arXiv Science⌕ Search

arXiv · 2609.37431

Query Complexity of Testing Structured Parenthesis Languages

Abstract

We study the query complexity of testing membership in structured string languages, focusing on Dyck languages and natural generalizations. A tester receives query access to a word and must distinguish valid inputs from words that are far in Hamming distance, while inspecting only a sublinear number of positions. Our results sharpen the boundary between constant-query testability and polynomial query complexity. First, we prove an $Ω(n^{2/5})$ lower bound for testing Dyck languages $D_m$ with any fixed number $m\ge2$ of parenthesis types, improving the previous $Ω(n^{1/5})$ lower bound of Fischer, Magniez, and Starikovskaya (SODA `18) and nearly matching their upper bound of $O(n^{2/5+o(1)})$. Furthermore, we show that all nonadaptive algorithms for these problems require $Ω(n^{1/2})$ queries. Our lower bounds use a Pólya-urn process to construct the hard distribution; Second, we identify a broad class of weighted-parenthesis languages, which we call {\em excursion languages,} that remain constant-query testable. These languages encode bounded-step walks that stay nonnegative and return to zero. For every fixed excursion language, we give a nonadaptive tester with query complexity $O(1/\varepsilon^2)$, and we prove this dependence on $\varepsilon$ is optimal, even for adaptive algorithms. As a special case, we obtain the tight $Θ(1/\varepsilon^2)$ query complexity of testing $D_1$, improving the previous $O(\log(1/\varepsilon)/\varepsilon^2)$ upper bound and giving the first matching two-sided-error lower bound. Third, we construct a simple hard language, Hidden String, that is generated by a deterministic linear grammar but nevertheless requires $Ω(n^{2/5})$ adaptive queries and $Ω(n^{1/2})$ nonadaptive queries to test. This shows that polynomial query complexity appears even for highly restricted string languages.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Tim Jackman, Diptaksho Palit, Sofya Raskhodnikova. 2026-09-29. Query Complexity of Testing Structured Parenthesis Languages. https://arxiv.org/abs/2609.37431

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Min-Sum Set Cover on Parallel Machines

We consider a generalization of the Min-Sum Set Cover to the setup with $m$ set-sequences, or in scheduling terminology, $m$ parallel machines. We call this problem Parallel Min-Sum Set Cover. To obtain approximation algorithms for its numerous variants we use a crucial sub-problem called Parallel Densest Subfamily. We prove that an $α$-approximation algorithm for this task gives a $4\cdotα$-approximation for the Parallel Min-Sum Set Cover, which yields $\frac{4\cdot e}{e-1}+ε$ and $4\cdot \frac{e}{e-1}^2+ε$-approximation ratios for identical and unrelated machines, respectively. To obtain the latter result we give a new $\frac{e}{e-1}^2+ε$-approximation algorithm for the Maximum Coverage Multiple Knapsacks problem which is of independent interest. If the sets are precedence-constrained, for unit cost sets we give an $\mathcal{O}(k^{2/3})$ approximation ($k$ is the number of sets). For the case of out-forest precedence constraints we improve this bound to $\mathcal{O}(\log k)$ via a reduction to the Group Steiner Orienteering problem, and show this is tight, unless $NP\subseteq ZTIME(n^{\mathcal{O}(\text{poly}(\log n))})$.

cs.DS↗

Learning Latent Algebraic Structure from Ambiguous Set Observations

We study when statistically learnable latent structure can also be recovered efficiently, and how membership queries change the answer. An unknown support $A\subseteq\mathbb F_2^n$ has small additive doubling and is observed through a fixed set $B$ satisfying $|A\triangle B|\leη|A|$. We seek one linear subspace $V$ such that every compatible support $A$ is covered by few $V$-cosets and satisfies $|V|\le|A|$. For every $η<1$, polynomially many uniform samples suffice statistically, with cost polynomial in the doubling constant and proportional to $(1-η)^{-1}$; this radius dependence is sharp. Under a specified hardness assumption for learning parities with noise (search-LPN), however, no polynomial-time sample-only learner achieves even constant covering cost, including when the latent support is unique. At fixed structural parameters and the same constant covering budget, adding exact membership queries to $B$ permits polynomial-time recovery. The general query learner constructs a short structural list and uses fresh samples to select one common output through a majority-coverage rule. Persistent structured cores make this candidate construction possible. At doubling one, a complementary distinction appears at $η=1/3$: coarse recovery remains polynomial time, while exact recovery requires exponentially many accesses in the worst case when latent cardinality is unknown.

cs.DS↗

Testing the Binary Rank with Polynomial Query Complexity

We design an adaptive two-sided error testing algorithm for the binary rank of a $0,1$ matrix $M$ with query complexity $O(d^3\log(d+1)/ε^2)$, where $d$ is the tested binary rank bound and $ε$ is the distance parameter. This answers an open question posed by Parnas, Ron and Shraibman~\cite{parnas2021property}, who asked if the binary rank can be tested with query complexity polynomial in $d$ and $1/ε$. Furthermore, our testing algorithm can be used to find an approximate binary decomposition of $M$ with an additional $d(n+m)$ queries. That is, under the promise that the binary rank of $M$ is at most $d$, we show how to find, with probability at least $5/6$, two $0,1$ matrices $A',B'$ such that $M' = A' \cdot B'$ is a $0,1$ matrix which differs from $M$ on at most an $O(ε)$ fraction of its entries. Our results also imply a testing algorithm with polynomial query complexity for the equivalent problem of testing if the edges of a bipartite graph can be partitioned into at most $d$ bicliques.

cs.DS↗