arXiv Science⌕ Search

arXiv · 2610.01007

Settling the Pass Complexity of Streaming Set Cover

Abstract

In the streaming set cover problem, $m$ sets from a universe of size $n$ are arriving one by one in a stream, and the algorithm is allowed to process the stream using one or a few passes and a space of $o(mn)$, which is sublinear in the input size. The goal is to determine the minimal (or approximately minimal) number of sets that cover the universe at the end of the last pass. This problem has been studied extensively over the years with rapid progress that led to several $O(\log{n})$-approximation algorithms in $\tilde{O}(mn^{1/p})$ space and $O(p)$ passes. However, progress on this front has largely stagnated over the past decade, despite the absence of any lower bounds that rule out even an $O(\log{n})$-approximation in $O(m)$ space and just two passes. We provide a simple explanation for this lack of progress by establishing an optimal three-way space-pass-approximation tradeoff for this problem: any $α$-approximation algorithm for streaming set cover requires $$ \widetildeΩ\Big(\frac{m}α \cdot \big(\frac{n}α\big)^{1/p}\Big) $$ space in $p$ passes whenever $α\ll n^{1/(p+1)}$. In light of prior work, this result is optimal up to constant factors in $p$ and logarithmic factors in $n,m$ for any $α\geq p$. Our bound is optimal with respect to the range of $α$ also, and fully settles the complexity of this fundamental problem in the streaming model. The proof of this result is (surprisingly) simple and non-technical and relies on a randomized reduction from a variant of the standard pointer chasing problem in communication complexity, using elementary properties of random sets.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Sepehr Assadi, Janani Sundaresan. 2026-10-01. Settling the Pass Complexity of Streaming Set Cover. https://arxiv.org/abs/2610.01007

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Instance-Optimality of Bidirectional PageRank Estimation

We study the problem of estimating a vertex's PageRank within a constant relative error, with constant probability. We prove that an adaptive variant of the simple classic bidirectional algorithm is instance-optimal up to a polylogarithmic factor for all directed graphs of order $n$ whose maximum in- and out-degrees are at most a constant fraction of $n$. In other words, there is no correct algorithm that can be faster than our algorithm on any such graph by more than a polylogarithmic factor. We further extend the instance-optimality to all graphs in which at most a polylogarithmic number of vertices have unbounded degrees. This covers all sparse graphs with $\tilde{O}(n)$ edges. In addition, we provide a counterexample showing that the bidirectional algorithm is not instance-optimal for graphs whose degrees are mostly equal to $n$. We also consider weighted graphs and multigraphs. We show that the bidirectional algorithm is instance-optimal on \emph{all} multigraphs, but for weighted simple graphs, we have almost the same limitations as for unweighted simple graphs.

cs.DS↗

Solving Hypergraph Laplacian Systems in Almost-Linear Time

For a connected weighted hypergraph, we give a randomized almost-linear-time solver for the Poisson problem for the cut-based hypergraph Laplacian in the natural input size $P=\sum_{e\in E}|e|$, the sum of hyperedge sizes. For every fixed constant $C>0$, our randomized algorithm runs in $P^{1+o(1)}$ time and, with high probability over its internal randomness, returns a primal point and a dual certificate, with additive optimality gap at most $\exp(-\log^C P)$. A key step is to rewrite the Fenchel dual as a convex-flow problem on an auxiliary $O(P)$-arc graph, yielding a near-optimal dual flow. The main difficulty is primal recovery, because this flow does not by itself determine a primal potential. Our main new ingredient is a recovery theorem showing that, for primal recovery, the detailed routing of the dual flow inside each hyperedge gadget can be discarded: one nonnegative scalar per hyperedge is enough. After the necessary finite-precision rounding, these scalars define a linear-cost min-cost-flow instance on the auxiliary graph, and solving it exactly recovers a primal potential. Finally, a ground-vertex reduction from regularized objectives to the Poisson solver gives randomized almost-linear-time resolvent/proximal primitives for the same cut-based hypergraph Laplacian.

cs.DS↗

Efficiently Listing Projected Trees, and Equivalence of Listing and Enumeration

The subgraph isomorphism problem and its generalizations, such as conjunctive queries where some nodes are projected, are among the most fundamental problems in graph algorithms and database theory. In this paper, we study the listing and enumeration variants of these problems and present two main results. The first result is an algorithm for enumerating projected trees with preprocessing time $\widetilde{O}(n^{17.42})$ and delay $\mathrm{polylog}(n)$. Prior to this work, for trees on $k$ nodes all algorithms in the literature required preprocessing time $n^{Ω(k)}$ or delay $n^{Ω(1)}$ or assumed $ω=2$. Our result generalizes to arbitrary projected hypergraphs, achieving enumeration in preprocessing time $\widetilde{O}(m^{17.42 \, \mathrm{subw}(H)})$ and polylogarithmic delay, where $\mathrm{subw}(H)$ is the submodular width of the pattern hypergraph $H$. We heavily rely on fast (rectangular and output-sensitive) matrix multiplication, which we complement by fine-grained lower bounds indicating that any algorithm beating preprocessing time $n^{Ω(k)}$ with polylogarithmic delay must rely on fast matrix multiplication. The second result is a generic enumeration-to-listing reduction, establishing that listing and enumeration are equivalent under natural assumptions. For (colored) subgraph isomorphism, our reduction transforms any listing algorithm running in time $O(f(n,m) + t \cdot g(n,m))$ into an enumeration algorithm with preprocessing time $O\left( (f(n,m)+g(n,m)+n+m) \log^2 n \right)$ and delay $O(g(n,m))$. We utilize this reduction to prove our first main result, and we expect that our generic reduction will find many future applications.

cs.DS↗