arXiv Science⌕ Search

arXiv · 1212.0297

The Geometry of Differential Privacy: the Sparse and Approximate Cases

Abstract

In this work, we study trade-offs between accuracy and privacy in the context of linear queries over histograms. This is a rich class of queries that includes contingency tables and range queries, and has been a focus of a long line of work. For a set of $d$ linear queries over a database $x \in \R^N$, we seek to find the differentially private mechanism that has the minimum mean squared error. For pure differential privacy, an $O(\log^2 d)$ approximation to the optimal mechanism is known. Our first contribution is to give an $O(\log^2 d)$ approximation guarantee for the case of $(\eps,δ)$-differential privacy. Our mechanism is simple, efficient and adds correlated Gaussian noise to the answers. We prove its approximation guarantee relative to the hereditary discrepancy lower bound of Muthukrishnan and Nikolov, using tools from convex geometry. We next consider this question in the case when the number of queries exceeds the number of individuals in the database, i.e. when $d > n \triangleq \|x\|_1$. It is known that better mechanisms exist in this setting. Our second main contribution is to give an $(\eps,δ)$-differentially private mechanism which is optimal up to a $\polylog(d,N)$ factor for any given query set $A$ and any given upper bound $n$ on $\|x\|_1$. This approximation is achieved by coupling the Gaussian noise addition approach with a linear regression step. We give an analogous result for the $\eps$-differential privacy setting. We also improve on the mean squared error upper bound for answering counting queries on a database of size $n$ by Blum, Ligett, and Roth, and match the lower bound implied by the work of Dinur and Nissim up to logarithmic factors. The connection between hereditary discrepancy and the privacy mechanism enables us to derive the first polylogarithmic approximation to the hereditary discrepancy of a matrix $A$.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Aleksandar Nikolov, Kunal Talwar, Li Zhang. 2012-12-03. The Geometry of Differential Privacy: the Sparse and Approximate Cases. https://doi.org/10.1145/2488608.2488652

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Min-Sum Set Cover on Parallel Machines

We consider a generalization of the Min-Sum Set Cover to the setup with $m$ set-sequences, or in scheduling terminology, $m$ parallel machines. We call this problem Parallel Min-Sum Set Cover. To obtain approximation algorithms for its numerous variants we use a crucial sub-problem called Parallel Densest Subfamily. We prove that an $α$-approximation algorithm for this task gives a $4\cdotα$-approximation for the Parallel Min-Sum Set Cover, which yields $\frac{4\cdot e}{e-1}+ε$ and $4\cdot \frac{e}{e-1}^2+ε$-approximation ratios for identical and unrelated machines, respectively. To obtain the latter result we give a new $\frac{e}{e-1}^2+ε$-approximation algorithm for the Maximum Coverage Multiple Knapsacks problem which is of independent interest. If the sets are precedence-constrained, for unit cost sets we give an $\mathcal{O}(k^{2/3})$ approximation ($k$ is the number of sets). For the case of out-forest precedence constraints we improve this bound to $\mathcal{O}(\log k)$ via a reduction to the Group Steiner Orienteering problem, and show this is tight, unless $NP\subseteq ZTIME(n^{\mathcal{O}(\text{poly}(\log n))})$.

cs.DS↗

Learning Latent Algebraic Structure from Ambiguous Set Observations

We study when statistically learnable latent structure can also be recovered efficiently, and how membership queries change the answer. An unknown support $A\subseteq\mathbb F_2^n$ has small additive doubling and is observed through a fixed set $B$ satisfying $|A\triangle B|\leη|A|$. We seek one linear subspace $V$ such that every compatible support $A$ is covered by few $V$-cosets and satisfies $|V|\le|A|$. For every $η<1$, polynomially many uniform samples suffice statistically, with cost polynomial in the doubling constant and proportional to $(1-η)^{-1}$; this radius dependence is sharp. Under a specified hardness assumption for learning parities with noise (search-LPN), however, no polynomial-time sample-only learner achieves even constant covering cost, including when the latent support is unique. At fixed structural parameters and the same constant covering budget, adding exact membership queries to $B$ permits polynomial-time recovery. The general query learner constructs a short structural list and uses fresh samples to select one common output through a majority-coverage rule. Persistent structured cores make this candidate construction possible. At doubling one, a complementary distinction appears at $η=1/3$: coarse recovery remains polynomial time, while exact recovery requires exponentially many accesses in the worst case when latent cardinality is unknown.

cs.DS↗

Testing the Binary Rank with Polynomial Query Complexity

We design an adaptive two-sided error testing algorithm for the binary rank of a $0,1$ matrix $M$ with query complexity $O(d^3\log(d+1)/ε^2)$, where $d$ is the tested binary rank bound and $ε$ is the distance parameter. This answers an open question posed by Parnas, Ron and Shraibman~\cite{parnas2021property}, who asked if the binary rank can be tested with query complexity polynomial in $d$ and $1/ε$. Furthermore, our testing algorithm can be used to find an approximate binary decomposition of $M$ with an additional $d(n+m)$ queries. That is, under the promise that the binary rank of $M$ is at most $d$, we show how to find, with probability at least $5/6$, two $0,1$ matrices $A',B'$ such that $M' = A' \cdot B'$ is a $0,1$ matrix which differs from $M$ on at most an $O(ε)$ fraction of its entries. Our results also imply a testing algorithm with polynomial query complexity for the equivalent problem of testing if the edges of a bipartite graph can be partitioned into at most $d$ bicliques.

cs.DS↗