arXiv ScienceSearch

arXiv subjects

Adam Smith

Publications and source records attributed to Adam Smith.

At least 55 records · Page 3Linked to original sources

Covariance-Aware Private Mean Estimation Without Private Covariance Estimation

We present two sample-efficient differentially private mean estimators for $d$-dimensional (sub)Gaussian distributions with unknown covariance. Informally, given $n \gtrsim d/α^2$ samples from such a distribution with mean $μ$ and covariance $Σ$, our estimators output $\tildeμ$ such that $\| \tildeμ- μ\|_Σ \leq α$, where $\| \cdot \|_Σ$ is the Mahalanobis distance. All previous estimators with the same guarantee either require strong a priori bounds on the covariance matrix or require $Ω(d^{3/2})$ samples. Each of our estimators is based on a simple, general approach to designing differentially private mechanisms, but with novel technical steps to make the estimator private and sample-efficient. Our first estimator samples a point with approximately maximum Tukey depth using the exponential mechanism, but restricted to the set of points of large Tukey depth. Its accuracy guarantees hold even for data sets that have a small amount of adversarial corruption. Proving that this mechanism is private requires a novel analysis. Our second estimator perturbs the empirical mean of the data set with noise calibrated to the empirical covariance, without releasing the covariance itself. Its sample complexity guarantees hold more generally for subgaussian distributions, albeit with a slightly worse dependence on the privacy parameter. For both estimators, careful preprocessing of the data is required to satisfy differential privacy.

cs.LG

Private Gradient Descent for Linear Regression: Tighter Error Bounds and Instance-Specific Uncertainty Estimation

We provide an improved analysis of standard differentially private gradient descent for linear regression under the squared error loss. Under modest assumptions on the input, we characterize the distribution of the iterate at each time step. Our analysis leads to new results on the algorithm's accuracy: for a proper fixed choice of hyperparameters, the sample complexity depends only linearly on the dimension of the data. This matches the dimension-dependence of the (non-private) ordinary least squares estimator as well as that of recent private algorithms that rely on sophisticated adaptive gradient-clipping schemes (Varshney et al., 2022; Liu et al., 2023). Our analysis of the iterates' distribution also allows us to construct confidence intervals for the empirical optimizer which adapt automatically to the variance of the algorithm on a particular data set. We validate our theorems through experiments on synthetic data.

cs.LG

Entanglement Transitions in Unitary Circuit Games

Repeated projective measurements in unitary circuits can lead to an entanglement phase transition as the measurement rate is tuned. In this work, we consider a different setting in which the projective measurements are replaced by dynamically chosen unitary gates that minimize the entanglement. This can be seen as a one-dimensional unitary circuit game in which two players get to place unitary gates on randomly assigned bonds at different rates: The "entangler" applies a random local unitary gate with the aim of generating extensive (volume law) entanglement. The "disentangler," based on limited knowledge about the state, chooses a unitary gate to reduce the entanglement entropy on the assigned bond with the goal of limiting to only finite (area law) entanglement. In order to elucidate the resulting entanglement dynamics, we consider three different scenarios: (i) a classical discrete height model, (ii) a Clifford circuit, and (iii) a general $U(4)$ unitary circuit. We find that both the classical and Clifford circuit models exhibit phase transitions as a function of the rate that the disentangler places a gate, which have similar properties that can be understood through a connection to the stochastic Fredkin chain. In contrast, the "entangler" always wins when using Haar random unitary gates and we observe extensive, volume law entanglement for all non-zero rates of entangling.

quant-ph

Analogue Quantum Simulation with Fixed-Frequency Transmon Qubits

We experimentally assess the suitability of transmon qubits with fixed frequencies and fixed interactions for the realization of analogue quantum simulations of spin systems. We test a set of necessary criteria for this goal on a commercial quantum processor using full quantum process tomography and more efficient Hamiltonian tomography. Significant single qubit errors at low amplitudes are identified as a limiting factor preventing the realization of analogue simulations on currently available devices. We additionally find spurious dynamics in the absence of drive pulses, which we identify with coherent coupling between the qubit and a low dimensional environment. With moderate improvements, analogue simulation of a rich family of time-dependent many-body spin Hamiltonians may be possible.

quant-ph

A rigorous benchmarking of methods for SARS-CoV-2 lineage abundance estimation in wastewater

In light of the continuous transmission and evolution of SARS-CoV-2 coupled with a significant decline in clinical testing, there is a pressing need for scalable, cost-effective, long-term, passive surveillance tools to effectively monitor viral variants circulating in the population. Wastewater genomic surveillance of SARS-CoV-2 has arrived as an alternative to clinical genomic surveillance, allowing to continuously monitor the prevalence of viral lineages in communities of various size at a fraction of the time, cost, and logistic effort and serving as an early warning system for emerging variants, critical for developed communities and especially for underserved ones. Importantly, lineage prevalence estimates obtained with this approach aren't distorted by biases related to clinical testing accessibility and participation. However, the relative performance of bioinformatics methods used to measure relative lineage abundances from wastewater sequencing data is unknown, preventing both the research community and public health authorities from making informed decisions regarding computational tool selection. Here, we perform comprehensive benchmarking of 18 bioinformatics methods for estimating the relative abundance of SARS-CoV-2 (sub)lineages in wastewater by using data from 36 in vitro mixtures of synthetic lineage and sublineage genomes. In addition, we use simulated data from 78 mixtures of lineages and sublineages co-occurring in the clinical setting with proportions mirroring their prevalence ratios observed in real data. Importantly, we investigate how the accuracy of the evaluated methods is impacted by the sequencing technology used, the associated error rate, the read length, read depth, but also by the exposure of the synthetic RNA mixtures to wastewater, with the goal of capturing the effects induced by the wastewater matrix, including RNA fragmentation and degradation.

q-bio.GN

Control, Confidentiality, and the Right to be Forgotten

Recent digital rights frameworks give users the right to delete their data from systems that store and process their personal information (e.g., the "right to be forgotten" in the GDPR). How should deletion be formalized in complex systems that interact with many users and store derivative information? We argue that prior approaches fall short. Definitions of machine unlearning Cao and Yang [2015] are too narrowly scoped and do not apply to general interactive settings. The natural approach of deletion-as-confidentiality Garg et al. [2020] is too restrictive: by requiring secrecy of deleted data, it rules out social functionalities. We propose a new formalism: deletion-as-control. It allows users' data to be freely used before deletion, while also imposing a meaningful requirement after deletion--thereby giving users more control. Deletion-as-control provides new ways of achieving deletion in diverse settings. We apply it to social functionalities, and give a new unified view of various machine unlearning definitions from the literature. This is done by way of a new adaptive generalization of history independence. Deletion-as-control also provides a new approach to the goal of machine unlearning, that is, to maintaining a model while honoring users' deletion requests. We show that publishing a sequence of updated models that are differentially private under continual release satisfies deletion-as-control. The accuracy of such an algorithm does not depend on the number of deleted points, in contrast to the machine unlearning literature.

cs.CR

CD-27 11535: Evidence for a Triple System in the $β$ Pictoris Moving Group

We present new spatially resolved astrometry and photometry of the CD-27 11535 system, a member of the $β$ Pictoris moving group consisting of two resolved K-type stars on a $\sim$20-year orbit. We fit an orbit to relative astrometry measured from NIRC2, GPI, and archival NaCo images, in addition to literature measurements. However, the total mass inferred from this orbit is significantly discrepant from that inferred from stellar evolutionary models using the luminosity of the two stars. We explore two hypotheses that could explain this discrepant mass sum; a discrepant parallax measurement from Gaia due to variability, and the presence of an additional unresolved companion to one of the two components. We find that the $\sim$20-year orbit could not bias the parallax measurement, but that variability of the components could produce a large amplitude astrometric motion, an effect which cannot be quantified exactly without the individual Gaia measurements. The discrepancy could also be explained by an additional star in the system. We jointly fit the astrometric and photometric measurements of the system to test different binary and triple architectures for the system. Depending on the set of evolutionary models used, we find an improved goodness of fit for a triple system architecture that includes a low-mass ($M=0.177\pm0.055$\,$M_{\odot}$) companion to the primary star. Further studies of this system will be required in order to resolve this discrepancy, either by refining the parallax measurement with a more complex treatment of variability-induced astrometric motion, or by detecting a third companion.

astro-ph.SR

Time Evolution of Uniform Sequential Circuits

Simulating time evolution of generic quantum many-body systems using classical numerical approaches has an exponentially growing cost either with evolution time or with the system size. In this work, we present a polynomially scaling hybrid quantum-classical algorithm for time evolving a one-dimensional uniform system in the thermodynamic limit. This algorithm uses a layered uniform sequential quantum circuit as a variational ansatz to represent infinite translation-invariant quantum states. We show numerically that this ansatz requires a number of parameters polynomial in the simulation time for a given accuracy. Furthermore, this favourable scaling of the ansatz is maintained during our variational evolution algorithm. All steps of the hybrid optimization are designed with near-term digital quantum computers in mind. After benchmarking the evolution algorithm on a classical computer, we demonstrate the measurement of observables of this uniform state using a finite number of qubits on a cloud-based quantum processing unit. With more efficient tensor contraction schemes, this algorithm may also offer improvements as a classical numerical algorithm.

quant-ph

Model-Independent Learning of Quantum Phases of Matter with Quantum Convolutional Neural Networks

Quantum convolutional neural networks (QCNNs) have been introduced as classifiers for gapped quantum phases of matter. Here, we propose a model-independent protocol for training QCNNs to discover order parameters that are unchanged under phase-preserving perturbations. We initiate the training sequence with the fixed-point wavefunctions of the quantum phase and then add translation-invariant noise that respects the symmetries of the system to mask the fixed-point structure on short length scales. We illustrate this approach by training the QCNN on phases protected by time-reversal symmetry in one dimension, and test it on several time-reversal symmetric models exhibiting trivial, symmetry-breaking, and symmetry-protected topological order. The QCNN discovers a set of order parameters that identifies all three phases and accurately predicts the location of the phase boundary. The proposed protocol paves the way towards hardware-efficient training of quantum phase classifiers on a programmable quantum processor.

quant-ph

Node-Differentially Private Estimation of the Number of Connected Components

We design the first node-differentially private algorithm for approximating the number of connected components in a graph. Given a database representing an $n$-vertex graph $G$ and a privacy parameter $\varepsilon$, our algorithm runs in polynomial time and, with probability $1-o(1)$, has additive error $\widetilde{O}(\frac{Δ^*\ln\ln n}{\varepsilon}),$ where $Δ^*$ is the smallest possible maximum degree of a spanning forest of $G.$ Node-differentially private algorithms are known only for a small number of database analysis tasks. A major obstacle for designing such an algorithm for the number of connected components is that this graph statistic is not robust to adding one node with arbitrary connections (a change that node-differential privacy is designed to hide): every graph is a neighbor of a connected graph. We overcome this by designing a family of efficiently computable Lipschitz extensions of the number of connected components or, equivalently, the size of a spanning forest. The construction of the extensions, which is at the core of our algorithm, is based on the forest polytope of $G.$ We prove several combinatorial facts about spanning forests, in particular, that a graph with no induced $Δ$-stars has a spanning forest of degree at most $Δ$. With this fact, we show that our Lipschitz extensions for the number of connected components equal the true value of the function for the largest possible monotone families of graphs. More generally, on all monotone sets of graphs, the $\ell_\infty$ error of our Lipschitz extensions is nearly optimal.

cs.DS

Numerical simulation of non-abelian anyons

Two-dimensional systems such as quantum spin liquids or fractional quantum Hall systems exhibit anyonic excitations that possess more general statistics than bosons or fermions. This exotic statistics makes it challenging to solve even a many-body system of non-interacting anyons. We introduce an algorithm that allows to simulate anyonic tight-binding Hamiltonians on two-dimensional lattices. The algorithm is directly derived from the low energy topological quantum field theory and is suited for general abelian and non-abelian anyon models. As concrete examples, we apply the algorithm to study the energy level spacing statistics, which reveals level repulsion for free semions, Fibonacci anyons and Ising anyons. Additionally, we simulate non-equilibrium quench dynamics, where we observe that the density distribution becomes homogeneous for large times - indicating thermalization.

cond-mat.str-el

On the `Semantics' of Differential Privacy: A Bayesian Formulation

Differential privacy is a definition of "privacy'" for algorithms that analyze and publish information about statistical databases. It is often claimed that differential privacy provides guarantees against adversaries with arbitrary side information. In this paper, we provide a precise formulation of these guarantees in terms of the inferences drawn by a Bayesian adversary. We show that this formulation is satisfied by both "vanilla" differential privacy as well as a relaxation known as (epsilon,delta)-differential privacy. Our formulation follows the ideas originally due to Dwork and McSherry [Dwork 2006]. This paper is, to our knowledge, the first place such a formulation appears explicitly. The analysis of the relaxed definition is new to this paper, and provides some concrete guidance for setting parameters when using (epsilon,delta)-differential privacy.

cs.CR

Improved Differential Privacy for SGD via Optimal Private Linear Operators on Adaptive Streams

Motivated by recent applications requiring differential privacy over adaptive streams, we investigate the question of optimal instantiations of the matrix mechanism in this setting. We prove fundamental theoretical results on the applicability of matrix factorizations to adaptive streams, and provide a parameter-free fixed-point algorithm for computing optimal factorizations. We instantiate this framework with respect to concrete matrices which arise naturally in machine learning, and train user-level differentially private models with the resulting optimal mechanisms, yielding significant improvements in a notable problem in federated learning with user-level differential privacy.

cs.LG

Differentially Private Sampling from Distributions

We initiate an investigation of private sampling from distributions. Given a dataset with $n$ independent observations from an unknown distribution $P$, a sampling algorithm must output a single observation from a distribution that is close in total variation distance to $P$ while satisfying differential privacy. Sampling abstracts the goal of generating small amounts of realistic-looking data. We provide tight upper and lower bounds for the dataset size needed for this task for three natural families of distributions: arbitrary distributions on $\{1,\ldots ,k\}$, arbitrary product distributions on $\{0,1\}^d$, and product distributions on $\{0,1\}^d$ with bias in each coordinate bounded away from 0 and 1. We demonstrate that, in some parameter regimes, private sampling requires asymptotically fewer observations than learning a description of $P$ nonprivately; in other regimes, however, private sampling proves to be as difficult as private learning. Notably, for some classes of distributions, the overhead in the number of observations needed for private learning compared to non-private learning is completely captured by the number of observations needed for private sampling.

cs.LG

Fully Adaptive Composition for Gaussian Differential Privacy

We show that Gaussian Differential Privacy, a variant of differential privacy tailored to the analysis of Gaussian noise addition, composes gracefully even in the presence of a fully adaptive analyst. Such an analyst selects mechanisms (to be run on a sensitive data set) and their privacy budgets adaptively, that is, based on the answers from other mechanisms run previously on the same data set. In the language of Rogers, Roth, Ullman and Vadhan, this gives a filter for GDP with the same parameters as for nonadaptive composition.

cs.CR

Instance-Optimal Differentially Private Estimation

In this work, we study local minimax convergence estimation rates subject to $ε$-differential privacy. Unlike worst-case rates, which may be conservative, algorithms that are locally minimax optimal must adapt to easy instances of the problem. We construct locally minimax differentially private estimators for one-parameter exponential families and estimating the tail rate of a distribution. In these cases, we show that optimal algorithms for simple hypothesis testing, namely the recent optimal private testers of Canonne et al. (2019), directly inform the design of locally minimax estimation algorithms.

math.ST

Methods for simulating string-net states and anyons on a digital quantum computer

Finding physical realizations of topologically ordered states in experimental settings, from condensed matter to artificial quantum systems, has been the main challenge en route to utilizing their unconventional properties. We show how to realize a large class of topologically ordered states and simulate their quasiparticle excitations on a digital quantum computer. To achieve this we design a set of linear-depth quantum circuits to generate ground states of general string-net models together with unitary open string operators to simulate the creation and braiding of abelian and non-abelian anyons. We show that the abelian (non-abelian) unitary string operators can be implemented with a constant (linear) depth quantum circuit. Our scheme allows us to directly probe characteristic topological properties, including topological entanglement entropy, braiding statistics, and fusion channels of anyons. Moreover, this set of efficiently prepared topologically ordered states has potential applications in the development of fault-tolerant quantum computers.

quant-ph

Finite-depth scaling of infinite quantum circuits for quantum critical points

The scaling of the entanglement entropy at a quantum critical point allows us to extract universal properties of the state, e.g., the central charge of a conformal field theory. With the rapid improvement of noisy intermediate-scale quantum (NISQ) devices, these quantum computers present themselves as a powerful tool to study critical many-body systems. We use finite-depth quantum circuits suitable for NISQ devices as a variational ansatz to represent ground states of critical, infinite systems. We find universal finite-depth scaling relations for these circuits and verify them numerically at two different critical points, i.e., the critical Ising model with an additional symmetry-preserving term and the critical XXZ model.

cond-mat.str-el