arXiv ScienceSearch

arXiv subjects

Marc Coram

Publications and source records attributed to Marc Coram.

6 recordsLinked to original sources

An AI system to help scientists write expert-level empirical software

The cycle of scientific discovery is frequently bottlenecked by the slow, manual creation of software to support computational experiments\cite{hannay2009how}. To address this, we present Empirical Research Assistance (ERA), an AI system that creates expert-level scientific software whose goal is to maximize a quality metric. The system uses a Large Language Model (LLM) and Tree Search (TS)\cite{silver2016mastering} to systematically improve the quality metric and intelligently navigate the large space of possible solutions. ERA achieves expert-level results when it explores and integrates complex research ideas from external sources. The effectiveness of tree search is demonstrated across a diverse range of tasks. In bioinformatics, ERA discovered 40 novel methods for single-cell data analysis that outperformed the top human-developed methods on a public leaderboard. In epidemiology, ERA generated 14 models that outperformed the CDC ensemble and all other individual models for forecasting COVID-19 hospitalizations. ERA also produced expert-level software for geospatial analysis, neural activity prediction in zebrafish, and numerical solution of integrals, and a novel rule-based construction for time series forecasting. By devising and implementing novel solutions to diverse tasks, ERA represents a significant step towards accelerating scientific progress.

cs.AI

Quantum Optimization with a Novel Gibbs Objective Function and Ansatz Architecture Search

The Quantum Approximate Optimization Algorithm (QAOA) is a standard method for combinatorial optimization with a gate-based quantum computer. The QAOA consists of a particular ansatz for the quantum circuit architecture, together with a prescription for choosing the variational parameters of the circuit. We propose modifications to both. First, we define the Gibbs objective function and show that it is superior to the energy expectation value for use as an objective function in tuning the variational parameters. Second, we describe an Ansatz Architecture Search (AAS) algorithm for searching the discrete space of quantum circuit architectures near the QAOA to find a better ansatz. Applying these modifications for a complete graph Ising model results in a $244.7\%$ median relative improvement in the probability of finding a low-energy state while using $33.3\%$ fewer two-qubit gates. For Ising models on a 2d grid we similarly find $44.4\%$ median improvement in the probability with a $20.8\%$ reduction in the number of two-qubit gates. This opens a new research field of quantum circuit architecture design for quantum optimization algorithms.

quant-ph

A general spline representation for nonparametric and semiparametric density estimates using diffeomorphisms

A theorem of McCann shows that for any two absolutely continuous probability measures on R^d there exists a monotone transformation sending one probability measure to the other. A consequence of this theorem, relevant to statistics, is that density estimation can be recast in terms of transformations. In particular, one can fix any absolutely continuous probability measure, call it P, and then reparameterize the whole class of absolutely continuous probability measures as monotone transformations from P. In this paper we utilize this reparameterization of densities, as monotone transformations from some P, to construct semiparametric and nonparametric density estimates. We focus our attention on classes of transformations, developed in the image processing and computational anatomy literature, which are smooth, invertible and which have attractive computational properties. The techniques developed for this class of transformations allow us to show that a penalized maximum likelihood estimate (PMLE) of a smooth transformation from P exists and has a finite dimensional characterization, similar to those results found in the spline literature. These results are derived utilizing an Euler-Lagrange characterization of the PMLE which also establishes a surprising connection to a generalization of Stein's lemma for characterizing the normal distribution.

stat.ME

Two Dimensional Density Estimation using Smooth Invertible Transformations

We investigate the problem of estimating a smooth invertible transformation f when observing independent samples X_1, ..., X_n ~ P \circ f, where P is a known measure. We focus on the two dimensional case where P and f are defined on R^2. We present a flexible class of smooth invertible transformations in two dimensions with variational equations for optimizing over the classes, then study the problem of estimating the transformation f by penalized maximum likelihood estimation. We apply our methodology to the case when P \circ f has a density with respect to Lebesgue measure on R^2 and demonstrate improvements over kernel density estimation on three examples.

stat.ME

Improving population-specific allele frequency estimates by adapting supplemental data: an empirical Bayes approach

Estimation of the allele frequency at genetic markers is a key ingredient in biological and biomedical research, such as studies of human genetic variation or of the genetic etiology of heritable traits. As genetic data becomes increasingly available, investigators face a dilemma: when should data from other studies and population subgroups be pooled with the primary data? Pooling additional samples will generally reduce the variance of the frequency estimates; however, used inappropriately, pooled estimates can be severely biased due to population stratification. Because of this potential bias, most investigators avoid pooling, even for samples with the same ethnic background and residing on the same continent. Here, we propose an empirical Bayes approach for estimating allele frequencies of single nucleotide polymorphisms. This procedure adaptively incorporates genotypes from related samples, so that more similar samples have a greater influence on the estimates. In every example we have considered, our estimator achieves a mean squared error (MSE) that is smaller than either pooling or not, and sometimes substantially improves over both extremes. The bias introduced is small, as is shown by a simulation study that is carefully matched to a real data example. Our method is particularly useful when small groups of individuals are genotyped at a large number of markers, a situation we are likely to encounter in a genome-wide association study.

stat.AP

Consistency of Bayes estimators of a binary regression function

When do nonparametric Bayesian procedures ``overfit''? To shed light on this question, we consider a binary regression problem in detail and establish frequentist consistency for a certain class of Bayes procedures based on hierarchical priors, called uniform mixture priors. These are defined as follows: let $ν$ be any probability distribution on the nonnegative integers. To sample a function $f$ from the prior $π^ν$, first sample $m$ from $ν$ and then sample $f$ uniformly from the set of step functions from $[0,1]$ into $[0,1]$ that have exactly $m$ jumps (i.e., sample all $m$ jump locations and $m+1$ function values independently and uniformly). The main result states that if a data-stream is generated according to any fixed, measurable binary-regression function $f_0\not\equiv1/2$, then frequentist consistency obtains: that is, for any $ν$ with infinite support, the posterior of $π^ν$ concentrates on any $L^1$ neighborhood of $f_0$. Solution of an associated large-deviations problem is central to the consistency proof.

math.ST