arXiv ScienceSearch

arXiv subjects

Peter Frazier

Publications and source records attributed to Peter Frazier.

14 recordsLinked to original sources

A geometric basis for materials families in inorganic solids

The thermodynamic stability of inorganic solids spans a vast compositional space, yet materials scientists have long organized their intuition around a manageable number of materials families. Here we show that this organization has a precise geometric basis. The formation-energy convex hull of all inorganic compounds from the Materials Project, spanning 92-dimensional elemental composition space, is captured to near DFT accuracy by a polyhedron with only seven facets. Each facet corresponds to a family of materials sharing similar chemical potentials. This low-dimensional structure is not merely an economical description of energies: without retraining or structural input, the same framework reproduces trends in DFT-calculated defect energies and elemental spatial correlations in high-entropy nanoparticles. These results reveal that a small number of material families, corresponding to geometric features of composition-energy space, govern bulk stability, defect energetics, and elemental mixing, and provide a unified, interpretable framework for rapid screening across diverse materials systems.

cond-mat.mtrl-sci

Experimentation Under Non-stationary Interference

We study the estimation of the ATE in randomized controlled trials under a dynamically evolving interference structure. This setting arises in applications such as ride-sharing, where drivers move over time, and social networks, where connections continuously form and dissolve. In particular, we focus on scenarios where outcomes exhibit spatio-temporal interference driven by a sequence of random interference graphs that evolve independently of the treatment assignment. Loosely, our main result states that a truncated Horvitz-Thompson estimator achieves an MSE that vanishes linearly in the number of spatial and time blocks, times a factor that measures the average complexity of the interference graphs. As a key technical contribution that contrasts the static setting we present a fine-grained covariance bound for each pair of space-time points that decays exponentially with the time elapsed since their last ``interaction''. Our results can be applied to many concrete settings and lead to simplified bounds, including where the interference graphs (i) are induced by moving points in a metric space, or (ii) follow a dynamic Erdos-Renyi model, where each edge is created or removed independently in each time period.

math.ST

Effective Atom Theory: Gradient-Driven ab initio Materials Design

We introduce Effective Atom Theory (EAT), a framework that transforms combinatorial materials design into a smooth, gradient-driven optimization within density functional theory (DFT). Atoms are represented as probabilistic mixtures of elements, enabling gradient-based optimizers to converge to a physically realizable material in about 50 energy evaluations -- far fewer than combinatorial optimization methods. Applied to Co-Cr-Ni-V oxides for the alkaline oxygen evolution reaction (OER), EAT leads to a final recommended composition of Co0.19Cr0.06V0.31Ni0.44O.

cond-mat.mtrl-sci

Multi-Armed Bandits with Interference

Experimentation with interference poses a significant challenge in contemporary online platforms. Prior research on experimentation with interference has concentrated on the final output of a policy. The cumulative performance, while equally crucial, is less well understood. To address this gap, we introduce the problem of {\em Multi-armed Bandits with Interference} (MABI), where the learner assigns an arm to each of $N$ experimental units over a time horizon of $T$ rounds. The reward of each unit in each round depends on the treatments of {\em all} units, where the influence of a unit decays in the spatial distance between units. Furthermore, we employ a general setup wherein the reward functions are chosen by an adversary and may vary arbitrarily across rounds and units. We first show that switchback policies achieve an optimal {\em expected} regret $\tilde O(\sqrt T)$ against the best fixed-arm policy. Nonetheless, the regret (as a random variable) for any switchback policy suffers a high variance, as it does not account for $N$. We propose a cluster randomization policy whose regret (i) is optimal in {\em expectation} and (ii) admits a high probability bound that vanishes in $N$.

cs.LG

Context, Composition, Automation, and Communication -- The C2AC Roadmap for Modeling and Simulation

Simulation has become, in many application areas, a sine-qua-non. Most recently, COVID-19 has underlined the importance of simulation studies and limitations in current practices and methods. We identify four goals of methodological work for addressing these limitations. The first is to provide better support for capturing, representing, and evaluating the context of simulation studies, including research questions, assumptions, requirements, and activities contributing to a simulation study. In addition, the composition of simulation models and other simulation studies' products must be supported beyond syntactical coherence, including aspects of semantics and purpose, enabling their effective reuse. A higher degree of automating simulation studies will contribute to more systematic, standardized simulation studies and their efficiency. Finally, it is essential to invest increased effort into effectively communicating results and the processes involved in simulation studies to enable their use in research and decision-making. These goals are not pursued independently of each other, but they will benefit from and sometimes even rely on advances in other subfields. In the present paper, we explore the basis and interdependencies evident in current research and practice and delineate future research directions based on these considerations.

cs.CE

Group Testing Enables Asymptomatic Screening for COVID-19 Mitigation: Feasibility and Optimal Pool Size Selection with Dilution Effects

Repeated asymptomatic screening for SARS-CoV-2 promises to control spread of the virus but would require too many resources to implement at scale. Group testing is promising for screening more people with fewer test resources: multiple samples tested together in one pool can be excluded with one negative test result. Existing approaches to group testing design for SARS-CoV-2 asymptomatic screening, however, do not consider dilution effects: that false negatives become more common with larger pools. As a consequence, they may recommend pool sizes that are too large or misestimate the benefits of screening. Modeling dilution effects, we derive closed-form expressions for the expected number of tests and false negative/positives per person screened under two popular group testing methods: the linear and square array methods. We find that test error correlation induced by a common viral load across an individual's samples results in many fewer false negatives than would be expected from less realistic but more widely assumed independent errors. This insight also suggests that false positives can be controlled through repeated tests without significantly increasing false negatives. Using these closed-form expressions to trace a Pareto frontier over error rates and tests, we design testing protocols for repeated asymptomatic screening of a large population. We minimize disease prevalence by optimizing a time-varying pool sizes and screening frequency constrained by daily test capacity and a false positive limit. This provides a testing protocol practitioners can use for mitigating COVID-19. In a case study, we demonstrate the effectiveness of this methodology in controlling spread.

q-bio.PE

Bayesian Optimization of Risk Measures

We consider Bayesian optimization of objective functions of the form $\rho[ F(x, W) ]$, where $F$ is a black-box expensive-to-evaluate function and $\rho$ denotes either the VaR or CVaR risk measure, computed with respect to the randomness induced by the environmental random variable $W$. Such problems arise in decision making under uncertainty, such as in portfolio optimization and robust systems design. We propose a family of novel Bayesian optimization algorithms that exploit the structure of the objective function to substantially improve sampling efficiency. Instead of modeling the objective function directly as is typical in Bayesian optimization, these algorithms model $F$ as a Gaussian process, and use the implied posterior on the objective function to decide which points to evaluate. We demonstrate the effectiveness of our approach in a variety of numerical experiments.

stat.ML

Matching Queues, Flexibility and Incentives

Problem definition: In many matching markets, some agents are fully flexible, while others only accept a subset of jobs. For example, ridesharing drivers can specify on the platform the destinations they are willing to accept. Conventional wisdom suggests reserving flexible agents, but this can backfire: anticipating higher matching chances, agents may misreport as specialized, reducing overall matches. We ask how platforms can design simple matching policies that remain effective when agents act strategically. Methodology/results: We model job allocation as a bipartite matching queueing system and analyze equilibrium throughput performance under different policies when agents choose which queue to join. We show that flexibility reservation is optimal under full information but can perform poorly with private information, sometimes substantially worse than random assignment. To address this, we propose a new policy -- flexibility reservation with fallback -- that guarantees robust performance across settings, without requiring precise knowledge of system parameters or agent utility functions. Managerial implications: Our results underscore the importance of accounting for strategic reporting in the design of matching policies: the proposed fallback policy both preserves flexibility and exploits latent flexibility when explicitly flexible agents are exhausted. Its simplicity and parameter-free nature also make it practical to implement in platforms such as ridesharing and affordable housing allocation.

cs.GT

Information Design in Spatial Resource Competition

We consider the information design problem in spatial resource competition settings. Agents gather at a location deciding whether to move to another location for possibly higher level of resources, and the utility each agent gets by moving to the other location decreases as more agents move there. The agents do not observe the resource level at the other location while a principal does and the principal would like to carefully release this information to attract a proper number of agents to move. We adopt the Bayesian persuasion framework and analyze the principal's optimal signaling mechanism design problem. We study both private and public signaling mechanisms. For private signaling, we show the optimal mechanism can be computed in polynomial time with respect to the number of agents. Obtaining the optimal private mechanism involves two steps: first, solve a linear program to get the marginal probability each agent should be recommended to move; second, sample the moving agents satisfying the marginal probabilities with a sequential sampling procedure. For public signaling, we show the sender preferred equilibrium has a simple threshold structure and the optimal public mechanism with respect to the sender preferred equilibrium can be computed in polynomial time. We support our analytical results with numerical computations that show the optimal private and public signaling mechanisms achieve substantially higher social welfare compared with no information or full information benchmarks in many settings.

cs.GT

Mean Field Equilibria for Resource Competition in Spatial Settings

We study a model of competition among nomadic agents for time-varying and location-specific resources, arising in crowd-sourced transportation services, online communities, and traditional location-based economic activity. This model comprises a group of agents and a single location endowed with a dynamic stochastic resource process. Periodically, each agent derives a reward determined by the location's resource level and the number of other agents there, and has to decide whether to stay at the location or move. Upon moving, the agent arrives at a different location whose dynamics are independent and identical to the original location. Using the methodology of mean field equilibrium, we study the equilibrium behavior of the agents as a function of the dynamics of the stochastic resource process and the nature of the competition among co-located agents. We show that an equilibrium exists, where each agent decides whether to switch locations based only on their current location's resource level and the number of other agents there. We additionally show that when an agent's payoff is decreasing in the number of other agents at her location, equilibrium strategies obey a simple threshold structure. We show how to exploit this structure to compute equilibria numerically, and use these numerical techniques to study how system structure affects the agents' collective ability to explore their domain to find and effectively utilize resource-rich areas.

cs.GT

An Asymptotically Optimal Index Policy for Finite-Horizon Restless Bandits

We consider restless multi-armed bandit (RMAB) with a finite horizon and multiple pulls per period. Leveraging the Lagrangian relaxation, we approximate the problem with a collection of single arm problems. We then propose an index-based policy that uses optimal solutions of the single arm problems to index individual arms, and offer a proof that it is asymptotically optimal as the number of arms tends to infinity. We also use simulation to show that this index-based policy performs better than the state-of-art heuristics in various problem settings.

math.OC

Unbiased Comparative Evaluation of Ranking Functions

Eliciting relevance judgments for ranking evaluation is labor-intensive and costly, motivating careful selection of which documents to judge. Unlike traditional approaches that make this selection deterministically, probabilistic sampling has shown intriguing promise since it enables the design of estimators that are provably unbiased even when reusing data with missing judgments. In this paper, we first unify and extend these sampling approaches by viewing the evaluation problem as a Monte Carlo estimation task that applies to a large number of common IR metrics. Drawing on the theoretical clarity that this view offers, we tackle three practical evaluation scenarios: comparing two systems, comparing $k$ systems against a baseline, and ranking $k$ systems. For each scenario, we derive an estimator and a variance-optimizing sampling distribution while retaining the strengths of sampling-based evaluation, including unbiasedness, reusability despite missing data, and ease of use in practice. In addition to the theoretical contribution, we empirically evaluate our methods against previously used sampling heuristics and find that they generally cut the number of required relevance judgments at least in half.

cs.IR

Mean Field Equilibria for Competitive Exploration in Resource Sharing Settings

We consider a model of nomadic agents exploring and competing for time-varying location-specific resources, arising in crowdsourced transportation services, online communities, and in traditional location based economic activity. This model comprises a group of agents, and a set of locations each endowed with a dynamic stochastic resource process. Each agent derives a periodic reward determined by the overall resource level at her location, and the number of other agents there. Each agent is strategic and free to move between locations, and at each time decides whether to stay at the same node or switch to another one. We study the equilibrium behavior of the agents as a function of dynamics of the stochastic resource process and the nature of the externality each agent imposes on others at the same location. In the asymptotic limit with the number of agents and locations increasing proportionally, we show that an equilibrium exists and has a threshold structure, where each agent decides to switch to a different location based only on their current location's resource level and the number of other agents at that location. This result provides insight into how system structure affects the agents' collective ability to explore their domain to find and effectively utilize resource-rich areas. It also allows assessing the impact of changing the reward structure through penalties or subsidies.

cs.GT

A Hierarchical Distance-dependent Bayesian Model for Event Coreference Resolution

We present a novel hierarchical distance-dependent Bayesian model for event coreference resolution. While existing generative models for event coreference resolution are completely unsupervised, our model allows for the incorporation of pairwise distances between event mentions -- information that is widely used in supervised coreference models to guide the generative clustering processing for better event clustering both within and across documents. We model the distances between event mentions using a feature-rich learnable distance function and encode them as Bayesian priors for nonparametric clustering. Experiments on the ECB+ corpus show that our model outperforms state-of-the-art methods for both within- and cross-document event coreference resolution.

cs.CL