arXiv ScienceSearch

arXiv subjects

Youngtak Sohn

Publications and source records attributed to Youngtak Sohn.

At least 19 recordsLinked to original sources

Universality and sharp thresholds for ellipsoid fitting

We establish a sharp phase transition for fitting random vectors by an ellipsoid. The random vectors have independent subgaussian coordinates with mean zero, variance one, and a common fourth moment, and the number of vectors is proportional to the square of the dimension. We identify an explicit satisfiability threshold such that, with high probability, a positive definite ellipsoid passes through every data point below the threshold, whereas no positive semidefinite fit exists above it. We also determine the optimal squared fitting error throughout the unsatisfiable regime. In particular, the threshold depends on the coordinate distributions only through their common fourth moment, revealing a fourth moment universality phenomenon. For standard Gaussian data the threshold is $1/4$, resolving the ellipsoid fitting conjecture.

math.PR

Stochastic block models with many communities and the Kesten--Stigum bound

We study the inference of communities in stochastic block models with a growing number of communities. For block models with $n$ vertices and a fixed number of communities $q$, it was predicted in Decelle et al. (2011) that there are computationally efficient algorithms for recovering the communities above the Kesten--Stigum (KS) bound and that efficient recovery is impossible below the KS bound. This conjecture has since stimulated a lot of interest, with the achievability side proven in a line of research that culminated in the work of Abbe and Sandon (2018). Conversely, recent work by Sohn and Wein (2025) provides evidence for the hardness part using the low-degree paradigm. In this paper we investigate community recovery in the regime $q=q_n \to \infty$ as $n\to\infty$ where no such predictions exist. We show that efficient inference of communities remains possible above the KS bound. Furthermore, we show that recovery of block models is low-degree hard below the KS bound when the number of communities satisfies $q\ll \sqrt{n}$. Perhaps surprisingly, we find that when $q \gg \sqrt{n}$, there is an efficient algorithm based on non-backtracking walks for recovery even below the KS bound. We identify a new threshold and ask if it is the threshold for efficient recovery in this regime. Finally, we show that detection is easy and identify (up to a constant) the information-theoretic threshold for community recovery as the number of communities $q$ diverges. Our low-degree hardness results also naturally have consequences for graphon estimation, improving results of Luo and Gao (2024).

math.PR

Low-degree estimation thresholds in planted hypergraphs and tensor PCA

A central question in high-dimensional statistics is to understand statistical--computational gaps: regimes in which recovering a hidden signal is information-theoretically possible but conjectured to be computationally intractable. The low-degree framework offers a concrete way to study this gap by restricting attention to estimators that are polynomials of degree at most $D$ in the observed data. In this paper, we study low-degree estimation in planted dense subhypergraph, sparse tensor PCA, and tensor PCA with a general prior. For the planted dense subhypergraph model on $n$ vertices, we identify two regimes depending on whether the planted set is larger or smaller than $\sqrt{n}$. Above this scale, we identify a sharp threshold for low-degree estimation. Below this scale, we establish hardness in the regimes predicted by prior work, thereby resolving a question of Schramm and Wein (2022) and Sohn and Wein (2025). For sparse tensor PCA, we identify an analogous sharp phase transition. For tensor PCA with a general prior, we prove a low-degree estimation lower bound at the critical signal scale, matching the degree--signal tradeoff suggested by prior work. Our lower bounds apply to degree $D=n^δ$, where $n$ is the dimension and $δ>0$ is a constant, and we complement them with corresponding low-degree upper bounds. In addition, for planted dense subhypergraph and sparse tensor PCA above the $\sqrt{n}$ scale, we convert our upper bounds into polynomial-time algorithms that achieve almost exact recovery above the sharp threshold, yielding polynomial-time algorithms succeeding up to this threshold. Our proofs extend the framework of Sohn and Wein (2025) through a conditional variant that yields the correct signal-to-noise ratio in settings where the unconditional approach is insufficient.

math.ST

Sharp Phase Transitions in Estimation with Low-Degree Polynomials

High-dimensional planted problems, such as finding a hidden dense subgraph within a random graph, often exhibit a gap between statistical and computational feasibility. While recovering the hidden structure may be statistically possible, it is conjectured to be computationally intractable in certain parameter regimes. A powerful approach to understanding this hardness involves proving lower bounds on the efficacy of low-degree polynomial algorithms. We introduce new techniques for establishing such lower bounds, leading to novel results across diverse settings: planted submatrix, planted dense subgraph, the spiked Wigner model, and the stochastic block model. Notably, our results address the estimation task -- whereas most prior work is limited to hypothesis testing -- and capture sharp phase transitions such as the "BBP" transition in the spiked Wigner model (named for Baik, Ben Arous, and Péché) and the Kesten-Stigum threshold in the stochastic block model. Existing work on estimation either falls short of achieving these sharp thresholds or is limited to polynomials of very low (constant or logarithmic) degree. In contrast, our results rule out estimation with polynomials of degree $n^δ$ where $n$ is the dimension and $δ> 0$ is a constant, and in some cases we pin down the optimal constant $δ$. Our work resolves open problems posed by Hopkins & Steurer (2017) and Schramm & Wein (2022), and provides rigorous support within the low-degree framework for conjectures by Abbe & Sandon (2018) and Lelarge & Miolane (2019).

math.ST

Fast mixing in Ising models with a negative spectral outlier via Gaussian approximation

We study the mixing time of Glauber dynamics for Ising models in which the interaction matrix contains a single negative spectral outlier. This class includes the anti-ferromagnetic Curie-Weiss model, the anti-ferromagnetic Ising model on expander graphs, and the Sherrington-Kirkpatrick model with disorder of negative mean. Existing approaches to rapid mixing rely crucially on log-concavity or spectral width bounds and therefore can break down in the presence of a negative outlier. To address this difficulty, we develop a new covariance approximation method based on Gaussian approximation. This method is implemented via an iterative application of Stein's method to quadratic tilts of sums of bounded random variables, which may be of independent interest. The resulting analysis provides an operator-norm control of the full correlation structure under arbitrary external fields. Combined with the localization schemes of Eldan and Chen, these estimates lead to a modified logarithmic Sobolev inequality and near-optimal mixing time bounds in regimes where spectral width bounds fail. We complement these results by proving exponential lower bounds on the mixing time for low temperature anti-ferromagnetic Ising models on sparse random regular graphs and Erdös-Rényi graphs, based on the existence of gapped states as in the recent work of Sellke.

math.PR

Balanced multi-species spin glasses

We identify a special class of multi-species spin glass models: ones in which the species proportions serve to ''balance'' out the interaction strengths. For this class, we prove a free energy lower bound that does not require any convexity assumption, and applies to both Ising and spherical models. The lower bound is the free energy of a single-species model whose $p$-spin inverse-temperature parameter is exactly the variance of the $p$-spin component of the multi-species Hamiltonian. For the Ising case, this generalizes an inequality recently found by Issa in the context of vector spin models. We further demonstrate that this lower bound is actually an equality in many cases, including: at high temperatures for all models, at all temperatures for convex models, and at zero temperature for pure bipartite spherical models. When translated to a statement about the injective norm of a nonsymmetric Gaussian tensor, our lower bound matches an upper bound recently established by Dartois and McKenna.

math.PR

Crisanti-Sommers formula and simultaneous symmetry breaking in multi-species spherical spin glasses

There is a rich history of expressing the limiting free energy of mean-field spin glasses as a variational formula over probability measures on $[0,1]$, where the measure represents the similarity (or "overlap") of two independently sampled spin configurations. At high temperatures, the formula's minimum is achieved at a measure which is a point mass, meaning sample configurations are asymptotically orthogonal up to a magnetic field correction. At low temperatures, though, a very different behavior emerges known as replica symmetry breaking (RSB). The deep wells in the energy landscape create more rigid structure, and the optimal overlap measure is no longer a point mass. The exact size of its support remains in many cases an open problem. Here we consider these themes for multi-species spherical spin glasses. Following a companion work in which we establish the Parisi variational formula, here we present this formula's Crisanti-Sommers representation. In the process, we gain new access to a problem unique to the multi-species setting. Namely, if RSB occurs for one species, does it necessarily occur for other species as well? We provide sufficient conditions for the answer to be yes. For instance, we show that if two species share any quadratic interaction, then RSB for one implies RSB for the other. Moreover, the level of symmetry breaking must be identical, even in cases of full RSB. In the presence of an external field, any type of interaction suffices.

math-ph

Replica symmetry breaking in multi-species Sherrington-Kirkpatrick model

In the Sherrington-Kirkpatrick (SK) and related mixed $p$-spin models, there is interest in understanding replica symmetry breaking at low temperatures. For this reason, the so-called AT line proposed by de Almeida and Thouless as a sufficient (and conjecturally necessary) condition for symmetry breaking, has been a frequent object of study in spin glass theory. In this paper, we consider the analogous condition for the multi-species SK model, which concerns the eigenvectors of a Hessian matrix. The analysis is tractable in the two-species case with positive definite variance structure, for which we derive an explicit AT temperature threshold. To our knowledge, this is the first non-asymptotic symmetry breaking condition produced for a multi-species spin glass. As possible evidence that the condition is sharp, we draw further parallel with the classical SK model and show coincidence with a separate temperature inequality guaranteeing uniqueness of the replica symmetric critical point.

math.PR

Free energy in multi-species mixed $p$-spin spherical models

We prove a Parisi formula for the limiting free energy of multi-species spherical spin glasses with mixed $p$-spin interactions. The upper bound involves a Guerra-style interpolation and requires a convexity assumption on the model's covariance function. Meanwhile, the lower bound adapts the cavity method of Chen so that it can be combined with the synchronization technique of Panchenko; this part requires no convexity assumption. In order to guarantee that the resulting Parisi formula has a minimizer, we formalize the pairing of synchronization maps with overlap measures so that the constraint set is a compact metric space. This space is not related to the model's spherical structure and can be carried over to other multi-species settings.

math.PR

Rapid phase ordering for Ising and Potts dynamics on random regular graphs

We consider the Ising, and more generally, $q$-state Potts Glauber dynamics on random $d$-regular graphs on $n$ vertices at low temperatures $β\gtrsim \frac{\log d}{d}$. The mixing time is exponential in $n$ due to a bottleneck between $q$ dominant phases consisting of configurations in which the majority of vertices are in the same state. We prove that for any $d\ge 7$, from biased initializations with $ε_d n$ more vertices in state-$1$ than in other states, the Glauber dynamics quasi-equilibrates to the stationary distribution conditioned on having plurality in state-$1$ in optimal $O(\log n)$ time. Moreover, the requisite initial bias $ε_d$ can be taken to zero as $d \to \infty$. Even for the $q=2$ Ising case, where the states are naturally identified with $\pm 1$, proving such a result requires a new approach in order to control negative information spread in spacetime despite the model being in low temperature and exhibiting strong local correlations. For this purpose, we introduce a coupled non-Markovian rigid dynamics for which a delicate temporal recursion on probability mass functions of minus spacetime cluster sizes establishes their subcriticality.

math.PR

Exact Phase Transitions for Stochastic Block Models and Reconstruction on Trees

In this paper we continue to rigorously establish the predictions in ground breaking work in statistical physics by Decelle, Krzakala, Moore, Zdeborová (2011) regarding the block model, in particular in the case of $q=3$ and $q=4$ communities. We prove that for $q=3$ and $q=4$ there is no computational-statistical gap if the average degree is above some constant by showing it is information theoretically impossible to detect below the Kesten-Stigum bound. The proof is based on showing that for the broadcast process on Galton-Watson trees, reconstruction is impossible for $q=3$ and $q=4$ if the average degree is sufficiently large. This improves on the result of Sly (2009), who proved similar results for regular trees for $q=3$. Our analysis of the critical case $q=4$ provides a detailed picture showing that the tightness of the Kesten-Stigum bound in the antiferromagnetic case depends on the average degree of the tree. We also prove that for $q\geq 5$, the Kestin-Stigum bound is not sharp. Our results prove conjectures of Decelle, Krzakala, Moore, Zdeborová (2011), Moore (2017), Abbe and Sandon (2018) and Ricci-Tersenghi, Semerjian, and Zdeborová (2019). Our proofs are based on a new general coupling of the tree and graph processes and on a refined analysis of the broadcast process on the tree.

math.PR

Local geometry of NAE-SAT solutions in the condensation regime

The local behavior of typical solutions of random constraint satisfaction problems (CSP) describes many important phenomena including clustering thresholds, decay of correlations, and the behavior of message passing algorithms. When the constraint density is low, studying the planted model is a powerful technique for determining this local behavior which in many examples has a simple Markovian structure. The work of Coja-Oghlan, Kapetanopoulos, Müller (2020) showed that for a wide class of models, this description applies up to the so-called condensation threshold. Understanding the local behavior after the condensation threshold is more complex due to long-range correlations. In this work, we revisit the random regular NAE-SAT model in the condensation regime and determine the local weak limit which describes a random solution around a typical variable. This limit exhibits a complicated non-Markovian structure arising from the space of solutions being dominated by a small number of large clusters. This is the first description of the local weak limit in the condensation regime for any sparse random CSPs in the one-step replica symmetry breaking (1RSB) class. Our result is non-asymptotic, and characterizes the tight fluctuation $O(n^{-1/2})$ around the limit. Our proof is based on coupling the local neighborhoods of an infinite spin system, which encodes the structure of the clusters, to a broadcast model on trees whose channel is given by the 1RSB belief-propagation fixed point. We believe that our proof technique has broad applicability to random CSPs in the 1RSB class.

math.PR

Weak recovery, hypothesis testing, and mutual information in stochastic block models and planted factor graphs

The stochastic block model is a canonical model of communities in random graphs. It was introduced in the social sciences and statistics as a model of communities, and in theoretical computer science as an average case model for graph partitioning problems under the name of the ``planted partition model.'' Given a sparse stochastic block model, the two standard inference tasks are: (i) Weak recovery: can we estimate the communities with non trivial overlap with the true communities? (ii) Detection/Hypothesis testing: can we distinguish if the sample was drawn from the block model or from a random graph with no community structure with probability tending to $1$ as the graph size tends to infinity? In this work, we show that for sparse stochastic block models, the two inference tasks are equivalent except at a critical point. That is, weak recovery is information theoretically possible if and only if detection is possible. We thus find a strong connection between these two notions of inference for the model. We further prove that when detection is impossible, an explicit hypothesis test based on low degree polynomials in the adjacency matrix of the observed graph achieves the optimal statistical power. This low degree test is efficient as opposed to the likelihood ratio test, which is not known to be efficient. Moreover, we prove that the asymptotic mutual information between the observed network and the community structure exhibits a phase transition at the weak recovery threshold. Our results are proven in much broader settings including the hypergraph stochastic block models and general planted factor graphs. In these settings we prove that the impossibility of weak recovery implies contiguity and provide a condition which guarantees the equivalence of weak recovery and detection.

math.PR

One-step replica symmetry breaking of random regular NAE-SAT II

Continuing our earlier work in \cite{nss20a}, we study the random regular k-NAE-SAT model in the condensation regime. In \cite{nss20a}, the 1RSB properties of the model were established with positive probability. In this paper, we improve the result to probability arbitrarily close to one. To do so, we introduce a new framework which is the synthesis of two approaches: the small subgraph conditioning and a variance decomposition technique using Doob martingales and discrete Fourier analysis. The main challenge is a delicate integration of the two methods to overcome the difficulty arising from applying the moment method to an unbounded state space.

math.PR

One-step replica symmetry breaking of random regular NAE-SAT I

In a broad class of sparse random constraint satisfaction problems(CSP), deep heuristics from statistical physics predict that there is a condensation phase transition before the satisfiability threshold, governed by one-step replica symmetry breaking(1RSB). In fact, in random regular k-NAE-SAT, which is one of such random CSPs, it was verified \cite{ssz22} that its free energy is well-defined and the explicit value follows the 1RSB prediction. However, for any model of sparse random CSP, it has been unknown whether the solution space indeed condenses on O(1) clusters according to the 1RSB prediction. In this paper, we give an affirmative answer to this question for the random regular k-NAE-SAT model. Namely, we prove that with probability bounded away from zero, most of the solutions lie inside a bounded number of solution clusters whose sizes are comparable to the scale of the free energy. Furthermore, we establish that the overlap between two independently drawn solutions concentrates precisely at two values. Our proof is based on a detailed moment analysis of a spin system, which has an infinite spin space that encodes the structure of solution clusters. We believe that our method is applicable to a broad range of random CSPs in the 1RSB universality class.

math.PR

Parisi formula for balanced Potts spin glass

The Potts spin glass is a generalization of the Sherrington--Kirkpatrick (SK) model that allows for spins to take more than two values. Based on a novel synchronization mechanism, Panchenko (2018) showed that the limiting free energy is given by a Parisi-type variational formula. The functional order parameter in this formula is a probability measure on a monotone path in the space of positive-semidefinite matrices. By comparison, the order parameter for the SK model is much simpler: a probability measure on the unit interval. Nevertheless, a longstanding prediction by Elderfield and Sherrington (1983) is that the order parameter for the Potts spin glass can be reduced to that of the SK model. We prove this prediction for the balanced Potts spin glass, where the model is constrained so that the fraction of spins taking each value is asymptotically the same. It is generally believed that the limiting free energy of the balanced model is the same as that of the unconstrained model, in which case our results reduce the functional order parameter of Panchenko's variational formula to probability measures on the unit interval. The intuitive reason -- for both this belief and the Elderfield--Sherrington prediction -- is that no spin value is a priori preferred over another, and the order parameter should reflect this inherent symmetry. This paper rigorously demonstrates how symmetry, when combined with synchronization, acts as the desired reduction mechanism. Our proof requires that we introduce a generalized Potts spin glass model with mixed higher-order interactions, which is interesting it its own right. We prove that the Parisi formula for this model is differentiable with respect to inverse temperatures. This is a key ingredient for guaranteeing the Ghirlanda--Guerra identities without perturbation, which then allow us to exploit symmetry and synchronization simultaneously.

math.PR

Universality of max-margin classifiers

Maximum margin binary classification is one of the most fundamental algorithms in machine learning, yet the role of featurization maps and the high-dimensional asymptotics of the misclassification error for non-Gaussian features are still poorly understood. We consider settings in which we observe binary labels $y_i$ and either $d$-dimensional covariates ${\boldsymbol z}_i$ that are mapped to a $p$-dimension space via a randomized featurization map ${\boldsymbol ϕ}:\mathbb{R}^d \to\mathbb{R}^p$, or $p$-dimensional features of non-Gaussian independent entries. In this context, we study two fundamental questions: $(i)$ At what overparametrization ratio $p/n$ do the data become linearly separable? $(ii)$ What is the generalization error of the max-margin classifier? Working in the high-dimensional regime in which the number of features $p$, the number of samples $n$ and the input dimension $d$ (in the nonlinear featurization setting) diverge, with ratios of order one, we prove a universality result establishing that the asymptotic behavior is completely determined by the expected covariance of feature vectors and by the covariance between features and labels. In particular, the overparametrization threshold and generalization error can be computed within a simpler Gaussian model. The main technical challenge lies in the fact that max-margin is not the maximizer (or minimizer) of an empirical average, but the maximizer of a minimum over the samples. We address this by representing the classifier as an average over support vectors. Crucially, we find that in high dimensions, the support vector count is proportional to the number of samples, which ultimately yields universality.

math.ST

Agreement and Statistical Efficiency in Bayesian Perception Models

Bayesian models of group learning are studied in Economics since the 1970s. and more recently in computational linguistics. The models from Economics postulate that agents maximize utility in their communication and actions. The Economics models do not explain the ``probability matching" phenomena that are observed in many experimental studies. To address these observations, Bayesian models that do not formally fit into the economic utility maximization framework were introduced. In these models individuals sample from their posteriors in communication. In this work we study the asymptotic behavior of such models on connected networks with repeated communication. Perhaps surprisingly, despite the fact that individual agents are not utility maximizers in the classical sense, we establish that the individuals ultimately agree and furthermore show that the limiting posterior is Bayes optimal. We explore the interpretation of our results in terms of Large Language Models (LLMs). In the positive direction our results can be interpreted as stating that interaction between different LLMs can lead to optimal learning. However, we provide an example showing how misspecification may lead LLM agents to be overconfident in their estimates.

math.ST