arXiv ScienceSearch

arXiv subjects

George Wang

Publications and source records attributed to George Wang.

At least 19 recordsLinked to original sources

Patterning in Practice: Debiasing Reward Models with Susceptibilities

Reward models trained on human preferences are known to suffer from length, formatting, and other stylistic biases. In this paper we use patterning, which reweights each preference pair according to its measured effect on posterior expectation values of benchmark losses (its susceptibility), to debias a Gemma 2 9B Instruct reward model trained on Skywork-Reward-Preference v0.2. We obtain $+14.2 \pm 1.2$ pp on RM-Bench Hard, the split where style cues point against correctness (mean $\pm$ s.e.\ over 5 seeds), with overall RM-Bench accuracy preserved, comparable to the strongest Hard-split gain reported by the closest published comparator (SteerRM, $+13.2$ pp). We demonstrate in a simple case that the reweighting is interpretable by tracing a side effect of the intervention (a regression on a safety subset of RM-Bench) to a small class of training pairs, which we confirm by ablation. The weights also transfer: those computed on Gemma 2 9B debias Gemma 2 2B and 27B with no recomputation, and transfer partially to Llama 3.1 8B. This is the first application of patterning, a program grounded in singular learning theory, beyond small models and synthetic tasks.

cs.LG

Patterning: The Dual of Interpretability

Mechanistic interpretability aims to understand how neural networks generalize beyond their training data by reverse-engineering their internal structures. We introduce patterning as the dual problem: given a desired form of generalization, determine what training data produces it. Our approach is based on susceptibilities, which measure how posterior expectation values of observables respond to infinitesimal shifts in the data distribution. Inverting this linear response relationship yields the data intervention that steers the model toward a target internal configuration. We demonstrate patterning in a small language model, showing that re-weighting training data along principal susceptibility directions can accelerate or delay the formation of structure, such as the induction circuit. In a synthetic parentheses balancing task where multiple algorithms achieve perfect training accuracy, we show that patterning can select which algorithm the model learns by targeting the local learning coefficient of each solution. These results establish that the same mathematical framework used to read internal structure can be inverted to write it.

cs.LG

Towards Spectroscopy: Susceptibility Clusters in Language Models

Spectroscopy infers the internal structure of physical systems by measuring their response to perturbations. We apply this principle to neural networks: perturbing the data distribution by upweighting a token $y$ in context $x$, we measure the model's response via susceptibilities $\chi_{xy}$, which are covariances between component-level observables and the perturbation computed over a localized Gibbs posterior via stochastic gradient Langevin dynamics (SGLD). Theoretically, we show that susceptibilities decompose as a sum over modes of the data distribution, explaining why tokens that follow their contexts "for similar reasons" cluster together in susceptibility space. Empirically, we apply this methodology to Pythia-14M, developing a conductance-based clustering algorithm that identifies 510 interpretable clusters ranging from grammatical patterns to code structure to mathematical notation. Comparing to sparse autoencoders, 50% of our clusters match SAE features, validating that both methods recover similar structure.

cs.LG

An Overabundance of Radio-AGN in the SPT2349-56 Protocluster: Preheating the Intra-Cluster Medium

Following the detection of a radio-loud Active Galactic Nucleus (AGN) in the z=4.3 protocluster SPT2349-56, we have obtained additional observations with MeerKAT in S-band (2.4 GHz) with the aim of further characterizing radio emission from amongst the ~30 submillimeter (submm) galaxies (SMGs) identified in the structure. We newly identify three of the protocluster SMGs individually at 2.4GHz as having a radio-excess, two of which are now known to be X-ray luminous AGN. Two additional members are also detected with radio emission consistent with their star formation rate (SFR). Archival MeerKAT UHF (816 MHz) observations further constrain luminosities and radio spectral indices of these five galaxies. The Australia Telescope Compact Array (ATCA) is used to detect and resolve the central two sources at 5.5 and 9.0 GHz finding elongated, jet-like morphologies. The excess radio luminosities range from L1.4,rest = (1-20)x10^25 W/Hz, ~10-100x higher than expected from the SFRs, assuming the usual far-infrared-radio correlation. Of the known cluster members, only the SMG `N1' shows signs of AGN in any other diagnostics, namely a large and compact excess in CO(11-10) line emission. We compare these results to field samples of radio sources and SMGs. The overdensity of radio-loud AGN in the compact core region of the cluster may be providing significant heating to the recently discovered nascent intra-cluster medium (ICM) in SPT2349-56.

astro-ph.GA

A large thermal energy reservoir in the nascent intracluster medium at a redshift of 4.3

Most baryons in present-day galaxy clusters exist as hot gas ($\boldsymbol{\gtrsim10^7\,\rm}\mathrm{K}$), forming the intracluster medium (ICM). Cosmological simulations predict that the mass and temperature of the ICM rapidly decrease with increasing cosmological redshift, as intracluster gas in younger clusters is still accumulating and being heated. The thermal Sunyaev-Zeldovich (tSZ) effect arises when cosmic microwave background (CMB) photons are scattered to higher energies through interactions with energetic electrons in hot ICM, leaving a localized decrement in the CMB at a long wavelength. The depth of this decrement is a measure of the thermal energy and pressure of the gas. To date, the effect has been detected in only three systems at or above $z\sim2$, when the Universe was 4 billion years old, making the time and mechanism of ICM assembly uncertain. Here, we report observations of this effect in the protocluster SPT2349$-$56 with Atacama Large Millimeter/submillimeter Array (ALMA). SPT2349$-$56 contains a large molecular gas reservoir, with at least 30 dusty star-forming galaxies (DSFGs) and three radio-loud active galactic nuclei (AGN) in a 100-kpc region at $z=4.3$, corresponding to 1.4 billion years after the Big Bang. The observed tSZ signal implies a thermal energy of $\mathbf{\sim 10^{61}\,\mathrm{erg}}$, exceeding the possible energy of a virialized ICM by an order of magnitude. Contrary to current theoretical expectations, the strong tSZ decrement in SPT2349$-$56 demonstrates that substantial heating can occur and deposit a large amount of thermal energy within growing galaxy clusters, overheating the nascent ICM in unrelaxed structures, two billion years before the first mature clusters emerged at $\mathbf{z \sim 2}$.

astro-ph.GA

Embryology of a Language Model

Understanding how language models develop their internal computational structure is a central problem in the science of deep learning. While susceptibilities, drawn from statistical physics, offer a promising analytical tool, their full potential for visualizing network organization remains untapped. In this work, we introduce an embryological approach, applying UMAP to the susceptibility matrix to visualize the model's structural development over training. Our visualizations reveal the emergence of a clear ``body plan,'' charting the formation of known features like the induction circuit and discovering previously unknown structures, such as a ``spacing fin'' dedicated to counting space tokens. This work demonstrates that susceptibility analysis can move beyond validation to uncover novel mechanisms, providing a powerful, holistic lens for studying the developmental principles of complex neural networks.

cs.LG

MAATS: A Multi-Agent Automated Translation System Based on MQM Evaluation

We present MAATS, a Multi Agent Automated Translation System that leverages the Multidimensional Quality Metrics (MQM) framework as a fine-grained signal for error detection and refinement. MAATS employs multiple specialized AI agents, each focused on a distinct MQM category (e.g., Accuracy, Fluency, Style, Terminology), followed by a synthesis agent that integrates the annotations to iteratively refine translations. This design contrasts with conventional single-agent methods that rely on self-correction. Evaluated across diverse language pairs and Large Language Models (LLMs), MAATS outperforms zero-shot and single-agent baselines with statistically significant gains in both automatic metrics and human assessments. It excels particularly in semantic accuracy, locale adaptation, and linguistically distant language pairs. Qualitative analysis highlights its strengths in multi-layered error diagnosis, omission detection across perspectives, and context-aware refinement. By aligning modular agent roles with interpretable MQM dimensions, MAATS narrows the gap between black-box LLMs and human translation workflows, shifting focus from surface fluency to deeper semantic and contextual fidelity.

cs.CL

Structural Inference: Interpreting Small Language Models with Susceptibilities

We develop a linear response framework for interpretability that treats a neural network as a Bayesian statistical mechanical system. A small perturbation of the data distribution, for example shifting the Pile toward GitHub or legal text, induces a first-order change in the posterior expectation of an observable localized on a chosen component of the network. The resulting susceptibility can be estimated efficiently with local SGLD samples and factorizes into signed, per-token contributions that serve as attribution scores. We combine these susceptibilities into a response matrix whose low-rank structure separates functional modules such as multigram and induction heads in a 3M-parameter transformer.

cs.LG

You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation

In this position paper, we argue that understanding the relation between structure in the data distribution and structure in trained models is central to AI alignment. First, we discuss how two neural networks can have equivalent performance on the training set but compute their outputs in essentially different ways and thus generalise differently. For this reason, standard testing and evaluation are insufficient for obtaining assurances of safety for widely deployed generally intelligent systems. We argue that to progress beyond evaluation to a robust mathematical science of AI alignment, we need to develop statistical foundations for an understanding of the relation between structure in the data distribution, internal structure in models, and how these structures underlie generalisation.

cs.LG

Evidence for environmental effects in the $z\,{=}\,4.3$ protocluster core SPT2349$-$56

We present ALMA observations of the [CI] 492 and 806$\,$GHz fine-structure lines in 25 dusty star-forming galaxies (DSFGs) at $z\,{=}\,4.3$ in the core of the SPT2349$-$56 protocluster. The protocluster galaxies exhibit a median $L^\prime_{[\text{CI}](2-1)}/L^\prime_{[\text{CI}](1-0)}$ ratio of 0.94 with an interquartile range of 0.81-1.24. These ratios are markedly different to those observed in DSFGs in the field (across a comparable redshift and 850$\,\mu$m flux density range), where the median is 0.55 with an interquartile range of 0.50-0.76, and we show that this difference is driven by an excess of [CI](2-1) in the protocluster galaxies for a given 850$\,\mu$m flux density. Assuming local thermal equilibrium, we estimate gas excitation temperatures of $T_{\rm ex}\,{=}\,59.1^{+8.1}_{-6.8}\,$K for our protocluster sample and $T_{\rm ex}\,{=}\,33.9^{+2.4}_{-2.2}\,$K for the field sample. Our main interpretation of this result is that the protocluster galaxies have had their cold gas driven to their cores via close-by interactions within the dense environment, leading to an overall increase in the average gas density and excitation temperature, and an elevated [CI](2-1) luminosity-to-far-infrared luminosity ratio.

astro-ph.GA

Differentiation and Specialization of Attention Heads via the Refined Local Learning Coefficient

We introduce refined variants of the Local Learning Coefficient (LLC), a measure of model complexity grounded in singular learning theory, to study the development of internal structure in transformer language models during training. By applying these \textit{refined LLCs} (rLLCs) to individual components of a two-layer attention-only transformer, we gain novel insights into the progressive differentiation and specialization of attention heads. Our methodology reveals how attention heads differentiate into distinct functional roles over the course of training, analyzes the types of data these heads specialize to process, and discovers a previously unidentified multigram circuit. These findings demonstrate that rLLCs provide a principled, quantitative toolkit for \textit{developmental interpretability}, which aims to understand models through their evolution across the learning process. More broadly, this work takes a step towards establishing the correspondence between data distributional structure, geometric properties of the loss landscape, learning dynamics, and emergent computational structures in neural networks.

cs.LG

Loss Landscape Degeneracy and Stagewise Development in Transformers

Deep learning involves navigating a high-dimensional loss landscape over the neural network parameter space. Over the course of training, complex computational structures form and re-form inside the neural network, leading to shifts in input/output behavior. It is a priority for the science of deep learning to uncover principles governing the development of neural network structure and behavior. Drawing on the framework of singular learning theory, we propose that model development is deeply linked to degeneracy in the local geometry of the loss landscape. We investigate this link by monitoring loss landscape degeneracy throughout training, as quantified by the local learning coefficient, for a transformer language model and an in-context linear regression transformer. We show that training can be divided into distinct periods of change in loss landscape degeneracy, and that these changes in degeneracy coincide with significant changes in the internal computational structure and the input/output behavior of the transformers. This finding provides suggestive evidence that degeneracy and development are linked in transformers, underscoring the potential of a degeneracy-based perspective for understanding modern deep learning.

cs.LG

The Local Learning Coefficient: A Singularity-Aware Complexity Measure

The Local Learning Coefficient (LLC) is introduced as a novel complexity measure for deep neural networks (DNNs). Recognizing the limitations of traditional complexity measures, the LLC leverages Singular Learning Theory (SLT), which has long recognized the significance of singularities in the loss landscape geometry. This paper provides an extensive exploration of the LLC's theoretical underpinnings, offering both a clear definition and intuitive insights into its application. Moreover, we propose a new scalable estimator for the LLC, which is then effectively applied across diverse architectures including deep linear networks up to 100M parameters, ResNet image models, and transformer language models. Empirical evidence suggests that the LLC provides valuable insights into how training heuristics might influence the effective complexity of DNNs. Ultimately, the LLC emerges as a crucial tool for reconciling the apparent contradiction between deep learning's complexity and the principle of parsimony.

stat.ML

Brightest Cluster Galaxy Formation in the z=4.3 Protocluster SPT2349-56: Discovery of a Radio-Loud AGN

We have observed the z=4.3 protocluster SPT2349-56 with ATCA with the aim of detecting radio-loud active galactic nuclei (AGN) amongst the ~30 submillimeter galaxies identified in the structure. We detect the central complex of SMGs at 2.2\,GHz with a luminosity of L_2.2=(4.42pm0.56)x10^{25} W/Hz. The ASKAP also detects the source at 888 MHz, constraining the radio spectral index to alpha=-1.6pm0.3, consistent with ATCA non-detections at 5.5 and 9GHz, and implying L_1.4(rest)=(2.4pm0.3)x10^{26}W/Hz. This radio luminosity is about 100 times higher than expected from star formation, assuming the usual FIR-radio correlation, which is a clear indication of an AGN driven by a forming brightest cluster galaxy (BCG). None of the SMGs in SPT2349-56 show signs of AGN in any other diagnostics available to us (notably 12CO out to J=16, OH163um, CII/IR, and optical spectra), highlighting the radio continuum as a powerful probe of obscured AGN in high-z protoclusters. No other significant radio detections are found amongst the cluster members, consistent with the FIR-radio correlation. We compare these results to field samples of radio sources and SMGs, along with the 22 SPT-SMG gravitational lenses also observed in the ATCA program, as well as powerful radio galaxies at high redshifts. Our results allow us to better understand the effects of this gas-rich, overdense environment on early supermassive black hole (SMBH) growth and cluster feedback. We estimate that (3.3pm0.7)x10^{38} W of power are injected into the growing ICM by the radio-loud AGN, whose energy over 100Myr is comparable to the binding energy of the gas mass of the central halo. The AGN power is also comparable to the instantaneous energy injection from supernova feedback from the 23 catalogued SMGs in the core region of 120kpc projected radius. The SPT2349-56 radio-loud AGN may be providing strong feedback on a nascent ICM.

astro-ph.GA

Rapid build-up of the stellar content in the protocluster core SPT2349$-$56 at $z\,{=}\,4.3$

The protocluster SPT2349$-$56 at $z\,{=}\,4.3$ contains one of the most actively star-forming cores known, yet constraints on the total stellar mass of this system are highly uncertain. We have therefore carried out deep optical and infrared observations of this system, probing rest-frame ultraviolet to infrared wavelengths. Using the positions of the spectroscopically-confirmed protocluster members, we identify counterparts and perform detailed source deblending, allowing us to fit spectral energy distributions in order to estimate stellar masses. We show that the galaxies in SPT2349$-$56 have stellar masses proportional to their high star-formation rates, consistent with other protocluster galaxies and field submillimetre galaxies (SMGs) around redshift 4. The galaxies in SPT2349$-$56 have on average lower molecular gas-to-stellar mass fractions and depletion timescales than field SMGs, although with considerable scatter. We construct the stellar-mass function for SPT2349$-$56 and compare it to the stellar-mass function of $z\,{=}\,1$ galaxy clusters, finding consistent shapes between the two. We measure rest-frame galaxy ultraviolet half-light radii from our HST-F160W imaging, finding that on average the galaxies in our sample are similar in size to typical star-forming galaxies at these redshifts. However, the brightest HST-detected galaxy in our sample, found near the luminosity-weighted centre of the protocluster core, remains unresolved at this wavelength. Hydrodynamical simulations predict that the core galaxies will quickly merge into a brightest cluster galaxy, thus our observations provide a direct view of the early formation mechanisms of this class of object.

astro-ph.GA

Overdensities of Submillimetre-Bright Sources around Candidate Protocluster Cores Selected from the South Pole Telescope Survey

We present APEX-LABOCA 870 micron observations of the fields surrounding the nine brightest, high-redshift, unlensed objects discovered in the South Pole Telescope's (SPT) 2500 square degrees survey. Initially seen as point sources by SPT's 1-arcmin beam, the 19-arcsec resolution of our new data enables us to deblend these objects and search for submillimetre (submm) sources in the surrounding fields. We find a total of 98 sources above a threshold of 3.7 sigma in the observed area of 1300 square arcminutes, where the bright central cores resolve into multiple components. After applying a radial cut to our LABOCA sources to achieve uniform sensitivity and angular size across each of the nine fields, we compute the cumulative and differential number counts and compare them to estimates of the background, finding a significant overdensity of approximately 10 at 14 mJy. The large overdensities of bright submm sources surrounding these fields suggest that they could be candidate protoclusters undergoing massive star-formation events. Photometric and spectroscopic redshifts of the unlensed central objects range from 3 to 7, implying a volume density of star-forming protoclusters of approximately 0.1 per giga-parsec cube. If the surrounding submm sources in these fields are at the same redshifts as the central objects, then the total star-formation rates of these candidate protoclusters reach 10,000 solar masses per year, making them much more active at these redshifts than what has been seen so far in both simulations and observations.

astro-ph.CO

Locks fit into keys: a crystal analysis of lock polynomials

Lock polynomials and lock Kohnert tableaux are natural analogues to key polynomials and key Kohnert tableaux, respectively. In this paper, we compare lock polynomials to the much-studied key polynomials and show that the difference of a key polynomial and lock polynomial for the same composition is monomial positive. We also examine the conditions for which key and lock polynomials are symmetric or quasisymmetric. We accomplish these goals combinatorially using key Kohnert tableaux and lock Kohnert tableaux. In particular, for the difference of a key minus a lock, we focus on the behavior of crystal operators on Kohnert tableaux. The Type A Demazure crystal can be realized on the vertex set of key Kohnert tableaux, and we show with an explicit combinatorial definition that a similar crystal-like structure exists on the vertex set of lock Kohnert tableaux. Finally, we construct an injective, weight-preserving map from lock Kohnert tableaux to key Kohnert tableaux that intertwines the crystal operators.

math.CO

Some conjectures on the Schur expansion of Jack polynomials

We present positivity conjectures for the Schur expansion of Jack symmetric functions in two bases given by binomial coefficients. Partial results suggest that there are rich combinatorics to be found in these bases, including Eulerian numbers, Stirling numbers, quasi-Yamanouchi tableaux, and rook boards. These results also lead to further conjectures about the fundamental quasisymmetric expansions of these bases, which we prove for special cases.

math.CO