arXiv Science⌕ Search

arXiv · 2610.09709

Pretraining Shapes Spectral Structure: Architecture- and Strategy-Conditional Prediction of OOD Robustness in Foundation Models

Abstract

Can we determine whether a foundation model will generalize out-of-distribution (OOD) before any target data is available? Existing diagnostics require source or target data, which rules them out before a target domain exists. Those that use the weights alone apply one statistic to every architecture, and do not separate robust models from fragile ones. We show the answer is encoded in the spectral structure of pretrained weights. Two forces shape that structure. Architecture determines how information is stored in weight matrices. Pretraining strategy determines what is rewarded. Together they set a spectral geometry that governs OOD robustness. We prove that the OOD accuracy gap is bounded by how tightly the source representations concentrate. A statistic computed from the pretrained weights alone serves as a proxy for that concentration. The direction of that proxy reverses between architecture families. We operationalize it: the direction is stable within one (architecture X strategy) combination, the finest grouping we test, which we call a cell. Pooled over 116 models spanning 7 modalities, a single statistic ranks OOD robustness weakly, because cells of opposite direction cancel. Within a cell, the statistic selected for it orders 92% of model pairs by OOD robustness in-sample. The selection does not leak the target: for each model family outside the matrix we logged the cell, metric and sign before running its OOD evaluation, and the predicted direction held in every case: EEG, genomic and protein. Acting on spectral concentration narrows the OOD gap by 24% at 87.5% ID retention. The diagnostic operates on released weights alone, so OOD robustness becomes checkable at model-selection time, before data or compute is committed to a target domain.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Sangyoon Bae, Sk Miraj Ahmed, Shinjae Yoo, Jiook Cha. 2026-10-07. Pretraining Shapes Spectral Structure: Architecture- and Strategy-Conditional Prediction of OOD Robustness in Foundation Models. https://arxiv.org/abs/2610.09709

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Learning in the Recurrent State: Gradient Descent with Linear Recurrent Networks

In-context learning lets a sequence model adapt to a new task from examples in its input. A prominent line of work shows how self-attention can be constructed to implement gradient descent on a linear predictor fit to the in-context examples during the forward pass. State-space models (SSMs) and other linear recurrent networks (LRNNs) model sequences at linear time cost, but it is unclear how their recurrent update could carry out the same in-context gradient descent. We introduce Gradient-based Recurrent In-context Learner (GRIL), a diagonal LRNN that factorizes a supervised gradient step into a short-window cross-product write and a multiplicative readout of the next query. For linear regression, this construction accumulates the context gradient in a matrix state and applies it in a single forward pass, with $O(f^2)$ learned degrees of freedom. The same design extends to multi-step updates and cross-entropy classification, with a limited MLP-based extension to non-linear regression. We show empirically that trained GRILs recover the behavior and parameters analytically predicted by the construction on synthetic ICL tasks. Furthermore, the same architecture can be extended and trained on general-purpose benchmarks, including Long Range Arena, language modeling and associative recall. Together, these results establish windowed cross-product self-attention as a concrete inductive bias that lets LRNNs learn in context through gradient-descent-like updates, while remaining trainable on general-purpose tasks.

cs.LG↗

LLaTA: Unlocking Graph Structure Learning with Tree-Guided Large Language Models

The emergence of large language models (LLMs) has popularized text-attributed graphs (TAGs), creating an urgent need for graph structure learning (GSL) methods that effectively leverage textual information. However, existing GSL approaches are designed for traditional graphs without text, and adapting them to LLMs faces two challenges: defining a suitable optimization objective given LLMs' massive parameters, and designing an efficient architecture without costly fine-tuning. To address these, we propose LLaTA (Large Language and Tree Assistant), which reformulates GSL as a tree optimization framework---shifting from training edge predictors to designing a language-aware tree sampler. LLaTA constructs structural encoding trees via entropy minimization to capture topology, then leverages tree-guided LLM in-context learning to integrate textual semantics without fine-tuning. Extensive experiments on 11 datasets demonstrate LLaTA's flexibility with any backbone, superior scalability over LLM-based GSL methods, and state-of-the-art effectiveness across diverse domains.

cs.LG↗

Causal Posterior Estimation

We present Causal Posterior Estimation (CPE), a novel method for Bayesian inference in simulator models, where evaluating the likelihood function is intractable or computationally expensive, but generating outputs given parameter values is straightforward. CPE approximates the posterior distribution using flow matching while directly incorporating the conditional dependence structure induced by the model's graphical representation into the neural network architecture. Across extensive experiments, we demonstrate that hard-coding these conditional dependencies into the network, rather than requiring them to be learned from data, enables CPE to achieve highly accurate posterior inference that matches or outperforms state-of-the-art baselines.

cs.LG↗