arXiv Science⌕ Search

arXiv · 2609.34458

Understanding Generalization Requires Universal Induction

Abstract

Classical statistical theory is insufficient to explain the successes of general-purpose AI models, because it depends on handcrafted inductive biases that it cannot justify. No Free Lunch (NFL) theorems force any learner that beats chance on some environments to underperform on others. We might hope that past experience informs which environments to expect, but NFL applies equally to meta-learning. Thus, any method that makes meaningful predictions necessarily begins with an inductive bias external to the data. Choosing to bias toward short programs yields Solomonoff induction (SI), whose performance is competitive against all computable learners - albeit up to "constants" that become large when comparing against specialized methods that exploit background information. We therefore relativize SI to an information vantage point, biasing toward short programs with access to all preexisting information. This reframes the inductive bias: instead of seeking some absolute notion of simplicity, we favor accessibility with respect to our vantage point. An algorithm can only outpredict the relativized SI to the extent that its code contains additional information about the data, and no algorithm can generate such information. While SI is incomputable and hence not a practical algorithm, it provides a formal optimum for inference in the limit of infinite compute, and there is evidence to suggest that frontier AI systems roughly approximate it. Thus, the only known answer to meta-NFL is rooted in algorithmic information theory, which we should expect to play a fundamental role in explaining the generalization behavior of modern (and future) AI systems.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Aram Ebtekar, Marcus Hutter, Danica J. Sutherland. 2026-09-28. Understanding Generalization Requires Universal Induction. https://arxiv.org/abs/2609.34458

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

High-Dimensional Partial Least Squares: Spectral Analysis and Fundamental Limitations

Partial Least Squares (PLS) is a widely used method for data integration, designed to extract latent components shared across paired high-dimensional datasets. Despite decades of practical success, a precise theoretical understanding of its behavior in high-dimensional regimes remains limited. In this paper, we study a data integration model in which two high-dimensional data matrices share a low-rank common latent structure while also containing individual-specific components. We analyze the singular vectors of the associated cross-covariance matrix using tools from random matrix theory and derive asymptotic characterizations of the alignment between estimated and true latent directions. These results provide a quantitative explanation of the reconstruction performance of the PLS variant based on Singular Value Decomposition (PLS-SVD) and identify regimes where the method exhibits counter-intuitive or limiting behavior. Building on this analysis, we compare PLS-SVD with principal component analysis applied separately to each dataset and show its asymptotic superiority in detecting the common latent subspace. Overall, our results offer a comprehensive theoretical understanding of high-dimensional PLS-SVD, clarifying both its advantages and fundamental limitations.

stat.ML↗

Empirical Bayes 1-bit matrix completion

The problem of predicting unobserved entries in a binary matrix, known as 1-bit matrix completion, has found diverse applications in fields such as recommendation systems. In this study, we develop an empirical Bayes method for 1-bit matrix completion motivated by the Efron--Morris estimator, a matrix generalization of the James--Stein estimator that shrinks singular values toward zero. The proposed method exploits the underlying low-rank structure of binary matrices, drawing parallels with multidimensional item response theory. Simulation studies and real-data applications demonstrate that the proposed method achieves competitive predictive accuracy and favorable predictive calibration.

stat.ML↗

Certified Adaptive Refresh: Anytime-Valid Monitoring for Federated Conformal RAG

Question-answering services built on retrieval-augmented generation (RAG), in which a language model answers from retrieved documents, are inspected continuously and upgraded repeatedly, so their reliability guarantee must survive both. We study federated conformal RAG: nodes holding private corpora score candidate answers with a shared language model and send compressed scores to a hub that returns an answer set; a miss omits the true answer. We formulate monitoring as a sequential test: alarm when misses exceed a certified bound on the miss rate (the allowance). But a conformal certificate covers one frozen configuration at one look fixed in advance, so it gives no anytime-valid alarm; a threshold calibrated for one model certifies nothing about its replacement; and re-certifying each upgrade at full error level compounds failures. We propose Anytime-FC-RAG with three auditable rules: pick each configuration before its fresh calibration, charge every certificate attempt to one trajectory-wide budget $δ_{\rm cal}$, and fix threshold, allowance and bet before each query. One betting wealth $E_t$, growing in expectation only when misses exceed the allowance, never resets across deployments. Our certified adaptive-refresh theorem shows that, with evidence budget $δ_e$, alarming when $E_t$ reaches $1/δ_e$ has false-alarm probability at most $δ_e + δ_{\rm cal}$ under a random number of history-selected, evidence-driven refreshes of model, retriever, corpus, score or threshold; $δ_{\rm cal}$ pays for wrong certificates. With an exact order-statistic certificate the bound is distribution-free, architecture-agnostic and about misses, not per-input coverage or distribution shift as such. On Qwen2.5/MMLU-Pro the alarm separated audited-harmful from benign shifts completely at a 5% budget; in a model swap, only fresh calibration kept a downgrade's miss rate in check.

stat.ML↗