arXiv · 2609.26268
Decoupling Is Not Identification: Supervised Evidential Learning in Next-Token Prediction
Abstract
A next-token probability says what a model predicts, not how much training support lies behind it. A Dirichlet head can represent this distinction by separating mean $m$ from concentration $S$, but decoupling does not identify what $S$ means. Here we propose an Evidential Next-Token Prediction (ENTOP) framework to audit this gap on character-level Moby-Dick, using exact 8-gram count as a reproducible lexical-support label and withholding count regression from 20% of context types. Standard implicit evidential training carries essentially no count signal beyond confidence on held-out-label types (partial Spearman $ρ= 0.001 \pm 0.014$), whereas explicit supervision generalizes ($ρ= 0.201 \pm 0.010$; matched-pair win $= 0.822 \pm 0.021$). CE predictive entropy is at chance for unseen 8-grams (AUROC $= 0.490 \pm 0.004$), while supervised vacuity reaches $0.772 \pm 0.003$, comparable with an indexed CE-representation baseline ($0.769$) but below the tautological corpus oracle ($1.000$). Neither longest-suffix nor representation-distance strata explain where amortization succeeds. Increasing count weight under the digamma objective improves support fit only by sacrificing prediction. A constant predictor wins natural log-RMSE, and vacuity does not improve error deferral. These results motivate a minimum evidence protocol---confidence control, matched pairs, held-out labels, a constant baseline, and a decision test---and show that concentration can pass identification while failing calibration and utility.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ge Wang. 2026-08-14. Decoupling Is Not Identification: Supervised Evidential Learning in Next-Token Prediction. https://arxiv.org/abs/2609.26268
Cite the original work for its findings. Save a collection to share your selection of sources.