arXiv · 2609.01311
One-Layer Transformer Provably Learns Multiclass One-Nearest Neighbor in Context
Abstract
We extend recent work establishing an equivalence between one-layer transformers and nearest-neighbor classifiers in the binary setting to the multiclass case. By leveraging the simplex encoding, we show that one-layer transformers with an argmax classification head behave identically to a one-nearest-neighbor classifier in the multiclass setting. This closes a gap left by prior work, whose multiclass result relied on a non-standard rounding-based approach rather than the typical argmax head used in practice.
Explore related subjects
Keep this discovery
Skanda Athreya, Yutong Wang. 2026-09-01. One-Layer Transformer Provably Learns Multiclass One-Nearest Neighbor in Context. https://arxiv.org/abs/2609.01311
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.