Intrinsic Interaction Geometry Controls the Low-Rank Complexity of Softmax Attention
How much matrix rank is required to preserve every bounded value output of normalized softmax attention? We study the unrestricted maximum-row-\(\ell_1\) approximation rank \(r_\varepsilon(A)\), exactly the least rank achieving uniform error over all bounded vector-valued values. Row softmax exposes the intrinsic interaction \(C=P_m(\log A)P_N\), whereas invertible \(Q/K\) gauges leave \(A\) fixed while changing the Euclidean geometry of a chosen query/key factorization. We replace that coordinate-dependent description by a projective residual \(q(C-T)\) and an attained factor-radius size \(κ(T)\). For every rank-\(r\) retained interaction with \(τ(T)<\varepsilon\), we prove $$ r_\varepsilon(A)\le \min\left\{ N,\; C_r\left( 1+\frac{κ(T)} {(\varepsilon-τ(T))^2} \right)^{r/2} \right\}, $$ with the same unknown dimension constant as the underlying weighted Gibbs-row cover. The profile is gauge invariant, termwise no worse than native retained-subspace bounds at the same declared dimension, and has a worst-case sharp \(r/2\) size exponent at fixed \(r\) and \(\varepsilon\). We then measure \(r_\varepsilon(A)\) directly on learned attention using 9,978 certified brackets across BERT, GPT-2, Qwen2.5, and two ViT checkpoints; where certificates do not close, the optimum remains interval-valued. A pre-specified 2,302-cell held-out study further shows that the historical native-coordinate geometry block contains coarse, mostly head-level information but no detectable incremental information beyond a strong calibrated baseline. The new intrinsic descriptor is not evaluated in that study. Together, the theory and measurements distinguish an operator-intrinsic complexity control from a stronger empirical explanation that the learned-head evidence does not support.