Unbiased Gradients, Moving Stability Boundaries: Exact Mini-Batch Geometry in Linear Self-Attention
Unbiased stochastic gradients can match the full gradient in expectation while changing finite-step stability. We study this effect in one-layer linear self-attention for in-context linear regression, where shared-mode mini-batch training reduces exactly to a random two-factor map. The full-batch map preserves an elliptic region, whereas sampled curvature and target correlation create batch-dependent stability boundaries. Their combined fluctuation is centered at the parameter-update level but induces a strictly outward drift of the full-batch boundary coordinate. We derive an exact one-step crossing criterion, adaptive concentration bounds, and conditional drift identities, and prove almost-sure escape in a balanced switching model. For independent long prompts, mode leakage decays at the inverse square root of prompt length. For finite-token softmax attention, we derive an explicit boundary-crossing law governed by an effective context length that shrinks as attention scores concentrate. Experiments in ambient attention objectives verify the reduced predictions, and four trained causal softmax Transformers show larger mean curvature shifts for smaller first-step batches, consistent with the predicted nonlinear outward bias.