arXiv Science⌕ Search

arXiv · 2609.32319

Phase Space Attention:A Hairer Lift Circumvents the Single-Layer Induction Obstruction

Abstract

We circumvent the Sanford-Hsu-Telgarsky (SHT) single-layer induction obstruction within a linear, one-step, causal, bilinear, symplectically consistent design class on the post-RoPE substrate, by lifting attention onto a symplectic phase space, mirroring Hairer's lift of Stormer-Verlet. The lift exits the premise of the SHT counting argument rather than the bound itself. The reframing exhibits the obstruction as a filter-order gap: a one-layer bilinear score realises a $z$-transform of joint order $(0,0)$, whereas the induction discriminator requires key-side order $\geq 1$. Applying the symplectic upper shear $M_γ:(q,p)\mapsto(q+γp,p)$ to the post-RoPE query and key streams closes it. We prove this lift is unique within the factorised subclass, exactly symplectic at operator level, and requires post-RoPE placement; and in an explicit $T_4$-only Gaussian reduction we derive a closed-form two-branch induction phase transition, held out at $r=0.9876$ with zero fitted parameters. That law is an analytically solvable limit, not a robust prediction: restoring the $T_3$ channel moves $γ_c$ at $d_k=64$ from 1.030 to 0.569 and removes the crossover. Deployability follows by exact derivation: the KV cache is unchanged, prefix reuse and speculative decoding are preserved, overhead is $6d$ FLOPs per token per layer, INT8 headroom grows by at most $\log_2(1+2γ)$ bits, fused kernels are unmodified, and no parameters are added. At 91.3M parameters a supercritical sweep locates an emergence band: induction forms 3/3 seeds at $γ=0.80$ in a mean of 717 steps, against 2/3 seeds and 2700 steps at $γ=0$. Adverse results are reported as directly: a key-only half-lift reaches 0.949 against 0.811 for the symmetric operator, so if induction accuracy is the objective, the half-lift is the better construction. Forty-one notebooks and result files ship as ancillary material.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Kingsuk Maitra, Shagun Sood Morteza Hosseini, Suman Gunnala, Vikram Gupta. 2026-09-26. Phase Space Attention:A Hairer Lift Circumvents the Single-Layer Induction Obstruction. https://arxiv.org/abs/2609.32319

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Stochastic Engrams for Efficient Continual Learning

The ability to learn continuously in artificial neural networks (ANNs) is often limited by catastrophic forgetting, a phenomenon in which new knowledge becomes dominant. By taking mechanisms of memory encoding in neuroscience (i.e., engrams) as inspiration, we propose a novel approach that integrates stochastically-activated engrams as a gating mechanism for metaplastic binarized neural networks (mBNNs). This method leverages the computational efficiency of mBNNs combined with the robustness of probabilistic memory traces to mitigate forgetting and maintain the model's reliability. Previously validated metaplastic optimization techniques have been incorporated to further enhance synaptic stability. Compared to baseline binarized models and benchmark fully connected continual learning approaches, our method is the only strategy capable of achieving average accuracies over 70% in both class-incremental and domain-incremental MNIST benchmarks, matching full-precision state-of-the-art methods. Furthermore, we achieve a significant reduction in peak GPU and RAM usage, under 5% and 20%, respectively, as well as an ~8x reduction in memory footprint compared to full precision counterparts. Our findings demonstrate (A) an improved stability vs. plasticity trade-off, (B) reduced memory intensiveness, and (C) enhanced performance in binarized architectures. By uniting principles of neuroscience and efficient computing, we offer new insights into the design of scalable and robust deep learning systems.

cs.LG↗

DRAN: A Distribution and Relation Adaptive Network for Spatio-temporal Forecasting

Spatio-temporal forecasting remains challenging under non-stationary environments because both data distributions and spatial relations evolve over time. Temporal normalization and de-normalization are widely used to mitigate distribution shifts, but they may distort inter-node relationships and thereby impair spatial dependency modeling. To address these issues, we propose the Distribution and Relation Adaptive Network (DRAN) for spatio-temporal forecasting. DRAN incorporates a Spatial Factor Learner (SFL) module, which enables effective normalization and de-normalization while preserving spatial dependencies in spatio-temporal systems. To model evolving spatial interactions, DRAN further proposes the Dynamic-Static Fusion Learner (DSFL) module. DSFL decomposes features into static and dynamic components and adaptively fuses them according to input variability. Experiments on six benchmark datasets show that DRAN outperforms state-of-the-art baselines. Additional analyses demonstrate that SFL consistently reduces spatial-relation distortion across multiple normalization schemes, whereas DSFL captures complementary static and dynamic dependencies and adjusts their contributions according to temporal variability.

cs.LG↗

AYLA: Architecting a loss landscape in shallow neural networks to accelerate feature recovery

Feature learning in shallow neural networks exhibits rich yet fragile dynamics, including prolonged plateaus, abrupt phase transitions, and sensitivity to optimization hyperparameters. While recent theoretical work has characterized these behaviors through the geometry of loss landscapes, saddle escape mechanisms, and emergent scaling laws, practical methods for actively shaping these dynamics remain limited. In this paper, we introduce AYLA, a principled loss reparameterization framework that dynamically modulates gradient magnitudes during training without altering the location of stationary points or optimal solutions. AYLA applies a smooth, sigmoid-controlled power-law transformation to empirical loss, yielding a state-dependent effective learning rate that accelerates descent in flat or saddle-dominated regions while stabilizing late-stage optimization. Crucially, AYLA preserves all critical points of the original objective, acting solely as a monotone transformation that reshapes optimization trajectories rather than objectives. We evaluate AYLA in controlled teacher student settings using two-layer tanh networks trained on synthetic Gaussian data. Across stochastic gradient descent and multiple loss-exponent schedules, AYLA consistently improves feature recovery. This evidence is observed in terms of weight alignment, per-neuron cosine similarity, hidden-activation correlation, and spectral properties of learned representations, while AYLA maintains competitive or faster loss convergence. Spectral analyses further demonstrate that AYLA mitigates rank collapse and promotes richer internal representations, signaling a transition from lazy to active feature-learning regimes. AYLA offers a lightweight, theoretically grounded way to improve shallow-network optimization, especially in resource-limited or noise-sensitive settings.

cs.LG↗