arXiv Science⌕ Search

arXiv subjects

Pratyush Acharya

Publications and source records attributed to Pratyush Acharya.

3 recordsLinked to original sources

Concepts Whisper: Spectral Anti-Concentration and the Dual Geometry of Transformer Representations

We find that transformer concept representations systematically anti-concentrate in the spectral tail of the unembedding covariance, encoding word-level concepts in low-variance directions across a 17-model core suite and an expanded set of 22 semantic concept categories, with convergent replications from three independent extraction methods. Residual-stream difference-of-means vectors anti-concentrate in all 17 models (model-level one-sample t-test, p = 3.8e-9), and remain more tail-aligned than norm-matched random directions in 13 of 17; convergent support comes from sparse autoencoder (SAE) features (p = 4.5e-19 across concepts within a model) and linear probes on Llama and Qwen. We identify a dual geometry: activation-space concept directions anti-concentrate while static unembedding-row contrasts concentrate in high-variance directions (p < 10^-4). This investigation arose from testing whether the causal inner product of Park et al. (2024) aids cross-lingual concept transport; a matched-spectrum randomization across 17 models and four language pairs finds no evidence that Whitened Causal Alignment improves over spectral regularization alone (p = 0.95). Split-injection interventions, restricted to steering strengths at which both arms remain interpretable, show the predicted interference asymmetry in four of five models (paired Cohen's d_z up to 1.19) with no significant reversal inside that regime, and POS-tag probing across eight models shows syntax preferentially encoded in the high-variance subspace in six of eight architectures, with a significant reversal in the Qwen 2.5 family. These results suggest transformers rotate semantic content into spectrally quiet regions during contextualized processing, where, in some architectures, interventions may reduce grammatical disruption relative to high-variance steering.

cs.LG↗

Assessing the Pedagogical Readiness of Large Language Models as AI Tutors in Low-Resource Contexts: A Case Study of Nepal's K-10 Curriculum

The integration of Large Language Models (LLMs) into educational ecosystems promises to democratize access to personalized tutoring, yet the readiness of these systems for deployment in non-Western, low-resource contexts remains critically under-examined. This study presents a systematic evaluation of four state-of-the-art LLMs--GPT-4o, Claude Sonnet 4, Qwen3-235B, and Kimi K2--assessing their capacity to function as AI tutors within the specific curricular and cultural framework of Nepal's Grade 5-10 Science and Mathematics education. We introduce a novel, curriculum-aligned benchmark and a fine-grained evaluation framework inspired by the "natural language unit tests" paradigm, decomposing pedagogical efficacy into seven binary metrics: Prompt Alignment, Factual Correctness, Clarity, Contextual Relevance, Engagement, Harmful Content Avoidance, and Solution Accuracy. Our results reveal a stark "curriculum-alignment gap." While frontier models (GPT-4o, Claude Sonnet 4) achieve high aggregate reliability (approximately 97%), significant deficiencies persist in pedagogical clarity and cultural contextualization. We identify two pervasive failure modes: the "Expert's Curse," where models solve complex problems but fail to explain them clearly to novices, and the "Foundational Fallacy," where performance paradoxically degrades on simpler, lower-grade material due to an inability to adapt to younger learners' cognitive constraints. Furthermore, regional models like Kimi K2 exhibit a "Contextual Blindspot," failing to provide culturally relevant examples in over 20% of interactions. These findings suggest that off-the-shelf LLMs are not yet ready for autonomous deployment in Nepalese classrooms. We propose a "human-in-the-loop" deployment strategy and offer a methodological blueprint for curriculum-specific fine-tuning to align global AI capabilities with local educational needs.

cs.CY↗

Grokking as a Variance-Limited Phase Transition: Spectral Gating and the Epsilon-Stability Threshold

Standard optimization theories struggle to explain grokking, where generalization occurs long after training convergence. While geometric studies attribute this to slow drift, they often overlook the interaction between the optimizer's noise structure and landscape curvature. This work analyzes AdamW dynamics on modular arithmetic tasks, revealing a ``Spectral Gating'' mechanism that regulates the transition from memorization to generalization. We find that AdamW operates as a variance-gated stochastic system. Grokking is constrained by a stability condition: the generalizing solution resides in a sharp basin ($λ_{max}^H$) initially inaccessible under low-variance regimes. The ``delayed'' phase represents the accumulation of gradient variance required to lift the effective stability ceiling, permitting entry into this sharp manifold. Our ablation studies identify three complexity regimes: (1) \textbf{Capacity Collapse} ($P < 23$), where rank-deficiency prevents structural learning; (2) \textbf{The Variance-Limited Regime} ($P \approx 41$), where generalization waits for the spectral gate to open; and (3) \textbf{Stability Override} ($P > 67$), where memorization becomes dimensionally unstable. Furthermore, we challenge the "Flat Minima" hypothesis for algorithmic tasks, showing that isotropic noise injection fails to induce grokking. Generalization requires the \textit{anisotropic rectification} unique to adaptive optimizers, which directs noise into the tangent space of the solution manifold.

cs.LG↗