arXiv Science⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 1,153 records · Page 64Linked to original sources

The Construction of an Empirical Dataset of Incomplete Software Changes from Open Source Projects

During software development, a modification to a software component may propagate across the system, requiring precise identification and correct revision of all affected components. This is a complex task, and developers often (45.7%) miss related changes. To address this, several methods have been developed to extract co-change rules from files that are frequently changed together in the revision history. However, previous research evaluated the methods using artificially created incomplete changes, which may not be representative of real-world data. To solve this problem, we construct a dataset by mining incomplete changes from a collection of open-source software, using information about induced bugs and their respective fixes from an issue tracking platform. We also analyze the characteristics of incomplete changes using this constructed dataset and found that 89.4% of missed changes involved five or fewer files. Finally, we re-evaluate LCExtractor, an existing co-change rule extraction method, on our constructed dataset, and we identify the optimal sorting criterion and the impact of the number of used commits.

cs.SE↗

HyDra: Demystifying and Taming Dynamic Context Parallelism at Production Scale

Long-context training runs on sequences whose lengths span orders of magnitude, and dynamic context parallelism (DCP) gives each sequence its own CP degree. Existing DCP systems either do not scale or perform poorly on mainstream models, leaving Megatron-Core (Mcore) DCP as the only option at production scale. Mcore DCP, however, sizes each degree to fit memory, which grows linearly with length while attention grows quadratically, so comparable token counts hide unequal computation. In our production 256K-context training job on more than 11K GPUs under Mcore DCP, per-rank microbatch times differ by up to 5x at similar token counts. The skew leads to a 46% pipeline bubble and a 13% data-parallel bubble. We present HyDra, a scalable load-driven DCP system. Its scheduler balances computation by pulling every rank toward one load target, and balance in turn makes that target solvable in closed form. It places sequences with lazy heaps rather than whole-pool scans. That balance asks for CP degrees larger than memory requires, so its CP engine nests an inner Ulysses group in a shallow outer ring, letting a higher degree lower computation and communication together. Evaluation at both scales shows consistent gains. On a 512-GPU testbed, HyDra raises throughput over Mcore DCP by 1.18x on average at 32K context and 2.48x at 256K. On a 2,048-GPU production job, it shrinks the pipeline bubble from 36% to 14%, cuts scheduling time by 2.3x, and raises throughput by 1.10-1.43x (avg. 1.25x) over Mcore DCP and 1.33-1.90x (avg. 1.59x) over static CP.

cs.DC↗

Text-Vision Synergistic Token Caching: A Training-Free Framework for Efficient Vision-Language-Action Inference

Vision-Language-Action (VLA) models enable generalizable robotic control but remain computationally expensive. Token caching provides a training-free, plug-and-play acceleration alternative. However, existing VLA caching does not fully exploit a key inductive bias of VLA models: text-vision synergy, wherein textual semantics guide the precise visual grounding of task-relevant regions. In particular, existing designs insufficiently account for head-wise reliability in attention aggregation and layer-wise stability in cache reuse. To address this, we propose Text-Vision Synergistic Token Caching (TVCache), a training-free framework for efficient VLA inference. TVCache filters attention heads based on text-vision information focus to improve task-relevant and physically consistent visual grounding. Concurrently, we introduce a reuse-layer selection mechanism guided by text-vision entropy differences to avoid caching unstable representations and improve cache resource allocation. Extensive experiments across four representative VLA models, two simulation benchmarks, and real-world robotic tasks demonstrate the effectiveness and generality of TVCache. At matched token-retention ratios, TVCache consistently improves task success over existing VLA caching with comparable computational cost. On OpenVLA-OFT, it improves average success by up to 14.5 percentage points over VLA-Cache at 12.5% retention while reducing FLOPs by 2.45x relative to full-token inference.

cs.CV↗

Certified Selective Automation of LLM Agent Evaluation

Evaluating LLM agents still ends with a human reading trajectories, because automatic judges carry no guarantee on how often they are wrong. We ask the operational question: what fraction of agent evaluation can a judge take over, with a certificate that the error rate among auto-decided trajectories stays below a budget alpha? Agent corpora resist the standard answer: many agents attempt the same tasks, so trajectories arrive in correlated clusters, and the i.i.d. certificates of existing selective-judging methods can overstate what is safe: a naive certificate can claim 98% automation while its realized error exceeds the budget in 17.5% of task resamples. We introduce a task-level bootstrap certificate that is valid in every regime we test while matching the naive certificate's coverage; finite-sample cluster-valid alternatives certify nothing at realistic task counts. Under this certificate, a 4B logprob judge trained with SFT and reject-weighted GRPO certifies 0.30-0.59 of evaluation on tool-use and web corpora at alpha=0.1, the only judge, among strongly elicited frontier models, certifying on both headline corpora. Certified coverage is predictable before training from base rate and discrimination alone (leave-one-corpus-out R^2=0.96). Finally, the certificate doubles as a self-training filter: pseudo-labels harvested inside certified regions have contamination bounded by alpha by construction (realized 0.000-0.041 across six harvests), letting a judge enter an unseen domain at in-domain strength with zero target training labels.

cs.CL↗

One Rollout Is All You Get: Fully Test-Time Adaptation for GUI Agents

GUI agents are deployed with frozen weights and discard everything they experience on the job. Existing ways to update an agent's weights assume something deployment withholds: ground truth, rollouts beyond the single attempt (retries, samples, practice runs), or a learning phase other than deployment. Because GUI actions can be irreversible, a deployed agent gets one attempt per task occurrence, in arrival order, and every attempt counts. No ground truth is available at any point. We define fully test-time adaptation for GUI agents by these constraints and pair it with a minimal weight-space method, SOLO. Auxiliary models read each episode: a judge selects the episodes it deems successful, and a proposer-verifier pair relabels a failed episode's prefix with the subtask that prefix completed. Admitted episodes enter a short sliding window, and each admission updates a small adapter by top-K self-distillation on the agent's own predictions, provided the window holds a judged success. On recurring task streams built from WebArena, VisualWebArena and MobileWorld, SOLO improves on the frozen agent with both UI-TARS-7B and Qwen3-VL-8B, by three to six points of success rate, and exceeds two in-setting memory methods on the web streams.

cs.LG↗

Test-Time Scaling via Budgeted Multi-Attribute Verification

Verifying LLM-generated answers under a shared computational budget requires jointly deciding which candidates to inspect and which verification attributes to evaluate. We formulate this problem as multi-attribute good-arm identification under a global budget: each candidate is an arm evaluated along several costly attributes, and the goal is to certify as many candidates as possible whose mean scores exceed the prescribed thresholds on all attributes. We propose \textsc{BMA-GAI}, an algorithm that combines cost-aware arm selection with adaptive sampling of attributes. Every observation serves both to guide adaptive allocation and to support anytime-valid certification, which removes the need for a separate confirmation stage. We establish an asymptotic coverage guarantee for \textsc{BMA-GAI} and derive a matching information-theoretic converse that characterizes the intrinsic complexity of the problem, thereby proving that \textsc{BMA-GAI} is first-order optimal away from critical budget levels. Experiments on synthetic benchmarks and an LLM answer-verification task show that \textsc{BMA-GAI} allocates the verification budget more efficiently and certifies more high-quality candidates than competing methods.

cs.AI↗

Three-loop anomalous dimensions of leading-twist operators in the Gross-Neveu-Yukawa model

We compute the anomalous dimensions of the leading-twist operators at three-loop level in the $\overline{\text{MS}}$ scheme in the Gross-Neveu-Yukawa model. Calculations are done for a number of low-spin operators using computer algebra systems, and the analytical form is reconstructed by solving a set of Diophantine equations. Calculations are optimized using the emergent supersymmetry of the corresponding generalized theory. For the obtained results, we perform a comprehensive analysis at the critical point, including a comparison with the $1/N$ expansion, a check of the generalized Gribov-Lipatov reciprocity, and an analysis of the behavior next to the rightmost singularity. The latter allows us to extract the values of Regge intercepts for all the leading-twist trajectories in the form of an $ε$ expansion.

hep-th↗

Relic Neutrinos Probing Small Scale Primordial Non-Gaussianity

Relic neutrinos retain a record of electromagnetic entropy production after neutrino decoupling, probing primordial modes erased from the primary CMB. We derive the cubic photon number source from acoustic damping and its leading folded sensitive bispectrum window, with neutrino decoupling, thermalization, and acoustic coherence combined in a single response kernel. For excited initial states, this kernel filters the physical folded profile and yields a sharpness dependent transmission condition, directly connecting small scale primordial non-Gaussianity to relic neutrino capture rates.

gr-qc↗

DORA: Dynamic Online Reinforcement Agent for Token Pruning in Vision Transformers

Vision Transformers (ViTs) incur quadratic self-attention cost in the number of tokens. Most token-reduction methods adapt token identities within a prescribed layer-wise compression schedule, or search a static mask offline, and thus limit online adaptation of when and how much to prune. We propose DORA (Dynamic Online Reinforcement Agent), which learns an input-adaptive pruning policy itself for frozen ViTs. At each eligible block, a hierarchical actor decides whether to prune, how many tokens to remove, and which tokens to remove from each image's evolving representation. Because early deletions change the states observed by later decisions, DORA formulates pruning as a finite-horizon Markov decision process. Complete-prefix shadow evaluations convert final-prediction fidelity into localized per-step credit, while closed-loop accuracy feedback adjusts the fidelity penalty toward a shared accuracy-drop target. A privileged critic and all shadow computations are training-only. Deployment retains the frozen backbone and a lightweight actor that applies hard deletion and packed variable-length FlashAttention, converting token reduction into measured speedups. On ImageNet-1K with DeiT-Base, DORA reduces FLOPs by 38.4% relative to the uncompressed backbone within one percentage point of accuracy loss. Averaged across four ViT-type backbones at matched accuracy, DORA uses 13.2% fewer FLOPs and achieves 32.4% higher throughput than the corresponding per-backbone baseline means. Under zero-shot transfer to ImageNet-A, these gains widen to 20.3% and 45.6%, respectively.

cs.CV↗

Routing Without Embeddings: Fast And Interpretable Routing With Regular Expressions

Large Language Model (LLM) routers commonly rely on neural query embeddings, with larger encoders expected to better capture query intent and difficulty. Yet scaling Qwen2.5 encoders from 0.5B to 72B parameters brings little improvement in routing accuracy (Figure 1b), suggesting that small encoders may already capture the query properties needed for routing. We therefore investigate which properties matter and whether they can be extracted directly from text without a neural encoder. We introduce REGEXROUTE, a pipeline that uses sparse autoencoders (SAEs) to discover interpretable regular-expression (regex) features. Using unlabeled text, an LLM turns descriptions of grouped SAE latents into regex extractors and refines them to match latent activation patterns. These extractors supply numerical features to a lightweight routing head, eliminating neural encoding at inference (Figure 1a). Across four benchmarks, one fixed set of 128 features achieves 76.43% average routing accuracy, comparable to 76.41% for the strongest neural text encoder baseline, with much smaller latency and strong robustness. These findings establish explicit, interpretable text features as a practical basis for designing and understanding LLM routers.

cs.LG↗

Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge

Scaling test-time computation is a powerful way to improve language-model reasoning, and is particularly appealing for small reasoning models (sRMs) that are cheap to serve. However, is additional thinking always the right operation? By intervening at intermediate reasoning states across two model families and multiple scales, we find that self-refinement largely consolidates probability mass onto solutions already reachable from the current state, rather than making new ones reachable. These interventions reveal two failure regimes: execution bottlenecks, where the correct path is reachable and reflection can recover it, and knowledge bottlenecks, where relevant external information makes it reachable. Motivated by this distinction, we introduce FlyBy, a selective querying framework, and train 4B and 8B variants to reason first, diagnose what remains unresolved, and, at a knowledge bottleneck, query stronger models whose parametric knowledge extends beyond its own. Supervised fine-tuning bootstraps a multi-depth query action, and cost-aware reinforcement learning calibrates whether to query, what to ask, and how much to spend. On 1,158 hard problems across six benchmarks, FlyBy-4B achieves 45.96% pass@8, surpassing Qwen3-14B (41.64%) at 2.7 times lower serving cost, while also exceeding Qwen3-8B in pass@1 (16.85% vs. 15.31%). Scaling to FlyBy-8B further improves pass@8 to 51.81%.

cs.AI↗

From samples to shape: a finite approach to topological invariants

Given a metric space $Y$ and a compact subspace $X$, we present a method for constructing an inverse sequence of finite topological spaces by sampling points from neighborhoods of $X$ in $Y$. We show that these inverse sequences can be employed to define shape theory and to approximate shape invariants, such as Čech homology groups.

math.GN↗

A Study on the Evolution of NGC6822. I. Investigating the Star Formation History through Evolved Stellar Populations

NGC6822 is an isolated dwarf irregular galaxy in the Local Group at a distance of $\sim$ 490 kpc . In this paper, we derived the star formation history (SFH) of NGC6822 employing a method based on evolved asymptotic giant branch (AGB) stars, known as long-period variable (LPV) stars. We utilized a dataset of 329 LPVs in JHK$_s$ bands to estimate the star formation rate (SFR) over time and reconstructed the SFH of the galaxy. In addition to obtaining the SFH assuming constant metallicity values within the range of 0.0001 $<$ Z $<$ 0.012, we also adopted three distinct age-metallicity relations (AMRs) to account for variations in the chemical content of the galaxy throughout its lifetime. The SFH has been investigated in two regions. The bar region encloses a central area of 189 arcmin$^{2}$, and the outer region encircles a field beyond the bar region, extending to the radial distance of 3 kpc. In the bar region, we identified three episodes of star formation peaking at $t \sim$ 4.7 Gyr (log $t = 9.67$), $t \sim$ 1.6 Gyr (log $t = 9.21$), and $t \sim$ 48 Myr (log $t = 7.63$), where $t$ is the look-back time. In the outer region, we found one episode of star formation occurring at $t \sim$ 4.6 Gyr. The significant peak observed $\sim$ 4-5 Gyr ago further supports the hypothesis suggested by the previous research, indicating that NGC6822 passed through the virial radius of the Milky Way in the past.

astro-ph.GA↗

MiCo: Mutual Information Coverage Optimization through Semantic Erasure Modeling for Efficient MLLM Inference

Multimodal large language models (MLLMs) have demonstrated impressive performance in multimodal understanding, but processing large numbers of visual tokens results in high computational costs. While many methods have been proposed to reduce the number of visual tokens, most of them rely on heuristics and are prone to discarding substantial visual information during pruning, leading to degradation in model performance. In this work, by using a semantic erasure model, we derive a general mutual information coverage objective from task log-loss and propose MiCo, a training-free two-stage pruning method. MiCo first uses visual signals to select a representative candidate pool before visual tokens enter the language model, then performs task-aware subset selection within it. At each stage, suitable observable proxies instantiate the derived objective as a monotone submodular coverage function, which MiCo greedily optimizes under the token budget. MiCo is evaluated on diverse MLLMs ranging from 7B to 13B parameters across a broad range of image and video benchmarks spanning general visual reasoning, fine-grained OCR and grounding, hallucination detection, and long-video understanding. MiCo consistently achieves the best performance across nearly all evaluated models under all pruning ratios. On LLaVA-NEXT-13B, MiCo uses only 5.6% visual tokens, retains 97.5% of baseline performance, and achieves a 3.8-fold inference speedup. Our experiments demonstrate the effectiveness of MiCo and our mutual information coverage objective for visual token pruning.

cs.CV↗

Riesz transform on eventually Gaussian local trees

We study the Riesz transform $\mathcal{R}=\partial(-Δ)^{-\frac{1}{2}}$ on uniform local trees, metric measure spaces that are locally real trees and whose canonical Dirichlet form is built from weak derivatives along the skeleton. The reference measure $m$ may be singular with respect to the length measure $ν$, so boundedness of $\mathcal R$ is understood from $L^{p}(m)$ to $L^{p}(ν)$. In contrast with fractal-like manifolds and cable systems, the diffusion is sub-Gaussian at small scales and Gaussian at large scales. Under uniform volume growth, two-sided heat kernel estimates and a pointwise gradient estimate for the heat kernel, we prove that a local Dini condition on the scale function implies boundedness of $\mathcal{R}$ on $L^{p}$ for every $p\in[2,\infty)$, and hence the reverse Riesz inequality for every $p\in(1,2]$. Conversely, boundedness of $\mathcal{R}$ for some $p<2$, or a reverse Riesz inequality for some $p>2$, forces the space to be one-dimensional at small scales. We show that a reverse Hölder inequality for harmonic functions yields the gradient estimate, and verify all hypotheses for spaces carrying a geometric group action whose generators have bounded displacement. As an application, for the alternating Vicsek fractafold in $\mathbb Z^{d}$ we determine the exact ranges of $p$ for which the Riesz and reverse Riesz inequalities hold.

math.FA↗

Observation of a topological edge state among localized bulk states in the anisotropic quantum Rabi model

Topological phases are governed by discrete symmetries that protect boundary modes against local perturbations. When translational periodicity is absent, the bulk states also become localized, so that a topological edge state can no longer be distinguished from them by spatial localization alone. Here, we investigate the topological edge state (TES) and bulk eigenstates of the anisotropic quantum Rabi model (AQRM) in a trapped-ion quantum simulator. The AQRM hosts a topological phase in a one-dimensional synthetic lattice, whose translational symmetry is broken by the non-uniform couplings scaling with the site index. While both the TES and bulk states show localized distributions, we find that the TES exhibits well-defined chirality and near-complete spin--boson separability as signatures of the topological phase, in contrast to the bulk states. Phase-space tomography further reveals that the bosonic component of the TES is a squeezed vacuum state, with squeezing up to 6.45 dB. These results identify the TES through its intrinsic topological signatures and establish eigenstate-level characterization as a route to probing topological phenomena.

quant-ph↗

Stacked Intelligent Metasurface-Diffractive Deep Neural Networks for Onboard Terrain Classification from SAR Level-0 Raw Data

Real-time terrain classification directly from Level-0 raw Synthetic Aperture Radar (SAR) data remains restricted by traditional digital-centric paradigms, where the processing of noisy, high-dimensional patches is hindered by computationally intensive processors and significant downlink latency. To address these fundamental limitations, this work establishes a new research paradigm for autonomous on-board sensing by proposing a Stacked Intelligent Metasurface-Diffractive Deep Neural Network (SIM-D$^2$NN). This architecture leverages the physical wave-propagation medium to offload inference tasks from digital processors to a physical device. By executing in-wave feature mapping, the SIM-D$^2$NN facilitates a move toward an integrated `compute-while-transmitting' framework, providing an alternative to the traditional `digitize-then-process' sequence. The multi-layer metasurface is positioned at the forefront of the satellite communication module. The initial layer modulates the raw SAR data through both amplitude and phase adjustments, where a 90$^\circ$ phase rotation is introduced as a lightweight but effective augmentation strategy to enhance robustness against noise and Doppler distortions. Subsequent layers learn variable phase shifts, enabling advanced feature mapping for the classification task. The classification results at the terrestrial station can be directly obtained based on the signal amplitude received at each antenna. This design reduces reliance on downlink bandwidth and high-power terrestrial computing, achieving performance around 90% in the binary task directly from real raw SAR data in terms of accuracy, precision, recall, and F1 Score. Therefore, our method helps bridge the gap between next-generation remote sensing tasks and in-orbit processing needs, paving the way for computationally efficient remote sensing applications.

eess.SP↗

HOCCL: Offloading Collective Communication from GPU Cores to Accelerate Distributed Training

Large language model training involves massive computation on GPU streaming multiprocessors (SMs), the primary compute units of GPUs. Since SMs host specialized accelerators such as Tensor Cores, their efficient utilization is critical to training efficiency. Unfortunately, existing collective communication systems compete with computation for SMs, as they consume SMs for communication-related data movement and synchronization operations. We observe that communication can, in principle, be driven by DMA engines, thereby eliminating SM involvement in communication. Based on this insight, we propose HOCCL, a zero-SM collective communication framework consisting of three components: a stream manager, a point-to-point (P2P) executor, and a collective scheduler. The stream manager preserves operator-level temporal ordering with other GPU kernels. The P2P executor enables zero-SM point-to-point communication, while the collective scheduler orchestrates P2P transfers to maximize bandwidth. Experiments show that HOCCL preserves near-peak communication performance, achieving within 3% of the state of the art on average, while eliminating communication occupancy on nearly 10% of total GPU SMs. By freeing SM resources for computation, HOCCL improves end-to-end training throughput by up to 5%.

cs.NI↗