arXiv Science⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 1,513 records · Page 84Linked to original sources

FedLore: Communication and Memory Efficient Federated Learning via Shared Gradient Low-Rank Projection

Federated training of foundation models is constrained by client memory and communication costs. LoRA-based methods reduce these costs through low-rank adapters, but their fixed rank budget can limit adaptation. Gradient low-rank optimization offers greater flexibility, yet independently chosen client subspaces create a problem we term \emph{subspace fragmentation}: local projections interact with data heterogeneity to bias aggregated directions, while aggregation can increase update rank and communication cost. Thus, accurate local gradient compression need not preserve global descent. We propose \texttt{FedLore}, which shares a low-rank optimization basis within each round and refreshes it across rounds. The shared basis enables exact aggregation in low-rank coordinates and eliminates the identified projection bias. Subspace refresh allows the accumulated model update to exceed the per-round rank budget. We characterize the aggregation bias and establish an $O(T^{-1/2})$ stationarity bound for the projected-SGD variant under a global-gradient coverage condition and standard smoothness and variance assumptions, with bounded gradient heterogeneity. Experiments on vision and language tasks, including federated pre-training, show that \texttt{FedLore} outperforms the evaluated low-rank adapter baselines and matches or exceeds full-parameter training, while reducing communication and optimizer-state memory.

cs.AI↗

Learning to Classify Threading in Melts of Rings

Dense melts of nonconcatenated ring polymers exhibit anomalous dynamics and nonlinear rheology, in some cases linked to long-lived inter-ring threading entanglement. Detecting these geometric constraints remains computationally demanding, with different methods being based on different geometric features tuned to specific models. Here, we introduce a machine-learning framework that detects threading directly from the writhe of ring conformations and we demonstrate its broad generalisability. Trained on simulations of isolated, distance-constrained pairs of short rings, a one-dimensional convolutional neural network (CNN1D) operating on the writhe profile classifies threading states with over 98% accuracy on held-out ring pairs. Remarkably, this model generalises without retraining to equilibrium dense melts spanning a wide range of densities and chain lengths (N = 100 to 1600), maintaining true-positive rates above 91% with false-positive rates near zero. Misclassified conformations are limited to physically ambiguous, shallow threading events which are arguably not impacting the dynamics of the rings. Effectively, writhe captures the interaction between the two rings, allowing writhe-based classifiers to improve the topological analysis of large ring-polymer systems with a 3-fold speed-up over the state-of-the-art minimal-surface detection. These results establish the writhe-trained neural networks as an accurate, scalable tool for characterising entanglement in ring polymer melts.

cond-mat.soft↗

The Buchholz Algebra from the Universal Resolvent Algebra--Particle Structure, Gibbs States, and Closed-Path Expansions without Fock Representations

We construct the Buchholz algebra without a Fock representation, as the bounded inverse limit of intrinsic particle-cutoff quotients of the gauge-fixed resolvent algebra, with isometric particle-degree coordinates. Its finite matrix corners carry canonical traces, KMS states, and trace-norm product formulas, together with closed-path expansions that lay the groundwork for path-integral representations. The known Fock-space description and a proper quasilocal subalgebra arise as realizations.

math-ph↗

Open-Source Multi-Wire SPI Readout for Wearable Ultrasound Probes

Wearable ultrasound probes must transfer increasingly large acquisition payloads while maintaining compact, low-power electronics. In TinyProbe, the current bottleneck in data transfer occurs between the acquisition FPGA and the wireless system controller. This work presents an open-source, multi-wire SPI readout interface that uses serial command and address phases followed by a build-time-selectable dual- or quad-lane payload phase that is intended to address this bottleneck by increasing the potential bandwidth over the wifi limit while retaining compatibility with the Microcontroller-centric wearable US architecture. The interface emulates a serial flash memory, enabling compatibility with a broad range of microcontroller families and their existing peripheral interfaces. On the FPGA, the data path connects the existing acquisition FIFOs to the SPI interface through clock-domain crossing, sample reshaping, and packing into 32-bit words. Dual-SPI readout is integrated into the existing IGLOO2/SiWG917 TinyProbe architecture and verified at an SCLK frequency of 5 MHz. A separate Kria K26 testbed is used to characterize the FPGA SPI interface independently of the acquisition and wireless subsystems, demonstrating error-free transfers at SCLK frequencies up to 66 MHz. These measurements identify the SiWG917 multi-lane SPI implementation as the next bandwidth-limiting component and motivate a future upgrade of the system controller. The HDL and MCU implementations are released under a permissive open-source license.

cs.AR↗

Relativistic Hirshfeld atoms in a molecule: An information-theoretic view, with application to Drude oscillator dispersion models

Several ad hoc dispersion models for density-functional theory are based on the use of Hirshfeld (or "stockholder") partition of a molecular charge density, which provides an in situ definition of atomic size. We show that a recently introduced "optimized" quantum Drude oscillator model for dispersion admits a closed-form solution in terms of the Lambert $W$ function, whose branches identify the compact and diffuse oscillator solutions. The compact solution determines the $C_8$ (dipole-quadrupole) dispersion coefficient analytically from the free-atom polarizability, $C_6$ coefficient, and van der Waals radius, without any reference $C_8$ data. Next, we provide a formal basis for a relativistic version of the atoms-in-molecule Hirshfeld partition. Using four-component Dirac-Hartree-Fock densities for isolated atoms defines a strictly positive deformation field that carries the relativistic changes in atomic density into the Hirshfeld partition. A uniqueness theorem for the non-relativistic case is extended to relativistic Hirshfeld atoms and admits an asymptotic expansion through quadratic order in the fine-structure constant. Normalization requires the relativistic density correction to reshape the reference atom while preserving its population. Finally, four-component polarizabilities and C6 coefficients are reported for closed-shell atoms and ions, which supply the reference data required to extend atoms-in-molecules dispersion models into the heavy-element regime. Periodic trends are observable in a scalar contraction factor that measures relativistic effects.

physics.chem-ph↗

Beyond Domain-Level Adaptation: Margin-Oriented Semantic-Appearance Interaction Correction for Personalized Federated Vision-Language Models

Federated parameter-efficient fine-tuning enables distributed clients to adapt pretrained vision-language models without sharing raw data or updating the full backbone. Its effectiveness, however, is limited by domain heterogeneity across clients. Existing personalized methods separate globally shared knowledge from client-specific style, but they largely treat each domain as a class-agnostic transformation. We show that this abstraction is insufficient: the cross-domain displacement associated with a fixed domain varies across semantic classes, and only a subset of these class-domain residuals damages the image-text decision margin. We therefore propose Margin-Oriented Semantic-Appearance Interaction Correction (MOSAIC), which first constructs a decision-aware harmfulness score that measures whether a training-derived class-domain residual favors a competing text prototype over the true class. It then models fine-grained class-domain interactions with a low-rank residual adapter whose class factors and residual basis are globally shared while domain factors remain client-private. An image-conditioned gate further controls candidate-wise correction, and harmful-pair-aware reweighting prioritizes decision-relevant residuals during local optimization. Extensive experiments on Office31, OfficeHome, and DomainNet100 demonstrate that MOSAIC consistently improves macro-client top-1 accuracy across all evaluated domain-shift and joint domain-label-shift settings.

cs.CV↗

Measuring the Stability Assumption Behind Action Chunking

Action chunking improves the performance of policies learned by behavioural cloning, and several mechanisms have been proposed to explain why, including temporal consistency, horizon reduction, representation learning, and reduced error compounding. We instead study what happens to an action error once it enters the system. At each state, we inject a small action error and measure how fast it grows or shrinks under two execution regimes: open-loop, where the rest of the chunk is replayed without replanning, and closed-loop, where the policy replans after the perturbation. The fitted rate labels each state as contracting, expanding, or unresolved. Across twelve manipulation tasks from three benchmark suites, we find that confidently stable states are rare, while error amplification is common among states whose propagation rate can be resolved. We further find that the measured propagation rate depends strongly on the fitting horizon: amplification is typically front-loaded, so short windows can overestimate longer-horizon propagation. Finally, we train predictors on these labels and find that a state's open-loop regime can be recovered from camera frames and proprioception alone, while its closed-loop propagation is only partially recoverable because it also depends on how the policy acts after the perturbation. These results suggest that error-compounding arguments alone do not provide a complete account of action chunking: neither passive open-loop dynamics nor policy replanning consistently contracts an injected error, and replanning rarely turns open-loop amplification into confident contraction. This suggests that closed-loop reactivity should be trained explicitly, using perturbation- and tree-coverage-oriented training to expose policies to deviations they must recover from, rather than expected to emerge reliably from standard imitation learning.

cs.AI↗

What Makes Something Hard(er)? Explaining Question Difficulty in Natural Language

Difficulty is one of the most fundamental properties of a question: it determines whether the question can meaningfully discriminate between models of differing ability. Although a variety of methods can now estimate or predict difficulty automatically, they yield only a single descriptive number, with no account of the underlying factors that make a question difficult in the first place. In this work, we propose a data-driven approach that automatically generates and validates natural-language hypotheses explaining what makes one question harder than another. We first estimate each item's difficulty from the responses of a large pool of LLMs using Item Response Theory. We then sample contrasting sets of easy and hard questions and prompt an LLM to propose candidate explanations of the difference, which are subsequently validated and selected on held-out questions. Experimental results across three datasets spanning mathematical, logical, and commonsense reasoning show that our method produces interpretable and predictive hypotheses. On their own, they predict the difficulty of unseen questions competitively with, or better than, advanced black-box difficulty regressors; used as additional features, they further improve those regressors, implying that they discover difficulty signals that existing models fail to capture. Moreover, we demonstrate that editing questions according to a hypothesis can shift their measured difficulty in the expected direction, indicating that the discovered hypotheses are causally valid difficulty factors rather than post-hoc descriptions. Our approach thus turns a purely descriptive difficulty score into actionable statements.

cs.CL↗

Diagonal operators on Janson-Sobolev and Janson-Sobolev-Hardy spaces

We study the Banach space and operator factorization structure of Janson-Sobolev and Janson-Sobolev-Hardy spaces. This new class of martingale spaces is determined by a $q$-adic filtration, a subspace $V\subset \mathbb{R}_0^{l\times q}$, and a rearrangement invariant function space $X$. Our main result shows that, for every bounded diagonal operator $D$, the operator $S = \sum_{t=1}^s λ_{\mathcal U}^{k_t}(D)Q_{\mathcal K_t}^{\mathcal B}$ determined by the linear functionals $λ_{\mathcal U}^{k_t}(D)$ and the canonical projections $Q_{\mathcal K_t}^{\mathcal B}$, almost projectionally factors through $D$ with constant $1^+$. As consequences, we obtain factorization results for diagonal operators both under a natural boundedness condition on the canonical projections and for all spaces equipped with the $L^1$-norm.

math.FA↗

Overlapping Subcritical Bubbles: Free Energy and Lifetime

Subcritical bubbles can form an appreciable population during weak first-order phase transitions, but are usually treated as isolated fluctuations. This raises the question of whether spatial overlap between neighboring subcritical bubbles can modify their evolution. We address this question by combining analytic free-energy calculations for composite Gaussian profiles with Langevin simulations of overlapping configurations. We find that the overlap lowers the free-energy cost and generally increases the lifetime of subcritical bubbles, with the enhanced persistence potentially feeding back on their abundance. Such overlap can therefore generate collective effects and should be incorporated into kinetic descriptions of subcritical-bubble populations.

hep-ph↗

After Cooperation Is Learned: Gradient Routing and Optimizer-Dependent Maintenance in Multi-Agent Reinforcement Learning

Cooperative MARL is commonly evaluated through cooperation discovery from random initialization, leaving open whether continued optimization can destabilize learned cooperation. Actor-critic comparisons can also conflate critic presence with value gradients entering shared actor representations. We study cooperation maintenance, defined as the survival of a behaviorally verified cooperative policy under continued training. We formulate maintenance as a right-censored event-time problem and compare matched warm starts: X0 allows value loss gradients to update shared actor features, X1 retains the critic while blocking those gradients, and X5 removes the learned critic as a critic-free reference. This isolates direct value-gradient access while controlling initialization, critic computation, and evaluation. Positive reward scaling preserves strategic preferences and equilibria while perturbing learning dynamics. Gradient audits confirm the intended routing pathways, and frozen-policy torso perturbations probe whether route-induced updates align with local cooperation boundaries. In confirmatory MinEx and CleanUp-lite experiments, higher scales selectively increase maintenance sensitivity in X0; X1 remains near the censoring ceiling, and X5 has no confirmed events in the tested settings. In CleanUp-lite, route-by-scale displacement is associated with reduced local cooperation margins; MinEx shows a weaker, optimizer-dependent effect. These results identify a conditional, scale-sensitive maintenance risk associated with direct value-gradient routing rather than a universal failure of critics.

cs.MA↗

Joint Geometric and QoS-Aware Routing in Optical LEO Satellite Networks via DRL

Optical inter satellite links ISLs are becoming the backbone of modern LEO constellations offering high capacity and low latency but introducing stringent geometric and physical layer constraints Routing in such networks must therefore account for time varying topology jitter induced outage and the heterogeneous reliability of intra and inter plane optical links aspects that classical shortest path or existing learning based schemes do not fully capture This paper develops a joint geometric and QoS aware routing framework for optical LEO networks We derive a closed form outage expression under Gaussian beam propagation with pointing errors and obtain analytical maximum feasible link ranges for different ISL classes These relations remove beam divergence from the optimization variables and embed optical feasibility directly into the routing layer leading to a latency reliability capacity constrained routing formulation that is proved to be NP hard To enable scalable decision making we cast snapshot routing as a Markov decision process and introduce an angle constrained masked deep Q network AC MDQN that integrates optical feasibility masks potential based latency shaping and a geometry aware corridor filter around the source destination great circle path This design significantly reduces the effective action space complexity while preserving near optimal routing choices Simulations on a Starlink like constellation demonstrate that AC MDQN achieves end to end latency within 1 to 2 percent of constrained shortest path solutions remains robust under varying pointing jitter and supports controllable hop latency trade offs through reward design The results confirm that the proposed framework provides an efficient and physically consistent routing solution for large scale optical LEO networks

eess.SP↗

SL-RFSIM: Enabling Scalable Multi-Hop 5G NR Sidelink Mesh Networking in OpenAirInterface

Recent 3GPP releases have extended 5G New Radio (NR) Sidelink (SL) to support device-to-device (D2D) relay and multi-hop capabilities. This paper presents SL-RFSIM, a component-based experimentation framework extending OpenAirInterface (OAI) with scalable multi-hop NR SL capabilities. SL-RFSIM replaces the legacy OAI RF simulator with a broker-based publish/subscribe architecture enabling arbitrary peer-to-peer connectivity while preserving compatibility with the OAI protocol stack. The framework further integrates pluggable mobility, propagation, reception, and monitoring services, and supports Layer-2 mesh networking through BATMAN-adv. Experimental evaluation on the SLICES-RI research infrastructure validates the proposed architecture through representative mesh networking scenarios and identifies the current software bottlenecks limiting scalability. SL-RFSIM provides an open-source foundation for reproducible experimental research on 5G NR SL and future multi-hop cellular mesh networks.

cs.NI↗

Generalization in Neural Networks Through the Lens of Magnitude Potential

Explaining generalization and training dynamics in neural networks remains a challenge, and various approaches have been developed to study different aspects of these phenomena. In this paper, we introduce the idea of {\em magnitude potential} -- a quantity based on the theory of metric magnitude -- that reflects how well an arbitrary point is represented by a given set. We find that this basic quantity can be applied to examine various features in neural generalization. The ratio between the magnitude potential with respect to a class and with respect to the entire data, computed at the logit layer, is informative of the representation of the point. In experiments, these ratios for individual training points are found to be correlated with the Feldman memorization scores. Magnitude potential ratios aggregated across points detect structural changes in the decision boundaries and provide a geometric indicator of grokking in modular arithmetic. Although the magnitude potential ratio and neural collapse are both closely associated with intra-class and inter-class geometric structure, the magnitude potential ratio remains informative even when neural collapse is explicitly suppressed.

cs.LG↗

Yo-ByT5: Efficient and High-Fidelity Diacritic Restoration for Yorùbá

Yorùbá is a widely spoken tonal language that depends on diacritics to avoid lexical ambiguity. However, it is often written without these diacritics, thereby hindering downstream Natural Language Processing (NLP) tasks. In this paper, we introduce Yo-ByT5, a byte-level Automatic Diacritic Restoration (ADR) model fine-tuned from ByT5-small. We evaluate Yo-ByT5 alongside five publicly released Yorùbá ADR models and one open-weight large language model (LLM) on the YAD benchmark under a consistent protocol. Our results demonstrate that Yo-ByT5 matches the performance of the strongest existing model, mT5-base, with a DER of 10.14% and a CER of 3.48%. Furthermore, it exhibits superior text fidelity despite using approximately half the parameter count of mT5-base. We also release our training code and model outputs, as well as call for the development of a larger, purpose-built benchmark for Yorùbá diacritic restoration.

cs.CL↗

Low Overhead IMU Assisted Predictive Beam Management for Multiband LEO Direct to Device Links

Direct to device D2D connectivity from low Earth orbit LEO satellites is moving to Ku band, where a handheld terminal must obtain directional gain from several small phased array panels distributed around the chassis. Beam management then becomes a joint satellite panel beam selection problem with hundreds of candidates, and ordinary hand motion can change the best candidate within one decision interval. This paper proposes a low overhead predictive beam management scheme for a multiband LEO D2D downlink in which a low frequency anchor link carries control signalling and fallback traffic, and a Ku band link carries broadband data. The handset inertial measurement unit IMU reports attitude and angular rate with a known delay; the scheme extrapolates the delayed attitude to the current orientation, scores all satellite panel beam candidates analytically from the satellite ephemeris and panel geometry, and trains only a small candidate set built with a satellite panel diversity rule under a fixed pilot budget. The Ku band link is activated only when its predicted post training rate exceeds the anchor rate by a margin. Trace driven Monte Carlo simulations with a four panel handset and two visible satellites show that, with six pilots 0.3 percent training overhead, the proposed scheme improves mean goodput by 21.2 percent and reduces broadband outage by 57.6 percent at 90 degrees per second relative to an equal budget delayed attitude baseline, and operates within 3.8 percent of a zero overhead oracle.

eess.SP↗

Autocatalysis and Boundary Stability in Chemical Reaction Systems

The next-generation matrix method is a powerful tool for computing the basic reproduction number in compartmental mathematical models of infectious diseases. The method has been recently extended to mathematical biochemistry, where it can be used to establish parameter regions of stability and instability for boundary steady states. Several significant challenges in its application remain, however, particularly around establishing conditions under which the method is mathematically valid, computationally tractable, and biologically meaningful in the biochemical setting. In this paper, we address these challenges by shifting the interpretation from new infections in the epidemiological setting to autocatalysis in the biochemical setting. We introduce a graph, called the ALT-graph (autocatalysis-leak-transition graph), which decomposes the contribution of each reaction to the Jacobian as autocatalytic, leak, or transition edges. We then present a systematic algorithm for splitting the ALT-graph, which is guaranteed to produce a valid reproduction number, $ρ(FV^{-1})$, while also decreasing computational complexity by lowering the rank of $FV^{-1}$. We apply the method to models of both biochemical reaction networks and infectious disease spread.

math.DS↗

Fusing Visual and Textual Representations via Multi-layer Fusing Transformers for Vietnamese Visual Question Answering

In recent decades, artificial intelligence has made significant progress in understanding and interacting with images. One of the important applications of this technology is Visual Question Answering (VQA), a research field that requires computers to understand and answer questions about images in a natural manner. Despite extensive research and development in VQA for English, there have been very few similar efforts made for other languages, especially Vietnamese. This gap presents a significant challenge and opportunity for the advancement of VQA technology in the Vietnamese language context. By bridging this gap, the field of Vietnamese VQA not only enriches the diversity of research in artificial intelligence but also enables practical applications in various domains, such as education, healthcare, and entertainment, catering to Vietnamese-speaking populations worldwide. Thus, the exploration and development of Vietnamese VQA systems hold immense potential for advancing both research and practical applications in the intersection of computer vision and natural language processing. In this paper, we propose a Multi-layer Fusing Transformer model utilizing a cross attention module to combine multiple modality features of images and texts from different layers in an aggregated representation. Our architecture allows us extract information from low level to high level. Through detailed experiments and ablation studies, our model achieves promising results against the competitive baselines in ViVQA dataset for Vietnamese language.

cs.CV↗