arXiv Science⌕ Search

SEARCH · arXiv Science

Search arXiv Science

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,171 records · Page 65Linked to original sources

SkillPE: Creativity-Oriented Cinematic Skill Evolution for Text-to-Video Prompt Engineering

Achieving high-quality, cinematic results in text-to-video generation remains challenging for non-experts, whose prompts often lack professional narrative and creative design. We propose SkillPE, a prompt engineering (PE) framework that evolves reusable cinematic skills from expert-authored seeds. SkillPE represents shot logic, composition, lighting, sound design, and other filmmaking cues in a fine-grained format, and retrieves movie references categorized as resonators (good matches), dissonants (weak matches), and divergents (creatively useful near-misses). The first two refine when and how a skill should be applied, while divergents inspire alternative cinematic realizations at different degrees of modification while preserving the user intent. Candidate skills are assessed through generated videos along prompt fidelity, cinematic quality, narrative appeal, and creativity to construct the final skill libraries. Experiments on StoryEval and VBench show improvements of up to 1.40 points over the strongest external baseline and 0.51 points over seed skills on 7-point four-dimensional evaluation, while remaining competitive on benchmark-native metrics. Overall, SkillPE offers a practical approach to balancing fidelity and creativity in cinematic text-to-video generation. Code is available at https://github.com/Ais0n/SkillPE .

cs.CV↗

Toward Resolution Convergence for Star Formation and Galactic Winds in Galaxy Simulations

The predictive power of hydrodynamic simulations of galaxies is often hampered by the resolution dependence of sub-grid models for star formation (SF) and stellar feedback. We present three-dimensional hydrodynamic simulations of an M82-like low-mass starburst galaxy at resolutions of 4, 8, 16, 32 pc and develop optimized sub-grid prescriptions to mitigate these numerical inconsistencies. For star formation, we adopt a $\sim20$ pc physical-scale control volume, corresponding to the typical size of giant molecular clouds (GMCs), for gas accretion onto sink particles, reducing its dependence on numerical resolution. We further incorporate an empirical, resolution- and density-dependent star formation scheme that accounts for variations in gas surface density across spatial scales, as suggested by observations. Furthermore, we refine the supernova feedback coupling by anchoring the injection volume to the physical cooling radius of the supernova remnant ($R_{\rm cool}$) and utilizing a resolution-dependent thermal-to-kinetic energy partition. These improvements substantially enhance convergence across 4 to 32 pc: variations in the total stellar mass formed and kinetic energy are reduced to $\sim20\%$, while relative errors in the cool-phase mass outflow rate and integrated X-ray luminosity decrease to below $\sim40\%$. The warm and hot outflow phases remain less well converged, with variations of $\sim50\%$ and $\sim75\%$, respectively. The upgraded model also improves the convergence of normalized wind properties, with relative errors below $\sim50\%$, $\sim80\%$, and $\sim20\%$ for the time-averaged mass loading factor, energy loading factor, and X-ray emission efficiency, respectively. Our results provide a practical framework for achieving more robust simulations of starburst galaxies.

astro-ph.GA↗

Audio Tokens as a Budgeted Resource: Marginal-Utility Allocation for Scalable Audio Representations

Discrete audio tokens are widely used as a representation interface, yet fixed-depth RVQ tokenizers allocate equal capacity to every frame despite varying refinement value. We introduce UniAdapt, which learns marginal utility of RVQ refinements on a frozen codec and allocates them under exact serialized-bit budgets. A rate-independent causal controller predicts acoustic utility, while an optional semantic head supports speech-only utterance-level allocation; measured acoustic and semantic marginal gains on speech have a correlation of 0.42. For causal allocation, a primal-dual allocator selects prefix-valid depths, while an exact guard constrains each sequence prefix to its matched fixed-depth serialized budget. Under utterance-level allocation, UniAdapt reduces Log-STFT distortion by 1.07-4.39 percent across speech, music, and environmental audio without larger budgets. Causally, it improves three of four speech rates with zero violations across 800 utterance-rate evaluations and runs faster than real time. A 20-listener utterance-level MUSHRA study shows a significant 3.52-point speech improvement, with no significant differences on music or environmental audio. These results support separating utility prediction from budget enforcement for scalable, budget-conditioned audio representations.

eess.AS↗

Medium effects on neutron star modified and direct Urca cooling rates

We study the effects of in-medium modification of the elementary cooling processes on observable properties of isolated neutron stars. We then deduce the neutron star mass distributions compatible with the cooling analysis and compare with current theoretical models. We conclude that current cooling data require fast direct Urca (DU) cooling, moderated by proton superfluidity, to be active in most neutron stars, and that the DU onset threshold must lie below canonical masses. In that case medium modifications of modified Urca (MU) rates are practically insignificant, but the $nn$ Bremsstrahlung rate plays a dominant role.

nucl-th↗

The directed temporal exploration problem

We study the temporal exploration problem on temporal digraphs. We prove that a lifetime of $O(n^2)$ suffices to guarantee the existence of a temporal exploration on always-unilateral temporal digraphs. We complement this with a $Ω(n^2)$ lower bound, even in the case where each snapshot has maximum undirected degree 2; for always-strong temporal digraphs, the lower bound still holds even if the maximum undirected degree is 3. This stands in stark contrast with the undirected setting. For the large minimum degree setting, we show that a lifetime of $4n/3 - 1$ is sufficient and necessary for guaranteeing the existence of a temporal exploration on temporal digraphs where each snapshot is semicomplete. For always-strong temporal digraphs where each snapshot has minimum undirected degree at least $n - c - 1$, we prove that a lifetime of $O(cn)$ guarantees the existence of a temporal exploration, and we also prove that this is asymptotically tight. From a computational perspective, our results for temporal semicomplete digraphs also yield a polynomial-time, factor-$4/3$ algorithm for deciding if a temporal semicomplete digraph admits a temporal exploration within the first $\ell$ snapshots. We complement this showing that no polynomial-time, factor-$(4/3 - ε)$ approximation algorithm exists, even if every snapshot is a tournament, unless P$=$NP.

cs.DS↗

VarioPath: Workload-Aware All-to-All Communication for PCIe GPU Clusters

AlltoAllv communication is a critical primitive in distributed large-model inference, particularly for mixture-of-experts (MoE) models. The growing adoption of PCIe GPU systems for cost-efficient inference makes AlltoAllv performance on these systems increasingly important. Without a dedicated scale-up interconnect (e.g., NVLink or Infinity Fabric), PCIe GPU systems carry both intra-node and inter-node traffic through the PCIe hierarchy, where concurrent transfers can contend for PCIe link bandwidth. This link contention, compounded by skewed traffic distributions and dynamic traffic demand, makes efficient AlltoAllv scheduling challenging. Existing approaches are either poorly suited to PCIe GPU systems or incur substantial schedule synthesis overhead that reduces their practicality in real-world deployments. We present VarioPath, an efficient AlltoAllv scheduling framework for PCIe GPU systems. It combines an offline topology-aware analyzer with an online demand-aware scheduler. The analyzer records contention-free transfer patterns as AlltoAllv channels and exploits topology symmetry to build a compact catalog for efficient search. The online scheduler decomposes each AlltoAllv invocation's demand across a sequence of channels, adapting to rapidly changing and skewed traffic while incurring low planning overhead. Evaluation on four platforms (up to 256 GPUs) shows average AlltoAllv speedups of 5.88x over FAST and 1.72x over DeepEP. End-to-end experiments show that VarioPath reduces Qwen3 inference latency by up to 27.2% and Wan2.1 generation latency by 6.1%.

cs.DC↗

A Comparative Study on Robust Topology Optimization of Design-Dependent Pressure-Actuated Compliant Mechanisms with Quadrilateral Elements

This paper presents a comparative study of compliant mechanisms generated using a robust topology optimization technique involving design-dependent pressure loads. Design domains are parameterized using standard and higher-order quadrilateral elements. Both eroded and blueprint configurations are considered. A min-max optimization model combined with an output-spring method is employed to extremize the mechanisms' output displacements. A volume and a strain energy constraint are applied to the blueprint and the eroded designs, respectively. The optimization process is executed using the method of moving asymptotes. Numerical experiments are performed to optimize the pressure-actuated inverter and gripper mechanisms using Q4, Q8, and Q9 elements, and the results are compared. The research highlights how quadrilateral element selection influences both the resulting topologies and performance characteristics.

cs.CE↗

SAGE: Structured Strategic Reasoning for Efficient LLM Game Playing

A strong LLM strategic agent should reason prospectively over uncertain futures, adapt its strategy to opponents' behavioral tendencies, and continuously recalibrate its decision process from interaction experience. However, incorporating these sources in free-form reasoning could lead to unsupported strategic assumptions, inconsistent opponent estimates, and harmful interference from irrelevant historical interactions. To address these issues, we propose SAGE, a training-free inference-time framework that structures LLM strategic reasoning around three coordinated operations: anchor, adapt, and recalibrate. SAGE first anchors reasoning to an equilibrium policy that provides a strategically valid prior. It then conditions deviations from this anchor on a soft belief over opponent behavioral tendencies, enabling opponent-specific exploitation. Finally, SAGE distills strategically related interactions into counterfactual hypotheses about previously missing considerations, allowing past experience to recalibrate the model's reasoning. We evaluate SAGE on three repeated imperfect-information games: Leduc Hold'em, Liar's Dice, and Goofspiel, against various opponent types in each game. Compared with reasoning-intensive LLM agents, including Suspicion-Agent, ReTA, Agent-Pro, EMO, and Hypothetical Minds, SAGE achieves up to a 127.6% payoff improvement in Liar's Dice while reducing input and output token usage by up to 80% and 90%, respectively. In direct match-up play, it attains non-negative mean payoff against 5/10, 8/10, and 8/10 evaluated opponents in Leduc Hold'em, Liar's Dice, and Goofspiel, respectively, while using relatively fewer tokens. Code is available at https://github.com/chenzhwsysu57/SAGE.

cs.AI↗

The Composition Gap in Dataset Distillation

Dataset distillation compresses a training set into a small synthetic set, usually evaluated one at a time. In federated and data-governance settings, several parties distill their own data and a user trains on their union. We ask whether the union of separately distilled sets reproduces training on the union of the real data composability and show that it can fail even when every source is distilled exactly and the total budget admits an exact joint distillate. Compressing a training trajectory into fewer steps transforms the source statistics nonlinearly, so averaging compressed sources differs from compressing their average. For quadratic objectives we derive the exact composition error for two-to-one step compression in terms of the source-Hessian variance and the linear terms of the losses, and on a smooth network at small step sizes this prediction captures the local endpoint discrepancy in magnitude and direction. For learned synthetic sets, however, the composed error decomposes exactly into this local discrepancy and an aggregate source residual. Under endpoint matching the residual exceeds the structural term by more than an order of magnitude, and under distribution matching the two terms partly cancel. Joint distillation also retains an accuracy advantage when both sets are distilled from the same dataset, where the local discrepancy is exactly zero. Training fidelity and downstream accuracy are therefore distinct requirements, neither established by evaluating each set on its own.

cs.LG↗

Learning to Steer, Steering to See: Unveiling the Geometry of RLVR in Large Language Models via Trainable Vectors

Reinforcement learning (RL) has become a key paradigm for enhancing the reasoning of large language models, yet the high dimensionality of parameter updates makes its training dynamics hard to analyze. We study reinforcement learning with verifiable rewards (RLVR) and use vector steering to identify a low-dimensional effective manifold in activation space associated with RL-induced gains. We uncover two geometric properties. (1) Effective Manifold Capacity: the capacity needed to reproduce RL gains can be very small but is not infinitely compressible; at extremely low capacity, intervention dimensionality and input-dependent expressiveness become key constraints, and this requirement varies with injection depth. (2) Control Manifold Separation: effective control directions lie mainly in the low-variance complement of the activation principal subspace. Within a task and base model, the learned geometry stays largely consistent across training configurations, and across tasks geometric alignment correlates with capability transfer. Experiments on 5 LLMs and 6 verifiable-reward tasks support these findings. We then propose Alpha-Stabler, a plug-and-play framework with a Predictor that monitors principal-subspace intrusion for early collapse warnings, and a Controller that removes the principal-subspace component of activation gradients during backpropagation while preserving the orthogonal complement. Alpha-Stabler stabilizes training for 2,000 steps and consistently improves RL gains, offering practical insights for robust post-training. Code: https://github.com/caiyuchen-ustc/On_Policy_Vector_Training

cs.LG↗

CRISP: Cultural Reward Modeling for Implicit Situated Propriety

As large language models (LLMs) are increasingly deployed across countries and regions, the ability to recognize and respond appropriately to diverse cultural contexts becomes increasingly important. However, existing research has largely focused on cultural knowledge or tasks with predefined response spaces, while open-ended culturally situated behavior remains comparatively underexplored. In this work, we introduce CRISP-RM, a culturally situated reward model that assigns rewards according to cultural appropriateness in open-ended social scenarios. During policy optimization, we further introduce Norm Grounding Supervision (NGS), providing guidance that enhances the policy's sensitivity to relevant cultural norms. To construct culturally situated data, we employ a collaborative multi-agent framework that instantiates implicit cultural norms into diverse social scenarios and further curate NormCompass as a dedicated testbed. We conduct comprehensive experiments to evaluate the effectiveness of CRISP-RM in both reward modeling and policy optimization. Best-of-\(N\) experiments show that CRISP-RM consistently outperforms strong general reward models. During GRPO policy optimization, CRISP-RM generally improves culturally situated behavior, while incorporating NGS yields further gains. Further analyses demonstrate the advantages of CRISP-RM in distinguishing culturally appropriate behavior beyond superficial fluency and politeness, while NGS provides complementary gains during policy optimization by improving norm grounding.

cs.CL↗

E-WAVE: Event-based Continuous Optical Flow via Warping-Aligned Visual Encoding

Temporally dense optical flow is essential for dynamic perception in immersive VR/AR systems, where rapid head, hand, and object motion must be continuously captured and tracked. Existing frame-based optical flow estimation methods are constrained by the tradeoff between temporal resolution and computational cost; while event cameras, with their high temporal resolution and energy efficiency, serve as a natural solution to the dilemma. However, event-based approaches commonly rely on correlation volumes to capture pairwise voxel correspondences, which incur substantial memory and computation overhead. We present E-WAVE, a correlation-free framework for high-temporal-resolution (HTR) optical flow estimation from event streams. Instead of constructing all-pairs correlation volumes, E-WAVE employs global attention mechanism to model long-range feature dependencies and performs trajectory guided feature warping using Bézier curve. Through iterative updates, it predicts trajectories that allow for querying at arbitrary timestamps without repeated inference. Experiments on MultiFlow and DSEC-Flow demonstrate a 25% lower trajectory error and comparable endpoint flow estimation accuracy relative to state-of-the art baselines. Additional evaluations on self-captured data using a head-mounted prototype validate that E-WAVE remains robust under challenging real-world conditions.

cs.CV↗

SAIL: Spatial Audio Intelligence with Large Language Models via Disentangled Acoustic-Spatial Encoding and Dual-Stream Q-Former

Spatial audio large language models (LLMs) enable embodied agents, wearable assistants, and immersive systems to recognize sound events, localize sources, and reason about their spatial relationships. However, existing spatial audio LLMs often rely on early fusion of acoustic and spatial features and source-agnostic token representations. These designs make it difficult to preserve the correspondence between individual sound events and their spatial attributes, particularly in multi-source scenes. To address this limitation, we propose SAIL, a Spatial Audio Intelligence framework with LLMs that preserves acoustic-spatial structure and source-level correspondence from audio encoding to LLM alignment. SAIL introduces a Disentangled Spatial Audio Transformer that represents Mel-spectrogram and interaural phase difference features as separate acoustic and spatial streams. Source-discriminative task queries further learn event, direction, and distance information for each source. A Dual-Stream Q-Former then aligns the two streams with the LLM using acoustic and spatial queries organized by source slots. Compared with the early-fusion baseline, SAIL achieves consistent improvements in dual-source sound event detection, direction and distance estimation, and spatial reasoning. These results demonstrate the importance of structured, source-discriminative audio representations for multi-source spatial understanding and reasoning.

cs.SD↗

Commutator Memory: Sparse, Path-Local Reading and Steering in Language Models

Gradient updates on different data generally do not commute: training a language model on two data sources in opposite orders gives different weights, even with the same data and total exposure. Loss or benchmark deltas show that the models differ, not where. We ask whether this path dependence leaves a parametric training-history memory: a weight component that flips sign when the two sources are swapped, is localized in output space, changes the held-out loss gap between the two orders under targeted interventions, and reveals which trained model came from which order. For one small SGD step of size $η$ on each of sources $A$ and $B$, the weight difference $θ_{AB}-θ_{BA}$ is, to leading order, $η^2 b_{AB}$, where $b_{AB}=H_Bg_A-H_Ag_B$ is the Lie bracket of the two gradient fields at the base model. We define commutator memory by projecting the bracket through the logits into one score per vocabulary token; the scores sum to the bracket's prediction of the gap. The scores are localized: on three models, the same readout of the measured $θ_{AB}-θ_{BA}$, or of a bracket from disjoint batches, shares 82-99% of the original top-20 tokens, versus 35-49% for norm-matched random directions. They are causally actionable: in Qwen-3-4B SFT, downweighting the ten tokens with the largest predicted share of the gap closes a median 32% of the measured gap, while frequency-matched tokens with near-zero scores have almost no effect. The weights themselves carry the component: projecting the difference between the two trained models onto $b_{AB}$ identifies which came from which order in 92% of cases across four LLMs (chance 50%). Controlled tests also cover matched-batch DPO, a frozen-rollout GRPO-style objective, and an AdamW endpoint check. The memory is defined per source pair, not per example, and its projection on $b_{AB}$ decays with further training.

cs.LG↗

Electrons with Anomalous Energy Generated in Vacuum Diodes with Different Pulse Duration

This paper presents results from an experimental study in which anomalous energy electrons (AEE) appear in vacuum diodes under varying amplitude and pulse-duration conditions. AEE refer to high-energy discharge electrons with kinetic energy $T$ exceeding $eU$ values (where $e$ is the electron charge and $U$ is the amplitude gap voltage). Here, two experimental setups were used to generate electron beams with different pulse durations. The electron energy spectra were reconstructed from the beam attenuation curves using an AI-driven methodology for solving an ill-posed inverse problem for the Fredholm equation. This confirms that vacuum diodes generate electron beams with a significant AEE fraction (up to 25~\%) only when operating with nanosecond-long voltage pulses. It has been confirmed that AEE in vacuum is generated using nanosecond- or sub-nanosecond voltage pulses. Experiments also indicate that photoelectric absorption of bremsstrahlung cannot produce significant AEE in vacuum discharges.

physics.plasm-ph↗

Second largest eigenvalue does not bound stationary entropy production

We examine the spectral dissipation-coherence trade-off inequality conjectured by Oberreiter {\it et al}. [Phys. Rev. E 106, 014106 (2022)], stating that the coherent number, defined from the second largest eigenvalue, provides a lower bound on the stationary entropy production per unit oscillation. We disprove this conjecture by explicitly constructing a counterexample, which accompanies finite coherent number with vanishing stationary entropy production. This model also excludes a wide class of thermodynamic bound with the real and imaginary parts of the second largest eigenvalue of the transition rate matrix. Our result implies that the second largest eigenvalue does not necessarily characterize the degree of coherent oscillation.

cond-mat.stat-mech↗

SemRD-V2X: Closure-Guided Communication with Bounded Inference for Cooperative Perception

Vehicle-to-Everything (V2X) cooperative perception improves 3-D detection by sharing intermediate features, but dense remote features may repeat context that the ego agent can infer locally. Most communication-efficient designs optimize masks or codes empirically, leaving a more basic question open: which remote evidence is indispensable given the receiver's own observation? We introduce a closure-fidelity perspective on ego conditioned remote perception. Under a finite deductive abstraction and explicit conditions, its rate--distortion function decomposes over an irredundant core, and the exact zero-distortion rate becomes $P_A H(π_A)$. This analysis suggests a concrete design principle: transmit compact evidence and recover derivable context with bounded receiver-side inference. Guided by this principle, SemRD-V2X is an operational neural proxy that combines exact-budget BEV support selection, pointwise channel compression, and masked shared-weight reconstruction before standard fusion. Experiments on simulated V2XSet and real-world DAIR-V2X validate the resulting design. In a controlled five-run V2XSet comparison against a locally reproduced V2X-ViT-v1 baseline on one Tesla V100, SemRD-V2X reduces the analytical feature payload by $26.6\times$ while improving AP@0.5/AP@0.7 by 4.13/8.57 points, with 3.81\% additional mean compute latency. These results position closure fidelity as both an analytical lens and an actionable design principle for communication-efficient cooperative perception.

cs.AI↗

FAST-Brain: A Flow-Aligned Spatio-Temporal Surrogate Brain Model

Modeling resting-state functional magnetic resonance imaging (rs-fMRI) data is crucial for understanding brain-wide neural activity. However, traditional methods struggle to capture complex temporal dynamics over long horizons, to account for the brain's anatomical spatial structure, and to model high-dimensional ambient signals that lie on a low-dimensional intrinsic subspace. We propose FAST-Brain, a unified flow-aligned spatio-temporal surrogate brain model that addresses all three challenges. At its core is a flow-aligned generative framework that directly predicts the clean blood-oxygen-level-dependent (BOLD) signal, paired with a graph convolutional network that captures spatial structural constraints and a Transformer that models long-range temporal dependencies. Theoretically, we show that under a low-dimensional subspace assumption, the approximation error of our model scales with the intrinsic dimension rather than the ambient dimension, which justifies our direct modeling of the BOLD signal. Extensive experiments on synthetic and Human Connectome Project datasets demonstrate that FAST-Brain achieves state-of-the-art performance in recovering functional connectivity, effective connectivity, and the implicit low-dimensional signal subspace.

cs.LG↗