arXiv Science⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 217 records · Page 12Linked to original sources

A Local Characterization of Unique Equilibrium States

We study when a prescribed ergodic measure is the unique equilibrium state of a continuous potential. For continuous actions of countably infinite discrete amenable groups on compact metric spaces with finite topological entropy, we characterize this property by a local condition on entropy differences. Fixing an ergodic measure $μ$, we extend $h(ν)-h(μ)$ to the cone generated by the differences $ν-μ$, scaling the entropy difference by the same factor. We prove that $μ$ is uniquely attainable if and only if this entropy difference function is weak\nobreakdash-$*$ upper semicontinuous at the zero measure. Upper semicontinuity of entropy at $μ$ alone does not suffice. In fact, we construct a shift space over a compact alphabet for which entropy is upper semicontinuous at every ergodic measure and continuous at a measure supported on a fixed point, but this measure is not an equilibrium state of any continuous potential.

math.DS↗

MR.ScaleMaster: Scale-Consistent Collaborative Mapping from Crowd-Sourced Monocular Videos

Crowd-sourced cooperative mapping combines monocular sessions from different front-ends, each an independently reconstructed keyframe sequence. Each session has its own local coordinate frame and may follow an incompatible scale convention. We present MR.ScaleMaster, a backend that accepts image, Sim(3) pose, and point-map packets without requiring a common reconstruction model, camera intrinsics, or front-end-provided metric scale. Our Cross-Front-End Loop Factor (CFL) uses a shared matcher for image correspondences but retrieves matched 3D points from the input point maps, so its scale estimates the inter-session ratio. Scale Preconditioning (SPC) initializes session scales for a Sim(3) anchor-node graph, which then corrects residual scale and drift. For front-ends exposing an incremental scale trajectory, an agent-side Scale Collapse Alarm (SCA) rejects or rolls back false intra-session loops that would otherwise collapse the session scale. We evaluate seven front-ends on KITTI and five on CODa. Based on ground-truth path lengths, session scales in our heterogeneous KITTI setting differ by 57-103x. CFL reduces mean ATE by 37% and inter-session scale error by 60% relative to loop factors from independent pairwise reconstructions. Adding SPC raises the reductions to 74% and 77%, respectively. On CODa, five-session fusion improves mean ATE over single-session runs for all five front-ends, while joint fusion registers 15 sessions from three different front-ends in a single map. Code will be released.

cs.RO↗

PipeLive: Efficient Live In-place Pipeline Parallelism Reconfiguration for Dynamic LLM Serving

Pipeline parallelism (PP) is widely used to partition layers of large language models (LLMs) across GPUs, enabling scalable inference for large models. However, existing systems rely on static PP configurations that fail to adapt to dynamic settings, such as serverless platforms and heterogeneous GPU environments. Reconfiguring PP by stopping and redeploying service incurs prohibitive downtime, so reconfiguration must instead proceed live and in place, without interrupting inference. However, live in-place PP reconfiguration is fundamentally challenging. GPUs are already saturated with model weights and KV cache, leaving little room for new layer placements and necessitating KV cache resizing, at odds with systems like vLLM that preallocate for throughput. Moreover, maintaining KV consistency during execution is difficult: stop-and-copy introduces large pauses, while background synchronization risks inconsistency as states evolve. We present PipeLive, which enables live in-place PP reconfiguration with minimal disruption. PipeLive introduces a redesigned KV cache layout together with a co-designed extension to PageAttention, forming a unified mechanism for live KV resizing. It further adopts an incremental KV patching mechanism, inspired by live virtual machine migration, to synchronize KV states between source and target configurations and identify a safe switch point. PipeLive achieves a 2.5X reduction in time-to-first-token (TTFT) without KV cache overflow compared to disabling KV resizing. Furthermore, compared to a variant without KV patching, it reduces reconfiguration overhead from seconds to under 10ms, and improves TTFT and time-per-output-token (TPOT) by up to 54.7% and 14.7%, respectively.

cs.DC↗

HazardArena: Evaluating Semantic Safety in Vision-Language-Action Models

Vision-Language-Action (VLA) models inherit rich world knowledge from vision-language backbones and acquire executable skills from action demonstrations. Yet current evaluations primarily measure task completion, leaving the semantic safety of learned action policies underexplored. This gap creates a critical vulnerability: a policy may execute the intended action correctly while producing unsafe outcomes when the surrounding visual-linguistic context changes. We present HazardArena, a benchmark for stress-testing semantic safety in VLA systems. Its core design is a set of safe/unsafe twin scenarios: paired environments with matched objects, layouts, and action requirements, but different semantic risk contexts. This controlled contrast isolates safety judgment from motor capability and directly tests whether a VLA can recognize when an otherwise valid action becomes hazardous. HazardArena includes over 2,000 assets and 51 risk-sensitive tasks across seven safety categories grounded in robotic safety standards. Across four representative VLA backbones, we observe a consistent and alarming pattern: safe-only fine-tuning improves benign task success while also increasing hazardous execution on matched unsafe scenarios. Physical-world experiments confirm that this failure transfers beyond simulation. These results show that stronger action execution does not imply safer behavior, and motivate semantic-risk-aware evaluation and enforcement as first-class requirements for real-world VLA deployment. Code released at https://github.com/HazardArena-Team/HazardArena ; updated code availability information.

cs.RO↗

Toward Measuring Structural Drift in LLM Communication Loops

Large language models increasingly run in stateful pipelines that assemble each prompt from retrieval, memory, tools, and other agents. Such pipelines drift: information that should shape the next response is dropped, compressed, or misrouted while every component still reports success. Existing diagnostics miss this because they evaluate isolated prompts, responses, or task scores, whereas what decouples is the relation between a prompt and the response it draws. Here we show that treating the prompt to response to next prompt chain as the fundamental unit of analysis makes these relations measurable. We introduce structural communication coherence, quantified by two metrics: communication closure, which asks if what the pipeline returns at one turn matches what it faces next, and normalized conditional action contribution, which measures how much a sent message resolves the subsequent reply. Across 2,171 human to human, 58 human to LLM, and 8 LLM to LLM dialogues, these metrics reveal directional interaction structures; crucially, the measured contribution drops by 87 to 92% when a response is swapped for one from another turn, leaving surrounding prompts untouched. Because this approach requires no labels, healthy reference data, or predefined rules only the raw prompts and responses drift can be defined and measured directly from operational traffic, rather than inferred from eventual task failure. Establishing prospective detection performance is the next step.

cs.CL↗

EP250302a: violent shell collision in a soft-X-ray-selected GRB-like transient

The Einstein Probe opens a previously unexplored soft X-ray window onto gamma-ray bursts, filling a critical observational gap in the soft X-ray coverage of their prompt emission. In this letter, we present EP250302a, a soft-X-ray-selected, GRB-like transient at $z=1.131$ detected by the Einstein Probe. Follow-up observations from X-ray to radio reveal a narrow X-ray flare at $\sim 1.1$\, ks and subsequent achromatic optical and X-ray rebrightening. These features challenge a standard single-component afterglow model and indicate the need for multiple ejecta components. A violent collision between a late relativistic shell and the decelerated leading blast wave provides a plausible interpretation: the flare arises from internal dissipation of the late ejecta, while the rebrightening is powered by the shocked emission produced in the collision. Quantitative modeling constrains the kinetic energy ratio between the late shell and the initial ejecta to $E_{\rm k,iso,2}/E_{\rm k,iso,1} \sim 5$ (with $E_{\rm k,iso,2} \sim 10^{53}$~erg and $E_{\rm k,iso,1} \sim 2\times10^{52}$~erg), as well as the Lorentz factor contrast to $Γ_{2,0}/Γ_{1,0} \approx 0.98$--$2.27$, required to reproduce the observed flare luminosity and rebrightening amplitude. Such an energetic late shell can be launched in a radiatively inefficient second episode of central-engine activity. Thanks to the well-sampled, early-time multiband coverage facilitated by the EP trigger, EP250302a provides a valuable case to test the physical connection between central-engine activity and shell collisions.

astro-ph.HE↗

Vision-Based Safe Human-Robot Collaboration with Uncertainty Guarantees

Safe human-robot collaboration (HRC) requires accurate human pose estimation and motion prediction to prevent critical collisions. Existing certifiable safe HRC approaches are highly conservative or rely on marker-based motion tracking, while vision-based pose estimators lack the statistical guarantees required for certification in accordance with ISO 13849-1. Hence, we propose a pipeline that predicts 3D human motion and strong probabilistic bounds on the prediction error using conformal prediction. A gradient-based monitor detects out-of-distribution input poses and replaces them with poses from past predicted motions to maintain smooth operation. The resulting conformal prediction sets directly integrate into the provably safe HRC approach SARA shield. In experiments on the Human3.6M dataset and a real-world HRC setting, our conformal prediction sets have a 7.6 times smaller volume than model-based predictions, and we bound the probability of a dangerous failure per hour by 9.5E-7 with 99.999 % confidence under our test distribution, which is necessary but not sufficient for performance level d. All code and models are available at https://jakob-thumm.com/conformal_human_motion_prediction/.

cs.RO↗

Preregistered Belief Revision Contracts

Deliberative multi-agent systems allow agents to exchange messages and revise beliefs over time. While this interaction is meant to improve performance, it can also create dangerous conformity effects: agreement, confidence, prestige, or majority size may be treated as if they were evidence, producing high-confidence convergence to false conclusions. To address this, we introduce PBRC (Preregistered Belief Revision Contracts), a protocol-level mechanism that strictly separates open communication from admissible epistemic change. A PBRC contract publicly fixes first-order evidence triggers, admissible revision operators, a priority rule, and a fallback policy. A non-fallback step is accepted only when it cites a preregistered trigger and provides a nonempty witness set of externally validated evidence tokens. This ensures that every substantive belief change is both enforceable by a router and auditable after the fact. In this paper, (a) we prove that under evidential contracts with conservative fallback, social-only rounds cannot increase confidence and cannot generate purely conformity-driven wrong-but-sure cascades. (b) We show that auditable trigger protocols admit evidential PBRC normal forms that preserve belief trajectories and canonicalized audit traces. (c) We demonstrate that sound enforcement yields epistemic accountability: any change of top hypothesis is attributable to a concrete validated witness set. For token-invariant contracts, (d) we prove that enforced trajectories depend only on token-exposure traces; under flooding dissemination, these traces are characterized exactly by truncated reachability, giving tight diameter bounds for universal evidence closure. Finally, we introduce a companion contractual dynamic doxastic logic to specify trace invariants, and provide simulations illustrating cascade suppression, auditability, and robustness-liveness trade-offs.

cs.AI↗

Isospin-symmetry violation - kaons and beyond (ISO-BREAK 25: summary and outlook)

This report summarizes the presentations and discussions during the ISO-BREAK 25 Workshop ``Isospin symmetry violation: kaons and beyond'', which was held at Jan Kochanowski University in Kielce on October 23--25, 2025. We address the current status of the isospin-symmetry breaking discovered by NA61/SHINE in nucleus--nucleus collisions at the CERN SPS, its confirmation by other experiments and studies in \ee and deep inelastic scattering. In addition, we discuss the theoretical status as well as we outline experimental and theoretical priorities towards understanding this currently unexplained phenomenon.

nucl-ex↗

Automated Palynological Analysis System: Integrating Deep Metric Learning, Detection and Classification in Bright Field Microscopy

Traditional melissopalynology is a time-consuming and subjective process, often taking 4-6 hours per sample. We present an automated, high-throughput microscopy system that integrates H_\infty robust mechanical control with advanced deep learning pipelines for the precise counting, classification, and morphological analysis of pollen grains from Bio Bio region in south central territory in Chile. Our system employs U^2-Net for salient object detection and a DINOv2 Vision Transformer backbone trained via Deep Metric Learning for classification. By integrating Gradient-Weighted Attention, the model provides human-interpretable texture and diagnostic feature annotations. The system achieves a 95.8% classification recall and at least 6x processing speedup compared to manual expert analysis.

cs.CV↗

AdaGScale: Viewpoint-Adaptive Gaussian Scaling in 3D Gaussian Splatting to Reduce Gaussian-Tile Pairs

Reducing the number of Gaussian-tile pairs is one of the most promising approaches to improve 3D Gaussian Splatting (3D-GS) rendering speed on GPUs. However, the importance difference existing among Gaussian-tile pairs has never been considered in the previous works. In this paper, we propose AdaGScale, a novel viewpoint-adaptive Gaussian scaling technique for reducing the number of Gaussian-tile pairs. AdaGScale is based on the observation that the peripheral tiles located far from Gaussian center contribute negligibly to pixel color accumulation. This suggests an opportunity for reducing the number of Gaussian-tile pairs based on color contribution. AdaGScale efficiently estimates the color contribution in the peripheral region of each Gaussian during a preprocessing stage and adaptively scales its size based on the peripheral score. As a result, Gaussians with lower importance intersect with fewer tiles during the intersection test, which improves rendering speed while maintaining image quality. The adjusted size is used only for tile intersection test, and the original size is retained during color accumulation to preserve visual fidelity. Experimental results show that AdaGScale achieves a geometric mean speedup of 13.8x over original 3D-GS on a GPU, with only about 0.5 dB degradation in PSNR on city-scale scenes.

cs.CV↗

Perfect matchings and $A_α$-spectral radius in 1-binding graphs

Let $G$ be a graph with vertex set $V(G)$ and edge set $E(G)$. For $α\in[0,1)$, we use $A_α(G)$ and $ρ_α(G)$ to denote the $A_α$-matrix and the $A_α$-spectral radius of $G$, respectively. The binding number $\mbox{bind}(G)$ of $G$ is defined by $\mbox{bind}(G)=\min\left\{\frac{|N_G(X)|}{|X|}:\emptyset\neq X\subseteq V(G),N_G(X)\neq V(G)\right\}$. If $\mbox{bind}(G)\geq1$, then $G$ is called 1-binding. A perfect matching in $G$ is a set of nonadjacent edges covering every vertex of $G$. Tutte proved that a graph $G$ of even order has a perfect matching if and only if $o(G-S)\leq|S|$ holds for every $S\subseteq V(G)$ [W. Tutte, The factorization of linear graphs, J. Lond. Math. Soc. 22 (1947) 107--111]. In this paper, we use Tutte's result to prove that a connected 1-binding graph $G$ of even order $n$ with $n\geq n(α)$ has a perfect matching unless $G=K_1\vee(K_{n-5}\cup K_3\cup K_1)$ if $ρ_α(G)\geqρ_α(K_1\vee(K_{n-5}\cup K_3\cup K_1))$, where $n(α)$ is defined as follows: $n(α)=\max\{18,\frac{2+8α}{1-2α}\}$ if $α\in[0,\frac{1}{2})$, and $n(α)=18$ if $α=\frac{1}{2}$.

math.CO↗

HIVE: Hidden-Evidence Verification for Hallucination Detection in Diffusion Large Language Models

Diffusion large language models generate text through iterative denoising, exposing hidden trajectories that may contain reliability signals beyond the final output. We propose HIVE, which compresses trajectory hidden states, selects informative step-layer evidence, and conditions a verifier through continuous prefix embeddings to produce a hallucination score and structured diagnostics. Across two D-LLMs and three QA benchmarks, HIVE outperforms eight established baselines and a verifier-backbone-matched text-only control in all six settings. Relative to text-only verification, hidden-evidence conditioning improves AUROC by 1.73--4.60 points and AUPRC by 1.10--3.62 points, with average gains of 3.15 and 2.28 points, respectively. Ablations, evidence interventions, and cross-dataset transfer further support the complementary value of fine-grained hidden trajectory evidence.

cs.CL↗

Typical entanglement entropy with charge conservation

We consider a many-body Hilbert space with a fixed global charge and show that the typical entanglement entropy of a subsystem, at the leading and subleading order in the thermodynamic limit, can be expressed in terms of a single quantity which represents the local thermal entropy at fixed charge density. We find a general formula which applies both to abelian U(1) symmetry and non-abelian SU(2) symmetry, including the case of a local Hilbert space which transforms under a general reducible representation of the symmetry group. We illustrate the general formula with model systems and discuss the relevance of the results as a probe of quantum chaos for physical Hamiltonians.

quant-ph↗

Entanglement of multi-qubit quantum graph states and studies structural properties of tripartite graphs with quantum computing

We propose a method for constructing multi-qubit entangled quantum states that represent weighted tripartite graphs, and develop approaches for investigating their structural properties using quantum computing. In the general case of multi-qubit states corresponding to arbitrary tripartite graph structures, we derive an expression for the entanglement distance. We establish a connection between entanglement and properties of the corresponding tripartite graphs. Namely, we show that the entanglement of a qubit with the rest of the system in a quantum tripartite graph state depends on the weights of the arcs in the neighborhood of the corresponding vertex, as well as on its degree with respect to the vertex sets of the tripartite graph. As an illustrative example, we consider a tripartite graph forming a triangle and evaluate the entanglement distance with quantum computing. We also compute quantum correlators for the general case of tripartite quantum graph states and relate these quantities to structural features of the underlying graphs, including the number of non-overlapping neighbors, the number of common neighbors of the corresponding vertices, and the number of 4-cycles. Obtained relationships between the quantum properties of multi-qubit states and the structural features of tripartite graphs opens up the possibility of investigating classical systems, such as tripartite graphs, using quantum computing. It is worth emphasizing that tripartite graphs have applications in practical problems, including resource allocation, scheduling, and database and hypergraph modeling.

quant-ph↗

Anon: Extrapolating Adaptivity Beyond SGD and Adam

Adaptive optimizers such as Adam and non-adaptive methods like SGD exhibit distinct generalization capabilities across different architectures. Prior tunable optimizers attempt to bridge this gap by strictly interpolating between SGD and Adam, effectively confining adaptivity within the 0-to-1 bound. However, this restricted interpolation is fundamentally insufficient: we reveal that optimal adaptivity often requires extrapolation, such as negative adaptivity for classical CNNs and adaptivity of at least one ($γ\geq 1$) for Transformers. Extrapolating adaptivity theoretically violates the strict non-decreasing pre-conditioner assumption, often leading to divergence in existing methods. To break this barrier, we propose Anon, an optimizer that achieves fully continuous adaptivity extrapolation across the entire real-number spectrum. To guarantee provable stability in these out-of-bound regimes, we introduce Incremental Delay Update (IDU), a novel mechanism that bypasses hard max-tracking strategies. We theoretically establish Anon's convergence in both convex and non-convex settings. Empirically, by exploring previously unreachable adaptivity landscapes, Anon demonstrates highly competitive and scalable performance among state-of-the-art element-wise optimizers on representative image classification, diffusion, and large language modeling tasks.

cs.AI↗

ANO: Robust Policy Optimization via Bounded, Redescending Gain Fields

Proximal Policy Optimization (PPO) dominates reinforcement learning and LLM alignment, yet its hard-clipping mechanism and unconstrained alternatives (e.g., SPO) sit at two extremes of a stability-efficiency dilemma. We argue that this dilemma is best understood dynamically: a surrogate objective is a feedback law on the probability ratio, and its clipping/penalty shape defines a gain field that drives the update dynamics. PPO's clip induces a dead zone (zero feedback outside the trust region), leaving the policy to drift open-loop under momentum; SPO's quadratic penalty induces an unbounded, linearly growing gain that stiffens the dynamics and destabilizes under aggressive step sizes. Guided by this view, we derive Anchored Neighborhood Optimization (ANO), which designs the gain field directly: a $C^\infty$ shaping kernel that anchors the identity map at $r{=}1$, peaks exactly at a prescribed trust-region boundary $1{+}ε$, bounds the push on severely off-policy samples by a tunable $κ_{+}$, and exerts a bounded, redescending pull of tunable depth $κ_{-}$ on extreme outliers. The three hyperparameters have decoupled roles, and all internal constants are solved in closed form. Empirically, ANO ranks first on both Atari (40 games) and MuJoCo in IQM and Median of normalized scores. While the runner-up differs across domains (PAPO on Atari, SPO on MuJoCo), ANO is the only method consistently at the top. Under a learning-rate stress test ($3\times10^{-4}\!\to\!10^{-3}$), ANO degrades by only $0.9\%$ whereas PPO collapses by $54.5\%$, and the stressed ANO still outperforms PPO and PAPO at their best-tuned learning rates.

cs.AI↗

From Reach to Insert: Tactile-Augmented Precision Assembly under Sub-Millimeter Tolerances

High-precision assembly frequently involves tight-tolerance insertions, where even slight pose errors can cause jamming or excessive interaction forces, making robust and safe insertion policies difficult to obtain. This paper proposes a tactile-augmented two-stage method that combines Imitation Learning (IL) and Reinforcement Learning (RL) for precision insertion tasks. In the first stage, IL learns a reaching policy with position generalization that grasps the peg and brings it to the vicinity of the target region. In the second stage, RL executes the insertion and enables recovery from failures during contact-rich interactions. To better exploit tactile feedback, we introduce tactile group sampling to increase coverage of critical contact segments during training, and design a tactile critic to more accurately evaluate policy values, improving insertion performance while maintaining low contact forces. We conduct systematic experiments across five hole geometries and three clearance settings. Results show that our method substantially improves insertion performance across all settings; under the most challenging 0.05\,mm clearance, it achieves a 67\% success rate while keeping contact forces low, reducing the maximum interaction force by 60\% and torque by 44\%, thereby validating both effectiveness and safety for precision assembly.

cs.RO↗