arXiv Science⌕ Search

SEARCH · arXiv Science

Search arXiv Science

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Constraint Preferences: Inattention and Aggregation

We study robust decision problems when individuals have maxmin preferences whose belief sets are neighborhoods around reference models, commonly known as "constraint preferences." We first show that a more disciplined form of rationally inattentive behavior is equivalent to the behavior implied by a subclass of constraint preferences. We then introduce an aggregation principle that requires collective beliefs to satisfy every individual's constraint. This requirement links collective beliefs to the information-processing technologies that generate individual constraints. Applications reveal how these technologies determine asset prices, when prediction-market prices become self-confirming, and how much dynamic mechanisms can reduce information rents.

econ.TH↗

Higher Gauge Theory via Differential Nonabelian Cohomology

This is a streamlined introduction to the global (infrared) completion of Maxwell-type higher gauge fields (as in the higher gauge sectors of higher dimensional supergravity and its brane probes) by electromagnetic flux quantization in differential nonabelian cohomology, using cohesive homotopy theory. Applications include D/NS brane charge in (unstable) K-theory, M-brane charge in unstable Cohomotopy and geometric engineering of topological quantum order on probe M5-branes.

hep-th↗

ContactWorld: What Representations Matter for Vision-Tactile Latent World Models in Contact-Rich Manipulation

Contact-rich manipulation poses a fundamental challenge for world models: visual and tactile observations capture different aspects of physical interaction, and their utility depends critically on how this information is represented. We introduce ContactWorld, a controlled benchmark and empirical study of vision-tactile representations across 12 contact-rich manipulation tasks. Using a fixed world-model architecture, training procedure, and planning framework, we examine representation effects through three complementary properties: spatial fidelity, motion coherence, and predictive stability. Point clouds preserve task-relevant geometry and track physical motion more reliably than image observations, helping explain their higher average planning success of 32.1%, compared with 20.7% and 22.0% for wrist- and front-view RGB, respectively. Tactile observations generally reduce long-horizon prediction-error accumulation, but these gains do not translate uniformly into task success. Structured tactile force fields provide the most consistent downstream improvements, with PointCloud+TacFF achieving the highest overall simulated success rate of 36.1%. We further validate these trends through 900 physical trials spanning six tasks, two robot platforms, and three tactile sensing systems. Together, ContactWorld identifies sensory representation as a central design factor in vision--tactile world models and provides empirical guidance for predictive planning in contact-rich manipulation.

cs.RO↗

JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence

Many moments in the real world do not wait for a user to ask. A fire starts on a security monitor, an expression flickers across a video call, or a product a viewer wants flashes by in a livestream. Yet today's large models remain mostly turn-based by design: they answer only when addressed, and even video-call apps that appear interactive still operate as question-answer systems, reacting only when polled or prompted. We argue for a different paradigm: a model that is present in the world like a person. It continuously watches what is happening now, decides on its own whether to speak or stay silent, interacts in real time, and delegates to a background model when the problem is hard. To advance interaction models and their adoption across domains, we make two fully open-sourced contributions. First, we release JoyAI-VL-Interaction, an 8B-scale, vision-first VL-interaction model. The model makes the response decision internally, choosing each second to stay silent, respond, or delegate to a background model, and it excels at vision-triggered responsiveness and time awareness. We pair it with a transferable training recipe, from which capabilities we never trained for emerge, such as guiding a shopper through changing app screens or improvising a lecture from a slide deck. Second, we release a complete, deployable system built around that model. The system streams any ongoing video into the model, making it genuinely present in the world. All other components are pluggable, including ASR/TTS modules, memory, visualization UI, and a background brain that can connect to any API or agent. Across six real-world scenarios, human raters prefer JoyAI-VL-Interaction over the in-app video-call assistants of Doubao and Gemini by a wide margin. To our knowledge, this is the first open, vision-driven interaction model released together with its training recipe, data, and complete deployable system.

cs.CV↗

Non-unital monoidal category of contact manifolds and Legendrian correspondence

There are two purposes of the present paper which are interrelated. The first goal is to construct the structure of a non-unital monoidal category $\mathfrak{Cont}$ of contact manifolds, not necessarily coorientable, by developing the contact topology \emph{without contact forms}. The non-unital monoidal product is the functorial contact product $\star$, called star product, introduced in \cite{oh:shelukhin-conjecture}. We prove that the product $\star$ is associative and there exist a collection $α= \{α_{X,Y,Z}\}$ of the \emph{associator} isomorphisms $α_{X,Y,Z}: X \star (Y\star Z) \cong (X \star Y) \star Z$ for $X, \, Y, \, Z \in \mathfrak{Cont}$, that satisfy the pentagon axiom, i.e., that the triples $(\mathfrak{Cont}, \star, α)$ form a nonunital monoidal category. The second goal is to develop the calculus of Legendrian correspondences, which are by definition embedded Legendrian submanifolds of the contact product $Q \star Q'$. Legendrian correspondences will play the role of 1-morphisms in the $2$-categorical structure to be equipped with $\mathfrak{Cont}$ whose two morphisms are contact instanton cohomologies $HI(R_{ab},R'_{ab})$ associated to a pair of Legendrian correspondences $R_{ab}, \, R'_{ab} \in \mathfrak{Leg}(Q_a,Q_b)$. With this future application in mind, we define the composition of Legendrian correspondences and prove that the composition of a generic pair is again embedded and hence canonically becomes a Legendrian correspondence.

math.SG↗

The entropy of black hole under second-order deviation from equilibrium

We investigate the entropy of a dynamical black hole arising from second-order perturbations of a general stationary background with a bifurcate Killing horizon. Using Gaussian null coordinates, we study the geometry of the apparent horizon perturbatively up to second order. Within the covariant phase space formalism, to explore the contribution of matter fields, we introduce a new modified canonical energy, and establish a balance law relating the second-order variation of the entropy to the energy flux entering the black hole. We show that the entropy is given precisely by the area of the apparent horizon at second order when the null energy condition holds for the infalling matter, and that the variation of the entropy also obeys the second law. We also discuss the possibility that the area law continues to hold when the null energy condition is violated.

gr-qc↗

Does DESI prefer Damped Oscillating Dark Energy over Cosmological constant?

We investigate a dark-energy equation of state governed by a damped harmonic oscillator equation, admitting underdamped, critically damped, and overdamped solutions. Confronting the model with Planck CMB distance priors, DESI BAO, BBN, cosmic chronometers, and three Type~Ia supernova compilations, we find that the data select an underdamped solution yielding $H_0 = 70.9 \pm 1.1$ km/s/Mpc with DES-Dovekie and $H_0 = 72.0^{+1.4}_{-2.1}$ km/s/Mpc with Union3, without any local $H_0$ prior. These higher values of $H_0$ arise along the $Ω_{\rm m}$--$H_0$ degeneracy direction while the sound horizon remains nearly unchanged at $r_{\rm d} \simeq 145$~Mpc, indicating that the enhancement of the late-time expansion rate is a geometrical effect that does not address the early-time calibration of $r_{\rm d}$. In contrast, the Pantheon+ compilation selects a near-critically damped solution with a prior-limited positive $w_0$ and $H_0 = 66.23 \pm 0.85$ km/s/Mpc, highlighting the sensitivity of the model to the low-redshift distance information encoded in the different supernova compilations. The Bayesian evidence relative to $Λ$CDM is inconclusive for the DES-Dovekie and Union3 combinations, whereas Pantheon+ shows a strong preference for the damped-oscillator model, driven by the departure from $w=-1$ at $z\lesssim0.1$.

astro-ph.CO↗

A nonlinear theory for chemotactic fronts of mixed populations

Collective migration of heterogeneous cell populations is central to many biological and physiological processes, including development and immune response. Recent experimental and theoretical advances have shown how asymmetric interactions with self-generated chemical gradients shape the spatial distribution of distinct cell types within migrating collectives. However, the principles governing robust spatial organisation of heterogeneous cell populations remain poorly understood. Here, we use asymptotic analysis to systematically derive a nonlinear analytical theory for heterogeneous cell collectives guided by self-generated chemotaxis. Our theory disentangles how heterogeneity in cell diffusivity, chemoattractant consumption, and chemotactic sensitivity shape the density profiles of migrating heterogeneous collectives, revealing four distinct dynamical behaviours that together capture all possible regimes. We calibrate our framework to experimental data on the co-migration of dendritic and T cells. We predict that this system operates in a parameter regime that balances intercellular mixing with T-cell localisation at the leading front of the migrating collective. Our theory reveals that this behaviour is enabled by intermediate long-range chemoattractant signalling generated through strong chemoattractant consumption by dendritic cells. Overall, our framework provides general principles for understanding how non-reciprocal chemical interactions shape robust collective migration in heterogeneous cell populations.

q-bio.CB↗

Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems

Agentic AI systems increasingly rely on language-model components to interpret instructions, process external data, invoke tools, and coordinate with other agents. These capabilities make prompt-injection and jailbreak attacks more consequential, especially as attackers adopt model-guided automation to scale probing, prompt refinement, and response evaluation. This work analyzes the resulting attack-defense setting through a probabilistic model of a target system, its defense mechanism, and the attacker's automated judge. Our analysis shows that conventional detect-and-block defenses can allow attacker success rate (ASR) to approach one as the query budget grows, since predictable refusals provide useful feedback to automated search. We then examine detect-and-misdirect, where detected malicious interactions receive controlled, non-operational responses designed to induce false-positive errors in the attacker's judge. This strategy reduces the positive predictive value of attacker-selected candidates and yields a bounded asymptotic ASR. We evaluate a proof-of-concept realization of this strategy through Contextual Misdirection via Progressive Engagement (CMPE), a lightweight conversational misdirection method designed to replace predictable refusal text with safe but strategically misleading responses in automated jailbreak settings. On jailbreak benchmarks, CMPE reduces estimated ASR upper bounds by up to two orders of magnitude and nearly eliminates verified attack success in end-to-end experiments with PAIR, GPTFuzz, and AutoDAN-Turbo.

cs.CR↗

Induced Directional Switching of Platicon Microcombs in Photonic Crystal Ring Resonators

Microcombs in normal-dispersion photonic crystal ring resonators (PhCRs) are versatile building blocks for next-generation integrated photonic circuits, but their inherent backward-propagation bias necessitates optical circulators or complex filtering for comb extraction, creating a significant bottleneck for full on-chip integration and precluding self-injection locking schemes. In this work, we introduce Side-mode Induced Forward Forcing (SIFF), a robust method to control and reverse this directionality. By engineering auxiliary mode splittings on resonances adjacent to the pump, we steer the nonlinear dynamics to favor stable, forward-propagating platicon states. We identify an optimal coupling condition that ensures forward-comb dominance across a wide parameter range. Our findings, validated numerically and experimentally, enable circulator-free, integrated normal-dispersion microcombs compatible with self-injection locking, offering a scalable architecture for compact telecommunications and sensing systems.

physics.optics↗

Interfacial melting as a thermodynamic indicator of solid-state synthesizability

Computational materials discovery commonly ranks candidate materials by their thermodynamic stability on the formation energy convex hull, yet many predicted-stable phases resist synthesis. We propose that solid-state synthesizability through interfacial-melt-mediated routes requires an additional thermodynamic condition: the interfacial melt at the target composition must itself remain locally stable against spinodal decomposition. We examine this in the classical Fe--B system, where thermodynamically stable FeB$_4$ has been reported under high-pressure synthesis but not in low-pressure synthesis attempts. Using melt--quench molecular dynamics driven by a fine-tuned machine-learning interatomic potential, we find that, at ambient pressure, the B-rich interfacial melt near the FeB$_4$ composition develops a concave free-energy landscape, signaling a demixing instability that is corroborated by the concentration--concentration structure factor and correlated with low-energy icosahedral and pentagonal-pyramidal boron motifs. In contrast to FeB$_4$, metastable Fe$_3$B and Fe$_{23}$B$_6$ remain synthesizable because their corresponding melts are stable. Applied pressure introduces a convex $PV$ contribution that strongly suppresses this instability, reducing the curvature at the FeB$_4$ composition to within the uncertainty of our fit at 1800~K, consistent with the experimental synthesis boundary. Comparison with CrB$_4$ further shows that weaker melt instability correlates with easier experimental synthesis. Interfacial-melt stability, which atomistic simulations can assess via the low-$k$ concentration--concentration structure factor, is thus proposed as a practical thermodynamic screening descriptor of synthesizability for AI-assisted materials discovery.

cond-mat.mtrl-sci↗

Flow as Flow: Modeling Robot Velocity Fields as Probability Velocity Fields for Flow-Based Object Manipulation

Cross-embodiment data have become central to training robotic foundation models. To leverage such heterogeneous data, we focus on flow-based object manipulation, where robot flows (robot velocity fields) serve as embodiment-agnostic motion representations. Previous studies do not formulate robot flows as dense velocity fields, but as displacements of sparse keypoints, even though dense velocity fields better match the continuous-time nature of motions. To address this, we propose Flow as Flow, a framework that models robot flows as probability flows based on a flow matching formulation. By naturally modeling such velocity fields within this formulation, our method achieves efficient and high-quality robot flow generation. Across standard benchmarks, our method outperforms representative baseline methods on standard metrics, while achieving approximately 24$\times$ faster generation than standard flow matching. Furthermore, through real-world experiments evaluating 9 methods with 260 trials per method across 13 manipulation tasks, we show that our method achieves a higher average success rate than the baseline methods.

cs.RO↗

Enabling Robust Cloth Manipulation via Inference-Time Simulator-in-the-Loop Refinement

Simulator-in-the-loop optimization offers a promising inference-time mechanism for robot manipulation. It uses a physical simulator as a backend rollout engine to evaluate candidate trajectories in parallel and refine nominal actions online, a paradigm shown to be effective in rigid-body manipulation where state and contact are relatively tractable. We bring this paradigm to real-world cloth manipulation from a single RGB input through three pillars. (i) We design a scalable synthetic-data generation and inference-time rollout pipeline built on FLASH, a deformable-object simulator that provides a practical balance among physical fidelity, numerical stability, and rollout efficiency. (ii) We develop a real-to-sim module, trained purely on synthetic data, that maps a single RGB observation to simulation-compatible cloth state by fusing pretrained visual features with learnable canonical tokens. (iii) We perform online planning by coupling a sparse-mesh rollout backend with prior-guided MPPI, anchored at an offline-distilled policy trajectory, preserving manipulation-relevant deformation and contact while enabling sufficient parallel rollout batches. Real-robot experiments show higher success rates than baseline methods and closed-loop correction under mid-fold perturbations. Project page: https://silr-cloth.github.io/

cs.RO↗

Higgs Scattering and Entanglement in SMEFT

We regard the weak isospin of the Higgs doublet as a qubit and classify the entanglement measures for the Higgs scattering in the Standard Model Effective Field Theories (SMEFT) and their Ultra-Violet complete models. We consider Higgs scattering in the unbroken phase for electroweak symmetry. Treating the final state as a momentum-isospin bipartite system, we obtain von Neumann and linear entropies to quantify momentum-isospin correlation. From the momentum reduced state, we calculate the concurrence, which measures the entanglement between the two isospins. Both quantities are set by the isospin singlet and triplet scattering amplitudes, and hence by the Wilson coefficients of the dimension-6 and dimension-8 Higgs operators. We find that the von Neumann entropy grows as a function of the total energy in SMEFT as compared to the SM case, but it undergoes a cancellation in the medium energy below the cut-off scale due to the interference effects between the dimension-4 and dimension-8 operators, in particular, when the effective interactions stem dominantly from a massive graviton. Assuming the dominance of dimension-8 operators, we find the conditions for entanglement suppression in the forward or backward scatterings or across all the kinematics. We also show the correlations between the entanglement suppression and the positivity bounds in the forward limit.

hep-ph↗

Decoupled energy-stable Runge-Kutta schemes of arbitrary order for the anisotropic phase-field dendritic crystal growth model

In this paper, we construct an arbitrary-order scheme for the anisotropic phase-field dendritic crystal growth model by introducing a time dependent auxiliary variable. By employing an algebraically stable Runge-Kutta method, the proposed scheme satisfies an unconditional discrete energy dissipation law. To reduce the computational cost, a matrix diagonalization technique is applied to the coupled elliptic system at each time step. This transforms the original system into independent elliptic equations with constant coefficients, which can be solved separately or in parallel. After the decoupling, the auxiliary variable is obtained from a uniquely solvable $q\times q$ algebraic system. For a fixed Fourier-Galerkin space, we further prove $q$th-order convergence in time for the scheme based on a $q$-stage Runge-Kutta method. Numerical experiments in two and three dimensions confirm the theoretical convergence rates and the discrete energy dissipation, and demonstrate the computational efficiency of the decoupled schemes. The effects of anisotropy, latent heat, orientation angle, and initial nuclei on the dendritic morphology are also investigated numerically.

math.NA↗

MPC-Injection: Biasing Off-Policy Locomotion RL Toward Controller-Induced Behavior Basins

Reinforcement learning (RL) for locomotion frequently converges to locally optimal but undeployable behaviors, such as vibrating limbs or scooting on the torso, that maximize return without producing a usable gait. We present MPC-Injection, a low-overhead method that steers RL toward a designer-preferred behavior by inserting transitions generated in the same environment by a model predictive controller (MPC). Unlike reward shaping, MPC-Injection does not require redesigning the task reward, and unlike adversarial imitation learning, it adds no discriminator, no kinematic retargeting, and no auxiliary objective. We analyze how the injected transitions bias the learning, allowing the policy to converge to behaviors that pure RL may fail to reach under simple reward functions. On a 2D walker in simulation and with sim-to-real evaluation on a Go2 quadruped, we show that MPC-Injection produces gaits qualitatively comparable to those of reward shaping and adversarial motion priors. We also show that MPC-Injection can complete a barrel roll that pure RL fails to achieve under the same simple reward and can select between trotting and bounding gaits only through changing the injected MPC data.

cs.RO↗

Closing the Quality Gap in Low-Resource Text-to-Speech: LoRA Fine-Tuning of VoxCPM2 for Khmer and Korean

Large pretrained text-to-speech (TTS) models sound almost human for well-resourced languages, but much worse for languages that are rare in their training data. We study this quality gap for Khmer and Korean using VoxCPM2, a 2.4B parameter, tokenizer-free TTS model that joins a MiniCPM-4 language-model backbone with a flow-matching diffusion decoder. We build one shared, language-tagged corpus of 25.5 hours after cleaning and adapt VoxCPM2 with a single Low-Rank Adaptation (LoRA) adapter, trained on both languages at once and added to both the language model and the decoder. The adapter is zero-initialized, so training starts exactly at the original zero-shot model. In native-speaker listening tests, the Khmer Mean Opinion Score (MOS) rises from 3.85 to 4.23 with the best adapter, rank 64. This gain is highly significant under a paired Wilcoxon test with p < 0.001, and it is achieved while training only 0.19 to 3.03 percent of the parameters. Two findings stand out. First, the training loss and human ratings disagree on the best rank. The loss is lowest at rank 128, but MOS peaks at rank 64. Second, the same adapter gives no significant gain for Korean, which the base model already covers well, and a high rank even hurts quality. This shows that adaptation helps mainly where the base model is truly weak.

cs.CL↗

All you need is log

How different are several probability distributions from one another? For two distributions the standard answer is the family of Rényi divergences, singled out by two natural requirements: processing the data never makes distributions easier to tell apart, and independent repetitions add. Many problems in learning and statistics compare more than two distributions at once, such as testing among several hypotheses or bounding generalization against several priors. The same two requirements leave one kind of building block, built on a coincidence probability: how unlikely it is that independent samples, one from each distribution, all show the same empirical distribution. The logarithm is forced because repetitions add, which is already visible for a single experiment repeated. This characterization is known in greater generality, and this paper is about the meaning of its building blocks. On a finite alphabet, each building block indexed by a rational point of the simplex is the exponential rate of that coincidence as the samples grow in fixed proportions. Each is also the limiting free energy of Bayesian inference over distributions. At any amount of data, the free energy of the posterior is the coincidence measure plus two costs: the expected distance from a posterior draw to the most likely distribution, and the information gained per unit of data. Both costs vanish as data accumulate. When the comparison is conditioned on side information, every kind of building block has a conditional counterpart, and the coincidence ones alone do not suffice.

cs.IT↗