arXiv Science⌕ Search

SEARCH · arXiv Science

Search arXiv Science

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,243 records · Page 69Linked to original sources

The Interplay of Harness Design and Post-Training in LLM Agents

Tool-integrated LLM agents are often wrapped within a harness: the scaffolding that determines which tools are exposed, how they are described, and what auxiliary information accompanies each per-step observation. While agents are routinely post-trained, this scaffolding is typically treated as a fixed engineering detail, with design effort limited to the training-free regime. Moreover, existing post-training algorithms assume a static environment, even though tool environments and tasks often shift upon deployment. To address this gap, we extend $\texttt{ALFWorld}$ (i) to treat the harness as a controllable design dimension and (ii) to support evaluation under task and tool environment shifts. Building on this, we systematically analyze how the harness design influences post-training in both in-distribution and out-of-distribution (OOD) settings. We empirically show that harness-aware post-training not only improves in-distribution performance but also enables agents to robustly adapt to OOD settings. Under a harness with minimal design effort, post-training suffers a drastic performance drop under stronger tool environment shifts, further highlighting the importance of harness-aware post-training under such shifts.

cs.LG↗

Cliff Tokens: Analyzing Failure Trigger Tokens in LLM Mathematical Reasoning

Large language models reach high accuracy in mathematical reasoning, but individual traces on the same problem diverge; some arrive at the correct answer while others fail. Prior work localizes such failures at the step, chunk, or sentence level, or identifies tokens where failure has already occurred. These approaches leave open which token triggers failure. We introduce the cliff token, a token at which the estimated probability of reaching the correct answer (success probability) drops beyond an adaptive threshold. Across seven models and three mathematical reasoning benchmarks (GSM1K, MATH500, AIME 2025), cliff tokens act as failure triggers. For incorrect traces containing cliff tokens, we compare resampling immediately before and after the first cliff token. Resampling before it shows higher pass@$k$ at the same sample count. We further introduce a cliff taxonomy of deterministic, uncertain, and sampled-off cliffs, defined by greedy choice and token entropy. Additionally, we show that the three types differ as training signals. Using single-token preference optimization at cliff positions (Cliff-DPO), we find that uncertain and sampled-off cliffs show larger accuracy gains than deterministic cliffs on three evaluation benchmarks. We release token-level rollout data and source code to enable further analysis without regenerating costly rollouts: https://github.com/beaver-22/Cliff-token

cs.AI↗

USS: Unifying Spatial-Semantic Prompting for End to End Embodied Visual Tracking

Embodied Visual Tracking (EVT) requires an agent to continuously follow a designated target while moving through dynamic environments. Existing embodied tracking methods generally rely on either an implicit target-selection convention or a language description, leaving the target-specification interface largely fixed. However, different tracking scenarios naturally favor different forms of target specification: language can specify a target outside the robot's current view, whereas spatial prompts provide direct instance designation for visible targets and can be useful for selecting among similar looking people, designating hard-to-describe individuals, or specifying a target under time pressure. We therefore introduce unified spatial-semantic prompting for EVT, in which text, a point, a box, and a mask serve as complementary target specifications that can be selected according to different scenarios, and present USS, an end-to-end architecture that maps any of them to egocentric waypoints. A modality-specific prompt encoder feeds a common design comprising hybrid-attention fusion, temporal memory, cross-view aggregation, latent prediction, and waypoint decoding, with one policy instance trained for each interface under an identical recipe. Across 320 zero-shot real-robot trials with policies trained only in simulation, spatial prompts perform comparably to language in ordinary tracking scenarios while providing clear benefits when visually similar targets require precise instance designation, supporting our premise that different scenarios favor different target specifications. On simulation EVT-Bench under the standard language-prompt protocol, USS obtains the highest success rate among non-MLLM methods at 57 FPS, against the 4.8-10 FPS reported by MLLM trackers that are stronger on several metrics. Project site: https://arescheah.github.io/uss-project-page/.

cs.CV↗

Fast LeWorldModel

Joint-Embedding Predictive Architectures (JEPAs), including recent LeWorldModel (LeWM), have become a promising foundation for reconstruction-free visual world models. For visual planning, however, LeWM evaluates candidate action sequences by repeatedly applying a local one-step latent transition model. This autoregressive rollout makes planning computationally expensive and exposes the predicted trajectory to accumulated latent errors as the horizon grows. We propose Fast LeWorldModel (Fast-LeWM), a fast latent world model that replaces repeated local rollout with action-prefix prediction. Given the current latent and a candidate action sequence, Fast-LeWM encodes its prefixes and predicts the future latents reached after executing those prefixes in parallel. By making action prefixes the basic prediction unit, Fast-LeWM directly models action effects accumulated to different extents over multiple horizons. This prefix-level supervision forces the model to learn how states continuously evolve under different action prefixes, rather than only fitting one-step state transitions. During planning, the predictor can use the prefix token from the encoded action sequence to evaluate the corresponding future latent without explicitly rolling through each intermediate imagined state. Across multiple tasks, Fast-LeWM improves average success over LeWM while substantially reducing planning time, achieving lower open-loop latent loss whose growth becomes significantly slower as the rollout horizon increases.

cs.LG↗

Nonsmooth Bloch Oscillations in Non-Hermitian Systems

Bloch oscillations (BOs) represent a paradigmatic manifestation of force-driven band dynamics in lattices, conventionally characterized by continuous quasimomentum evolution and smooth periodic trajectories. Here we show that non-Hermiticity qualitatively changes this paradigm, introducing nonsmooth BOs characterized by cusp-like real-space trajectories and abrupt momentum jumps. Reversing the force switches the dynamics between smooth and nonsmooth BOs, enabling strongly nonreciprocal dynamics. We derive a universal relation for the critical force that governs the nonsmooth-to-smooth dynamical transition in general one-dimensional lattices, which is induced by a pitchfork bifurcation in momentum space. Our work establishes a unified framework for non-Hermitian BOs and uncovers a fundamental bifurcation mechanism that can generate nonsmooth wave transport in entirely linear systems.

quant-ph↗

LISA: Likelihood Score Alignment for Visual-condition Controllable Generation

The prevalent dual-branch paradigm, i.e., training a side network to encode visual conditions and fusing its intermediate-layer features to a frozen pretrained main network, has shown remarkable success in visual-condition controllable generation. Despite its widespread adoption, the role of the side branch and its training efficiency remain underexplored. In this paper, we first revisit this mainstream paradigm through the lens of score-based generative modeling: 1) The main network preserves visual perceptual quality by providing a prior unconditional score. 2) The side network steers conditional control by implicitly contributing a likelihood score. Guided by this perspective, we propose LIkelihood Score Alignment (LISA), an effective regularization method that explicitly aligns the intermediate feature of the side network with an approximated likelihood score. Specifically, we first hook features from a designated layer of the side network and project them into the score latent space by a lightweight decoder. Then, we construct an approximated likelihood score target and calculate the distance between the decoder's output and this target as an additional regularization loss. Finally, we jointly optimize the side network and decoder with both standard diffusion loss and our regularization loss. Experiments across various image/video tasks, architectures, and diffusion/flow models demonstrated that LISA can not only consistently accelerate the training convergence and improve final synthetic results, but also encourage the side network's features to be more disentangled for conditional modeling with negligible additional training cost and zero extra inference cost.

cs.CV↗

Exact subsystem dynamics in the deterministic Floquet-PXP model

The dynamics of local subsystems in a thermodynamically large quantum many-body system can be understood as effectively open as the system produces its own effective bath. The action of this bath can be characterised in terms of the so-called influence matrices. In generic situations, the complexity of these objects grows unfavourably with time, however, there exist solvable cases where influence matrices can be characterised exactly even in the presence of non-trivial interactions. Here we show that Rule 201, a deterministic version of the Floquet-PXP model, is one of these solvable instances. Indeed, it admits influence matrices given by a finite-dimensional matrix-product operator (MPO) that solves a finite set of algebraic conditions. We provide the solution, and characterise multi-time autocorrelation functions.

cond-mat.stat-mech↗

HPRO: Hierarchical Progressive Reward Optimization via Preference Extraction for Emotional Text-to-Speech

Recently, Large Language Model (LLM)-based Text-to-Speech (TTS) models have achieved remarkable naturalness. However, the standard Supervised Fine-Tuning paradigm often converges to statistically averaged prosody, limiting emotional expressiveness. While preference-driven optimization offers a promising alternative, existing approaches suffer from two structural mismatches: information conflict, where content and emotion in a shared latent space produce conflicting gradients, leading to reward hacking and semantic degradation; and scale gap, where sparse sentence-level rewards struggle to guide dense frame-level generation. To overcome these challenges, we propose HPRO, a hierarchical progressive reward optimization framework. Within HPRO, we introduce the HD-Emo codec as a novel differentiable reward model to mitigate the information conflict. It extracts speech into distinct content and style preference tokens, structurally isolating emotional optimization from semantic content. Building upon this structured preference space, HPRO bridges the scale gap by progressively aligning frame-, word- and sentence-level objectives. Experiments demonstrate that HPRO significantly enhances emotional expressiveness, while effectively preserving linguistic intelligibility. The code and audio samples are publicly available at https://xxh333.github.io/hpro-demo/.

eess.AS↗

A Gravitational Interpretation of Safety Reversion under Fine-Tuning

Safety alignment in large language models can degrade during post-training even when neither the data nor the objective is intentionally adversarial. Alignment rebound and reverse dynamics suggest that this degradation may reactivate behavior suppressed during safety alignment. Building on these ideas, we hypothesize that ordinary non-adversarial post-training follows a reversion direction: the activation-space displacement from the safety-aligned model toward a more permissive, earlier helpful-only state. We see that for Llama, every tested trajectory across references, tasks, and seeds exceeds a matched empirical null, while at aligned Llama and Qwen checkpoints, a vocabulary readout shows that the direction locally favors task-engaging over fixed refusal-like openings. Its geometric expression is behaviorally informative: as post-training proceeds, alignment with the direction and harmfulness increase together, yielding a strong descriptive correlation (Spearman r=0.958). To move beyond correlation, we test causal relevance during adaptation using objectives constructed from this coordinate. Across all tested Llama, Qwen, and Gemma settings from 3B to 14B, an optimizer-matched objective opposing positive motion reduces geometric alignment and harmfulness relative to ordinary fine-tuning, whereas a separately stabilized objective reinforcing that motion increases both. Every model and scale exhibits the same mean block-baseline-push ordering, showing that the causal relevance of the reversion direction is not tied to one architecture or model size. Finally, we show that a standard safety-rehearsal objective, built without access to the direction, independently opposes it and cuts cumulative reversion by about 30% in Llama and Qwen.

cs.LG↗

Second-Order Sensitivity of Efficient Solution Maps in Parametric Vector Optimization with Set Constraints

We study second-order fixed-ray Dini sensitivity of the efficient solution map \(S\) for the parametric vector problem \(\min_C f(p,x)\) subject to \(x\in H(p)\). Under a value-to-decision error bound (VDB) -- an inverse stability estimate -- the second-order semi-derivative of the marginal map \(Φ\) transfers to an exact second-order Dini formula for \(S\). The result separates value sensitivity from the inverse recovery of efficient decisions. This mechanism allows a nonsingleton efficient-value fiber and requires neither injectivity nor strict monotonicity of \(\nabla_x f\). For structured feasible maps \(H(p)=\{x\inΩ: g(p,x)\in D\}\), we derive primitive-data formulas for \(D^2H\), \(D^2Φ\), and \(D^2S\) under Robinson metric regularity, second-order regularity of \(Ω\) and \(D\), objective-aware recession conditions, and directional second-order semi-derivability of the data. Polyhedral inequality/equality systems give explicit specializations. A parametric multi-objective portfolio family verifies (VDB) with an explicit constant, obtained directly from the model's feasible directions, while a two-branch example shows how the formula operates on a nonsingleton decision fiber.

math.OC↗

Computational Complexity of Strong and Average Justified Representation

We study the approval-based multiwinner election problem where a set of $n$ voters cast approval-based ballots to a set of $m$ candidates, and we are to select a winner committee consisting of $k$ candidates. We consider two axioms: strong justified representation (SJR) and average justified representation (AJR). A winner committee satisfies SJR if the satisfaction for each voter in every $\ell$-cohesive group is at least $\ell$. AJR is a weaker axiom that requires the average satisfaction for each $\ell$-cohesive group to be at least $\ell$. It is well known that a winner committee satisfying AJR may not exist (and neither does SJR). In this paper, we study the computational complexity of the following decision problem: given an approval-based multiwinner election instance, decide if there exists a winner committee satisfying SJR/AJR. We prove that this problem is $Θ_2^p$-complete for SJR, and $Σ_2^p$-complete for AJR. As byproducts, we derive some results that are interesting in their own right. Firstly, we show that adding one more adaptive query to an NP oracle on top of polynomially many non-adaptive NP queries does not add more computational power, and the resulting complexity class is still $Θ_2^p$. Secondly, we construct a set system that can be useful in other applications, especially when doing reductions from typical satisfiability problems such as 3SAT.

cs.GT↗

Sharp hypercontractivity for free orthogonal quantum groups of Kac type

Let $N\geq2$, let $F\in GL_N(\mathbb C)$, and let $O_F^+$ be the associated free orthogonal quantum group of Kac type, with tracial Haar state $h$. We prove that, for $1<p\leq r<\infty$, the normalized heat semigroup $P_t=e^{-tL}$ on $O_F^+$ \cite{CFK14,BVY21} is hypercontractive with the optimal time \[ \|P_t:L_p(O_F^+)\to L_r(O_F^+)\|\leq1 \quad\Longleftrightarrow\quad t\geq\frac12\log\frac{r-1}{p-1}. \] In particular, this solves a conjecture of Brannan, Vergnioux and Youn \cite{BVY21}. To prove this hypercontractivity result, we establish the equivalent logarithmic Sobolev inequality \[ \operatorname{Ent}(|x|^2)\leq2\mathsf{E}(x,x), \qquad x\in D(\mathsf{E}), \] with sharp constant $2$, where $\operatorname{Ent}(y)=h(y\log y)-h(y)\log h(y)$ is the relative entropy and $\mathsf E(x,x)=\|L^{1/2}x\|_2^2$ is the Dirichlet form associated with $P_t$. The main ingredients are the cubic-majorant idea from the recent work of Frank--Ivanisvili \cite{FI26} and Xie--Zhang \cite{XZC26,XZ26}, and a centered third-moment estimate coming from the representation theory. We also prove the sharp Beckner inequalities \[ \frac{\|x\|_p^2-\|x\|_2^2}{p-2}\leq\mathsf E(x,x), \qquad 1\leq p\leq3,\quad p\ne2, \] whose limit as $p\to2$ recovers the logarithmic Sobolev inequality above. Moreover, when $N=2$, the results can be sharpened by replacing the heat semigroup with $\mathsf P_t=e^{-tΛ}$, where $Λ=\sqrt{I+3L}-I$. The key step is to use the graded-twist realization of $SU_{-1}(2)$ to transfer Beckner's sharp $p=3$ inequality on $SU(2)$ \cite{Beckner93}.

math.OA↗

Criteria for Weighted Homogeneity via Logarithmic Vector Fields

Recently in [6] the authors proved a lower bound for the $\GSV$ index of a logarithmic vector field along an isolated hypersurface singularity and proposed a conjecture that the weighted homogeneity of this hypersurface singularity can be detected by the existence of either a transverse or a non-degenerate holomorphic logarithmic vector fields. In this paper we prove this conjecture affirmatively. We also prove that the $\GSV$ index attains the lower bound if and only if the hypersurface singularity is weighted homogeneous.

math.AG↗

Influence of the KK graviton decay into hh on the triple Higgs measurement at LHC

Evidence for two KK graviton candidates has been previously reported at 380 GeV and 700 GeV. They would provide graviton factories in an e+e- collider. Following a Randall Sundrum interpretation, two extra resonances should appear at 1000 GeV and 1322 GeV. Recently ATLAS, in its search for triple Higgs coupling, has reported an excess in that mass region which comforts this prediction. Local cross sections are therefore clearly in excess of the standard predictions which implies muHH>1. While still marginally significant, this effect appears in a mass region with low background which allows to expect good prospects of discovery with RUN3 data. Such a result could therefore allow to confirm the existence of a series of KK graviton resonances observable at LHC. An optimised strategy, which includes H(95) is proposed to do so. Other opportunities appear in searches for heavier KK graviton recurrences decaying into ZZ/WW in semi-leptonic and fully hadronic modes. Prospects for measuring heavy KK graviton decays GKKi->G380G380 at HL-LHC are discussed. Indirect evidences for the RS model coming from precision measurements are also presented. KK lower mass limits for other KK particles are deduced from precision measurements at LEP/SLC and LHC.

hep-ex↗

FacePlex: Toward Natural Full-Duplex Conversational Avatars

Natural human conversation is inherently a real-time interaction in which speech and facial behavior continuously evolve. Enabling such interaction requires a conversational avatar to jointly generate speech and facial motion in real time, prepare facial motion for upcoming speech before the corresponding audio is emitted, and produce non-verbal reactions that reflect the ongoing dialogue. However, existing conversational avatar systems cannot address such requirements: audio-driven methods rely on pre-given speech, while joint streaming generation alone does not ensure anticipatory articulation or semantically appropriate reactions. We propose $\textbf{FacePlex}$, a unified framework for full-duplex speech-facial motion generation and real-time avatar rendering. FacePlex jointly coordinates speech, facial motion, and Gaussian splatting rendering on a shared streaming timeline. To prepare facial motion for upcoming speech, FacePlex predicts a short speech continuation and uses its future hidden states through asymmetric conditioning and denoising, while continuously updating them as new user audio arrives without observing future user input. For dialogue-grounded non-verbal behavior, we construct $\textit{SemReact}$, a dataset aligning dialogue context, reaction semantics, and facial motion, and introduce a semantic behavior router that guides continuous motion during both speaking and listening. Extensive experiments show improved audio-visual synchronization and facial articulation, natural dialogue-grounded non-verbal responses, and low-latency of 122 ms end-to-end avatar interaction.

cs.AI↗

Submission Responsibility Matters: Role-Aware Submission Quotas under Coauthorship

Author-level submission quotas are increasingly used to control growing peer-review load. Recent coauthorship-sensitive quota rules improve over fixed per-author limits by reducing the quota cost of multi-author submissions, often using harmonic authorship-credit models to prevent simple author-list padding. However, these rules conflate three distinct quantities: review burden, authorship credit, and submission responsibility. As a result, they can penalize genuine solo-authored work, treat all coauthors as equally responsible for a submission, and create bottlenecks for student-led papers when a faculty advisor appears on multiple unrelated submissions. We argue that submission quotas should be designed around the responsibility structure of a paper rather than only its number of coauthors. We formalize desiderata for quota rules, including venue-load control, padding resistance, role sensitivity, solo neutrality, and student non-blocking. We then propose a role-aware quota framework that assigns author-specific quota costs based on constrained roles such as lead author, regular coauthor, and designated advisor. The framework includes fixed, per-capita, and harmonic-style rules as special or limiting cases, while allowing venues to distinguish lead authors, corresponding authors, advisors, and peripheral contributors. We show how simple role constraints can preserve resistance to manipulation while avoiding several structural disadvantages of coauthor-symmetric quota rules. Our analysis suggests that role-aware quota mechanisms provide a more faithful and flexible foundation for managing peer-review load under modern collaborative authorship.

cs.DL↗

Why can genetic algorithms work in high-dimensional search spaces?

We show that the effective dynamics of the elitist $(1+M)$ genetic algorithm is, in the limit of small mutations, normalized gradient descent on the loss in the presence of anisotropic Gaussian white noise. In expectation, therefore, a simple mutation-selection genetic algorithm follows the gradient of the loss, without explicit calculation of gradients and without averaging over loss evaluations. The genetic algorithm is slower than gradient descent because of the noise that acts in directions transverse to the gradient. However, this slowdown is controlled not by the number of parameters of the search space but by the effective rank of the Hessian of the loss function. For the concentrated Hessian spectra observed in neural-network loss functions the effective rank can be far smaller than the number of parameters, which may explain why genetic algorithms can scale to large search spaces.

cond-mat.stat-mech↗

Reductive monoids over general base

We develop a theory of affine algebraic monoids over connected base schemes whose unit groups are split reductive groups. Our main result is a classification theorem for such objects, generalizing works of Vinberg and Rittatore over a field. As applications, we obtain combinatorial descriptions and normality properties of orbit closures, prove a Steinberg-type theorem on adjoint quotients of split reductive monoids, and construct finite type integral models of the Vinberg monoids.

math.RT↗