arXiv Science⌕ Search

SEARCH · arXiv Science

Search arXiv Science

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,495 records · Page 83Linked to original sources

AdSpark: A Large-Scale Dataset and Benchmark for Product-Centric Advertisement Video Generation

Product-centric advertisement video generation aims to create promotional videos that preserve fine-grained product identity while presenting selling points through coherent multi-shot narratives. However, this emerging task remains underexplored due to the lack of large-scale advertisement-specific datasets and comprehensive evaluation frameworks. To address this gap, we introduce \textbf{AdSpark}, a large-scale dataset and benchmark for product-centric advertisement video generation, based on data from a major e-commerce platform. \textit{AdSpark-300K} contains approximately 300K reference image--prompt--video triplets, comprising a real-world subset and a synthetic subset. Each sample provides structured advertisement annotations, including product identity annotations, selling-point descriptions, creative plans, and aligned audio scripts, enabling models to learn product preservation and advertisement-oriented visual storytelling. We further propose \textit{AdSpark-Bench}, a diagnostic benchmark that evaluates generated advertisements across six dimensions, including visual quality, product fidelity, instruction adherence, temporal coherence, audio alignment, and advertisement effectiveness. Based on AdSpark-Bench, we evaluate representative models, revealing key challenges in product preservation, multi-shot storytelling, and selling-point visualization. Experiments with AdSpark-300K-finetuned models further validate the effectiveness of our dataset. AdSpark provides a unified dataset and benchmark for future research, and we will release the dataset upon acceptance.

cs.CV↗

Coupled quantum critical states in a circuit simulator

Hybrid metal-semiconductor quantum circuits offer a controlled setting for realizing and probing strongly correlated quantum matter. A largely open question that may be addressed by such simulators is whether coupled critical states can combine into new non-Fermi liquid (NFL) phases. Inspired by the recent realization of a two-site charge-Kondo circuit [W. Pouse \textit{et al.}, \href{https://www.nature.com/articles/s41567-022-01905-4}{Nat. Phys. \textbf{19}, 492 (2023)}], we uncover the fate of coupled Kondo anyons emerging from the two-channel-Kondo critical point at each site. Bosonization and quantum Brownian motion methods yield the full quasiballistic phase diagram and low-temperature transport properties, while numerical renormalization group calculations resolve the complementary weak-coupling limit. Together, these provide a unified description, and reveal robust NFL phases with continuously tunable transport and susceptibility exponents. Stronger inter-site coupling ultimately produces a screened Fermi liquid.

cond-mat.mes-hall↗

The Long Road to the Same Answer: Cognitive Bias Under Escalating Reasoning Budgets in Large Language Models

Reasoning models allocate extra computation at inference time and present their answers as the product of deliberate thought. If this deliberation works the way dual-process accounts of human cognition suggest, longer thinking should weaken the classic decision biases that fast, intuitive judgment produces. Using 30 vignettes covering six biases (anchoring, framing, loss aversion, escalation of commitment, availability, confirmation) from an established benchmark, we run a dose-response study across four model families, pairing each reasoning model with a matched non-reasoning sibling and requesting thinking ceilings of 0, 1,024, 4,096, and 8,192 tokens, for 12,350 API calls. Because a requested ceiling is not the same as realized deliberation, we use the reasoning tokens each call consumed as the dose. First, reasoning models are not less biased than their siblings; the point estimate leans the other way in every family, but the item-level pooled contrast is not reliable (Delta = +0.031, t(29) = 1.45, p = .157). Second, bias magnitude does not reliably fall as realized deliberation grows: no slope is significantly negative, and where anything moves it is the signed score drifting further from the human direction. Third, anchoring is the only bias in the human direction (d = 1.89). Four of the other five lean the opposite way in all seven models; with five items per bias, that reversal is reliable for framing and directional for escalation of commitment, confirmation, and loss aversion, while availability is absent. A one-line instruction to restate the anchor before answering lowered anchoring on all five anchoring items, which no amount of additional thinking did, although the effect does not reach significance (p = .057). The results argue against treating test-time reasoning as a rationality guarantee and for auditing deployed models bias by bias.

cs.CL↗

Bounds on the maximum number of limit cycles of piecewise linear Lienard systems I. The continuous case

This paper concerns the planar Liénard system \(\dot x=y-F(x),\ \dot y=-x\), where \(F(x)\) is a continuous piecewise linear function with exactly \(n\) fold points. Tonnelier [SIAM J. Appl. Math. 63 (2002)] conjectured that the maximum number of limit cycles of the system is \(n\) when \(F\) has exactly \(n\) fold points. The conjecture was confirmed for \(n=1\) and \(n=2\) in [J. Nonlinear Sci. 25 (2015)] and [J. Lond. Math. Soc. 113 (2026)], respectively, whereas the case \(n\ge3\) remained open. In this paper, we show that the maximal number of limit cycles is at least \(|3n-4|\) for \(n\in\mathbb N^+\), thereby disproving Tonnelier's conjecture for \(n\ge3\). The proof reveals a unified multiscale perturbation mechanism underlying the creation and coexistence of multiple limit cycles. Moreover, we prove that the number of limit cycles admits the finite upper bound \(2^{28(n+1)^2}\).

math.CA↗

Almost-splitting of spacetimes under timelike sectional curvature bounds

We establish an almost-splitting theorem for spacetimes under timelike sectional curvature bounds in the spirit of the celebrated Riemannian result due to Cheeger and Colding. More concretely, we show that if a globally hyperbolic spacetime with timelike sectional curvature bounded below by a nonpositive constant close to zero contains a long maximizing timelike segment, then causal diamonds along that segment are close in the Lorentzian Gromov--Hausdorff sense of Minguzzi--Suhr to those in a product metric spacetime. Our result can be interpreted as a quantitative analogue of the Lorentzian splitting theorem for spacetimes with nonnegative timelike sectional curvature by Beem--Ehrlich--Markvorsen--Galloway.

math.DG↗

Consecutive Rankin-Cohen Bases, Full-Spark Periods, and Divisor-Tau Congruences

For $r\ge1$, let $K=2r+14$ and $d=\dim S_K$. We prove strict positivity for determinants of Mellin period functionals, yielding a full-spark theorem for periods in a fundamental half-range and, in particular, the mixed odd--even period independence conjecture of Xue. As a consequence, the consecutive first Rankin--Cohen brackets $[E_{2r+10-2j},E_{2j+2}]_1$, $1\le j\le d$, form a basis of $S_K$, and we determine the signs of all admissible ordered determinants of first Eisenstein brackets. This gives a uniform exact all-weight formula for the divisor--tau convolution $C_r(n)=\sum_{m=1}^{n-1}σ_{2r+1}(m)τ(n-m)$. Reducing the same coordinate identity modulo primes, we obtain canonical prime-wise reductions for infinitely many primes in every weight and characterize sparse Ramanujan-type specializations through vanishing Cramer coordinates. Finally, for homogeneous $f,g\in\mathbf Q[E_4,E_6]$, we prove $[f,g]_1/Δ=-3456\det((f_{E_4},f_{E_6}),(g_{E_4},g_{E_6}))$, reducing the Cramer system to a one-variable coordinate problem in $T=E_6^2/E_4^3$.

math.NT↗

On Bonart's interpretation of the Square-Root Impact Law

The square-root impact law (SRIL), $I = Yσ\sqrt{Q/V}$, bundles two facts that a single mechanism must explain at once: a shape (impact proportional to square-root of traded volume $Q$) and an amplitude ($Y=O(1)$, independent of the participation rate $φ$). Bonart has recently proposed an elegant solution: if realized and counterfactual prices are both diffusive, information-neutral impact must have white increments and the SRIL follows without the wart. We reformulate and simplify his argument, and foreground the ingredient he himself regards as essential --- that the market whitens the stationary stream of each participant's metaorders, not any isolated one. The Lillo--Mike--Farmer model satisfies his premises, yet gives an isolated metaorder the super-square-root impact $φ^{(1-γ)/2}Q^{(1+γ)/2}$, $γ\in(0,1)$, because single metaorders (i.e. not part of a stream) are statistically invisible and thus priced mechanically. We then construct an explicit multi-agent propagator that realizes Bonart's whitening, and show it reproduces the clean SRIL precisely when the market can attribute trades to their issuer. But correct re-attribution of anonymous trades is a formidable task, which requires markets to behave as implausibly efficient signal processing machines. The stream hypothesis makes a sharp, falsifiable prediction --- isolated metaorders, well separated in time, should not obey the SRIL --- that does not seem agree with empirical data known to us. Independently, a volume-conservation argument singles out the square-root law and lays bare the step the diffusivity route must assume: diffusivity fixes the variance of impact --- an additive, stationary, age-blind quantity --- whereas the SRIL is a law for the mean impact, which marginally decays as $1/\sqrt{\text{age}}$ and therefore demands an origin of time.

q-fin.TR↗

Transition Path Sampling Using Koopman Operators and Exit-Time Optimal Control

Sampling transitions between metastable states is a central problem in dynamical systems theory and molecular dynamics in particular. A key challenge is the existence of high free-energy barriers that separate the states, making transitions extremely rare. Recent machine learning-based methods cast transition path sampling (TPS) as an optimal stochastic control (OSC) problem over a fixed time horizon, and parameterize the drift bias via a neural network trained by simulation-in-the-loop, requiring repeated biased rollouts. To address computational and performance guarantee issues of these models, we propose a new approach for the problem based on Koopman operators. Because Koopman operators are linear, their leading eigenfunctions reveal the metastable sets and provide an estimate of the committor function with no transition path information required. Furthermore, we formulate TPS as an OSC problem up to an exit time. Our time horizon is the first hitting time of the target set, and our running cost penalizes time spent in nonreactive regions by encoding the estimated committor function. We derive the optimal controller in closed form and approximate it in a reproducing kernel Hilbert space (RKHS). This reduces the problem of constructing the optimal controller to solving a single equality-constrained quadratic program, whose solution can be characterized by a linear Karush-Kuhn-Tucker (KKT) system. On the two-channel double well and alanine dipeptide, our controller increases the fraction of trajectories reaching the target from 0% to 99.8% within 1000 steps, and from 0% to 93% within 1ps, respectively.

eess.SY↗

Human movement reconstruction through latent structural representations during assisted transitions and multidirectional hopping

The quantification of three-dimensional human kinematics is fundamental to neurorehabilitation and musculoskeletal research, yet atypical posture, physical assistance, and rapid movement create conditions of partial observability. Here, we evaluate a latent-structure-informed framework combining convolutional rank-reduction autoencoding with feedforward regression to reconstruct movement from synchronized video. Three case-based demonstrations span therapist-assisted sit-to-stand in a child with cerebral palsy, unassisted sit-to-stand in a typically developing child, and multidirectional single-leg hopping in a young adult. Reconstruction was examined through positional trajectories, projected angular descriptors, and hip-knee coupling. Across test repetitions excluded from model training, mean absolute errors in a predicted marker position during assisted sit-to-stand ranged from 12.02 to 57.11 mm. Sagittal knee and trunk descriptor errors were 5.32° and 5.38° in the typically developing case. Hip-knee cyclograms captured aspects of coupled-motion patterns in both cases, with closer trajectory correspondence in the typically developing case. Sixteen test hops yielded a vertical knee-marker mean absolute error of 15.58 mm. These findings connect latent-structure-informed reconstruction with biomechanically interpretable movement features, providing a foundation for video-derived assessment spanning clinically atypical movement and high-dynamic athletic movement.

cs.CE↗

A solid-state nuclear clock based on VUV absorption spectroscopy of $^{229}$Th

We demonstrate sustained operation of a solid-state thorium-229 nuclear clock using continuous-wave absorption at 148.4 nm. In a $^{229}\mathrm{Th}:\mathrm{CaF}_2$ crystal, we resolve four quadrupole components of the D center and a broader O-center resonance. Feedback on the strongest narrow component is updated approximately every 10 s and remains active throughout a 30-h record. A hydrogen-maser-referenced frequency comb measures the output. Its fractional frequency instability follows approximately $1.29\times10^{-11}(τ/\mathrm{s})^{-1/2}$ and reaches $1.24\times10^{-13}$ at $10^4$ s. A separate 6.6-h run demonstrates feedback on a second quadrupole component. These measurements connect the resolved absorption spectra with sustained nuclear-clock operation and frequency comparisons between crystal segments.

physics.atom-ph↗

WxFM-XL: Adapting Univariate Foundation Models to Multi-Station Weather Forecasting

With the rise of univariate time series foundation models (e.g., Sundial, Timer), initial efforts have been made to extend them to multivariate settings. However, these models mainly focus on modeling correlations among variables. When they are applied to multi-station weather forecasting, two important factors are often overlooked: (1) the spatial information of stations, and (2) different error priors of different stations relative to the foundation model. In this paper, we propose WxFM-XL, a model for adapting univariate time series foundation models to multi-station weather forecasting. WxFM-XL introduces a cross-station error correlation prior graph to capture stationwise error priors with respect to the foundation model. Building on this, we further propose a dynamic fusion mechanism that adaptively integrates a spatial correlation graph with the error correlation prior graph. Experiments on multiple datasets demonstrate that our model outperforms state of the art baselines.

cs.LG↗

Cache the Encoder Within:Compact, Reusable Memory across LLM Queries

Repeated queries over shared documents incur redundant encoding, while caching model states introduces persistent storage costs. Building on CoMem's intermediate-state interface, EncBank treats a pretrained LLM's lower layers as a reusable document encoder and compactly stores their outputs for an adapted upper-layer reader. A self-distilled suffix adapter is shared across storage precisions within each backbone, without quantization-specific retraining. Across five benchmark suites on three Qwen backbones spanning different sizes and full-attention and hybrid architectures, 4-bit storage keeps each reported benchmark aggregate within one score point of native-precision EncBank. In a fixed Qwen3-8B workload, it retains 28.1% of the native-precision persistent GPU store. Separate native-precision controls yield a 1.40x selected-pack prefill speedup over same-evidence, same-adapter text replay, at a 3.12-point RULER accuracy cost. A native-precision Qwen3.8-27B configuration also passes 70 of 89 Terminal-Bench 2.1 tasks. EncBank thus combines reusable computation with compact memory, while task fidelity and end-to-end benefits remain dependent on the workload, preparation costs, and reuse frequency.

cs.CL↗

A toy model for the quark pole mass

We study the pole mass for a heavy quark probe in the Gross--Neveu model, at the first non-trivial order in the $1/N$ expansion. We show that this mass suffers from a perturbative renormalon ambiguity structurally similar to the one in QCD. The exact large $N$ solution leads however to an explicit expression for the pole mass which can be decoded as a trans-series, and we find that non-perturbative corrections cure the renormalon ambiguity. We also calculate the self-energy of the heavy quark beyond perturbation theory by using the OPE method with condensates, and show that the results are in agreement with the large $N$ solution. The non-perturbative corrections to the pole mass turn out to be determined to a large extent by the condensate corrections to the self-energy. However, there is a class of corrections that has to be reorganized due to on-shell effects and cannot be obtained from the conventional condensate expansion. We discuss possible implications of this toy model for QCD.

hep-th↗

TRACK: Telemetry-Based Racing Analysis and Coaching Kit in Sim Racing Games

This paper presents TRACK (Telemetry-Based Racing Analysis and Coaching Kit), which is a framework for analyzing driving performance in sim racing and profiling how individual drivers behave behind the wheel. We report this framework together with its limitations: we calibrate each clustering result against a null, and when one does not separate from chance, we say so. Instead of restricting ourselves to scoring drivers or sorting them into preset labels, we represent each recording session as a compact geometry in a four-dimensional behavioral space (speed, braking, strategy, and consistency), and we group these fingerprints by their similarity using unsupervised clustering. Over time, we have developed and refined this framework on the open Assetto Corsa Gym (ACGym) dataset. Our study suggests that corner types differ along a behavioral dimension that was not used to define them. It also suggests that when the car changes, only speed and consistency carry over in the restricted population, while repeatability could not be shown there for any of the braking or strategy measures. Cluster separation becomes less distinct as the range of available telemetry widens. Until that repeatability is shown, grouping on the braking and strategy dimensions cannot treat the car as interchangeable, which divides an already small sample into smaller cells. It is also not clear whether a driver's grouping carries over from one corner type to the next. We also normalize each metric against a reinforcement-learning reference agent. The reference does not depend on the sample, so the scale does not shift when the sample does. We intend these results as an analytical foundation for a personalized improvement suggestion system. The sample is small. The cross-car result changes when the sample is defined more broadly. These outcomes are preliminary.

cs.GR↗

Loud Failures, Quiet Failures: Fault Detection and Recovery in Tool-Using Language Model Agents

Tool-using agents are usually scored on whether they finish a task while the tools work. Deployments are less forgiving: services time out, endpoints disappear, parameter names change, and results come back well formed but wrong. Prior work has shown that language models over-trust tool outputs that fail silently; we ask how that over-trust plays out across the stages of failure handling in multi-turn agents. Wrapping the executable environments of an established function-calling benchmark in a fault-injection layer, we inject one of four typed faults at a controlled point in the trajectory and record whether the agent notices, changes plan, recovers the task, or repeats itself. Six models from three families, half of them reasoning variants, ran 1,920 trials over 24 multi-step tasks. Agents treat a failure as a problem in 91.3% of trials when the tool returns an explicit error, but in 58.8% of trials when it returns a plausible wrong value, against a 26.8% rate of reporting problems when nothing was wrong. Reasoning models are not better placed: paired against instruct siblings, they notice less (-9.3 points, p < .001) and change plan more (+10.4 points, p < .001), and recovery is unchanged (p = .512). Because agents are stochastic, two fault-free runs of the same task end in the same state only 63.3% of the time; against that baseline, only a missing tool clearly lowers recovery (39.9%), while timeouts, schema drift, and corruption stay within run-to-run variation. After a fault, agents return to the same tool three or more times in a row in up to 22.2% of trials, though strictly identical repeats are rare. A prompt line asking the agent to check each result did not move detection. Agents respond to the error channel rather than to the content of what a tool returns, so failures that stay inside the expected format pass through.

cs.AI↗

Universal orbital-coupling rules for hydrogen-defect interactions in bcc metals

Hydrogen-defect interactions control the performance of body-centered-cubic (bcc) metals. However, the complex variations among defects lead to various empirical models that lack transferability across defects. Here, we propose an orbital-coupling model, built on coordination number and interatomic distance, that quantifies the H solution energetics across nanovoids, vacancy loops and grain boundaries in bcc metals, and even predicts the potential-energy surfaces of H at nanovoids. Our model reveals an unconventional s-d coupling rule in the confined environment of defects: H-metal interactions exhibit a unique coordination-dependent law, whereas H-H interactions, deviating from the usually speculated s-s coupling, acquire the distance-decay law of H-metal coupling. This unusual rule proves essential to reproduce the experimentally observed bimodal profile of H desorption. Our electronic-structure-origin, unified model is thus crucial to understanding the nature of chemical bonds under constraint and engineering the H-tolerant materials.

cond-mat.mtrl-sci↗

Comprehension Audits to Mitigate Risks from Automated AI Research

AI is already writing a majority of code for frontier AI labs. This creates a safety risk if there is insufficient human oversight. Existing work proposes minimum comprehension thresholds and unaided checks to mitigate this. To our knowledge, however, there is currently no published frontier-AI assurance regime that requires demonstrated evidence that the responsible humans understand what they are building as a precommitted condition for continuing development or usage. We propose comprehension audits, a novel development-process assurance mechanism in which the responsible people explain R&D contributions to auditors to demonstrate understanding. With independent administration and graded reports, they provide a gate: development of a contribution stops based on a failure to demonstrate human understanding until remediated, with escalating consequences for repeated failures. Our analysis of leading open-source AI projects finds increased output of code with reduced human review commentary rates per line of code, with far lower rates for automated fleet accounts. We advocate for labs to conduct them with embedded independent auditors.

cs.CY↗

Distributed Motion Planning for Multi-Robot Systems under Topological Constraints

Efficient and distributed coordination of mobile robots is one of the main challenges in multi-robot systems. Topological constraints, often expressed as topological braids, are a popular tool to encode complex coordination patterns between multiple mobile robots, as they offer a compact and abstract representation of the desired qualitative relation between the space-time trajectories of the robots. However, execution of joint motion plans encoded as braid-based topological constraints via distributed controllers is challenging, with existing approaches, generally based on the execution of one braid generator at a time, producing slow and suboptimal trajectories. We propose a distributed controller based on Model Predictive Control (MPC) to efficiently execute braid-based topological specifications. Rather than directly tracking the braid specification, we propose to use winding numbers, which are topological invariants for braids, as a proxy. This has the twofold benefit of converting braids into a continuous function, which can be easily tracked by an MPC controller through an appropriate term in the cost function, and of decoupling the global braid specification into a set of pairwise specifications, which can be tracked distributedly through the solution of only local MPC problems. To maintain global coordination, we propose a consensus-based progress estimation approach, which allows the robots to synchronize their motion toward the desired specification. We validate the proposed approach in simulation and in real-world experiments, where we demonstrate the effectiveness of the proposed approach and the improvement over existing approaches in terms of execution speed and control effort.

cs.RO↗