arXiv Science⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 1,117 records · Page 62Linked to original sources

Agentic High-Dimensional Bayesian Optimization with Hypothesis- and Evidence-Guided Search

High-dimensional Bayesian optimization (HDBO) seeks sample-efficient optimization when the number of variables is large relative to the evaluation budget. Recent LLM-based and agentic BO methods incorporate task knowledge and adapt search decisions during a run, but have primarily been evaluated on low- and moderate-dimensional problems. We ask whether this paradigm can transfer to the higher-dimensional regime. Our experiments show that these methods do not remain reliable in the high-dimensional regime, where the challenge is not only where to evaluate, but also which modeling assumption and search geometry to use when the objective's useful structure is unknown. We therefore introduce HERA, a Hypothesis- and Evidence-guided Research Agent that uses task context, optimization feedback, and structural diagnostics to revise search hypotheses, select and configure HDBO strategies, and determine their execution length. PRISM, its numerical optimization engine, generates and evaluates candidates sequentially within each search block, updating numerical models after each observation. HERA remains competitive with strong numerical HDBO baselines and outperforms the evaluated LLM-based and agentic methods on four metadata-free synthetic functions. Across eight real-world tasks, HERA achieves the best mean final objective among all evaluated systems on most benchmarks. Further analyses show that structural diagnostics change strategy use, metadata effects vary across tasks, and adaptive search blocks reduce inference cost.

cs.LG↗

Idempotents in the Temperley-Lieb Monoid and Other Categories

This paper examines idempotents in algebras and categories that arise from factorizations of the identity morphism. In the diagrammatic and combinatorial contexts considered here, these factorizations correspond to generalizations of meanders, where a meander is understood to be a curve in the plane that wanders transversely back and forth across a given straight line. By formulating the algebra of the Temperley-Lieb Monoid in terms of planar curve combinatorics, one can understand idempotents in the Temperley-Lieb Monoid in terms of meanders. Corresponding results are shown for the Brauer Monoid and for the Tangle Monoid and Tangle Category.

math.QA↗

Stable Constant Mean Curvature Hypersurfaces in $\mathbb{H}^n$

For every $n\geq 4$ and every $H$ in a neighborhood of $n-1$ (depending on $n$), we prove the existence of a properly embedded strongly stable CMC hypersurface in $\mathbb H^n$ with mean curvature $H$ and infinitely many ends, invariant under the action of a Schottky group. In particular, stable Bernstein-type rigidity fails in this range.

math.DG↗

Over-Personalization Is a Decision Failure: Generation-Induced Apply Bias in LLMs

Personalized LLMs must decide, for each stored preference, whether the current context calls for applying or suppressing it, which we call its applicability. They frequently over-personalize, applying preferences the context rules out, yet existing benchmarks score only the final response and cannot tell where this failure arises. We decompose preference handling into three stages and measure each separately: (1) knowing whether a preference applies, (2) deciding on an explicit Apply/Suppress label, and (3) generating a response consistent with that label. Using linear probes, we first show that this applicability signal remains decodable from hidden states during generation. By making the decision explicit, we then find that in most settings wrong decisions faithfully followed outnumber correct decisions lost in generation. We thus locate the failure in the decision, which breaks once the model is also asked to answer. To determine whether this reflects lost sensitivity or a response bias, we propose ABIDE (Apply-Bias Investigation via Decision-score), which adapts signal detection theory to Apply-vs-Suppress decision scores read directly from logits. ABIDE reveals a generation-induced Apply bias: merely stating an answer-generation objective shifts the decision score toward Apply while sensitivity is largely preserved, and the shift persists under controls for prompt structure, cascades across preference slots, and prompt wording. Finally, we show that subtracting a single bias scalar, estimated on a held-out split, from the decision score at decoding time reduces leakage while largely preserving fulfillment.

cs.CL↗

Epistemic Learning from Imprecise Annotation

Imprecise annotations may support several plausible labelling distributions, yet learning methods often resolve this ambiguity into a single predictive distribution. This can obscure what the annotation evidence leaves unresolved. We introduce epistemic learning from credal supervision, a framework that uses convex sets of plausible labelling distributions, called credal sets, as supervision and learns sets of predictive distributions. We instantiate the framework with the pessimistic--optimistic credal classifier (POCC), which combines a shared backbone with two classification heads trained to minimise worst-case and best-case losses over the supervision sets. Their outputs define a predictive credal set whose spread provides an uncertainty score. We also show how credal labels can be obtained through a simple relaxation of existing probabilistic labels, reducing commitment to their precise probability assignments. This construction admits closed-form inner optimisation under cross-entropy loss, enabling efficient training. Assuming the supervision sets contain the true conditional label distributions, and other regularity assumptions, we establish a finite-sample generalisation bound for the averaged predictor with an explicit penalty for supervision imprecision. We evaluate POCC using human annotator disagreement and teacher predictions, alongside label smoothing as a controlled proxy for annotation imprecision. Across these settings, POCC achieves a favourable balance of predictive accuracy, calibration, and uncertainty-based selective classification versus competitive baselines.

cs.LG↗

Dexterous Tactile World Model

World models for manipulation are typically trained from video, yet the events that determine how manipulation unfolds, such as making and releasing contact, are difficult to observe visually and are often easier to sense through touch. We present the Dexterous Tactile World Model (DTWM), a video world model for future-frame prediction of egocentric manipulation from both observed video and tactile signals from a glove worn on each hand. We condition a pretrained video diffusion transformer on each hand's tactile signal through a zero-initialized residual at the corresponding hand location in the video tokens, while a causal mask prevents predicted frames from accessing future information. Compared with a vision-only model matched in architecture, parameters, and training, DTWM reduces the underestimation of hand motion from 23% to 9%, while reducing the perceptual error in the hand region by 7.4% across three training runs per model. The benefit also increases over the prediction horizon, with the improvement in the later predicted chunks being about 4.1x larger than in the first. DTWM also outperforms other visual-tactile world models under the same setting, and training with touch improves future-frame prediction even when no touch is available at inference. Ablations show that the model benefits from both the magnitude and spatial location of force: replacing the tactile signal with binary contact states, either per hand or per location, increases prediction error. The observed course of the force indicates whether the interaction will persist or change.

cs.CV↗

ReScraper: Unified Scraping and Cleaning of Web Data for Effective LLM Pretraining

LLM pretraining corpora are normally cleaned by a stack of hand-written heuristics. A heuristic scraper extracts the main content from HTML, and dozens of rule-based filters then clean it, so corpus quality is capped by the coarseness and accuracy of the rules. In this work, we propose ReScraper, a unified language model of only 0.6B parameters that replaces this entire stack. To train ReScraper, we carefully curate supervised data from the outputs of three teacher models, so it learns to first extract the main content from raw data and then choose among four operations: keeping the page as extracted, editing out noisy lines and spans, deleting it entirely, or rewriting it when it is poorly written but informative. Based on the same crawled data pool, pretraining 400M, 1.4B, and 2.8B models on our curated data improves the DCLM Core score by a relative 3.8--4.7% over the strongest baseline at each scale, including the costly multi-agent curation. Our analyses show that each operation plays a distinct and complementary role, and that extracting and cleaning in one model outperforms a cascade of separate models. ReScraper also concentrates its operations on the pages that need them, raising the quality of poor pages the most while keeping the corpus diverse. These results demonstrate the feasibility and effectiveness of AI4AI for pretraining data curation, where a small learned model takes over an entire stage of the pipeline from hand-written heuristics. We open-source our code at https://github.com/cxcscmu/ReScraper

cs.CL↗

Linesearch-Free Nonlinearly Preconditioned GD with Semiautomatic Geometry Design

Nonlinear preconditioning makes it possible to adapt gradient updates to the growth of the objective's curvature, but combining its design with suitable stepsize selection is challenging. We propose adaptive dual preconditioned gradient descent (adaptive DPGD), which combines curvature-based preconditioner design with a linesearch-free stepsize rule based on local information. For convex objectives, we establish linear convergence under local strong convexity of the objective. By adapting the stepsize rule, we extend the method to weakly convex objectives and establish asymptotic stationarity without global Lipschitz smoothness. Moreover, our local assumptions give greater freedom in preconditioner design. We exploit this freedom to develop a semiautomatic recipe guided by the objective's curvature, yielding a family of preconditioners that includes those underlying normalized and hyperbolic gradient descent. Experiments on convex and weakly convex problems with real data demonstrate the effectiveness of the preconditioner design and adaptive stepsize selection.

math.OC↗

Higher orthogonality and truncation in canonical pretriangulated quotients

Canonical ideal quotients of extriangulated categories admit natural one-sided triangulated or pretriangulated structures. A fundamental question is whether these quotient structures still retain enough information to recover higher orthogonality properties of the subcategory being factored out. We show that, for a strongly functorially finite $n$-rigid subcategory, such information is encoded by the nilpotency of the canonical suspension and, equivalently, of the canonical loop. More precisely, this nilpotency characterizes two-sided maximal $n$-orthogonality. We further establish an objectwise recognition criterion that combines one-sided higher orthogonality with the vanishing of the $n$th suspension or loop. In the triangulated setting, our results give a converse to the known truncation construction and show that the quotient is $n$-truncated if and only if the subcategory is $(n+1)$-cluster tilting. Moreover, in the nonsplit case, the nilpotency index is exactly $n$. These results turn truncation from a consequence of higher orthogonality into a sharp criterion for recognizing it.

math.CT↗

Beyond Polytopes: Damped Newton Frank--Wolfe Methods over Compact Convex Sets

We study convex optimization over compact convex sets for twice differentiable objective functions with positive definite Hessians, assuming access to the feasible set through a linear minimization oracle. Existing second-order Frank--Wolfe methods either provide only a local linear rate over general convex sets or achieve global linear and local quadratic convergence only over polytopes. We propose damped Newton Frank--Wolfe methods that approximately solve constrained damped Newton subproblems using Frank--Wolfe or away-step Frank--Wolfe inner iterations. We develop residual backtracking, using a root Newton stepsize as a safe reference, together with a switching rule that eventually takes full steps. Under Hölder smoothness assumptions for $p$th derivatives, we establish global linear convergence and local convergence with Q-order at least $1+ν$ ($p=2$) and local quadratic convergence ($p=3$). Experiments on matrix sensing and ridge-regularized logistic regression show that the residual-backtracking variant is consistently robust and often achieves the best performance among the tested first- and second-order baselines.

math.OC↗

Counterexamples to preservation of flat normal bundles under mean curvature flow

We construct an embedded torus and an entire graph in $\mathbb{R}^4$ whose normal bundles are initially flat but lose this property instantaneously under mean curvature flow. We also give an example showing that the parallel principal normal condition is not preserved under mean curvature flow, even though the normal bundle remains flat along the flow.

math.DG↗

Pre-registered tests of solid-state-physics-inspired LLM compression: a cluster-level negative result at small-language-model scale

We report a three-month autonomous research-agent program testing five solid-state-physics-inspired compression mappings on pretrained language models, with predictions committed to git before any pilot data and a 3-sigma gate deciding PASS or SHELVE. The common anchor -- area-law / Kohn-nearsighted decay of the one-particle density matrix -- has a distance face (P001 Wannier, P002 tight-binding) and a rank face (P003 DMRG-truncated MLPs, P005 Wilson-RG, P011 tensor-train embeddings). P005 was pre-empted at Phase 1; three of four Phase-3 pilots were falsified. On the attention face, GPT-2-medium attention-versus-distance is best fit by a stretched exponential in 12 of 16 median-layer heads once probe padding is excluded, and a tight-binding cutoff costs +96% perplexity (P002); on Pythia-160M the Wannier sparsity 0.054 +/- 0.004 is indistinguishable from PCA, random-Haar and identity baselines (P001). On the rank face, per-token tensor-train bond dimension does not track surprisal (r = 0.016 vs a pre-registered 0.65) and the format inflates rather than compresses (P011). P003 is mixed: its scaling claim shelved (r = -0.434), its MPO premise died at stage-0, and its cross-paper check, r = 0.523 as first written, collapses to 0.047 under the same correction, leaving both cross-paper checks null. The results invert the pre-registered prediction that most attention heads behave like Kohn-nearsighted insulators, pointing instead to critical, glassy or heavy-tailed regimes; the inversion is specific to the <= 350M scale tested, while the rank-face no-gain result held to 7-8B. We contribute the pre-registration + 3-sigma + cluster-framing + append-only-catalogue discipline -- including why our own enforcement gate was designed but not deployed -- four pre-registered negative results with full data release, and the inversion. The catalogue holds eighteen concluded studies, seventeen negative.

cond-mat.dis-nn↗

Optimal Quantum Algorithms for Ordered Search

Ordered search is the problem of locating a target element in a sorted list of size $n$ using comparison queries. Classically, binary search requires $\lceil \log_2 n\rceil$ queries, which is optimal. Quantum algorithms offer a constant-factor speedup, but the precise constant has been a longstanding open question. We close this gap by exhibiting two new quantum algorithms for ordered search, each using the optimal $\frac{1}π\ln n+o(\log n)$ queries. The first, discovered by Claude Fable 5, is a simple zero-error algorithm derived from a continuum relaxation of the problem. The second, discovered by GPT-5.6-Sol (informed by Claude's zero-error algorithm), is an exact algorithm based on an analytic solution of the polynomial program of Farhi, Goldstone, Gutmann, and Sipser.

quant-ph↗

Semantic Modality Compensation for Unsupervised Visible-Infrared Person Re-identification under Unpaired Settings

Unsupervised visible-infrared person re-identification (USL-VI-ReID) learns person representations that can be compared across modalities without identity annotations. In the unpaired setting, however, identity correspondences between modalities are often incomplete, leaving many identities without an observed counterpart in the other modality. Existing unpaired methods bridge this gap by generating or mapping features for the other modality, mainly by exploiting the statistics of visual features without explicitly separating content that is discriminative for identity from style that is specific to modality. Consequently, the generated features may distort identity cues or inherit bias from the source modality, undermining the reliability of supervision across modalities. We formulate unpaired learning across modalities as a semantic compensation problem and propose Semantic Modality Compensation (SMC), a framework based on prompt composition that decouples identity semantics from modality style within a shared visual semantic space. SMC first constructs a discriminative ReID space through augmented dual contrastive learning, yielding pseudo labels, cluster prototypes, and memory banks for each modality. It then learns visible and infrared modality prompts in the CLIP semantic space and maps clusters obtained from pseudo labels to identity semantic tokens. For each cluster lacking a reliable match in the other modality, SMC combines its identity token with the prompt for the target modality to synthesize a semantic counterpart in the missing modality. The synthesized counterpart is then projected back into the ReID space and injected into a compensation memory through confidence gating. Extensive experiments under both paired and unpaired settings demonstrate that SMC consistently outperforms state-of-the-art methods, with particularly large gains when identity mismatch is severe.

cs.CV↗

Monotone Sobolev functions: approximation, critical points, and level sets

We give an affirmative answer to the planar local smoothing problem in Question~1.7 of D.~Ntalampekos and positive and negative answers to the basic approximation and level-set parts of his higher-dimensional Question~1.8. In every dimension $n\ge2$, each continuous Lebesgue monotone function in $W^{1,p}$ on a bounded open set admits uniform and strong $W^{1,p}$ approximation by monotone $C^{1,α}_{\loc}$ functions, with unchanged Sobolev boundary values and no increase of the $p$-Dirichlet energy, for $1<p<\infty$. The energy can be made strictly smaller unless the original function is $p$-harmonic. In the plane we obtain smooth local replacement at every isolated $p$-harmonic critical point, with arbitrarily small $C^1$ and Sobolev error. A point of gradient index $-m$ can be resolved into exactly $m$ nondegenerate saddles. Together with Ntalampekos's planar theorem, this gives smooth monotone density with fixed boundary values; at a nonsmooth $p$-harmonic minimizer the energy increase is unavoidable but can tend to zero. In dimensions $n\ge3$, an explicit Lipschitz monotone function has a nonmanifold point on every level in an interval, although it admits smooth monotone approximation. A product extension of a planar homogeneous $7$-harmonic function also rules out a general discrete exceptional set under the exact boundary and energy constraints. The answer to Question~1.8 thus distinguishes the five basic approximation properties from the stronger topological and exceptional-set conclusions. Applications include constrained integral functionals, strict $BV$ convergence of superlevel sets, nonlinear flux convergence, and stability of persistence diagrams under tameness assumptions.

math.CA↗

Dr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research Agents

Rubric-based tasks are increasingly addressed through reinforcement learning (RL), with rubric scores used as training rewards. However, these rewards typically supervise final answers without distinguishing the contributions of intermediate decisions. Many existing credit assignment methods rely on ground-truth answers to define process rewards, limiting their applicability to open-ended tasks without canonical solutions. To address this limitation, the proposed rubric-grounded credit uses task requirements as a shared reference for final answer evaluation and process supervision. The information returned by tools is assessed for the additional support it provides toward satisfying each rubric relative to that rubric's history of accepted support. By referencing these histories, credit distinguishes new support from evidence already present in the trajectory while recognizing partial support for each rubric. Dr.Credit uses rubric-grounded credit to supervise intermediate tool turns in an RL framework for deep research agents. The resulting process advantages are combined with GRPO outcome advantages to guide research decisions while retaining supervision of final-report quality. Evaluations on four in-domain and out-of-domain benchmarks show that Dr.Credit outperforms the evaluated open deep research baselines on every primary metric and submetric. Meanwhile, with an 8B-parameter backbone, the trained agent achieves average performance competitive with the evaluated frontier proprietary models. Further analyses suggest more efficient evidence acquisition and higher-quality reports under limited research-turn budgets, motivating the extension of rubric-grounded process supervision to a broader range of rubric-based tasks.

cs.CL↗

TLC-DiT: Task-Aligned Local Visual Conditioning for Robust Multitask Robot Manipulation

Language-conditioned robot policies have made clear progress in multitask manipulation, but task-relevant local visual evidence usually stays hidden inside a visual backbone or attention layers. This leaves the policy difficult to inspect and fragile under visual change, two symptoms of a missing explicit, task-aligned local visual channel. We present TLC-DiT, a plug-in extension of the Multitask Diffusion Transformer (DiT) policy that adds explicit task-guided local visual feature maps without changing the diffusion objective or the action-generation process. For each camera view, frozen DINOv2 patch features are modulated by the CLIP task embedding through FiLM and refined by a lightweight CoordConv CNN adapter into smooth spatial maps, which are concatenated with the original global image, language, joint-state, and timestep conditions. On LIBERO, TLC-DiT reaches a 93.5% average success rate, compared with 86.5% for Multitask DiT and 79.25% for SmolVLA. On LIBERO-plus, the total success rate improves from 54.07% to 57.24%, with larger gains under camera, background, and sensor-noise changes. In real-world bimanual tasks, TLC-DiT raises Teabag Putting completion from 44% to 89% while maintaining comparable Match Box Opening performance. Feature-map visualizations confirm that the model attends to task-relevant regions across views and perturbations, providing a direct way to inspect the visual evidence.

cs.RO↗

Nuclear scattering phase shifts with factorized geometric-time RODEO on a quantum processor

We reconstruct elastic $s$-wave neutron-proton ($np$) phase shifts from trapped spectra measured on IBM Aachen using a compressed RODEO circuit. For a schematic square-well interaction represented by classically constructed four-dimensional effective Hamiltonians, four positive-energy levels at five trap strengths supply twenty inputs to classical modified effective range expansion (MERE) extrapolation to free space. The six-cycle factorized geometric-time RODEO implementation (FG-R6) combines ancilla reuse, exact query-phase separation, post-compilation binding, and six numerically optimized geometric evolution times. Numerical cycle-count tests support this choice within the adopted local spectral tolerances. At the same three-qubit width and equal shot budgets, FG-R6 reduces median compiled depth and two-qubit-gate count by 33.1% and 37.2% relative to ten-cycle dynamic direct RODEO (direct R10), increases the median fitted amplitude, and lowers median finite-shot energy uncertainty from 0.610 to 0.534 keV. For this dataset's primary fit on the 0.1-30.0 MeV grid, the maximum central phase-shift deviation from exact-energy MERE decreases from $0.842^\circ$ to $0.163^\circ$, despite a larger root-mean-square (RMS) trapped-energy deviation. The full-grid gain arises mainly in the low-energy extrapolation region; direct R10 has slightly smaller residuals relative to the same reference on the 10.0-30.0 MeV common energy-interpolation subset. The FG-R6 central curve differs from the analytical square-well solution by at most $0.156^\circ$. Fit-form sensitivity remains appreciable and is assessed separately from finite-shot uncertainty. This reduced-space benchmark connects RODEO circuit compression to nuclear continuum observables and shows why circuit performance must be assessed through the scattering reconstruction and its energy range, not spectral RMS errors alone.

nucl-th↗