arXiv Science⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 1,135 records · Page 63Linked to original sources

PSM: Dataset Distillation Based on Precise Statistical Matching by Difficulty

Dataset distillation (DD) condenses a large original dataset into a small distilled dataset with high training utility. Decoupled statistical matching methods substantially reduce distillation time and memory overhead while achieving strong performance. However, they typically supervise all distilled samples using running statistics estimated from the entire original dataset. These statistics mainly capture the average feature distribution while overlooking differences in sample difficulty, limiting their ability to characterize the difficulty structure of the original data. To address this issue, we propose Precise Statistical Matching (PSM) by difficulty. After pretraining, PSM uses the Global Precision Score (GPS) to estimate image difficulty, ranks the samples within each class, and partitions each class into IPC (images per class) difficulty groups. During distillation, Statistics Updated Again (SUA) updates the teacher's batch normalization (BN) running statistics through forward passes on original samples from each group, providing difficulty-specific supervision for the corresponding distilled batch. Meanwhile, Initial Sample Screening (ISS) initializes distilled samples using original images from the corresponding difficulty group, providing an effective starting point for precise matching. Experiments across multiple datasets and model architectures demonstrate that PSM broadens the difficulty range of distilled samples and improves downstream performance in most evaluated settings. Code will be released.

cs.CV↗

When World Models Lie: Adaptive Safety Analysis Under Wrong Imaginations

World models offer a powerful substrate for safety reasoning in high-dimensional robotic systems, but they are also fallible: their predictions can be biased, miscalibrated, or confidently wrong. This creates a central challenge for latent-space safety filters, which often learn Hamilton-Jacobi safety value functions on the dynamics of a world model. If the world model is incorrect, the resulting value function can inherit its errors and produce overconfident safety estimates. Existing latent safety filters often rely on auxiliary signals such as ensemble disagreement or value-target consistency residuals for adaptation, but these signals can remain small even when the world model's predictions deviate from observations. We propose an adaptive latent safety filter that calibrates safety reasoning using directly observed world-model error. Our method uses Adaptive Conformal Inference to construct online uncertainty sets from discrepancies between predicted and observation-inferred latent states, then evaluates safety pessimistically by minimizing the learned value function over these sets. This allows the filter to remain minimally conservative when the world model is accurate, while becoming more cautious when observations reveal model mismatch. We provide a finite-time coverage guarantee for the adaptive uncertainty radius. Through simulation and hardware experiments, we show that our method significantly reduces failures relative to state-of-the-art latent safety filters while preserving task completion.

cs.RO↗

One Sequence, Many Decodings: CAGenMol-2 Recasts Drug Design as Masked Molecular Inference

Drug design couples property evaluation, conditional generation, structure-based design, and local optimization, yet machine learning systems typically address these capabilities with separate task-specific models. We introduce CAGenMol-2, a masked diffusion molecular language model that represents molecules, continuous scalar properties, and 3D protein pockets within a single wrapped sequence. Within this pretrained interface, downstream operations are selected by which sequence regions are observed or masked at inference, allowing one checkpoint to perform property prediction, property- and pocket-conditioned generation, and partial-constraint design without task-specific architectures or backbone fine-tuning. We further propose Adaptive Fragment Optimization (AdaFO), a gradient-free mask-and-refill search that turns the masked decoder into an iterative local molecular optimizer. On CrossDocked2020, AdaFO increases Success Rate from 30.2\% to 70.8\%, the best reported under this protocol, while largely preserving drug-likeness and diversity. Finally, scaffold-preserving directional editing and CRBN/VHL case studies demonstrate its use in compound design workflows spanning local molecular editing, structure-based prioritization, and downstream simulation-based screening.

cs.LG↗

CAT-Free: Multi-View Pedestrian Localization without Calibration, Annotations, or Target-Scene Training via Adaptive Geometric Filtering

Multi-camera pedestrian localization is useful for wide-area monitoring in public and commercial spaces. However, deploying these systems often requires considerable setup for each new environment. Existing methods typically require camera calibration, position annotations, or target-scene training. CAT-Free removes all three requirements. It uses synchronized RGB video as its only scene-specific input. Camera configuration is estimated directly from the video. Pedestrian locations are then estimated by combining observations from multiple cameras. Automatic camera estimation is not always accurate. This can produce unreliable pedestrian locations. CAT-Free therefore introduces two adaptive geometric filters. They remove unreliable position estimates. Their thresholds are estimated from each input sequence. CAT-Free achieves 82.5, 84.5, and 65.7 MODA on WildTrack, MultiviewX, and GMVD. It uses no supplied calibration, position annotations, or target-scene training. Published methods using such scene-specific information report 88.2--95.0 MODA on WildTrack and 83.9--96.5 on MultiviewX under their respective protocols. CAT-Free also transfers without retuning. It reaches 74.9 MODA on four additional sequences and 78.6 on an unseen 8-camera installation. Finally, localization uncertainty predicts MODA with $r=-0.98$. This provides a label-free estimate of localization reliability.

cs.CV↗

Asymptotic Lower Bounds for Continuous Optimization

Nonasymptotic convergence rates for optimization problems have been extensively studied across a wide range of settings, with carefully constructed worst-case instances establishing matching lower bounds for many algorithm classes. Recent work has shown, however, that these rates can often be improved in the asymptotic regime, while the corresponding asymptotic lower bounds remain largely unknown. We propose a general technique for converting existing nonasymptotic lower bound constructions into asymptotic lower bounds on Hilbert spaces. This allows us to show that several known asymptotic convergence upper bounds in Hilbert spaces are tight. In particular, we recover tightness of the $o(n^{-2})$ suboptimality guarantee for accelerated gradient descent on smooth convex functions \cite{attouch2016rate}, as well as the $o(n^{-1/2})$ suboptimality guarantee for gradient descent on smooth nonconvex functions \cite{gratton2025refining}. For stochastic convex optimization, we also develop a method that achieves an $o(n^{-1/2})$ suboptimality guarantee in finite dimensions, and prove a matching one-dimensional lower bound.

math.OC↗

EXO-200 Public Data Release for AI/ML Applications

We present a public release of a subset of calibration data from the EXO-200 experiment, a liquid xenon time projection chamber designed to search for neutrinoless double-beta decay in $^{136}$Xe. The released dataset consists of approximately one million events collected during $^{228}$Th calibration campaigns performed near the end of Phase-II operations between February 26 and March 2, 2018. For each event, the dataset includes raw detector waveforms from the U-wire, V-wire, and avalanche photodiode (APD) readout channels together with reconstructed event quantities including charge energy, light energy, rotated energy, and charge-cluster information. The dataset is distributed in HDF5 format and is intended to support data preservation, educational activities, and the development of modern machine-learning techniques for rare-event physics. This paper describes the detector, dataset contents, file structure, and public access mechanism.

physics.ins-det↗

Improving Lens Modelling from Ground-Based Imaging with Deconvolution

Accurate and precise mass modelling of strong lensing systems is essential for extracting insights into cosmology, galaxy environments, and the nature of dark matter. The Hyper Suprime-Cam Subaru Strategic Program (HSC-SSP), which has mapped $\sim$1000 square degrees of the sky, discovered over 1000 definite or probable galaxy-scale gravitational lens candidates. As the resolution of ground-based imaging poses challenges for accurate lens modelling, high-resolution imaging is traditionally sought after for detailed analyses of these systems. In this work, we instead propose applying the STARRED algorithm, which leverages starlet wavelet regularisation, as a preprocessing step before lens modelling to maximise the scientific output of the large HSC-SSP sample. We test this approach on seven HSC-like mock lensing systems with different lensing configurations. Our results show that the resolution boost from $\sim$0.6" to $\sim$0.15" enables improved accuracy and sub-3% precision measurements of key lensing parameters such as the Einstein radius and the mass axis ratio for ground-based images. These results pave the way for accurate measurements of key strong lensing quantities from a large sample of ground-based observations.

astro-ph.GA↗

SPRINT: Single-Step Generative Recommendation via Average Probability Velocity

Semantic ID (SID) based generative recommendation represents each item as a sequence of discrete tokens, and recommends by generating the SID of the item a user would like to interact with. Both dominant paradigms in this domain generally pay for generation token by token: autoregressive models decode the tokens left-to-right, while non-autoregressive models decode in parallel yet still need multiple rounds of refinement to stay competitive. Therefore, both generally spend multiple forward passes per item, a cost that is prohibitive in latency-sensitive recommender systems. We ask whether an item can be generated in a single forward pass, and answer it through a new perspective which we call average probability velocity. We view SID generation as a flow of token generation probabilities and characterize it by its average velocity over the whole generation process. We prove that this average velocity is fully determined by the average generation probability of each token. Therefore, we directly parameterize and learn the probabilities of all tokens in a single forward pass with a bidirectional Transformer. As these probabilities are generated independently across positions and the coherence among tokens is lost, we further design a dual-level flow contrastive objective to restore the coherence among an item's tokens. It contrasts the target SID against negative SIDs at both the token and SID levels. The token level ranks the generation probabilities of the target tokens above those of negative SIDs, while the SID level scores the tokens of each SID as a whole item for capturing token coherence of each item. Extensive experiments show that our model not only generates recommendations far more efficiently ($8.39-10.04\times$ speedup over the second-fastest AR/NAR method) but also attains superior recommendation accuracy ($7.77\%$ average improvement over the second-best.

cs.IR↗

On Representations of $\mathrm{GL}_n(\mathrm{D})$ admitting a generalized linear period

Let $\mathrm{D}$ be a quaternion division algebra over a non-Archimedean local field $\mathrm{F}$ of characteristic zero, and let $\mathrm{G}_n=\mathrm{GL}_n(\mathrm{D})$. Let $\mathrm{H}_{1,n-1}=\{\mathrm{diag}(g_1,g_2)\in\mathrm{G}_1,\ g_2\in\mathrm{G}_{n-1}\}$ and, for $s \in \mathbb{R}$, define $χ_s(\mathrm{diag}(g_1,g_2))=ν(g_1)^{2s}ν(g_2)^{-2s}$. We classify, for $n=3,4$, the irreducible smooth representations of $\mathrm{G}_n$ admitting a generalized linear period with respect to $(\mathrm{H}_{1,n-1},χ_s)$. Motivated by these results, we conjecture a complete classification for all $n>2$. Assuming this conjecture, we characterize such representations in terms of Langlands parameters: an irreducible smooth representation $π$ of $\mathrm{G}_n$ admits such a period if and only if $\mathfrak{L}(π)$ contains a Weil-Deligne subrepresentation isomorphic to $\mathfrak{L}(ν^{-2s})$, the $(2n-4)$-dimensional parameter of $ν^{-2s}$ on $\mathrm{G}_{n-2}$, and the four-dimensional quotient is the Langlands parameter of either the trivial representation of $\mathrm{G}_2$, with $s=\pm\frac{n-2}{2}$, or an irreducible infinite-dimensional $\mathrm{H}_{1,1}$-distinguished representation of $\mathrm{G}_2$. We also verify the Lapid-Prasad conjecture in this setting: the $L$-packet of an irreducible representation of $\mathrm{G}_n$ admitting a linear period with respect to $\mathrm{H}_{1,n-1}$ is invariant under $ρ\mapsto\widetildeρ^θ$.

math.RT↗

$\mathcal F$-Transitivity of Backward Shifts on Lattice Graphs

We study $\mathcal F$-transitivity of backward shifts on finite- and infinite-width directed lattice graphs. For the unilateral and bilateral finite-width lattices on weighted $\ell^p$- or $c_0$-spaces, we obtain exact weight characterizations of $\mathcal F$- and $\widetilde{\mathcal F}$-transitivity for an arbitrary Furstenberg family $\mathcal F$. If $\mathcal F$ is finitely invariant, these conditions also characterize topological $\mathcal F$-recurrence. For the infinite quadrant lattice with radial weights on weighted $\ell^2$-spaces, a finite-dimensional reduction and sharp smallest-singular-value estimates for Pascal-type transfer matrices yield equivalent characterizations of $\mathcal F$-transitivity, hypercyclicity and mixing. This provides a partial answer, in the radial $\ell^2$-setting, to the open problem posed by Baranov, Lishanskii and Papathanasiou concerning hypercyclicity on the infinite lattice graph. Examples at the critical exponential growth rate illustrate the finer dynamical information provided by our criteria.

math.FA↗

MaLiang-Harness: A Programmable Path to Image and Video Generation

Executable programs offer explicit control over how images and videos are constructed, but generating runnable code is only the beginning of visual creation. A program can execute correctly while violating the requested composition, appearance, or motion. We define this discrepancy as the Program-to-Visual (P2V) gap and introduce MaLiang-Harness, a unified framework for organizing MLLM-driven visual generation into a persistent process of construction, inspection, and revision. Its central design is to make the evolving visual program, its construction history, and its verification share a common revision reference. We define the Persistent Executable Generation (PEG) state as preserving programs and task context. Traceable Generation Process (TGP) connects edits to rendered evidence, and Revision-aware Editing and Verification (REV) supports restoration and checks the current revision before completion. Together, these mechanisms coordinate planning, execution, and visual feedback across rendering backends. We evaluate 11 powerful closed-source MLLMs on MaLiang-IBench and four on MaLiang-VBench, measuring generation success, visual quality, and computational cost. GPT-6-Astra achieves 100% generation success on both benchmarks, with 96.0% of image tasks and 76.9% of video tasks meeting all quality thresholds. The comparison also reveals a mismatch between general capability scores and visual generation performance, with similarly scored models differing substantially in their ability to satisfy visual requirements. MaLiang-Harness provides a systematic basis for studying how MLLMs translate executable code into visual outcomes, exposing both the potential of programmable generation and the limitations of general benchmarks as predictors of this ability. The project is available at https://github.com/gulucaptain/MaLiang-Harness.

cs.CV↗

Riemannian Difference-of-Convex Optimization for K-Means Clustering

K-means is a widely adopted clustering approach in signal processing and machine learning. In this paper, we study K-means clustering through a cardinality-constrained formulation on a compact embedded submanifold. We replace the cardinality constraint with a difference-of-convex (DC) penalty and establish a global error bound to prove that the penalized and constrained formulations share the same global minimizers whenever the penalty parameter exceeds a finite threshold. To solve the resulting nonsmooth Riemannian DC problem, we reformulate it as a minimax problem and propose RADA-DC, a Riemannian alternating descent ascent method combining dual regularization with DC linearization. Under standard assumptions and suitable parameter choices, RADA-DC finds an $ε$-Riemannian critical point within $O(ε^{-3})$ iterations. We conduct experiments on synthetic and real-world datasets to demonstrate that the proposed method outperforms the tested baselines, including K-means++, in solution quality at competitive computational cost when the number of clusters is large.

cs.LG↗

Weak Endpoint Estimate for Commutators of Rough Maximal Singular Integral operators

Let $d\geq 2$ and $T^*_{Ω,b}$ be the commutator of the rough maximal singular integral on $\mathbb{R}^d$ defined by $$T^*_{Ω,b}f(x)=\sup_{\varepsilon>0}\left|\int_{|x-y|>\varepsilon}\bigl(b(x)-b(y)\bigr)\frac{Ω(x-y)}{|x-y|^d}f(y)\,dy\right|.$$ Assume that $Ω\in L(\log L)^2(\mathbb{S}^{d-1})$ has mean zero and that $b\in\operatorname{BMO}(\mathbb{R}^d)$. We prove that, for every $λ>0$, \[\bigl|\{x\in\mathbb{R}^d:T^*_{Ω,b}f(x)>λ\}\bigr|\lesssim_{Ω,b}\int_{\mathbb{R}^d}\frac{|f(x)|}λ\log\!\left(e+\frac{|f(x)|}λ\right)\,dx.\] The argument starts from a Calderón-Zygmund decomposition of $f$ and a dyadic linearization of the maximal truncation. A decomposition of $Ω$ by size, followed by regularization of the spatial cutoffs and a microlocal decomposition, reduces the problem to a family of estimates with quantitative decay. The required decay follows by interpolating localized $L^1$ and $L^3$ bounds and using an analytic-family argument for the commutators.

math.CA↗

Gauss--Manin connections extend to holonomic coadmissible D-cap-modules

We prove that the pushforwards of Gauss--Manin connections on smooth rigid analytic varieties along Zariski-open immersions are coadmissible and holonomic over Ardakov--Wadsley's sheaf D-cap of infinite order differential operators. This may be viewed as a rigid analytic variant of the classical theorem of Griffiths, Deligne, and Katz that Gauss-Manin systems have regular singularities after compactification. The proof combines p-adic de Rham comparison and Diao--Lan--Liu--Zhu's logarithmic Riemann--Hilbert correspondence with extendability results for log-connections with nilpotent residues. In the algebraic snc case, we also establish weak holonomicity via a p-adic Bernstein-Sato criterion of Bitoun and the first author.

math.AG↗

ControlScope: Workflow Revision and Reliability in LLM Agents

How much of a running workflow should a language model agent revise? ControlScope compares continuing generated code, editing the next tool call's data arguments, and replacing the unfinished workflow from the same public execution state. The nested permissions separate available repairs from the actions an agent selects. We evaluate one-time and repeated reviews across filesystem tasks, ALFWorld, and AppWorld. Across two source programs per task and three reasoning-reviewer draws on 20 filesystem tasks, FULL completes 15-16 tasks versus 13 for KEEP; across four fast draws it completes 10-13 versus 13. Fresh student-record confirmation reproduces a batch-read repair. ALFWorld fast panels yield KEEP/ARG/FULL scores of 85/86/87 on 87 tasks across 52 scenes and 134/134/127 on 134 tasks across four scenes; reasoning on the 87-task cohort also yields 85/86/87 with substantial review cost. An AppWorld V1 official-test panel of 585 task instances from 195 scenario templates shows small net differences. Frozen replays expose viable agent-written replacements interrupted by later revision in two failed file-organization runs. An offline source-trajectory midpoint comparison shows later reviews completing an insufficient repair. Five-call protection saves 19.4% of logged model output and loses one success across 20 fresh source runs. An argument-only shortcut shows that the broader sampled policy can overlook a cheaper successful edit available in both operation sets. These outcomes tie repair access to actual choices and subsequent execution.

cs.AI↗

PlaylistEval: Can Video-Language Judges Be Trusted at Day Scale and Beyond?

Video-language models are increasingly used as judges of video understanding, both for evaluating model outputs and for training reward models. Whether their judgments remain reliable when the evidence is buried in day-long videos has yet to be established. Existing benchmarks cannot answer this. Their videos are typically only a few minutes long, many answer pairs can be separated from the transcript alone, and collecting human judgments does not scale to ultra-long videos. We introduce PlaylistEval, an agentic framework that builds video-language judge benchmarks over 100-hour playlist collection without human annotation. It automatically generates questions with paired answers whose differences are controlled by causal degradation, so that every pair demands retrieval across the collection. The resulting benchmark contains 630 pairs across seven domains spanning both static and dynamic knowledge, and on a stratified subset of 152 pairs it agrees with human judgments 93.0% of the time (IAA 0.781). Evaluating 17 omnimodal and multimodal models from eight families reveals that frontier judges reach only 75.4% pairwise accuracy, while open-source judge models perform far behind. We further show that both retrieval and final judgment depend on using multiple modalities, and that judge accuracy degrades as the playlist set grows. We release our pipeline, benchmark, and evaluation code at https://playlisteval.github.io.

cs.CV↗

Augmented James--Stein estimation for leading eigenvectors and eigenspaces in high dimensions

Building on the James--Stein approach to leading eigenvector estimation (Goldberg and Kercheval, Proc. Natl. Acad. Sci. USA 120, e2207046120, 2023), we develop a data-adaptive augmented James--Stein shrinkage framework for estimating leading eigenvectors and eigenspaces under a generalized spiked population model, in the high-dimensional regime where the dimension $p$ and sample size $n$ grow proportionally. For each spiked eigenvector, we construct an augmented target subspace that combines auxiliary information, either from domain knowledge or prior information, with the remaining sample spiked eigenvectors. This augmentation allows information shared across the sample spiked components to be exploited while retaining a fully data-driven shrinkage rule. We show that the resulting eigenvector estimator strictly improves upon standard PCA whenever the target subspace contains nonvanishing information about the population eigenvector, while asymptotically reverting to PCA when the target is uninformative. The individual estimators further yield a nested sequence of estimators for all leading spiked eigenspaces, with analogous dominance properties. The proposed estimator also strictly improves upon the existing HDLSS-motivated shrinkage estimator in the proportional high-dimensional regime. Simulation studies demonstrate substantial finite-sample gains and robustness to misspecification of the number of spikes.

math.ST↗

Dynamical Parameters: An Interpretability Framework for Time-Series Foundation Models

This work studies a central gap in interpreting time-series foundation models (TSFMs): a dynamical property may be accessible in a hidden state even when the forecast fails to respond correctly as that property changes. We formalize these properties as Dynamical Parameters, including trend slope, oscillation frequency, and autoregressive dependence. We compare their representation accessibility, measured by recovery from hidden states, with their forecast response, measured by agreement with the expected forecast change. Across nine frozen TSFMs and thirteen laws, 42 of 63 model-parameter cells achieve accessibility above 0.95, whereas their median reference-aligned response relative to the conditional reference is only 0.46. To explain this gap, causal geometry compares the hidden-state change required to produce the reference response with the change induced by the parameter intervention. Directly modifying the hidden state recovers the reference response, but the parameter intervention often moves the state in a different direction. These results show that accessible parameter information need not be expressed in forecasts when input changes miss the required hidden-state direction.

cs.AI↗