arXiv Science⌕ Search

SEARCH · arXiv Science

Search arXiv Science

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,423 records · Page 79Linked to original sources

Catastrophic Forgetting in Sequential Thermal Anti-UAV Detection: The Role of Scale-Conditioned Gradient Imbalance

Counter-UAV systems based on thermal infrared detection must stay accurate as operational datasets evolve, yet sequential fine-tuning causes catastrophic forgetting of prior tasks, a problem that remains insufficiently characterized in this domain. This continual-learning study measures the stability-plasticity trade-off in YOLOMG, a YOLOv5-based detector run as a single thermal-infrared stream with the motion channel disabled, trained sequentially across three anti-UAV benchmarks of rising scale difficulty: Anti-UAV-RGBT, Anti-UAV410, and CST Anti-UAV. Naive fine-tuning on CST yields a Forgetting Measure of -0.605 against the Stage 1 ceiling, corresponding to a 90% capability loss, with -0.572 occurring in Stage 3 alone. In contrast, knowledge distillation from a frozen teacher is associated with FM = -0.033 +/- 0.004 across three seeds, corresponding to 95% retention. Because no Stage 2 no-KD control is included, this result establishes retention under KD training rather than a causal KD effect. Per-stratum analysis shows large-target detection collapsing to near zero within the first epoch, despite an inter-stage cosine similarity of 0.987 over the gradient-updated weights, pointing to scale-conditioned gradient imbalance, rather than weight drift, as a candidate mechanism. Scale-Stratified Herding (SSH), a 300-exemplar buffer balanced across four UAV size strata, roughly halves the forgetting (FM = -0.605 to -0.311) and keeps large-target detection non-zero. An ablation attributes the gain primarily to scale stratification rather than herding: random-stratified replay performs at least as well (FM = -0.221 versus -0.311 for SSH). These replay results are single-seed and should therefore be treated as preliminary.

cs.CV↗

MARCO: The Radioactive Watermark for Protein Generative Models

Protein Generative Models (PGMs) have revolutionized structural biology by enabling the design of complex 3D protein structures from sequence data. However, this breakthrough introduces a dual-use challenge, exposing high-value PGMs to economic risks like unauthorized model extraction and biosecurity threats such as biohazard synthesis. To mitigate these threats, we propose \textbf{MARCO} (\textsc{COnformation waterMARk}), the first radioactive watermarking framework specifically tailored for PGMs. MARCO establishes a Dual-Layer defense that simultaneously protects intellectual property and ensures the forensic traceability of potential biosecurity misuses. (i) To preserve efficiency, MARCO iteratively embeds watermarks during diffusion reverse denoising via an auxiliary encoder-decoder, allowing the original PGM parameters to remain frozen for broad compatibility. (ii) To preserve biophysical fidelity and maximize robustness, we employ specialized loss functions targeting $C_α$-atom pairwise distances and torsion angles ($ψ, ϕ$) within an adversarial training framework integrated with stochastic attack simulations. (iii) Crucially, MARCO exhibits ``radioactivity'' where the watermark automatically transfers to the outputs of any pirate models trained on the watermarked data, effectively countering model extraction attacks. Comprehensive experiments demonstrate that MARCO achieves superior fidelity and robustness while successfully validating watermark transferability.

cs.CR↗

Bayesian heterogeneous copula mixtures with nonparametric margins: consistency, identifiability and tail asymmetry in physical fitness data

Finite mixtures of Clayton, Gumbel, Frank and Gaussian copulas can describe dependence that differs between the upper and lower tails. We study a two-stage Bayesian analysis in which ranks or kernel estimates replace the margins and the copula likelihood is evaluated at the resulting pseudo-observations. This pseudo-posterior is strongly consistent for the copula density whenever the log density admits a logarithmic boundary envelope. Both stages may use the same data, margins may be standardized within observed strata, and neither smoothness nor identifiability is required. The envelope holds for finite mixtures of Gaussian, Student, Clayton, Gumbel and Frank copulas in any fixed dimension; tail-dependence coefficients and conditional tail probabilities are therefore consistently estimated. We further prove that Clayton, Gumbel, Frank and Gaussian copulas are jointly finitely linearly independent, which makes mixture weights and components identifiable and consistently estimated. In simulations the pseudo-posterior matches multi-start maximum pseudo-likelihood in large samples. It is more stable in small samples and near independence (every component close to the independence copula), where tail coefficients are learned long before the weights and a marginal Metropolis sampler is up to twice as efficient as data augmentation. Two physical fitness datasets show mirror-image asymmetries. Among 8772 university students, sprint and jump performance are coupled mainly at the top. Among 5336 adults in a national health survey, low grip strength and low daily activity cluster together, increasingly with age, whereas high values do not. A Gaussian copula misses both patterns in held-out data.

stat.ME↗

SCOPE: Certified Theorem Proving with a Language Model as the Policy Planner

In proof assistants such as Lean, a generated proof must pass machine compilation checks, so evaluation needs no human scoring. Direct generation fails on multi-step numeric propositions: a proof is valid only if every content integer is correct, so the pass rate is bounded by the k-th power of the per-integer accuracy. Controlled corruption across 2,617 reference proofs confirms this power law. SCOPE (State-Conditioned Operator Planning and Execution) enforces the natural division of labor: the model plans over an operator vocabulary, a symbolic engine executes the numerics, and a compiler renders the proof. On a 218-problem suite it certifies 191/218 (87.6%) with a 135M backbone; the 7B DeepSeek-Prover-V1.5-RL certifies 18/218 at 27.5 times the tokens and 37.5 times the wall-clock, and DeepSeek-Prover-V2-7B certifies zero on a bidirectional dual suite. Multi-step thinking costs 6.12 discrete decision actions per problem and produces no natural-language thinking text. Replacing the lagged engine state in the decision frame with the current one lifts the pass rate from 117/218 to 191/218, while up-weighting the chain-end loss hurts. On the public Lean-Workbook library, 2,132 of 3,536 gradeable admissible problems certify (60.29%) with zero regression on the main suite. All readings come from a version-frozen review with independent rechecks and reverse verification. Restricting free generation and keeping decision-time information visible is a more direct route than enlarging the model.

cs.AI↗

Humanoid Horizon: Extending Task Horizon in Whole-Body Loco-Manipulation via Parallel Training, Dynamic Starting, and Reward Gating

Cluttered indoor environments, where large and heavy objects are scattered across diverse surfaces, require humanoid robots to sequentially navigate, grasp, transport, and accurately place each item at its target location within a single uninterrupted episode. This long-horizon, whole-body loco-manipulation task remains a significant challenge for current methods. Previous approaches often suffer from two main issues: easy-reward bias, where training overemphasizes early transport stages at the expense of later ones, and catastrophic forgetting, where focusing on later stages leads to a decline in earlier-stage performance. In this work, we introduce Humanoid Horizon, a unified policy framework designed to overcome these limitations through three interrelated mechanisms. The Parallel Training Strategy organizes $N$ scenes into $S$ concurrent stage streams governed by a shared policy, ensuring all transport stages receive continuous gradient updates and removing the bottleneck of sequential optimization. The Dynamic Starting Mechanism updates each environment's initial state with terminal states from upstream rollouts, gradually broadening transition coverage and enhancing robustness at stage boundaries. Reward Gating sets the reward to zero for the rest of the episode in later-stage streams when the immediately preceding object is displaced beyond a set threshold, so the shared policy learns not to disturb a just-placed object and earlier placements are preserved throughout the episode. Collectively, these strategies achieve per-stage success rates exceeding 80\% on the two-object LHM-Humanoid benchmark (350 training scenes, 66 held-out scenes). As the number of sequentially transported objects grows beyond two, success declines with the horizon, but the degradation is graceful relative to the sharp drop seen in all baselines.

cs.RO↗

Covariate-dependent Nonparametric $g$-modeling for regression via infinite Mixture-of-Expertizing class

Empirical Bayes $g$-modeling captures unit-level heterogeneity by estimating a latent prior distribution from observed data. In the existing formulations, however, the prior is shared by all units. In this paper, we develop a covariate-dependent g-modeling framework for regression in which the entire prior distribution of the regression coefficients is allowed to depend on covariates. We formulate the estimation of the prior as nonparametric maximum likelihood estimation (NPMLE) of the covariate-dependent prior, and show that the unrestricted problem is ill-posed. To resolve this, we introduce the infinite Mixture-of-Expertizing class of conditional priors, under which the NPMLE is precisely a softmax-gated Mixture of Experts (MoE) whose number of experts is not fixed in advance but is determined by the data. Building on a first-order optimality condition, we propose two exemplar-based estimation algorithms that select experts automatically, together with a post-hoc aggregation of experts for interpretation. On the theoretical side, we show that every conditional prior in the class is Lipschitz continuous in the covariates, and that aggregated softmax gates can approximate any continuous gate function. The effectiveness of the proposed NPMLE is shown through application to synthetic datasets and real datasets.

stat.ME↗

Structure-Aware Graph Abstention for Reliable Selective Forecasting

Selective forecasting abstains on high-risk test windows under a retained-coverage budget. Existing gates such as TEM (Brusokas et al., 2025) score each forecast as a whole; for multivariate outputs, trajectories can look plausible while violating dependencies among variables. We treat instance-level plausibility and relational consistency as distinct reliability axes and operationalize the latter via a learned sparse graph and a Dirichlet-style structural energy E_struct, trained with error-weighted graph regularization and score-error alignment. On seven long-horizon benchmarks and four backbones, structural gating often reduces selective MSE versus TEM at matched coverage, with the largest gains where cross-variable structure appears more informative in our benchmarks; gains are not universal, indicating a complementary abstention signal. Table 1 is a Protocol A ranking diagnostic (seed 2024); three-seed deployable Protocol B on an aligned subset is in Table 3 (full validation-to-test grids: Appendix A).

cs.LG↗

Communication-Free Obstacle Localization from Aggregate Wrench Measurements in Leader--Follower Cooperative Transport

We consider obstacle localization for a team of robots cooperatively transporting a rigid payload without explicit inter-robot communication. A leader robot directs the payload's motion, while follower robots assist and react to locally detected obstacles. The leader measures the followers' aggregate wrench, i.e., the combined force and torque they exert on the payload, but cannot directly distinguish their individual reactions. We design a follower control law that allows the leader to recover obstacle locations from these measurements. Each follower resists motion toward nearby obstacles, resulting in a piecewise-linear relationship between the payload's translational and angular velocity and the aggregate wrench. Changes between adjacent linear regions reveal an obstacle's bearing and distance and identify the responding follower. We give sufficient conditions for exact recovery at a fixed payload configuration and develop an adaptive probing procedure in which the leader applies translational and rotational inputs to the payload to obtain the required measurements. We demonstrate the performance of the proposed method in simulations.

cs.RO↗

Synthesizing free-electron wave functions by stimulated near-field interactions

Stimulated electron-optical-near-field interactions imprint a coherent, phase-coherent modulation on the wave function of a free electron, of which the free-space dispersion subsequently converts into a train of attosecond density peaks. Light thereby becomes a tool for synthesizing the electron wave function itself, setting when, where, and how narrowly the electron density concentrates. Here we study photon-induced near-field electron microscopy (PINEM) to gain control over that synthesis in two stages: first for a single PINEM interaction, whose design parameters we obtain in closed form, and then for multiple PINEM interactions acting in parallel, which extends the control across space and time. The design rules follow from decomposing the propagated density into temporal harmonics, which we classify into three: (i) the arrival time of the attosecond train, fixed by the optical phase of the coupling; (ii) the distance at which the train forms, set by diffraction at the point where all classical trajectories converge; and (iii) the pulse duration, fixed by the imprinted energy spread. We evaluate these parameters for a scanning-electron-microscope (SEM) configuration at 10~keV, where the compression completes within tens of micrometers from the interaction region, with results 20% better than previous approaches. Letting the electron then interact with several spatially separated near fields in parallel, normal to the trajectory, each pathway carrying its own coupling strength and optical phase, we show that retaining or erasing which-path information changes both the propagated density and the measurable energy spectrum, the latter providing in turn a measure of the mutual coherence of the driving fields. These results establish PINEM as a quantitative synthesis tool for free-electron wave functions in compact, low-voltage electron microscopes.

physics.optics↗

Motif: A Modular Finite-Volume Framework for Transient Incompressible Flow

This document presents the theory, implementation, and verification of Motif, a two-dimensional incompressible flow solver developed as a teaching and research framework. The governing equations are discretized using the finite-volume method on staggered Cartesian grids. Convection is treated explicitly using the second-order Adams--Bashforth method, diffusion implicitly using the Crank--Nicolson method, and pressure--velocity coupling using projection methods. Both the Standard projection and the approximate second-order projection of Perot (1993) are considered. Particular attention is given to the discrete operators, boundary conditions, temporal treatment of pressure, conservation properties, and numerical dissipation and dispersion. The implementation is verified using a sequence of periodic, wall-bounded, and inlet--outlet flow problems on uniform and smoothly stretched grids. Spatial and temporal convergence, discrete mass conservation, kinetic-energy behavior, enstrophy, and energy spectra are examined. The results demonstrate the expected accuracy of the spatial discretization and distinguish the temporal behavior of the two projection approaches. The document is intended both to describe Motif in sufficient detail for independent implementation and to examine the numerical properties relevant to its continued development toward direct numerical simulation.

physics.flu-dyn↗

MedZERO: Self-Evolving Agents for Open-Ended Medical Reasoning Through Controlled Knowledge Accumulation

Large language models (LLMs) have shown promise in medical question answering and clinical reasoning, yet their improvement remains constrained by static parametric knowledge and costly expert supervision. Self-evolving agents offer a promising alternative by enabling models to improve through iterative task generation and problem-solving. However, most existing self-evolving methods are designed for easily verifiable domains such as mathematics and coding, where solutions can be checked by exact answers or executable programs. Medical reasoning is fundamentally different: it is open-ended, knowledge-intensive, and often only partially verifiable. We present MedZERO, a self-evolving framework for open-ended medical reasoning. MedZERO couples an Examiner that generates frontier medical question-option pairs with a Reasoner that solves them through evidence-grounded multi-turn reasoning with external knowledge tools. To support reliable, continual improvement, MedZERO adopts controlled knowledge accumulation, which maintains temporary exploratory knowledge and curated persistent knowledge in reasoning. We evaluate MedZERO on five public medical reasoning benchmarks using 4B- and 8B-scale base models under open-ended evaluation. Across all settings, MedZERO consistently outperforms the underlying base models and prior self-evolving baselines, achieving up to 13.7 average accuracy-point gains over the next-best self-evolving baseline.

cs.AI↗

Enhancing LLMs with Cognitive-Affective Personality Inference for Simulating Human Social-Psychological Behavior

Large language models are increasingly used to simulate human participants in social and behavioral studies, yet static persona prompting typically maps a participant profile and an experimental scenario directly to a response, entangling stable dispositions with situation-specific interpretations. To address this limitation, we introduce \textbf{SPIN}, a cognitive-affective personality system-inspired inference pipeline for simulating human social-psychological behavior. Specifically, SPIN implements this structured inference process through three zero-shot LLM calls that compile a task-blind participant core, elicit condition-specific cognitive-affective states, and read out decisions from those states, thereby reusing stable personality structure while routing each trial-specific response through an explicit state representation. We evaluate SPIN on two reconstructed social-psychological study families spanning uncertainty reasoning and pluralistic ignorance, across four base LLMs. Compared with blank, demographic, narrative, and chain-of-thought prompt variants, SPIN consistently delivers the strongest overall alignment performance across base LLMs and study families. Ablations and state analyses further show that both personality compilation and structured state elicitation contribute to the gains, and that the elicited states shift interpretably across informational and normative conditions. These results suggest that structured personality-state inference can improve benchmark-level behavioral alignment beyond richer persona descriptions or generic multi-step reasoning.

cs.SI↗

An AI-Assisted Formalization of the Poincaré Conjecture

We present an AI-assisted Lean 4 formalization of the Poincaré conjecture. The project began with limited reusable formal infrastructure for the geometric analysis behind the proof. To organize this work, we combined a proof blueprint prepared by mathematicians with explicit milestone statements. These milestones enabled parallel agent work and gave mathematicians clear points to locate blockers and provide effective mathematical guidance. Our analysis identifies the human interventions and organizational choices behind this workflow. The project provides a starting point toward reusable infrastructure for future formalization projects; such infrastructure, once developed, could eventually reduce the cost of verifying mathematical results in geometric analysis.

cs.AI↗

MoF: Preference-Aware Mixture Modeling for Black-Box LLM Personalization

Proprietary Large Language Models (LLMs) have demonstrated remarkable capabilities across a wide range of tasks, yet aligning their outputs with diverse user preferences remains challenging. Existing personalization approaches for black-box LLMs often rely on user-specific scoring heads, causing the number of personalized parameters to grow linearly with the number of users and requiring additional adaptation for unseen users. To address these limitations, we propose Mixture-of-Facets (MoF), a scalable personalization framework for black-box LLMs that models user preferences as compositions of shared latent preference facets rather than dedicated user-specific parameters. MoF performs personalization through history-conditioned routing over shared facet heads, enabling personalization for users unseen during training without additional parameter updates. Across diverse personalization tasks, MoF delivers stronger personalization performance while maintaining a more scalable and parameter-efficient design than prior approaches. Additional analysis indicates strong generalization to unseen users.

cs.AI↗

Transferable Spatial Temporal Coherence Adversarial Attack on Black-Box Vision Language Models for Autonomous Driving

The rapid integration of Vision Language Models (VLMs) into sensitive systems introduces critical safety vulnerabilities that remain unexplored in exist studies. While adversarial attack robustness has been extensively studied for image-based models, the susceptibility of VLMs to temporally-aware adversarial attacks against video in driving context poses a distinct and under examined threat. In this paper, we introduce novel adversarial attack against video targeting VLM models used for autonomous driving scenes named Spatial Temporal Coherence Adversarial Attack (STCA). Our attack comprise from three stages: modalities expansion, Spatial attack, and STCA attack. In modalities expansion, we propose caption-guided frame selection method in order to ensure that adversarial perturbation target the most semantically significant frames. Secondly.In spatial attack, we craft effective perturbation and preserve high similarity. Then the perturbed video generated fed into STCA stage that disrupt cross-frame temporal coherence using motion guided mask. Our method operate under black box threat model against victim target VLMs, relying solely on transferability from white-box surrogate model.We conduct our experiments on the BDD100K and nuScenes autonomous driving datasets across three VLM models: Video LLaVA-7B, Qwen2.5-VL-7B, and Dolphin. Experimental results demonstrate spatial attack achieves an ASR with high SSIM. Our finding reveal that existing video language model, remain highly susceptible to adversarial attack in autonomous driving scenarios, underscoring the urgent need for robust defense for VLM models.

cs.CV↗

SC3BF: Shifted Collision Cone Control Barrier Function for Dynamic Obstacle Avoidance

The collision cone used by velocity-space control barrier functions is conservative: it rejects every relative velocity aimed into an obstacle, however slow. We propose the \emph{shifted collision-cone CBF} (SC3BF), which adds a state-dependent \emph{allowance} to the cone condition, so the robot may approach the obstacle at a rate that grows with distance and with its own speed. SC3BF is enforced by an ordinary quadratic program, and its safe set is forward invariant under bounded inputs without a minimum forward speed or a clearance margin. We prove that a nonzero allowance preserving safety always exists, and derive one in closed form. Against three velocity-space baselines on a kinematic bicycle among up to $100$ moving obstacles, SC3BF reaches the goal more often and modifies the nominal input less than half as much.

cs.RO↗

Perfect state transfer on mixed graphs: complete classes and transfer times

For perfect state transfer (PST) on unweighted mixed graphs, we classify the normalized transfer times of complete PST classes. A finite set $Λ\subset\mathbb R/\mathbb Z$ containing zero occurs at a nonstationary periodic vertex if and only if $\cos(2π(x-y))\in\mathbb Q$ for all $x,y\inΛ$. Every admissible set has a connected oriented realization. We also classify the possible return phases of oriented realizations at the minimum vertex period. Transfers at rational multiples of the common minimum vertex period partition a complete class into sets of size at most six, or at most three in an oriented graph with return phase $-1$; both bounds are sharp. We construct complete classes of every finite size, including classes in which all transfers between distinct vertices occur at irrational multiples of the period and no switching automorphism maps a class vertex to a distinct class vertex. We also characterize simultaneous realization in connected oriented graphs with prescribed relative minimum vertex periods and return phases. The proof combines an imaginary quadratic field restriction with an unweighted construction that selects the complete target set.

math.CO↗

On constructibility of klt type varieties and simple $D$-module

In this paper, we prove that the combination of dense klt type and $\mathcal{D}$-simplicity property establishes the klt type property. The result paves the way to connect the theory of algebraic differential operators with the area of birational geometry.

math.AG↗