arXiv Science⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 1,639 records · Page 91Linked to original sources

Broadly Applicable Approximate MCMC for Switching Stochastic Differential Equations Using Uniformization and Time-Conditioned Factorized Neural Likelihood Estimation

Switching stochastic differential equations (SSDEs) describe continuous-time dynamics whose parameters switch according to a latent regime process that follows a continuous-time Markov chain (CTMC). By allowing dynamics to change between regimes, SSDEs represent heterogeneous system behavior and have been applied across diverse fields. However, Bayesian inference for SSDEs remains difficult, and existing SSDE inference methods have limited applicability, with restrictions such as noise-free observations, univariate states, linear drift, or state-independent diffusion. In this study, we propose an approximate Markov chain Monte Carlo sampler for SSDEs using uniformization and factorized neural likelihood estimation (FNLE), a simulation-based inference method. Uniformization provides an exact representation of the CTMC but requires SDE transition densities over arbitrary time intervals. We approximate these densities by training a time-conditioned FNLE model. The resulting sampler is broadly applicable to SSDEs without requiring analytically tractable transition densities. In synthetic-data experiments, our method recovered regime paths and parameters for three SSDE models for which previous methods have limited applicability. We also applied our method to a real dataset and detected a regime transition.

stat.ML↗

Effective Neutrino Mass and Apparent Phantom Crossing

Recent cosmological constraints on the sum of neutrino masses have reached, and in some analyses fallen below, the minimum value implied by the normal mass ordering, suggesting a possible tension with neutrino-oscillation measurements. We investigate whether this discrepancy can be alleviated in a mass-varying-neutrino (MaVaN) cosmology in which a canonical pseudo-Nambu-Goldstone boson (pNGB) drives late-time cosmic acceleration and controls the masses of massive neutrinos through an exponential coupling. Using current cosmic microwave background, baryon acoustic oscillation, and type Ia supernova data, we reconstruct the neutrino-mass evolution and show that the cosmologically inferred effective neutrino mass can fall below the present-day normal-ordering floor while remaining consistent with oscillation constraints. The same interaction produces a distinctive late-time signature, with 96% of the posterior expansion histories exhibiting a crossing of the apparent phantom divide, $w_{\rm app}=-1$, at $z_{\rm cross}=0.90^{+0.28}_{-0.13}$ at 68% C.L. The underlying scalar field nevertheless remains canonical, with $w_ϕ\geq -1$. Our results therefore show that neutrino mass variation can simultaneously alleviate the apparent neutrino-mass tension and generate the phantom-crossing behavior favored by current cosmological data, without introducing a phantom field or violating the null energy condition.

astro-ph.CO↗

HuLiGen: Human LiDAR Generation from Parametric Body Models

LiDAR point clouds of humans are extremely expensive to collect and annotate, thus represent a scarce resource that hinders the development of human analysis using this modality. To alleviate this scarcity, prior work relies on simulated human LiDAR, but such samples do not fully reflect the geometry and sensing characteristics of real observations. In contrast, we introduce HuLiGen, a generative model that generates human LiDAR point clouds from a parametric body model, using a point transformer trained with a flow-matching objective. We show that our generated point clouds are closer to the real capture distribution. Using HuLiGen to generate synthetic data, we propose a synthetic-only pretraining scheme for LiDAR-based HPE that achieves state-of-the-art performance, with even larger gains in low-annotation and low-data regimes, where MPJPE is reduced by up to 50%. Code, models and generated samples are available at https://github.com/valeoai/HuLiGen.

cs.CV↗

VolCo: Volumetric Contact for High-Fidelity Human Grasp Generation

Accurate contact modeling is fundamental to understanding hand-object interaction, yet existing contact representations are typically restricted to object surfaces and rely on hand-crafted rules to recover contact details, leading to severe penetrations and implausible results. To better exploit the rich detail in motion-capture data, we introduce Volumetric Contact (VolCo), a representation that expands surface points to a set of 3D volumetric grids. VolCo encodes 3D contact that allows precise hand part recovery, and is organized in an inherent hierarchy: local contact details within each volume and global hand geometry across all volumes. Our framework, VolCoDiff, employs two modules to capture local and global features following this hierarchy. For local contact details, we use a 3D variational autoencoder to model the possible hand configurations conditioned on the local object signed distance field (SDF). For global hand geometry, we design a prior-guided diffusion model that learns the distribution of compressed latent features aggregated from the volumetric grids. We evaluate our method on two benchmark datasets and demonstrate state-of-the-art performance in penetration and stability, indicating the capability to generate tight grasps with much less severe penetrations. Our code is available at https://github.com/chzh9311/volco.

cs.CV↗

Benchmarking Behavioral Steerability in Behavior Foundation Models

Behavior Foundation Models (BFMs) are emerging as a paradigm for translating human intentions into executable humanoid behaviors. As these models evolve beyond behavior generation toward general-purpose behavioral systems, a fundamental question arises: can they be reliably steered according to user intentions? In this paper, we introduce the concept of behavioral steerability, defined as the ability of BFMs to faithfully generate behaviors that satisfy user-specified intentions. To study this capability, we present RoboSteer, the first benchmark for behavioral steerability in BFMs. RoboSteer organizes behavioral steerability into a three-level hierarchy-Conditional Steering, Constraint Steering, and Compositional Steering-and establishes a unified evaluation framework supported by a large-scale multimodal motion corpus. Using RoboSteer, we conduct the first large-scale empirical study of behavioral steerability across 9 existing BFMs. We view behavioral steerability as more than a capability for controlling motion: it concerns how embodied systems translate human intentions into purposeful actions. We hope RoboSteer will advance research on intention realization as a foundation for general-purpose embodied intelligence.

cs.RO↗

Robust Decentralized Fairness Auditing

Emerging legislation requires large language models (LLMs) to be audited for compliance with regulatory standards, particularly fairness. Such black-box audits typically assume a single auditor with access to a large, representative set of queries. In practice, it can be difficult for an auditor to obtain such a query set, but multiple auditors can together cover the relevant demographic groups by auditing the LLM collaboratively with their individual query sets. However, relying on multiple auditors raises a fundamental trust problem, as they may act on behalf of the LLM provider to portray a misleading appearance of fairness, i.e., fairwashing. We propose Auditopus, a novel approach for robust decentralized fairness auditing. In Auditopus, auditing proceeds in rounds without a central server. In each round, every auditor issues a fixed number of queries to the LLM, and sends only cumulative statistics vectors of its query results to other auditors instead of sensitive queries in clear. The fairness of the audited LLM is then estimated by aggregating all the vectors. We show theoretically and empirically that even a single adversarial auditor in the network can steer this estimate by fabricating the vectors it sends, making an unfair LLM appear fair. To address this threat, Auditopus has each honest auditor locally down-weight any auditor whose cumulative statistics vectors are statistically inconsistent with previous ones. We implement Auditopus and compare it to robust aggregation baselines on two datasets with two pre-trained LLMs. Against an attacker that optimizes the vectors it sends to make the LLM appear fair, Auditopus reduces audit error by up to 78% on average relative to no defense and at least 62% relative to the robust aggregation baselines. Even when 49% of the auditors are adversarial, Auditopus never lets a very unfair or moderately unfair LLM pass as fair.

cs.LG↗

On an ideal membership problem

Suppose $f_1, \dots, f_{n+1}$ are elements of a regular ring $R$, where $n := \dim R$. The first author proved that $f_1^n \cdots f_{n+1}^n \in (f_1^{n+1}, \dots, f_{n+1}^{n+1})R$; this uses the Brian\c con-Skoda theorem --- no simpler proof is known, as far as we are aware. We show here that this result is optimal in many respects. On the other hand, when the $f_i$ are homogeneous general polynomials of a fixed degree in $R := \mathbb{F}[x_1,\dots,x_n]$, for $\mathbb{F}$ a field of characteristic zero, we prove that $f_1 \cdots f_{n+1} \in (f_1^2, \dots, f_{n+1}^2)R$.

math.AC↗

GAGR-Lab: Evaluating Joint Spatial-Geometric and Analytic Function Reasoning

Joint spatial-geometric and analytic function reasoning requires translating a perceived spatial configuration into a symbolic function whose executed curve satisfies geometric constraints. We present GAGR-Lab, a framework for measuring this capability through Cartesian game scenes, explicit function semantics, and authoritative Rust trajectory execution. It distinguishes spatial perception, metric grounding, geometric relations, function interpretation, function construction, and constrained synthesis. We specify four configurable scene-difficulty presets and a prospective 24-cell diagnostic design, while reporting only the subset actually evaluated. A bounded pilot of one hosted model (Llama 3.2 11B Vision Instruct) using two API credentials as execution replicas yields 72 balanced games with 432 attempts, 429 valid provider responses, and no target hits; exploratory ordinary-function prompt variants also fail to hit, while the structured localization interface yields no scoreable outputs. A privileged analytic search control independently succeeds on 600 directional cases from 300 generated scenes, with exact repeatability and 1,200 successful vertical-reflection or translation checks. The framework separates serving reliability, symbolic compliance, and geometric success, and preserves exact model-visible inputs and realized paths. A staged protocol outlines diagnostic calibration, held-out replication, multi-model comparison, and paired robustness tests. The contribution is an operational research framework with an executed pilot and a clearly identified prospective study plan; the full difficulty matrix and comparative model results remain untested.

cs.AI↗

Chiral Extensions of Cotangent Yangian

We construct chiral extensions of the cotangent Yangian using vertex algebroids. At zero level, the construction yields a vertex algebra with a compatible coproduct, whose Zhu reduction recovers the cotangent Yangian of Abedin and Niu. We also construct vertex algebra extensions at non-zero levels and study level-additive coproducts. The zero-level vertex algebra admits two natural quasi-conformal structures. We compute the corresponding sphere one-point conformal blocks and equip each block space with an associative product induced by the coproduct. For the first quasi-conformal structure, the blocks space can be identified with the dual of a certain Lie algebroid homology; for the second, we describe the block algebra and its classical limit in terms of mini-twistor geometry.

math.QA↗

How to train your model organism

Model organisms of alignment-relevant behaviors (e.g., backdoors, sycophancy, spurious correlations) have emerged as a key tool for evaluating whitebox interpretability techniques. We argue that the prevailing practice of training model organisms to a single objective of installing the target behavior is insufficient and propose validating model organisms with respect to three objectives with associated metrics: target-behavior installation, general-capability preservation (i.e., parametric knowledge, chat quality), and output naturalness (i.e., CoT and activations). We re-visit two publicly released organism suites using this validation framework and show that (1) chat quality and CoT naturalness degrade substantially across training recipes, and (2) validation metrics predict how well interpretability methods recover the installed behavior, e.g., a logit lens readout covaries with an organism's general capabilities. We introduce a multi-objective training approach based on model merging to train more realistic model organisms. Finally, on a new suite of model organisms targeting demographic biases in clinical reasoning, we compare training recipes and find that DPO training stays closer to the base model than supervised finetuning, and the proposed model optimization approach better preserves capabilities and naturalness. Auditing this suite with an investigator agent, we again observe validation metrics tracking bias recovery. In sum, training methods shape the interpretability conclusions an organism supports, and we argue that one should consider multiple objectives to draw generalizable conclusions about interpretability methods using (realistic) model organisms.

cs.LG↗

Three-dimensional Lagrangian ecosystems: carbon dynamics and potential for artificial fertilization

Transient supplies of nutrients to the surface ocean, both natural and artificial, stimulate blooms of phytoplankton and the formation of organic matter, driving air-sea gradients and uptake of CO$_2$. However, quantifying the associated carbon budget remains challenging, as it requires tracking the coupled biophysical evolution of water masses while they are transported, stretched and diluted. Here, we present an idealized three-dimensional model that describes biomass production and carbon dynamics within a Lagrangian patch in the ocean. The framework reproduces observed biogeochemical patterns from an artificial fertilization experiment and provides integrated metrics for the local carbon budget. Utilizing large ensembles of simulations, we examine the sensitivity of patch-scale primary production and carbon uptake to biochemical and physical factors. Our results show that patch dilution can enhance the ecosystem response, and how carbon uptake is sensitive to initial injected area, horizontal divergence and vertical diffusivity. In the context of renewed interest in ocean fertilization strategies for climate mitigation, our approach can thus provide quantitative tools to assess their efficacy and potential.

physics.ao-ph↗

Breaking the Odd Cycle: Regularization and First-Order Optimization for the Maximum $s$-Plex Problem

Building on a new regularized continuous formulation which avoids spurious solutions of the maximum $s$-plex problem, we develop a tailored block-coordinate first-order method that combines away-step Frank-Wolfe updates on the standard simplex with linear optimization over the fractional $b$-matching polytope. We establish finite stabilization of the auxiliary variables, convergence to stationary points, and finite support identification under suitable conditions. Numerical experiments on standard benchmark graphs show that, also in practice, regularization avoids the support-identification failures observed with the original cubic formulation and enables the proposed method to quickly find high-quality $s$-plexes.

math.OC↗

Computations of the slice genus and the unknotting number of links via machine learning

Links are disjoint unions of circles smoothly embedded in $S^3$. We use reinforcement learning and Bayesian optimisation to obtain new upper bounds on several link invariants that are not known to be algorithmically computable: the slice genus and the unknotting number for links, and the strong slice genus for algebraically split links. We also compute lower bounds using known invariants. Combining the upper and lower bounds, we obtain new exact values in many cases. Our unknotting agents can reproduce the non-additivity of the unknotting number for several counterexamples due to Brittenham and Hermiller, in some cases finding new unknotting trajectories.

math.GT↗

Identifiable Preference Types and Optimal Menu-Choice Experiments

A planner wants to elicit partial information about a subject's preference relation. Preferences are grouped into ``types" relevant to the planner's purpose. Can the planner identify the subject's type through menu-choice data or experiment, without distinguishing preferences within the same type? We provide a complete answer by characterizing identifiable type spaces. The characterization uses a novel geometric condition, called face consistency, which governs distinctions within and across square and hexagonal faces of the permutohedron. For every identifiable type space, we explicitly construct the unique irreducible identifying experiment from its boundary. This experiment is optimal among all identifying experiments under any cost criterion for which adding menus cannot reduce cost. Using the equivalence between identifiability and elicitation with acts, as shown in Azrieli et al. (2021), we study symmetric policies requesting unranked shortlists that contain only acceptable products, include every product preferred to a selected one, and are nonempty whenever an acceptable product exists. Only two such policies are elicitable with acts: requesting the favorite acceptable product or the entire acceptable set.

econ.TH↗

CARES: A Controlled Synthetic Benchmark of Speaker Reactions to Sound

Automatic audio scene description turns a recording into a text account of a situation. One difficulty is deciding which elements of the audio should be kept, since a description cannot include them all. Annotators disagree about this, making a ground truth hard to obtain. In this work, we first define the ground truth, then generate the data. We focus on audio events and define sound salience with a simple rule: a sound is salient when a speaker audibly reacts to it. For scale and variety, a controlled set of scenarios fixes the ground truth, and a language model writes the dialogues. The resulting corpus, CARES, contains 10,000 two-speaker scenes. We then benchmark six audio-language models on three tasks: identifying the scene, tagging the sounds present, and classifying reactions. We show that these models hear the sounds but miss how the speakers react to them.

cs.SD↗

Points and their multiples on curves in powers of simple abelian varieties

Let $G$ be a simple abelian variety of dimension $g \in \mathbb{N}$ defined over $\mathbb{Q}^\mathrm{alg}$ and let $C_1, C_2 \subseteq G^N(\mathbb{C})$ be irreducible closed algebraic curves with $N \geq 3$. Further assume that at least one of $C_1$ and $C_2$ is not defined over $\mathbb{Q}^\mathrm{alg}$. Suppose that there does not exist an algebraic subgroup $G \subseteq G^N(\mathbb{C})$ of dimension $g$ such that $C_1 \subseteq G$ and that there does not exist an algebraic subgroup $H \subseteq G^N(\mathbb{C})$ of dimension $2g$ such that $C_1 \cup C_2 \subseteq H$. Denoting $\mathcal{N} = \{n \in \mathbb{N} \ | \ [n]C_1 \subseteq C_2\}$, we prove that $\bigcup_{n \in \mathbb{N} \setminus \mathcal{N}}\{x \in C_1 \ | \ x^n \in C_2\}$ is finite.

math.LO↗

OrthoGen: A Generative Orthogonal Learner for Time-Varying Treatments

Estimating conditional distributional potential outcomes (CDPOs) over time is important in medicine (e.g., to estimate patient-specific risks under different treatment sequences). However, this task is challenging because of time-varying confounding, yet existing adjustment strategies for this task are limited. In this paper, we aim to learn CDPOs under time-varying treatments using flexible generative models. Our contributions are two-fold. (1) We introduce a tailored adjustment strategy for our setting, namely, generative recursive g-computation. Our adjustment strategy recursively propagates full conditional outcome distributions rather than conditional means, modeling the variables of interest directly rather than full trajectories. Building on our adjustment strategy, we formulate simple generative learners for CDPO estimation. However, these learners can be sensitive to nuisance estimation errors, which motivates an orthogonal learner. (2) We thus introduce OrthoGen, a Neyman-orthogonal and doubly robust generative learner. Importantly, we show that OrthoGen further achieves rate double robustness and quasi-oracle efficiency under suitable conditions. Our learners are flexible and can be instantiated with different generative backbones (e.g., normalizing flows and diffusion models). Across experiments with synthetic, semi-synthetic and real-world datasets, we find that OrthoGen is highly effective. To the best of our knowledge, we are the first to propose a generative orthogonal learner for estimating CDPOs under time-varying treatments.

cs.LG↗

Sequential Random Sampling PIR with Multiple Colluding Servers in DNA-Based Data Storage

As DNA-based data storage evolves, protecting user privacy during data retrieval has become increasingly important. We study sequential random sampling DNA private information retrieval (SRS DNA PIR) with multiple colluding random sampling servers, where the database is partitioned into servers of equal size. We investigate the tradeoff between the download cost, defined as the expected number of queries, and the privacy leakage, measured by mutual information. We derive lower bounds on this tradeoff, including a bound given by an optimization problem. This bound is tight when each server stores two files, and we construct schemes that attain it. For servers of any size, we construct schemes that apply a single-server scheme to a randomly selected subset of servers.

cs.IT↗