arXiv Science⌕ Search

SEARCH · arXiv Science

Search arXiv Science

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,279 records · Page 71Linked to original sources

Complete parameterization and parameter-space topology of discrete Wigner representations on $d \times d$ phase space

Toroidal discrete Wigner functions (DWFs) for finite-dimensional systems are non-unique. We classify the labeled single-qudit ${d\times d}$ representations whose phase-point operators are Hermitian, unit-trace, Hilbert-Schmidt orthogonal, and Weyl-Heisenberg covariant. Our previously established stencil theorem expresses this family as descents of a doubled ${2d\times2d}$ parent. Here, we solve its projected-stencil admissibility conditions explicitly. In the symplectic-Fourier representation, admissibility fixes the modulus, leaving phase data on a ${d\times d}$ base cell satisfying a parity-dependent twisted-oddness relation. This gives the parameter space ${(S^1)^{(d^2-1)/2}}$ for odd ${d}$ and ${(S^1)^{(d^2-4)/2}\times\mathbb{Z}_2^3}$ for even ${d}$. Canonical horizontal and vertical marginals reduce these spaces to ${(S^1)^{(d-1)^2/2}}$ and ${(S^1)^{d(d-2)/2}\times\mathbb{Z}_2}$, respectively. The same phase data parameterize validity-preserving rephasings of the doubled Weyl-Heisenberg displacement operators. Requiring these operators to have order dividing the Hilbert-space dimension ${d}$ replaces each ${S^1}$ factor by a discrete ${\mathbb{Z}_d}$ factor. The associated symplectic-Fourier characteristic functions encode the same operator information, with pointwise magnitudes independent of the valid stencil convention. This classification separates the freedom intrinsic to DWF validity from that selected by marginal and displacement-algebra requirements and makes explicit the structural distinction between odd and even dimensions.

quant-ph↗

Is Better Teacher Supervision Enough? Unlocking Student-side Learning in Multimodal On-Policy Distillation

On-policy distillation (OPD) improves reasoning by providing token-level supervision from a teacher on a student's own trajectories. Existing methods primarily focus on enhancing this teacher-side guidance (e.g., by enriching teacher inputs and refining teacher feedback), yet we find that limited student perception is another critical bottleneck in multimodal OPD. By providing oracle visual facts, the performance of OPD-trained students can still be substantially improved for both weak and strong teachers. To address this bottleneck, we propose S-OPD, a simple multimodal on-policy distillation framework that explicitly strengthens student perceptual learning through two objectives. Specifically, Teacher-calibrated Policy Contrast separates student policies under original and masked images with teacher-based token-level gating, strengthening the student's reliance on visual evidence during reasoning. Policy Agreement aligns student policies under original and noise-perturbed images, further improving perceptual robustness to visual noise. Notably, our method can be seamlessly plugged into existing OPD frameworks, requiring no additional data annotations, model parameters or inference operations. Extensive experiments on eight benchmarks across student scales and distillation paradigms demonstrate consistent performance improvements, with gains of up to 4.25 points on LogicVista. When combined with existing teacher-side supervision methods, our method can yield further gains. Code is available at https://github.com/Sirilaw/S-OPD.

cs.CV↗

Bivariate Bicycle Codes and Metachecks: Syndrome Repair, Measurement-Fault Ambiguity, and Logical Obstructions

Faulty syndrome measurements can corrupt an otherwise correct quantum-error correction step. Bivariate bicycle (BB) codes contain dependent stabilizer checks, so their measured syndromes obey parity constraints that can be used as metachecks. We study when this built-in redundancy actually supports syndrome repair and how it interacts with the logical structure of the code. Annihilator quotients identify single-block logical classes and show when a mixed-block search is unavoidable. For syndrome repair, the metasyndrome is equivalent to reduction modulo the valid-syndrome ideal. The translation group therefore maps into the unit group of a quotient algebra of dimension $k/2$, with kernel $K_M$. This gives a distance-two characterization, a unit-group bound on distinguishable single faults, and a family-level obstruction: for bounded-$k$ BB families with growing block length, metasyndrome-only exact repair fails with probability tending to one at any fixed measurement-error rate. With no intervening data fault, the same orbit count also gives the minimum partial second measurement needed to remove every single-fault ambiguity. Exact calculations separate the standard examples. The $[[72,12,6]]$ code has syndrome distance three and minimum-weight repair corrects every single measurement fault, whereas Gross $[[144,12,12]]$ has 36 indistinguishable single-fault pairs. Sustained phenomenological experiments show the same qualitative contrast and favor joint data--measurement decoding over a separated repair stage on the codes with stronger ambiguity. The analysis extends our earlier coprime-period treatment to general two-block BB codes and separates measurement ambiguity from the data error it can induce.

quant-ph↗

Duality functors for Coulomb branches I

To a quiver of finite type one can associate two abelian categories: the category of finite dimensional modules over the corresponding affine quiver Hecke algebra and the Coulomb category of Koszul-perverse coherent sheaves. In this paper we define an exact functor from the former to the latter. This functor sends a simple to a simple or zero and is (almost) essentially surjective on simples. This implies that simple Koszul-perverse sheaves categorify the dual canonical basis.

math.RT↗

Prequential E-Values for Selected-GP Near-Optimality Certificates

When optimizing an expensive black-box function sequentially, as in hyperparameter optimization, we may want to stop once the best evaluated value is certified within $\varepsilon$ of the global optimum. Such a certificate needs two ingredients: a lower confidence bound for the selected value and an upper confidence envelope over the domain, typically supplied by a Gaussian process (GP). GP-UCB-style stopping rules are valid when the kernel and constants defining this envelope are fixed before the run, but the practical temptation is to tune the envelope from the same adaptive evaluations and then certify as if it had been fixed. We use prequential e-values to make this selection auditable: starting from a predeclared set of fully specified GP/RKHS envelopes, each candidate is tested by its own one-step-ahead e-process, contradicted candidates are deleted, and certification uses the largest upper bound among the survivors. With a valid selected-point lower bound and one declared candidate having valid latent coverage and noise calibration, the rule is anytime-valid. On a 512-seed noisy RBF stress sweep, it roughly halves false-certification risk at comparable power versus fit-then-certify. Relative to random fixed GP precommitment on smooth $d=3,4$ objectives, each additional false certificate is accompanied by 3.0 and 13.5 additional correct certificates, respectively.

cs.LG↗

CDMD: A Cross-Dataset Mixed-Type Diffusion Model for Tabular Data

Generative models for tabular data are typically trained separately for each dataset, limiting knowledge transfer and requiring the storage of many specialized models. In this paper, we introduce CDMD, a tabular diffusion model trained jointly across heterogeneous datasets with different schemas and variable numbers of numerical and categorical features. Unlike existing cross-dataset tabular diffusion models that operate in continuous representation spaces, CDMD defines diffusion directly over the mixed-type feature space and is trained end-to-end. To accommodate heterogeneous categorical domains, we introduce a schema-restricted reverse-process parameterization for masked diffusion models, in which the output space dynamically adapts to each feature's vocabulary. We then compose numerical and categorical feature-level diffusion processes into a schema-dependent row-level process. A shared schema-aware Transformer denoiser captures dependencies between features and parameterizes the reverse process across varying schemas. On seven real-world datasets, a single jointly trained CDMD achieves the highest average generation quality among strong single-dataset and cross-dataset baselines, while using substantially fewer total parameters than the collection of separately trained models. Furthermore, pre-training on a corpus of 337 datasets improves generation on previously unseen datasets under both limited target data and limited adaptation epochs. These results demonstrate the potential of direct mixed-type diffusion for shared and transferable tabular data generation. Our code is available at https://github.com/ketatam/cdmd.

cs.LG↗

LBDU-VIO: Learned Bias Dynamics and Uncertainty for Visual-Inertial Odometry with Unreliable Vision

Visual-inertial odometry (VIO) for aerial robots relies on high rate inertial measurement unit (IMU) propagation between visual updates. However, conventional multi state constraint Kalman filters (MSCKFs) use random walk bias assumptions and fixed noise parameters, which can limit robustness when visual information is unreliable. To address this problem, we propose LBDU-VIO, a learning-augmented MSCKF with learned continuous time bias dynamics and an IMU uncertainty model. A neural ordinary differential equation (ODE) models continuous time bias dynamics to propagate the filter's bias states, replacing their random walk model. The IMU uncertainty model predicts motion adaptive measurement noise covariances for covariance propagation. Both models are trained with pose supervision without direct labels. Experiments on real world EuRoC and TUM-VI benchmarks show lower errors than representative visual-inertial baselines, including a 25.1% reduction in mean relative position error compared with S-MSCKF on EuRoC sequences with 10s visual outage.

cs.RO↗

An uncertainty principle for entanglement

Consider two orthonormal bases of a bipartite Hilbert space such that states in one basis are weakly entangled and states in the other are strongly entangled. We establish a quantitative version of the following uncertainty principle: every state that is localized in one basis must be delocalized in the other. More specifically, we lower bound the sum of the Shannon entropies of any state's coefficient distributions in the two bases by the difference between the entanglement Rényi entropies of their basis states. The result follows from new lower and upper bounds on the entanglement of superpositions that may be of independent interest. As an application, our result limits the power of weakly-entangled states in representing ground states of Hamiltonians with strongly-entangled energy eigenstates. We illustrate this application numerically for the Sachdev-Ye-Kitaev model.

quant-ph↗

How to Reduce Localization Ambiguity? Geometry-Semantic Constrained BEV Representation Learning for Satellite-Ground Localization

Satellite-ground localization estimates the planar position and yaw orientation of a ground camera within a geo-referenced satellite image. Most recent methods map ground and satellite features into a shared bird's-eye-view (BEV) space and establish spatial correspondences. However, insufficient depth constraints can assign one ground feature to different distances along a viewing direction, creating geometric ambiguity in BEV feature placement. Similar appearances at different locations can also create descriptor matching ambiguity, while existing descriptor learning lacks explicit semantic supervision to distinguish them. We propose GeoSem-BEV, a geometry-semantic constrained BEV representation learning method. Radial depth supervision constrains distance assignment, and vertical height supervision constrains height aggregation. Shared explicit semantic supervision promotes consistent semantic predictions across views and helps distinguish locations with similar semantics. These constraints improve feature placement and descriptor discriminability, enhancing state-of-the-art BEV localization models. On VIGOR with unknown orientation, GeoSem-BEV reduces mean orientation error by 37.2% and 38.1% in the cross-area and same-area settings, respectively. The corresponding errors are reduced by 10.8% and 15.6% on DReSS-D. On KITTI-CVL, it reduces same-area mean orientation error by 26.8% under 10 degree orientation noise.

cs.CV↗

GeoGAT: Bidirectional Temporal Sampling Meets Hierarchical Graph Attention for Global Video Geo-localization

Global video geo-localization aims to infer the geographic location of a video worldwide, evaluating performance across four geographic hierarchies: city, state/province, country, and continent. Existing methods typically employ one-way uniform sampling to process video frames and train independent classifiers for each hierarchy, which leads to the loss of key geographic cues and prediction conflicts between hierarchies, especially for complex multi-shot edited videos. To address these limitations, we propose GeoGAT, which integrates bidirectional temporal sampling with graph attention networks (GATs). Specifically, GeoGAT extracts forward and offset-reversed frame sequences to construct complementary spatiotemporal features. These fused features are then fed into a predefined geographical hierarchy graph, where GATs perform structure-aware message passing, while a dual-constraint mechanism prunes predictions to eliminate cross-hierarchy conflicts. We construct GeoGAT10k, comprising 9,720 multi-shot edited videos from 166 cities worldwide, specifically to benchmark generalization ability on complex video structures. Experimental results on CityGuessr68k and GeoGAT10k demonstrate that GeoGAT eliminates hierarchical conflicts entirely and achieves state-of-the-art performance across all four geographic hierarchies. On CityGuessr68k, GeoGAT outperforms the strongest baseline, evaluated under both classification and retrieval protocols, by 2.6 percentage points at the city level. On the more challenging GeoGAT10k with multi-shot edited videos, the accuracy improvement exceeds 24 percentage points, validating strong generalization to complex real-world scenarios.

cs.CV↗

Drape-Compatible Tool-Tip Localization for Hand-Held Laparoscopic Instruments via UWB Carrier-Phase Ranging and Trocar-Constrained Geometry

A surgical robot policy needs to know where each instrument's working end sits relative to the camera and to the other instrument, yet a laparoscopic operation leaves only the endoscope video, and a draped instrument hides every optical path on itself. We estimate the tool tips from distances and an IMU alone. The distances are measured by ultra-wideband carrier phase between antenna nodes on the handle side of the instruments: one per hand-held instrument, one on the endoscope, all outside the sterile barrier. Take the endoscope antenna as the reference. The two instrument antennas carry six coordinates while the three pairs give three distances, so the problem is short by three, and averaging cannot close that gap because what is missing is information, not precision. Carrier phase adds one more unknown per pair, an integer that leaves distance fixed only modulo lambda/2 = 23.1 mm. The trocar settles both difficulties: each shaft passes through a port whose position is known and whose direction the on-board IMU measures, so its antenna keeps a single degree of freedom, the insertion depth. Two unknowns now face three distances -- two fix the depths and the third is left over, and that spare equation exposes a wrong integer or a wrong mounting offset instead of absorbing it. The same solve returns both tool tips in the endoscope frame, without the second carrier frequency that integer resolution usually needs, so the link never leaves a single PLL lock. In an 11-hole phantom among metal instruments the three distances hold sigma = 0.10-0.33 mm at 35-46 Hz, a sterile drape pressed onto the antennas leaves the link unchanged where an optical path would be blocked, and the logger was carried into a live rabbit appendectomy.

cs.RO↗

Perceptual Color Difference Modeling Using Machine Learning and Human Similarity Judgments

Accurate assessment of color differences is essential for applications ranging from digital design to quality control. While existing color difference metrics, such as CIEDE2000, aim to approximate human perception, they may still exhibit inconsistencies with perceptual judgments. In this study, we investigate a data-driven approach to color-difference estimation based directly on human evaluations. We collect similarity judgments for 2,000 systematically generated color pairs, each rated by seven observers using a four-point ordinal scale. These judgments are then used to train regression models using different color representations, including RGB channel differences, HSI differences, and COLIBRI fuzzy linguistic categories. Experiments with five regression algorithms show that the choice of color model has a greater influence on prediction performance than the choice of regression algorithm. Using COLIBRI features alone, linear regression achieves an R2 of 0.595, outperforming RGB and HSI representations, which achieve R2 values of 0.479 and 0.493, respectively. The best performance is obtained by LightGBM using the combined representation, reaching an R2 of 0.703. The results indicate that human perceptual color differences are better captured when numerical color coordinates are complemented by graded perceptual categories, highlighting the potential of data-driven models for perceptually aligned color-difference estimation.

cs.CV↗

Characterizing High Bandwidth Flash for LLM Serving

Large language model (LLM) serving requires substantial memory to store model weights and KV caches. As models grow larger and contexts become longer, memory capacity and bandwidth increasingly become bottlenecks for serving performance. Agentic workloads compound this pressure through repeated interactions over growing contexts, making it increasingly important to retain KV state for reuse. High-bandwidth flash (HBF) offers a way to expand accelerator memory capacity for large language model (LLM) serving, but its access costs and limited write endurance complicate its use. We evaluate HBF for high-throughput agentic serving across system design and scheduling choices to understand when additional capacity improves serving performance and energy efficiency. We introduce an HBM-HBF-host hierarchical storage system and buffered cache-aware scheduling, and use trace-driven simulations to analyze their effects on performance, energy consumption, and HBF write lifetime. Across the evaluated workloads, the fastest HBF-augmented systems reduce completion time by 36.1-87.0% relative to HBM-only systems. Modeled energy savings reach 55.8%, although HBF increases energy consumption on some light workloads. Buffered cache-aware scheduling extends estimated HBF write lifetime from 4.77 to 14.82 years in the evaluated configuration. These results demonstrate the importance of coordinating data placement and scheduling to improve serving efficiency while sustaining a practical HBF write lifetime.

cs.LG↗

Uncertainty-Aware Consistency Distillation for Few-Step Video Generation

We study few-step video generation, i.e., distilling a multi-step video generator, which typically requires tens of sampling steps, incurring substantial latency and compute, into a few-step student. Consistency distillation is a common recipe, in which a multi-step teacher provides the consistency targets for a few-step student. However, these teacher-guided targets are not equally trustworthy, and the content is harder to learn where it varies rapidly over time, e.g., moving foliage shadows or flowing water. We observe that supervision reliability follows the local difficulty of the content rather than semantic complexity: regions that change little yield consistent endpoint predictions, whereas regions with large temporal variation produce larger discrepancies that coincide with the largest perceptual errors. Motivated by this observation, we propose Uncertainty-Aware Consistency Distillation (UACD), which reweights consistency supervision at each spatiotemporal region using a local, parameter-free uncertainty estimate. Specifically, we construct two independently perturbed teacher-guided consistency paths, whose student endpoint predictions provide a consensus target; the discrepancy between the student's direct prediction and this target is the uncertainty proxy. We then relax the consistency penalty on high-uncertainty regions through an exponential weight, while keeping the full penalty elsewhere, since the student cannot be expected to match targets that are hard to learn. To preserve perceptual quality under aggressive step reduction, we integrate feature-space adversarial training with semantic alignment. With parameter-efficient LoRA adaptation of the 50-step Wan model, our method achieves state-of-the-art 4-step generation on VBench 2.0 (0.556 mean score) and is preferred over competing methods in a user study.

cs.CV↗

Limit F-invariants of the simple singularities and the Fermat hypersurfaces

In this article, we compute the limit Hilbert--Kunz multiplicities and limit F-signatures of all the simple singularities and the Hilbert--Kunz multiplicities of the Fermat hypersurfaces, generalizing the well-known sec z + tan z formula by Gessel and Monsky for the A1 singularities, using the theory of h-functions. We also obtain a functional equation relating the phi-function of a hypersurface to its reflection. Using limit F-invariants, we generalize both a characterization of regular local rings through F-invariants and the known cases of the Watanabe--Yoshida conjecture to characteristic zero.

math.AC↗

Feature-Aware Token Attack for Compression-Triggered Stealthy Failures in Large Vision-Language Models

Visual-token compression improves the efficiency of large vision-language models, but can expose failures that full-token evaluation misses. We study adversarial images that preserve full-token correctness yet induce errors after compression, even when both inference paths succeed on the clean image. Creating such failures is challenging because perturbing token importance can also damage the visual content needed for full-token inference. We propose Feature-Aware Token Attack (FATA), which couples attention suppression with cosine-based feature preservation on a fixed set of salient clean-image tokens. In the primary LLaVA-1.5-7B setting, FATA uses only vision-encoder gradients, without access to the deployed compressor, token budget, or downstream task. Across four visually dependent task subsets and four compressors under a controlled reconstruction protocol, FATA achieves SR = 96.3% full-token accuracy retention and CBR = 22.1% conditional blinding, compared with 89.8% and 15.7% for CAA. Ablations support the role of both objectives in balancing compressed-path failure against full-token preservation. FATA also has the lowest measured detection rate among four attacks across three evaluated detectors at a 5% false-positive rate. These findings motivate assessing adversarial robustness jointly across full-token and compressed inference.

cs.CV↗

Asking the World: Generalist Physical Reasoning through Agentic World Modeling and Probing

Physical reasoning from video requires inferring latent physical properties and dynamics beyond direct observation. Direct VLM inference remains unreliable on complex physical tasks without explicit modeling and validation, while predefined tool pipelines rely on task- and domain-specific priors that limit generalization across materials, dynamics, and reasoning tasks. We introduce Asking the World (ATW), a generalist agent that constructs and interrogates task-relevant executable worlds through two adaptive stages: World Modeling calibrates a world from video, while World Probing queries, simulates, and intervenes on it to obtain question-relevant evidence. Rather than prescribing the operations in either stage, ATW determines how to model and probe according to the scene and question. We develop PolyWorld Engine, a lightweight and highly programmable Warp-based multiphysics simulator for constructing and probing worlds with rigid bodies, soft bodies, cloth, ropes, fluids, and their coupled interactions. CEM-based system identification recovers task-relevant dynamics during World Modeling. The resulting world becomes an active workspace for question-directed physical experiments rather than a predetermined downstream tool. We evaluate ATW on CLEVRER, ContPhy, and three real-world scenarios. Using Gemini-3-Flash as its base VLM, ATW achieves 80.82% overall per-question accuracy on CLEVRER, improving direct Gemini-3-Flash by 46.50 points, GPT-5.5 by 13.58 points, and PhysMind by 8.27 points. On ContPhy, it reaches 70.56% overall accuracy, surpassing Gemini-3-Flash by 28.10 points and GPT-5.5 by 3.53 points. Across the three real-world scenarios, ATW achieves 71.67% accuracy, 28.33 points above GPT-5.5. These results establish agentic world modeling and probing as an effective, execution-grounded approach to generalist physical reasoning.

cs.CV↗

Characterization of Sobolev regularity of plurisubharmonic functions

Let $f$ be a nonzero holomorphic germ at $0 \in \mathbb C^n$ with $f(0)=0$, and let $χ$ be a $C^2$ non-decreasing convex function on the left half-line. We establish sharp necessary and sufficient conditions for the local Sobolev regularity of the plurisubharmonic function $v=χ(\log|f|).$ The criteria for the $L^p$-integrability of the classical Laplacian and for $W^{1,p}$-regularity are given by a weighted integral involving $χ''$ and $χ'$, respectively, and depend on $f$ only through the smallest multiplicity of $\operatorname{Div}(f)$. We also obtain the $W^{2,p}_{\mathrm{loc}}$ criterion for $1<p<\infty$. At the endpoint $p=1$, we prove that \[ v\in W^{2,1}_{\mathrm{loc}} \quad\Longleftrightarrow\quad χ'\in L^1((-\infty,A)), \] equivalently, $v$ is locally bounded. As applications, we obtain counterexamples to Calderón-Zygmund theory and characterize a class of functions in the local Monge--Ampère domain. % and exhibit a natural family in $W^{2,1}_{\mathrm{loc}}$ whose limit fails to belong to $W^{2,1}_{\mathrm{loc}}$.

math.CV↗