arXiv Science⌕ Search

SEARCH · arXiv Science

Search arXiv Science

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,441 records · Page 80Linked to original sources

From Demailly's Inequality to Harbourne--Huneke Containments for General Points

Let $I$ be the defining ideal of a general set of $s$ points in projective $N$-space over an algebraically closed field. Harbourne and Huneke conjectured the strengthened symbolic containment $$ I^{(Nr)}\subseteq \mathfrak{m}^{r(N-1)}I^r $$ for every $r\ge 1$. For general points this was known in low dimension and for small numbers of points, while in arbitrary dimension and arbitrary cardinality the best general result was the stable version, valid on one dense Zariski-open set for all sufficiently large $r$. We show that the recent proof of Demailly's conjecture for arbitrary finite point sets by Hà and Sivakumar removes the stability threshold. The key observation is that the strict generic Waldschmidt estimate of Bisui and Nguyen allows one to select a single finite symbolic level $m_0$. Equality of the corresponding interpolation number is a Zariski-open condition. Demailly's inequality then propagates this one finite condition to lower bounds for every symbolic power, and a standard regularity criterion converts those bounds into the Harbourne--Huneke containment for every $r$. Consequently, for every $N\ge2$ and every $s\ge1$, one dense Zariski-open family of $s$-point configurations in $\mathbb{P}^N$ satisfies the Harbourne--Huneke containment simultaneously for all $r\ge 1$. The argument also clarifies a quantifier issue in the earlier literature: before the full Demailly theorem, estimates on the generic fibre naturally yielded either all $r$ on a very general set, or one open set only for $r\gg 0$. The new theorem replaces infinitely many symbolic conditions by a single finite interpolation condition.

math.AC↗

Quantum geometry of collective pairing fluctuations in the superfluid weight of multiband superconductors

Beyond the frozen-pairing contribution, the superfluid weight of a multiband superconductor contains a correction from the self-consistent relaxation of the pairing amplitudes under a phase twist. We show that this correction has a quantum-geometric origin, not in the space of Bloch states but in the space of collective pairing fields. For an attractive multiband Hubbard model at the mean-field level, the thermodynamic Hessian governing the pairing response coincides with the static Gaussian pair-fluctuation kernel at zero momentum at every temperature below $T_c$. Since gauge invariance relates a phase twist to a finite center-of-mass momentum of the pair field, the superfluid weight is set by the small-momentum curvature of the eigenvalue branch that evolves from the Goldstone mode, and the relaxation correction is a stiffness-weighted quantum metric of the corresponding soft eigenvector on the manifold of pairing configurations. The same construction yields the Ginzburg-Landau gradient coefficient with the identical geometric correction, hence the coherence-length tensor, and, with the low-frequency dynamics of the kernel, the pair-mass tensor. Lieb-lattice calculations verify the equivalence of the thermodynamic, fluctuation-kernel, and soft-mode formulations and the independence of the superfluid weight from the orbital embedding.

cond-mat.supr-con↗

Machine Learning for German Redispatch Forecasting under Data Delays and Temporal Distribution Shift

Public redispatch records provide empirical data for grid congestion forecasting, but delayed reporting, zero-inflated distributions, and temporal shift present major modeling challenges. We assess the accuracy and reliability of probabilistic machine-learning forecasts using published German transmission records under experimentally imposed information-age constraints. The benchmark evaluates eight daily series of upward and downward intervention energy across four German transmission system operators from 2021 to 2024 (48,242 eligible records; 354 evaluation dates in 2024). We compare seasonal empirical, regularized autoregressive (ARX), quantile LightGBM, GRU, and Transformer models under a minimum seven-day target-latency constraint. Neural architectures use a zero-censored output head to accommodate exact-zero outcomes. Static, rolling, and adaptive delayed-feedback calibration are evaluated using normalized weighted interval score (nWIS), empirical coverage, and block-bootstrap inference. Raw LightGBM achieved nWIS 0.7952, outperforming ARX (1.0604) and the seasonal baseline (0.8739) by 25.0% and 9.0%, respectively (Holm-adjusted p<0.005). Rolling calibration improved LightGBM to nWIS 0.7767 versus 0.8251 for static calibration (p=0.0092), with 91.81% coverage for nominal 90% intervals. The zero-censored Transformer achieved nWIS 0.8161, with no significant difference from LightGBM (p=0.260). However, aggregate coverage concealed substantial undercoverage during high-volume interventions (61.91% coverage among above-threshold events). These results show that boosted-tree models with rolling calibration provide accurate probabilistic forecasts of aggregate redispatch volumes under target delays, while nominal aggregate validity does not ensure reliability during extreme congestion events.

eess.SY↗

Simulated Annealing for Antenna Placements

The 2026 ACM SIGSPATIAL GIS Cup posed the following geometric challenge: Given a set of simple, pair-wise disjoint polygons, compute $k$ antennas (points) on polygon boundaries, maximizing the number of polygons with at least a fraction $τ$ of their perimeter visible from the antennas. Contestants had 24 hours to produce the best possible solutions. We describe an approach that separates geometric visibility preprocessing from a portfolio of incremental combinatorial searches. A construction pool provides initial solutions with antennas restricted to polygon vertices; simulated annealing explores one-for-one antenna swaps, complemented by exact discrete one-swap descent, ruin-and-recreate, and elite crossover. A secondary objective rewards progress toward the service threshold without overriding the primary score. Candidate antennas are later expanded to include points in polygon-edge interiors. Our submitted solutions cover between 261 and 16,273 buildings across the nine parameter combinations, and the team was invited to present as one of the top entries. The code is available at https://github.com/JacobusTheSecond/giscup.

cs.CG↗

Digital Twin-Driven Real2Sim2Real: Simulator-Conditioned Generation via Paired Driving-Scene Reconstruction

Camera-based 3D perception for autonomous driving relies heavily on large annotated datasets, and deploying such a system to a new target region typically requires data collection and annotation. Generative augmentation has been proposed to reduce this cost, but existing approaches face a fundamental trade-off: label-conditioned methods consume the very annotations they aim to replace, while simulator-conditioned methods offer free annotations but lack visual grounding to specific real environments. This work investigates the extent to which a digital-twin-driven Real2Sim2Real pipeline (DT-R2S2R) can substitute for target-region real data. By reconstructing recorded driving clips inside a georeferenced digital twin (DT-R2S), we condition a diffusion model on geometrically aligned simulator renderings, establishing a digital twin-grounded Sim2Real model (DT-S2R). As a result, DT-S2R synthesizes photorealistic driving images given low-cost yet georeferenced simulator data across both reconstructed and novel simulator scenes within digital-twin coverage. The efficacy of generated data is verified on diverse 3D detectors. DETR3D, especially, reports 93.18% of mAP obtained by a target-region real-data oracle, without employing target images for detector training. Furthermore, simple co-training with existing out-of-target real data outperforms the oracle. Thus, DT-R2S2R can substantially reduce the cost of manual on-site data collection and annotation in digital twin-available districts, providing a practical foundation for scaling 3D perception.

cs.CV↗

Compositeness of near-threshold exotic hadrons in systems with Coulomb and short-range interactions

We study the universal properties of near-threshold $s$-wave eigenstates in systems with Coulomb plus short-range interactions by analyzing their compositeness. Using a nonrelativistic effective field theory, we derive a relation between the compositeness of near-threshold states and the Coulomb scattering length, the Coulomb effective range, and the Bohr radius. We find that the long-range Coulomb interaction modifies the threshold behavior expected from the low-energy universality of purely short-range systems. When the Coulomb interaction is relatively weak compared with the short-range interaction, a remnant of the low-energy universality associated with the short-range interaction survives, leading to composite-dominant near-threshold eigenstates. In contrast, in the strong-Coulomb regime, this remnant disappears.

hep-ph↗

DIPrune: Task-Aware Token Pruning with Dual Importance for Efficient Multimodal Language Models

Recent training-free pruning approaches for Multimodal Large Language Models (MLLMs) effectively cut computational overhead by exploiting visual redundancy or text-vision attention. However, they frequently suffer from semantic degradation due to their task-agnostic design or unreliable attention estimates. Based on our empirical analysis, we have found that this issue arises because salient tokens in shallow layers persistently suppress emerging semantic ones through numerical inertia, leading to premature discarding of signals crucial for deep reasoning. To address the aforementioned issue, from the task-oriented aspects, we first reformulate training-free pruning as a minimization of the distortion in the final task loss and derive a tractable, token-wise upper bound to serve as a surrogate objective. Specifically, this formulation inherently reveals a previously neglected inter-layer term that accounts for gradients across layers. Accordingly, for the implementation, we propose DIPrune, a rank-based framework that employs a dual importance scoring mechanism to jointly optimize intra-layer static feature saliency and inter-layer dynamic semantic evolution. Extensive experiments on LLaVA and Qwen-VL demonstrate that DIPrune consistently achieves state-of-the-art results.

cs.CV↗

Perfect Bound States and non-Markovian Dynamics for Imperfect Giant Atoms in a Semi-Infinite Waveguide

Bound states in structured reservoirs require exact interference conditions and are therefore sensitive to imperfections in the coupling geometry. In this paper, we study a two-level giant atom coupled at multiple points to a mirror-terminated semi-infinite waveguide. The exact atom-photon bound state requires both cancellation of the outgoing radiation amplitude and dispersive matching to the atomic detuning. Generalized delay-equation dynamics, continued-pole spectra, and real-space photon fields verify this criterion for coupling-strength, position, and coupling-phase imperfections. For representative realizations, retuning restores one-, two-, and three-root responses after strength errors and one- and two-root responses after position errors. The specified position-imperfect three-root configuration fails because its $2π$ root is no longer dark, consistent with a $58.77\%$ spectral mismatch. Ensembles of 1000 realizations at each error amplitude show near-unity target-state recovery for strength errors, unity recovery of the specified one- and two-root targets but no three-root recovery for position errors, and no nontrivial exact recovery for independent phase errors. Optimized phase-imperfect responses instead retain finite linewidths. Root matching with realization-dependent control retuning therefore constitutes a compensation procedure, rather than generic passive robustness against arbitrary disorder.

physics.optics↗

Event-Driven Proactive Robot Assistance through Vision-Language Reasoning

Assistance in collaborative manipulation is often initiated by user instructions, making high-level reasoning request-driven. In fluent human teamwork, however, partners often infer the next helpful step from the observed outcome of an action rather than waiting for instructions. Motivated by this, we investigate an event-driven formulation of proactive assistance, where human--object interaction outcomes initiate assistive reasoning without user-provided task specifications at inference time. To this end, we propose an event-driven framework that monitors workspace state changes with an event monitor and, upon event completion, extracts stabilized pre/post snapshots that characterize the resulting state transition. A frozen pretrained Vision-Language Model (VLM) then uses its semantic priors to infer the task context, decide whether assistance is appropriate, and, when needed, generate a sequence of assistive actions from the observed transition. To make outputs executable and verifiable, we restrict actions to a set of action primitives and reference objects via integer IDs.We evaluate the same framework across three distinct real world tabletop collaboration tasks without task-specific training or fine-tuning. The event-driven framework achieves performance comparable to variants given user instructions.

cs.RO↗

High-Dimensional Statistical Inference for Sparse Support Vector Machines

Using a replica-symmetric high-dimensional characterization, we develop an inferential framework for sparse support vector machines when the sample size and number of features grow proportionally. The main challenge is the nonsmooth hinge loss, which prevents direct application of debiasing arguments developed for smooth classification losses. We overcome this difficulty by representing the $L_1$-penalized support vector machine (SVM) as a linear program and identifying the hinge-loss subgradient through its dual variables. This yields a computationally accessible debiased estimator whose coordinates are asymptotically Gaussian under the proportional asymptotic regime. The resulting distributional characterization provides confidence intervals and hypothesis tests for individual features and enables false-discovery-rate-controlled variable selection. Extensive simulations examine calibration, power, and variable-selection performance under a range of covariance structures, including strongly correlated designs. An analysis of high-dimensional breast cancer gene-expression data illustrates how the proposed inference can distinguish statistically significant features from variables selected by the original sparse SVM.

stat.ML↗

PolarScale: A Physics-Grounded Benchmark for Radiometrically Consistent RGB-to-Stokes Estimation

Polarization imaging provides physical cues beyond intensity imaging but typically requires specialized hardware. Recent methods infer polarization from RGB-like inputs, yet predict only normalized Stokes components or relative descriptors, from which the radiometric scale needed for full Stokes reconstruction has been divided out. We introduce PolarScale, a benchmark that makes this scale an explicit prediction and evaluation target. Built on existing trichromatic full-Stokes measurements, PolarScale takes the per-scene normalized total-intensity image $s_0$ (a scene-referred linear image, not a consumer sRGB photograph) and asks models to predict normalized Stokes components, AoLP/DoLP/DoCP, and a per-scene scale. Because the scale is divided out of the input, it is not physically identifiable; PolarScale therefore evaluates dataset-conditioned semantic scale estimation against a constant-scale control, together with angular, self-consistency, and physical-bound metrics. Across seven restoration-based and generative backbones and three prediction strategies, the strongest restoration models estimate the scale with 3.6-4.3% mean relative error versus 5.7% for the constant control and violate physical bounds on fewer than 0.25% of pixels, whereas two generative baselines collapse to a near-zero scale; explicit descriptor supervision improves descriptor accuracy (23.66 vs. 18.88 dB PSNR for MAE). Predicted full-Stokes representations improve diffuse/specular separation, material segmentation, and glare classification, although in diffuse/specular separation the learned scale performs only on par with the constant control.

cs.CV↗

Network Intervention by Polling Strategic Agents

A planner in a network of strategic agents faces three entangled challenges: the optimum depends on agents' private information, queried agents may misreport to steer the outcome, and exact computation does not scale. We study these challenges in multi-activity network games with heterogeneous private technologies, in which the planner sets non-discriminatory prices. We show that the optimal prices admit a centrality-based decomposition of the welfare kernel: each agent's contribution scales with its squared centrality in a network reweighted by agents' preferences across activities. This decomposition motivates Poll, a polling algorithm in which the planner samples one agent per round, walks briefly through the agent's neighborhood, and updates the price from a local report. From the same decomposition flow three forms of efficiency: computationally, Poll uses significantly fewer operations than exact computation and other distributed methods, requiring up to three orders of magnitude less communication on a real-world network with over 300,000 agents; statistically, its query complexity scales with topology and preference heterogeneity rather than explicitly with population size; and economically, it converges to welfare-maximizing prices while admitting behavior-specific implementations that induce truthful reports and detect adversarial deviations.

cs.GT↗

Can one hear the shape of a lattice random walk?

We construct distinct high-dimensional mean-zero finite range lattice random walks having pairwise-equal return probabilities for all step counts. The same examples provide pairwise-distinct shapes of discretizations of the standard Laplacian with pairwise-equal density state functions. The main contribution is a reconstruction theorem: to a colored trivalent graph one associates a quantum Clebsch--Gordan polytope, and this association is a full functor, in particular from the polytope one can uniquely recover the original graph. These polytopes appear as moment polytopes of toric degenerations of character varieties (moduli spaces of rank-2 bundles on curves), yielding a combinatorial non-abelian Torelli theorem. In symplectic geometry, the reconstruction theorem implies that monotone Lagrangian tori on odd character varieties associated with these degenerations are pairwise non-Hamiltonian isotopic. These results, and some of the applications, arose from the study of mirror symmetry for moduli spaces of vector bundles, and of the related Laurent phenomenon for mutations of graph potentials.

math.PR↗

Optimal Conversion Bandwidth for MDS Convertible Codes in the Split Regime with $r^F<k^F<r^I$

Erasure codes are widely used in distributed storage systems to provide fault tolerance. An $[n,k]$ erasure code encodes $k$ data symbols into $n$ coded symbols and distributes them across $n$ storage nodes. Once the code parameters are fixed, the achievable fault tolerance is also fixed. However, the failure rates of storage nodes may vary over time, and dynamically adapting the code parameters to these variations can substantially reduce storage overhead. Motivated by this observation, Maturana and Rashmi introduced convertible codes, which allow an $[n^I,k^I]$ initial code with redundancy $r^I=n^I-k^I$ to be transformed into an $[n^F,k^F]$ final code with redundancy $r^F=n^F-k^F$, while preserving the required code properties. Convertible codes whose initial and final codes are both MDS codes are called MDS convertible codes. They are of particular interest because MDS codes provide the maximum erasure tolerance for a given amount of storage overhead. Several works have established lower bounds and constructions for the bandwidth cost of MDS convertible codes in the split regime. However, the tight bound in the parameter range $r^F<k^F<r^I$ remains unknown. In this work, we establish a family of new rank inequalities for stable MDS convertible codes with linear conversion procedures and derive an improved lower bound on the bandwidth cost consisting of three cases for this remaining range. We prove that the bound is tight by presenting three explicit constructions, one for each case.

cs.IT↗

How Much Planning Is Enough? Reducing Search and Computation in World-Model Planning

Visual world models enable goal-directed control through decision-time action search, but their deployment efficiency is often limited by conservatively large planning budgets. We show that competitive task performance can be achieved without agreement with the Full-budget action, that sufficient budgets vary across model--task pairs, and that iterative planners repeatedly encode solve-invariant context. To address these inefficiencies, we propose {SufficientPlan}, a simple deployment framework that requires no modification to pretrained world models or planner updates. Its {Paired Sequential Budget Certification (PSBC)} component uses paired closed-loop evidence to search for and certify a reduced model--task-specific budget within a predefined Full-performance tolerance. Its {Static-Context Reuse (SCR)} component caches observation and goal representations across search iterations while preserving candidate-dependent planning and selected actions. Experiments across multiple world-model backbones and visual-control tasks show that SufficientPlan substantially reduces search budgets and planning latency while maintaining competitive control performance.

cs.RO↗

Regularity of Free Boundary for Monge-Ampère Obstacle Problem up to the Boundary

In this paper, we study the regularity of free boundaries for Monge-Ampère obstacle problem up to the boundary. We mainly focus on how the growth order of boundary value near the contact point impact the regularity. We also find an interesting phenomena that under quadratic growth boundary value, the free boundary enjoys better regularity in high dimensions than in two dimension.

math.AP↗

6G NeXt - Towards 6G Split Computing Network Applications: Use Cases and Architecture

The definition of the sixth generation mobile communication is in the full swing. Mobile communication aims for the convergence of physical, human and digital world. Research project 6G NeXt is considering two demanding use cases, holographic communications and drones anti-collision system, which set heterogeneous requirements on the communication as well as the computing infrastructure. In both use cases, the clients are widely spread in the network and are cooperatively interacting with each other. Especially for holographic communication, also high processing power is required. This makes a high-speed distributed backbone computing infrastructure, which realises the concept of split-computing, inevitable. Furthermore, tight integration between processing facilities and wireless network is required in order to provide adequate quality of service to the users. This paper illustrates the use case scenarios and their requirements. Afterwards, an appropriate solution approach to realise those is elaborated. Here, the novel technological approaches are discussed based on the developed overall communication and computing architecture.

cs.NI↗