arXiv Science⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 1,423 records · Page 79Linked to original sources

Exact Distinguishability in Non-Markovian Decision Processes

Non-Markovian environments are often modeled as Regular Decision Processes (RDPs), where dynamics depend on the interaction history through a finite automaton. Existing offline guarantees for RDPs rely on a distinguishability assumption on the behaviour policy but provide no means of verifying it. When the assumption is violated, distinct models may explain the data equally well. We study when data collected under a fixed behaviour policy can distinguish two candidate RDPs. We prove that the posterior odds between observationally equivalent candidates remain equal to the prior odds at every sample size, even when the policy visits every automaton state, and verify both results formally in Lean 4. We then characterize this equivalence exactly and derive PEC, an algorithm that decides it in time linear in the size of the product automaton. The distinguishability assumption of prior work fails on three of our four test environments, and the experiment identified by PEC restores it in each case.

cs.LG↗

A Nonlinear Two-Sheath Circuit Model for Low-Pressure Symmetric and Asymmetric Capacitively Coupled Radio-Frequency Plasmas

We develop a nonlinear self-consistent two-sheath circuit model for low-pressure capacitively coupled plasmas that applies to geometrically symmetric as well as asymmetric discharges. The quasineutral plasma bulk is represented by an inductive-resistive element and coupled to stationary particle and electron-energy balances. Both boundary sheaths are treated dynamically using a lncosh sheath charge-voltage model with bounded differential elastance, based on a Riccati closure for the differential sheath width. The model recovers the quadratic depletion-sheath relation in the small-charge limit, while the characteristic sheath scales are determined from the RF-averaged sheath voltages using a collisionless Child-Langmuir/Bohm closure. The resulting four-variable RF subsystem contains the two sheath charges, the blocking-capacitor voltage, and the discharge current. In the symmetric monofrequent limit, the two sheath nonlinearities compensate strongly and the dc self-bias vanishes, whereas geometrical or electrical asymmetry breaks this compensation and enhances harmonic generation and plasma-series-resonance oscillations.

physics.plasm-ph↗

Determining $W^{1,n}$ conductivities in Calderón's problem for $n\geq4$

We establish a global uniqueness result at the critical Sobolev regularity for the isotropic Calderón problem in all dimensions $n\ge 4$. More precisely, for every $n\ge4$ and every bounded Lipschitz domain $Ω\subset\mathbb{R}^n$, any real-valued, uniformly elliptic scalar conductivity $γ\in W^{1,n}(Ω)$ is uniquely determined by its full Dirichlet-to-Neumann map. This settles a long-standing open problem in the field. To prove uniqueness at this endpoint, we develop an averaged trace norm method, which combines complex geometrical optics with trace norm estimates averaged over wave orientations.

math.AP↗

Calibrating Prediction Timeliness Through Multi-Objective Hyperparameter Optimization for Remaining Useful Life Prediction

In predictive maintenance, early and late RUL prediction errors carry asymmetric consequences, yet hyperparameter optimization typically targets a single accuracy metric that treats both directions equally. This study treats the optimization objective itself as a design variable. Five architectures (MLP, LSTM, XGBoost, TCN, and Transformer) are evaluated under three regimes: single-objective maximization of $R^2$, single-objective minimization of the NASA scoring function, and a multi-objective formulation that jointly optimizes both criteria. The multi-objective search employs NSGA-II with Entropy-CRITIC weighting for Pareto selection. Seventy-five model-dataset-strategy combinations are assessed on the NASA C-MAPSS turbofan and BackBlaze hard-disk drive benchmarks. On C-MAPSS, all strategies achieve comparable accuracy ($R^2 \approx 0.89$), yet multi-objective optimization reduces directional imbalance by approximately 33%, improving calibration of early versus late predictions. Model rankings prove configuration-dependent, with simpler architectures frequently outperforming deeper temporal models. On BackBlaze, the objectives shift from complementary to conflicting, producing divergent Entropy-CRITIC weights and a substantial generalization gap (best $R^2 \approx 0.34$). These results demonstrate that the optimization objective materially shapes prognostic behavior and that multi-objective search provides a practical mechanism for calibrating prediction timeliness in RUL modeling.

cs.LG↗

Towards Reliable Vision-Language Models for Autonomous Driving

Vision-Language models (VLMs) are increasingly being explored in autonomous driving for tasks such as scene understanding, driving reasoning, decision-making, and end-to-end driving. As their role becomes more prominent, ensuring their robustness and reliability is increasingly important. In real-world conditions, visual inputs may be degraded by sensor imperfections and environmental conditions, potentially affecting both model predictions and their associated confidence. Such degradation is especially concerning in autonomous driving, where safety-critical decisions require models to make accurate predictions and recognize when their predictions may be unreliable. In this work, we evaluate five VLMs (Qwen3.5-9B, Gemma4-E4B, LLaVA-OneVision-7B, DriveFusion/DriveFusionQA-4B, and NVIDIA Alpamayo-1.5-10B) across four driving-related QA datasets with different visual input settings, including single-frame, multi-view, multi-frame, and monocular inputs. Our results show that the effects of visual corruption vary across models, datasets, and input settings, with changes in accuracy and confidence reliability and also differing across conditions. We then apply Visual Evidence Augmentation ($\mathrm{V}{\scriptstyle \mathrm{EA}}$), a recent inference-time method to examine whether it can improve model reliability under degraded visual conditions. We find that $\mathrm{V}{\scriptstyle \mathrm{EA}}$ improves performance for some models and datasets, although the gains are not consistent across all settings.

cs.AI↗

Efficient Bayesian inference for multiple network data

We investigate distributional properties of the centered Erdős--Rényi distribution (Lunagòmez et al., 2021) and propose a semi-conjugate Bayesian approach to multiple network data. In simulations, both Gibbs sampling and empirical Bayes accurately recover network summaries, with the latter scaling efficiently with network size. As a companion to this note, we provide the R package BayesCER, which implements the proposed methodology.

stat.ME↗

Gate Control Improves Routing Efficiency for Periodic Traffic

Traffic movement on highways and data packet transmission across the Internet serve as prototypical examples of collective dynamical behavior in complex systems. Although enhancing traffic flow in complex networks has been extensively investigated, most existing studies focus on single-queue models. In this paper, we introduce a simple yet effective two-queue scheme with gate control that regulates vehicle entry and integrates it with the routing protocol. When vehicles arrive at a on-ramp gate, they first join a waiting queue and are permitted to enter the main flowing queue once it has emptied. This two-queue strategy is especially advantageous when highway inflow varies periodically with large fluctuations, as is commonly observed in real traffic. Moreover, it modifies the system's failure characteristics, replacing sudden, discontinuous jamming transitions with gradual, continuous ones. The approach is based on a simple rule using only local queue information, making it easy to implement and widely applicable. Beyond the specific setup studied here, it can be applied to other transportation networks under different conditions or extended to other queuing systems, such as Internet data packet routing.

cond-mat.stat-mech↗

False Floors: LLM Safety Routing Evaluations Break Under Distribution Shift

Safety routers send each request to one of several models and are judged against the best single model. A major routing benchmark picks that comparator on the evaluation data. In the benchmark's own setting this is harmless, but under distribution shift it is not. On HELM Safety the selection cost is 0.003-0.030 of harm under random splits and 0.045-0.113 under held-out categories, comparable to the whole deficit attributed to routing, with its direction holding under either published judge alone. It rises seven- to ninefold on AgentDojo when suites are held out. Across seven safety corpora chosen by rules fixed in advance, three meet a registered interval test and four beat a later permutation null, and three of the four interval misses are corpora where some models have zero observed harm. Prior work proves the direction of this bias. We size it on harm and accuracy, show that it is larger under the held-out splits we measure, and bound it by optimism plus a shift-dependent regret. Scored honestly under shift, routing buys little on these benchmarks. In most pool cells the nested router serves the honest baseline's model, and on the nearly saturated AgentDojo corpus a perfect pre-dispatch router is worth at most two points of harm. We also find a model's expressed recognition of a late injection steerable. On held-out reruns an attacker who knows which model it faces lowers GPT-5.4's judged recognition by 19.6 points, confirmed by an independent label. In an offline counterfactual composition into a controller, the same attack raises or lowers estimated harm depending on the fallback model. Safety routing should be evaluated under shift, against a baseline chosen without the test labels, and recognition-based defences should be scored on harm against an attacker who chooses what the model sees.

cs.CR↗

An Explicit Polynomial Counterexample to Connes' Embedding Conjecture

We construct an explicit Hermitian polynomial $ f $ with integer coefficients, of degree $ 12 $ in $ 65 $ selfadjoint variables, whose normalized trace is at least $3/4$ on every tuple of selfadjoint matrix contractions, in every dimension, and equals $-1$ at a specified tuple of selfadjoint unitaries in a group von Neumann algebra. Consequently, $f+\varepsilon$ lies outside the contraction quadratic module modulo commutators for $0\le\varepsilon<1$, giving an explicit counterexample to the algebraic formulation of Connes' embedding conjecture. Combining the group construction of Kun and Thom with the normalization argument of Thom and the spectral correction theorem of Alekseev, Liu, and Thom, we determine an explicit positive integer $μ$ for which $f=1-(P Q)^2+μ\sum_{ν=1}^{825}E_ν^*E_ν+μ\sum_{j=1}^{65}(1-X_j^2)^2$. Here $P,Q$ encode conjugate involutions, and the $E_ν$ encode relation defects.

math.OA↗

FedFit: Federated Fine-Tuning of LLMs via Vector-Bank Parameterization and Quantization

Federated Learning (FL) enables privacy-preserving fine-tuning of Large Language Models (LLMs), yet the massive communication overhead remains a critical bottleneck. Furthermore, applying Low-Rank Adaptation (LoRA) in FL faces a fundamental "aggregation dilemma" between the accurate Sum-of-Products (SoP) and the communication-efficient Product-of-Sums (PoS) implementations. To tackle these challenges, we propose FedFit. First, to significantly reduce communication overhead, we introduce a disjoint shared vector-bank parameterization that reconstructs high-dimensional adapter matrices from two compact and disjoint global vector banks. Second, to address the aggregation dilemma, we devise an alternating optimization schedule. By cycling between decoupled single-bank updates (which allow for accurate aggregation) and joint updates corrected by a Residual Spectral Aggregation mechanism, we resolve the conflict between SoP and PoS. Additionally, we integrate blockwise quantization with client-side error feedback to further compress the transmitted vectors. Furthermore, we establish theoretical convergence guarantees for the proposed algorithm. Extensive experiments on Qwen2.5 models demonstrate that FedFit achieves perplexity performance comparable to standard federated LoRA methods, while providing compression ratios up to 100x higher.

cs.LG↗

Efficient learning of quantum interactions from thermal metastable states

Learning quantum interactions from finite-temperature many-body systems is a central task in emerging quantum platforms. Recently, the problem of learning from lattice quantum Gibbs states has found rigorous, efficient protocols. Nevertheless, exact Gibbs states, as the input premise, are in fact computationally intractable to prepare and may not faithfully represent generic finite-temperature quantum systems. In contrast, a system coupled to a heat bath can be stuck at an approximate stationary state (metastable state) long before it truly equilibrates. Here, we formulate a physically and algorithmically consistent alternative: learning from such metastable states of detailed-balanced master equations (Lindbladians) arising from system-bath interactions. We distill the algorithmic mechanism and structural condition underlying Gibbs-state learning and extend it in full to metastable states, attaining nearly optimal sample and computational complexity (in the system size and the precision). More broadly, we sharpen notions of metastability and develop a unified framework for finite-temperature learning.

quant-ph↗

The AI Assessment Sandbox Configurator: A Framework to Support Technical Assessment in AI Regulatory Sandboxes

The EU's Artificial Intelligence Act requires all Member States to establish AI Regulatory Sandboxes (AIRS) by August 2027: supervised environments bringing together national Competent Authorities, technical experts, and the organisations under assessment. When AIRS engagements include structured technical testing, running such testing at scale demands dedicated infrastructure, yet the tooling ecosystem remains structurally fragmented, with heterogeneous tools producing outputs that are difficult to compare, trace, and reuse. From the procedural conditions of AIRS engagements and the AI Act obligations for high-risk systems, we derive 11 architectural and governance requirements for the infrastructure that operationalises technical testing within an AIRS. In response to these requirements, we introduce the AI Assessment Sandbox Configurator, an open-source framework combining a curated Catalogue of tests and controls accessed through a stable plug-in API, a shared data model that harmonises heterogeneous outputs, role-specific dashboards for multi-disciplinary interpretation, and audience-segmented reporting. We describe the architecture and current release, and report an early-stage pilot that exercised the harmonisation and reporting layers within a live AIRS engagement and contributed to an official Exit Report. We discuss the roadmap, the governance questions raised by the Catalogue's tiered contribution model, and the institutional pathways through which an open-source assessment ecosystem could emerge across Member States.

cs.AI↗

Measuring Geometric Phase based on Indefinite Causal Order in a Sagnac Interferometer

Sagnac interferometers are important tools for precision measurements. Cancellation of common-path noise ensures a high degree of stability against external perturbations, while their sensitivity to rotations provides the basis of laser gyroscopes. Here we highlight another, less explored feature of Sagnac interferometers: the fact that optical elements are traversed in reverse order for clockwise and counterclockwise propagating light components. This provides the opportunity to explore a classical realization of indefinite causal order, where the order of events within a sequence is not fixed. We explore this concept to identify the geometric phase associated with two non-commuting polarization operations with a single measurement. This idea may have applications for the rapid determination of polarization manipulations or optical activity within chiral media, but foremost it provides a geometric illustration of indefinite causal order.

physics.optics↗

Weighing Galaxies Inside-Out: Small-Scale Lensing and the Stellar Mass Problem

Small-scale galaxy-galaxy weak lensing provides an independent way to probe the stellar initial mass function by constraining the matter distribution within galaxies. We forecast the potential of upcoming \textit{Euclid} observations using realistic mock catalogues designed for a DESI-like Bright Galaxy Survey, which is expected to cover approximately $9000\,\mathrm{deg}^2$ of the \textit{Euclid} wide survey area. We assess the impact of foreground lens light, which introduces significant scale-dependent biases in source detection, photometry, and shape measurements. Subtracting the lens light with \textsc{Galfit} reduces these biases by nearly an order of magnitude at scales down to four times the half-light radius. We model the predicted lensing signal around central galaxies with stellar masses between $9.5 \leq \log(M_*/h^{-2}M_\odot) \leq 11.5$ over projected separations of $10$ to $100\,h^{-1}\mathrm{kpc}$. The resulting signal-to-noise ratios reach $\sim 150$ for intermediate-mass galaxies. Assuming an NFW profile for the dark matter contribution, we constrain the stellar mismatch parameter $α$ to a precision of approximately $8.5$ percent for the full sample and $8.6$ percent for massive galaxies. These results highlight the potential of small-scale galaxy--galaxy weak lensing as a robust probe of the IMF and the connection between stellar and dark matter in galaxies using upcoming \textit{Euclid} data.

astro-ph.GA↗

Synthetic training for long-tail haemorrhagic lesion segmentation in data-scarce settings

Cerebral microbleeds (CMBs) and cortical superficial siderosis (cSS) are imaging markers of cerebral small vessel disease, but their automated segmentation is limited by the scarcity of positive cases and voxel-level annotations. We propose a synthetic training framework for long-tail haemorrhagic lesion segmentation that requires no real lesion annotations for training and leverages radiological description of the lesions. Starting from anatomical brain parcellations, the framework applies spatial augmentation and voxel resampling, procedurally inserts cSS and CMB labels using clinical priors on lesion location and morphology, and synthesises images through randomised intensity assignment, blurring, and Rician noise simulation. Models were trained on dynamically generated image-label pairs and evaluated against manual delineations in 10 cSS cases and 13 CMB cases. The proposed configurations outperformed classical filter baselines. For cSS, the hypointensity constrained model achieved higher AUPRC and AUROC than the Frangi filter (AUPRC: 0.284 vs 0.083; AUROC: 0.907 vs 0.731). For CMBs, explicit synthesis of blood vessels as lesion mimics improved performance over the classical baseline (AUPRC: 0.538 vs 0.004; AUROC: 0.999 vs 0.968). These results support our proposal as a feasible strategy for data-scarce haemorrhagic lesion segmentation.

cs.CV↗

A groupoid model for Rørdam's finite-infinite $C^*$-algebra

We construct a groupoid model for Rordam's example of a simple, nuclear, separable $C^*$-algebra containing a finite and an infinite projection. This groupoid model arises from a (crossed product of a) carefully chosen model for the inductive limit construction in Rordam's original approach, which we show to yield a Cartan subalgebra of the limit. In particular, this procedure yields a locally compact, second countable, Hausdorff, étale, amenable, minimal and topologically principal groupoid whose unit space is compact, and whose $C^*$-algebra contains an infinite and a non-zero finite projection. Hence, this groupoid is amenable, and yet does not have comparison.

math.OA↗

Revisiting Cross-Reconstruction for Generalizable Deepfake Detection

Existing image forgery detectors often suffer from generalization to unseen manipulation methods due to the limited ability to capture transferable forensic cues. Recent cross-reconstruction based methods attempt to improve generalization through semantic-artifact disentanglement, but typically align heterogeneous artifacts across generators and exclude artifact representations during reconstruction, which may overlook the inherent diversity and visual cues of manipulation artifacts. In this work, we revisit cross-reconstruction and introduce an artifact-oriented disentanglement framework for robust image forgery detection. We argue that \textbf{artifact diversity}, i.e., the intrinsic variations of manipulation artifacts introduced by different generation processes, contains complementary forensic cues rather than undesirable domain variations. Instead of enforcing explicit artifact alignment, our framework preserves diverse artifact characteristics through semantically aligned cross-generator reconstruction. Furthermore, we incorporate artifact representations into the reconstruction process and introduce a masked frequency-aware reconstruction strategy to emphasize manipulation-related residuals while reducing semantic interference. This design enables the model to learn transferable forensic representations from diverse artifacts. Extensive experiments on multiple benchmark datasets demonstrate improvements under both cross-dataset and cross-generator evaluation settings. Further analysis and ablation studies validate the effectiveness of artifact diversity preservation and artifact-aware cross-reconstruction.

cs.CV↗

Vector- and operator-valued backward stochastic equations with finite-variation drivers and a maximum principle for singular stochastic control in infinite dimensions

We study a mixed regular--singular control problem for stochastic evolution equations in a Hilbert space with possibly unbounded random linear operators, a nonconvex regular-control domain, and a state-dependent singular coefficient. The singular control is an adapted nondecreasing càdlàg process whose terminal value need not be bounded. We prove well-posedness and weighted moment estimates for the forward and backward equations, and characterize the second-order adjoint by a conditionally expected operator-valued backward stochastic integral equation. An Itô-type formula for the quadratic form of this adjoint, together with spike and convex variations, yields a second-order Hamiltonian condition for the regular control, as well as nonnegativity and a contact condition for the optional singular Hamiltonian.

math.OC↗