arXiv Science⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 541 records · Page 30Linked to original sources

Investigation of Widefield NV-Center Magnetic Imaging for Non-Destructive Materials Testing

Due to progressing miniaturization, material fatigue is expected to gain importance on microscopic scales. Nevertheless, most non-destructive testing techniques are optimized for macroscopic scales. Therefore, new approaches applicable to miniaturized samples need to be explored. High spatial resolution imaging of a material's magnetic stray field enables sensitive detection of microstructural changes due to the local interplay of magnetic properties, strain, and defects. Up to now, this approach has rarely been used for materials testing since high-resolution magnetic sensing is traditionally challenging. Novel quantum sensing techniques offer the potential to close this gap. One of the most prominent quantum sensors, the nitrogen vacancy center in diamond, is investigated for non-destructive testing applications within the scope of this work. An experimental system based on a widefield sensing approach was developed to image magnetic stray field distributions within seconds to minutes over areas up to 1 x 1 $mm^2$. Magnetic sensitivities below 10 $μ$T/$\sqrt{\text{Hz}}$ and a spatial resolution of 1-2 $μ$m were achieved. The technique offers high mechanical stability, which is crucial for non-destructive testing applications. Comparison with magneto-optical Kerr effect measurements showed that the magnetic field maps are strongly correlated with the material's surface domains while containing additional information from deeper inside the sample. Characteristic changes in the magnetic stray field distribution were detected in an electrical steel sample after cyclic loading. A potential marker for early fatigue damage was deduced from the splitting gradient distribution, while 2D Fourier transform analysis provided additional insight.

physics.ins-det↗

A Bayesian framework for multilevel data under model mis-specification

We propose a Bayesian framework for uncertainty quantification from the perspective that the working model is mis-specified in settings of a multilevel data-generating process. We focus on settings in which the mis-specification fails to match the functional form of the mean structure, and discuss Bayesian estimation of target parameters under dependence induced by a mismatch between working and data-generating models. The proposal represents a Bayesian semi-parametric procedure aimed at estimating population-level parameters while accounting for cluster- and unit-level variation in the estimating function. The proposal extends the regular Bayesian bootstrap to account for cluster- and unit-level variation using multilevel weights from an enriched Dirichlet model. Simulation studies indicate that the proposed approach has good frequentist properties when the data-generating process and the proposed model induce a partially exchangeable sequence associated with the unknown quantity of interest. Applications to radon (Gelman and Hill, 2007), Programme for International Student Assessment 2022 (OECD, 2023), and tuberculosis (Nobre et al., 2023) datasets are presented for illustrative purposes. The results demonstrate that the proposed method is competitive with variations of multilevel models, with major differences observed in the range of credible intervals, which are justified by the nonparametric assumptions underlying the proposed method.

stat.ME↗

Scaling of Capability and Efficiency at Inference Time in Large Reasoning Models

Capability and efficiency are two key dimensions of reasoning in large language models (LLMs). Capability refers to the ability to solve a given problem correctly, whereas efficiency refers to the ability to do so with limited resources. When LLMs use Chain-of-Thought (CoT) reasoning to solve problems of controlled hardness, both the number of problems solved correctly and the number of tokens required to reach a correct answer depend on problem hardness and model size. However, how these factors jointly shape capability and efficiency remains poorly understood. Here, we use hierarchical Bayesian models to evaluate the capability and efficiency of LLMs from the DeepSeek-R1-Distill model family across four classes of arithmetic and algorithmic reasoning problems. At a fixed model size, the probability of correctly solving an instance decays approximately exponentially with instance size, our proxy for problem hardness. The decay scale grows sublinearly with model size, indicating that larger models are more capable, but that capability gains diminish with scale. Output length grows as a power law with instance size, which serves as a proxy for difficulty. However, the parameters of this power law do not vary systematically with model size, suggesting that larger models do not become more efficient. Together, these findings reveal potential limitations of naive scaling as a strategy for developing more capable AI systems: capability improves with diminishing returns, while efficiency shows little to no improvement.

cs.LG↗

LOCKR: A Hidden-State Trajectory-Guided Planner for Detecting and Repairing Stable-but-Wrong Lock-In in Diffusion Language Models

Diffusion language models generate text through iterative denoising, exposing intermediate trajectories before final answers are produced. We identify a recurring reasoning failure, stable-but-wrong lock-in, where an answer stabilizes early around an incorrect value while substantial denoising remains. Surface-level decoding signals such as confidence, entropy, margin, and answer stability are insufficient to reliably distinguish correct from erroneous lock-in. We formulate selective reasoning repair as a lightweight test-time planning problem and propose LOCKR, a hidden-state trajectory-guided planner that decides when to allocate additional computation, expands a structured set of targeted repair branches, and selects the most promising continuation using trajectory-aware verification. Across two diffusion language models and three mathematical reasoning benchmarks, hidden-state trajectories consistently outperform surface signals and single hidden snapshots for both wrong-lock-in detection and repair selection. On natural evaluation distributions, LOCKR yields absolute accuracy gains of 2.21--5.37 percentage points across all five evaluated settings, with repair rates ranging from 22% to 41%. These results establish hidden diffusion trajectories as actionable signals for selective test-time reasoning repair.

cs.CL↗

One-Step Voice Conversion by Learning kNN Transport in WavLM Space

Voice conversion (VC) systems fall into two families: non-parametric embedding-space methods, which need no trained model but degrade on short target utterances, and spectrogram-based neural architectures, which achieve strong quality via multi-module pipelines with tens of millions of parameters. We propose kNN-FM-VC, a single conditional flow-matching network that learns to approximate the kNN-VC mapping between WavLM embedding distributions of source and target speakers, replacing explicit pointwise kNN matching with a neural regressor trained on kNN-generated pairs. The model is conditioned on the target speaker via cross-attention and FiLM, and trained under three Gaussian conditional paths (Schrödinger bridge, straight line, and constant-variance Gaussian tube), enabling few-step sampling. Unlike Phoneme Hallucinator, which uses an upsampling stage followed by kNN matching, our 13M-parameter model performs conversion with a single learned network and supports one-step inference. On LibriSpeech, the one-step Gaussian Bridge achieves lower WER and higher estimated speech quality than FreeVC and Phoneme Hallucinator. Relative to kNN and kDOT, it substantially reduces WER.

eess.AS↗

What fidelity metrics miss: a structural check on synthetic educational data

Secondary use of educational records is increasingly mediated by platforms that share a differentially private synthetic version of a dataset and validate specific findings against the real data on request. The synthetic version is evaluated by comparing summary statistics of each variable, yet reported confirmation rates suggest that such comparisons do not predict which findings survive. We propose a structural check: the number of connected components of a weekly proximity graph over learners, tracked across a term. Across four annual cohorts of lower-secondary study-habit logs, the synthetic versions reproduced the level of this quantity and the shape of the weekly partition, but its variation across the term was between 2.6 and 4.9 times smaller than in the real data at a common working point, without exception, and those changes fell in different weeks: the synthetic cohorts single out the term's examination weeks and the real cohorts do not. We also show that a routine rule for setting the graph threshold makes naive comparisons between two datasets invalid, and illustrate this with an error of our own. The real curves are also distinguishable from marginal-preserving surrogates of themselves in all four cohorts, where three of the four synthetic ones are not, a comparison that needs no real data; these differences trace to what the generator was given.

cs.CY↗

A note on bistability of a two-gene competitive system

Self-regulation together with mutual promoter competition provides a simple mechanism for bistability in gene-regulatory models. We study a two-gene system with regulatory terms of Hill exponent one, allowing distinct basal production rates and distinct degradation rates. Each gene product, when bound to its own promoter, may enhance or reduce production relative to the basal rate, while the two products compete through promoter occupancy. We show that the system has at least one and at most three equilibria in the positive quadrant. Exactly two positive equilibria can occur only if one nullcline intersection is degenerate; consequently, a configuration in which all positive nullcline intersections are transverse has either one or three positive equilibria. If there are exactly three distinct positive equilibria, then all three are automatically hyperbolic: the two outer equilibria are asymptotically stable nodes and the middle equilibrium is a saddle. Moreover, every positive solution converges to an equilibrium. Hence the positive quadrant is the disjoint union of the basins of attraction of the two stable nodes and the one-dimensional stable manifold of the saddle, yielding global bistability.

math.DS↗

The Power of Recruiting the Smaller Side: Two Additional Traders Suffice in Two-Sided Markets

We study Bulow-Klemperer-style competition complexity in two-sided double auctions with $m$ unit-demand buyers drawn i.i.d. from $F_B$ and $n$ unit-supply sellers drawn i.i.d. from $F_S$. When $m \ge n$ and buyer valuations first-order stochastically dominate seller costs ($F_B \succeq_{\mathrm{FSD}} F_S$), we prove that recruiting just two additional sellers enables Seller Trade Reduction (STR), a prior-independent mechanism, to achieve expected Gains From Trade (GFT) at least the first-best GFT of the original market. When the buyer side is the smaller side of the market ($m \le n$), an analogous result holds for Buyer Trade Reduction with 2 additional buyers. This resolves open questions of Babaioff, Goldner, and Gonczarowski (SODA 2020) and Cai, Liaw, Mehta, and Zhao (STOC 2024). We complement our upper bound by showing that this uniform bound is optimal: already for $m = n = 1$, no prior-free mechanism (deterministic or randomized) that is dominant-strategy incentive-compatible, individually rational, and weakly budget-balanced can match the first-best GFT by recruiting only one additional seller.

cs.GT↗

Maximal Monotone Differential Inclusions with Volterra and One-Sided Lipschitz Perturbations under Nonlocal Initial Conditions and Applications

We investigate a class of differential inclusions governed by non-autonomous and autonomous maximal monotone operators, involving set-valued perturbations with a Volterra integral term and subject to a nonlocal condition. Under suitable assumptions on the governing operator, the set-valued perturbation, and the Volterra kernel, we establish existence results for solutions. In particular, the set-valued perturbation is assumed to satisfy a one-sided Lipschitz condition, while the Volterra kernel is required to be Lipschitz continuous. The existence of a trajectory is established by means of an iterative construction and an application of Zorn's lemma. Finally, several examples are presented to illustrate the applicability of the abstract results.

math.OC↗

A Bulletproof Business? Towards Detecting Infrastructure-as-a-Service Offerings on Telegram

Cybercriminal operations increasingly depend on reusable digital infrastructure---including hosting, proxies, and virtual private networks (VPNs)---rented through Cybercrime-as-a-Service markets and advertised on platforms such as Telegram. We present a taxonomy for identifying Telegram messages advertising cybercriminal Infrastructure-as-a-Service (IaaS). The taxonomy comprises six service categories across compute, network, and communication infrastructure, together with three trust attributes: Bulletproof, Payment Security, and Transparency. Using 261 human-annotated messages, we evaluate keyword-based and TF--IDF classifiers and examine prompt-based large language models as exploratory baselines. We select a TF--IDF pipeline and apply it to 1,116,071 messages from 167 cybercrime-related Telegram communities. The pipeline assigns at least one infrastructure category to 207,244 messages (18.57%) spanning 113 communities. Classified advertising is highly concentrated: a single community accounts for 50.3% of infrastructure-positive messages, while the trust-attribute classifiers identify Bulletproof claims in 37.66% of those messages. These findings characterize the scale, composition, and concentration of infrastructure advertising on Telegram and can inform the prioritization of communities and actors for monitoring and investigation.

cs.CR↗

Latent evolving World Action Model

World Action Models (WAMs) jointly model action generation and environment dynamics and are mostly built on pretrained Video Diffusion Models (VDMs). In VDM-based WAMs, observations are first encoded by a VAE, and the resulting compressed latents are then processed by large video diffusion backbones to extract effective features for action generation. However, this paradigm ties WAM performance and training cost to large-scale video generation pretraining, limiting WAM efficiency and scalability. In this paper, we theoretically and empirically investigate how visual representations affect action generation in WAMs. Our results show that predictive embeddings from Joint-Embedding Predictive Architecture (JEPA) encoders better support action generation than compressed VAE latents, with I-JEPA performing best in our encoder comparison. Based on these findings, we propose LeWAM, which conditions action generation on JEPA embeddings and models environment evolution by predicting future embeddings in the same space, without relying on a video diffusion backbone. We further find that imitation learning matches demonstrated actions but does not distinguish better actions from worse ones, even though small action deviations can greatly affect task success. To address this limitation without additional environment interaction or the human oversight required for resets and safety, we introduce Demonstration-Guided DPO (DemoDPO), an offline preference refinement stage that derives preference supervision directly from demonstrations. With only 0.4B trainable parameters, LeWAM achieves an average success rate of 92.28\% on RoboTwin 2.0, comparable to that of state-of-the-art VLAs and WAMs, and maintains practical effectiveness on real-world manipulation tasks.

cs.CV↗

Landau singularities and convex geometry

We revisit the classic 1959 paper of L. D. Landau on possible singularities of the integral associated to a Feynman graph. Focusing on the real setup, we relate it to convex geometry by representing a collection $p$ of incoming momenta as the weighted normals of a convex polytope $Q$ via the Minkowski problem. Then, a regular polyhedral subdivision $\mathcal{P}$ of $Q$ exhibits $p$ as a Landau singularity labelled by the dual graph of $\mathcal{P}$ with masses being the areas of the faces. Positivity of the Landau/Feynman/Schwinger multipliers is interpreted as strict convexity of a PL-function. This gives an interesting class of ``polyhedral'' Landau singularities. In the planar case going back to the original 1959 paper, Landau graphs can also be identified with plane webs of Gaiotto-Moore-Witten that provide a language dual to that of regular polygonal subdivisions.

math.AG↗

ProCredit: From Outcome Rewards to Progress Credit in Agentic Reinforcement Learning

Long-horizon agentic tasks require an agent to modify an environment through a sequence of tool calls, with success determined by the final state. The standard recipe assigns a single outcome reward at the end and compares trajectories sampled for the same task. As a result, a group with no successful trajectory yields no training signal, failed attempts cannot be told apart by how close they came to completion, and turns that advance the task receive the same credit as turns that only query the environment. Prior work refines the unit of comparison from the trajectory to the step, or trains a reward model to supply intermediate signal: the former still derives its signal from final success alone, and the latter estimates it with a model. We observe that the acceptance checks that decide success can also be run on intermediate states, so progress is as verifiable as the outcome. We propose ProCredit, which turns this verified progress into credit: it reruns the acceptance checks after each turn, rewards the turn by its change in progress, and uses these rewards to assign credit both across attempts at the same task and across the turns within a trajectory. Starting from Qwen3.5 base models at three scales on AppWorld, ProCredit outperforms outcome-reward baselines and progress-based baselines in task completion rate at every scale on both test sets, exceeding the strongest outcome-reward baseline by 4.1 percentage points at 4B, and results in a second environment show the same direction of improvement. Ablations show that adding the final progress to the trajectory score alone does not improve performance: the gain comes from crediting progress to the turn where it occurs.

cs.LG↗

CCR: Towards a Common, Quality-Gated CACAO Integrations Registry for European Cybersecurity Automation

Standardised, machine-readable cybersecurity playbooks provide a basis for portable, shareable, and reusable incident-response logic. OASIS CACAO provides a vendor-neutral representation for such playbooks, but not the product-specific integration artefacts needed to invoke external products and services. We introduce the Common CACAO Registry (CCR), an open, provenance-aware registry of CACAO HTTP-API connector envelopes. Each envelope captures an API operation's command, inputs, target, authentication-related information, provenance, validation evidence, and maturity metadata. CCR is quality-gated, with acceptance requiring both CACAO v2 schema validity and a mean back-validation score of at least 0.8 against the source OpenAPI operation, while a six-level maturity model records progressively stronger evidence and distinguishes gate acceptance from operational readiness. To seed CCR, we develop a hybrid OpenAPI-to-CACAO pipeline. Deterministic code extracts source-derived interface facts, generates identifiers, wires cross-references, and validates structure, while a constrained LLM provides bounded semantic enrichment, including action naming, authentication interpretation, and CACAO activity annotation. Evaluation across eight security APIs yields 713 CACAO-schema-valid envelopes with a mean back-validation score of 91.5%, of which 675 produce well-formed, dispatchable HTTP requests in a local harness. Comparison with a deterministic rule-based baseline shows that mechanical API structure is preserved more reliably through rule-based translation, while the LLM contributes bounded semantic enrichment, most notably CACAO activity annotation. Together, these results support CCR as reusable integration infrastructure for CACAO action steps and as an initial foundation for a broader common European registry.

cs.CR↗

SHRAV: State-Hypothesis-Reason-Action-Verify Framework for Physical Modeling and Inverse Design

Physical modeling and inverse design require computation that can continue from reusable state. We introduce SHRAV, an architecture-independent computational framework organized around State, Hypothesis, Reason, Action, and Verify. Its central mechanism is a state-continuation core with declared reuse boundaries and explicit roles for learned evolution and numerical quantities. Forward configurations evolve predictive state and read out physical responses; inverse-design configurations additionally generate target-directed modifications and consume evaluator feedback. Electromagnetic world-model studies are mapped to forward configurations, with selected readout and reuse diagnostics reported here. Computational lithography demonstrates an inverse-design configuration: four fixed-weight design updates improve thresholded aerial-image intersection-over-union from 0.5313 to 0.8153 under independent scalar-pupil replay, with a maximum absolute IoU difference of approximately 0.000824 between predictor estimates and independent replay.

cs.AI↗

Brain-to-Language Decoding: Tasks, Signals, Methods, Evaluation, Practical Use and Beyond

Brain-to-language decoding translates neural activity associated with language production, internal speech and perception into linguistic or expressive outputs. It offers a route to restoring communication after speech loss and a means of studying how the brain represents language. Advances in neural recording and representation learning have expanded the field from constrained recognition and acoustic reconstruction to text generation, streaming personalised speech and facial animation. This survey synthesises these developments across invasive and non-invasive measurements, drawing on a search without a lower year limit and source-led updates through September 2026. We connect Articulated, Inner and Perceived tasks to the neural populations they engage, the representations available to decoders and the outputs those representations can support. We examine model development, public resources and the evolution of evaluation, and compare published performance and communication costs within their reported protocols. The synthesis identifies complementary routes to progress: phonetic, acoustic and semantic targets preserve different aspects of a message; shared representations support reuse across recording conditions and tasks; and online communication increasingly depends on calibration, feedback and user control alongside decoding accuracy. Shared benchmarks enable algorithmic comparisons, while longitudinal studies reveal the demands of sustained use. We discuss these developments and their remaining limitations, then outline a prospective five-level trajectory from commands and language to meaning, scenarios and bidirectional cognitive exchange

cs.CL↗

Mixing profile for Glauber dynamics of the discrete Gaussian Free Field starting from super-harmonic functions

We study the convergence rate of the heat-bath Glauber dynamics for the Discrete Gaussian Free Field on arbitrary connected finite graphs. We show that, when starting from super-harmonic initial conditions, the evolution enjoys a strong form of monotonicity. This allows us to get a sharp mixing profile as the size of the graphs diverges. More precisely, we show that mixing occurs at time $\frac{1}{2λ}\log(\mathcal{E})$ with window $\mathcal{O}(1/λ)$, where $λ$ is the spectral gap of the graph Laplacian and $\mathcal{E}$ is the energy of the super-harmonic initial condition. This result holds for arbitrary graphs that do not exhibit extreme connectivity properties (one way or the other). In particular, it holds for finite boxes of the grid $\mathbb{Z}^d$, in dimension $d\geq 3$.

math.PR↗

Consequential Behaviour and Representational Fairness in the Validation of Synthetic Research

Researchers in industry and academia use synthetic survey respondents powered by large language models as substitutes for human samples. These synthetic populations require validation against real-world data, so researchers often address them using ad hoc comparisons with human surveys. Inspired by the intention-behaviour gap in behavioural science, we argue that these validations test the wrong thing for most applied cases where decision makers commission synthetic research to anticipate consequential behaviour. To address this problem, we propose a validation framework with two requirements. First, every validity claim must state its level of correspondence with human data: does the sample predict what the represented people do, which of four diagnostics (location, dispersion, response process and structure) does the validation address, and does the validation compare against experimental effects? Second, researchers must report validity claims for subgroups, since these groups are often the most affected by consequential decisions and aggregate accuracy hides their misrepresentation. Our validation framework operationalises three justice dimensions (distributional, procedural, and recognition) as measurable quantities and defines within-persona counterfactual experiments as a validation requirement. We then apply the framework to electric vehicle charging tariffs, before closing with a reporting checklist that researchers can use to make convincing validity claims.

cs.CL↗