arXiv Science⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 343 records · Page 19Linked to original sources

CONSISTRE: A Unified Consistency-Aware Framework for Document-Level Relation Extraction with Large Language Models

Document-level relation extraction (DocRE) aims to extract relations among multiple entities across extended contexts while maintaining consistency across predicted triples. Although large language models (LLMs) show remarkable reasoning capabilities in information extraction, their predictions are typically generated independently for each candidate triple and may violate fundamental relational constraints such as transitivity, symmetry, and functional uniqueness, leading to contradictory and unreliable outputs. We propose CONSISTRE, a unified consistency-aware framework for DocRE that addresses this limitation through two complementary tracks. The first operates at inference time for black-box LLMs, combining constraint-aware prompting, constraint-based verification, and iterative self-reflection to refine predictions without task-specific fine-tuning. The second injects consistency knowledge into smaller open-source models via a knowledge distillation and reinforcement learning pipeline: reasoning traces from a powerful teacher are distilled into a student via supervised fine-tuning, followed by GRPO alignment using a composite reward that jointly optimizes extraction performance and relational consistency. Together, the two tracks cover both API-accessible and locally deployable scenarios under a unified consistency formulation. Experiments on DocRED show that both tracks outperform their baselines, with the inference-time track achieving competitive F1 using off-the-shelf black-box LLMs and the training-time track substantially narrowing the gap between 7--8B open-source models and state-of-the-art proprietary LLMs at a fraction of their inference cost. Ablation studies confirm that explicit consistency modeling mitigates relational contradictions and enhances the reliability of LLM-based DocRE across both deployment paradigms.

cs.CL↗

A Bogoliubov-ratio framework for quantum-information diagnostics of time-dependent two-mode Boson System

We formulate a compact dynamical representation of quantum-information diagnostics for time-dependent two-mode bosonic systems in terms of the Bogoliubov ratio $λ_k(η)=β_k(η)/α_k(η)$. For vacuum-evolved pure Gaussian pair states generated by a Hermitian quadratic Hamiltonian, $λ_k$ obeys a closed complex Riccati equation. After one partner mode is traced out, the reduced-state spectrum is determined by $q_k=|λ_k|^2$, from which the purity, linear entropy, Rényi-2 entropy, and von Neumann entropy follow directly. These Gaussian results are mathematically equivalent to those obtained from the standard density-matrix or covariance-matrix formalism; the advantage of the ratio representation is a compact separation between model-dependent dynamics and universal state diagnostics. Writing $λ_k=\sqrt{q_k}e^{iθ_k}$ further exposes how pair production drives radial growth, whereas frequency rotation and detuning act through the phase and can suppress coherent squeezing accumulation. We also establish an exact extension to non-Gaussian $SU(1,1)$ sectors. For a lowest-weight Fock seed with conserved number difference $d$, the same ratio equation governs the evolution, while the reduced spectrum becomes a negative-binomial distribution determined by $(q_k,d)$. The Gaussian formulas are recovered at $d=0$. Cosmological perturbations and a chirped-pulse optical parametric amplifier illustrate the common dynamical mechanism and its information-theoretic consequences.

quant-ph↗

Global Convergence of DGM and PINN Algorithms for Solving Nonlinear PDEs

The Deep Galerkin Method (DGM) and Physics Informed Neural Networks (PINNs) have become widely-used methods for solving partial differential equations (PDEs) in the rapidly growing field of scientific machine learning. In these methods, a neural network is trained to approximate the PDE solution by using (stochastic) gradient descent to minimize the PDE residual of the neural network. Due to the non-convexity of the PDE residual objective function, the trained neural network may, in principle, only converge to a local minimizer of the objective function (which would not be a solution of the PDE). Therefore, there is a longstanding question regarding the mathematical foundations of these algorithms, and it is highly valuable to establish that the trained neural network will converge to the PDE solution. In this paper, we consider a class of semilinear PDEs with nonlinearities in the solution and its first derivative. For this class of PDEs, we prove that neural networks trained with gradient descent to minimize the PDE residual objective function will converge to the PDE solution as the network width and training time $\rightarrow \infty$.

cs.LG↗

Femtoscopy Measurement with S$π$RIT TPC in Radioactive BeamHeavy-ion Collisions

Femtoscopy is a powerful tool for exploring the dynamic emitting structure in heavy-ion collisions, while radioactive beam heavy-ion collisions enable the investigation of nuclear matter under extreme isospin conditions. Here, we successfully perform femtoscopy measurements using the S$π$RIT Time Projection Chamber (TPC). A dedicated correction scheme for track merging and splitting is proposed, which is well applicable to rectangular TPCs housed inside dipole magnets and effectively improves the reconstructed correlation functions at small relative momenta. Focusing on the proton-proton (p-p) correlation function in the 270 MeV/u $^{132}\text{Sn}+^{124}\text{Sn}$ system, we successfully apply the track merging and splitting correction; additionally, the TPC angular acceptance exhibits a negligible impact on the correlation function. A systematic uncertainty quantification framework is established. The experimental results of the p-p correlation function confirm the feasibility of the S$π$RIT TPC for femtoscopy measurements and provide technical support for high-precision femtoscopy studies using rectangular TPCs in radioactive beam heavy-ion collisions.

physics.ins-det↗

Chart-Supported or Model-Supplied? Examining MLLM-Generated Claims for Accessible Visualization

Multimodal large language models (MLLMs) can connect visualization patterns to external causes, consequences, and domain knowledge, but the evidential basis of these interpretations is often unclear. We present an exploratory study of 102 visualizations from four sources, three MLLMs, and four input conditions that vary access to the image, accessible chart context (non-image artifacts such as data tables, captions, alt text, and screen-reader structures), and withheld-context framing. Across 1,224 descriptions, we analyze model-attributed DIRECT, DERIVED, and SPECULATIVE labels and conduct an automated audit of numeric agreement. Accessible chart context shifted Gemini and GPT toward DIRECT claims and improved numeric agreement for some models. Adding the image to the full context did not yield a consistent numeric benefit, and the withheld-context prompt did not reliably increase cautious language. The prompt-defined Real-World Significance section remained predominantly SPECULATIVE. These results motivate accessible description systems that distinguish claims supported by supplied evidence from model-supplied interpretation.

cs.AI↗

Three Failures of Pain Location: Why Its Diagnostic Utility Is Three Quantities, Not One

Patient-reported pain location is diagnostically decisive for some presentations and nearly uninformative for others. The prevailing account treats this as one gradient of diagnostic utility set by anatomical complexity. That explanation conflates three epistemically distinct failures, each with its own mathematics, its own optimal instrument, and its own public-health consequence. In anatomical multiplexing, many structures share one location: a non-identifiable inverse problem. In delocalized amplification - clinically, central sensitization or nociplastic pain - a centrally driven pain-behaviour pattern replaces the peripheral generator: a change of generative model. In referred and atypical displacement, location is hypothesized to shift in a systematic, person-dependent way: a group-conditional bias whose direct evidence is still open. The three are one Bayesian inference problem failing at different nodes - the likelihood, the model class, and the group-conditional prior - with a fourth node at the report itself. The formal development is in a companion paper; this paper states what each model shows and what follows clinically. Re-examination finds that the published "high-utility" accuracy band leans on overstated specificity (Lipton et al., 2003; Bruyninckx et al., 2008; Devillé et al., 2000), so the gradient is real but flatter than drawn. The well-evidenced finding that better detection alone does not improve outcomes when treatment uptake lags is about the care pathway, not perception - a distinction the three-way split makes visible and a single utility number hides. The paper organizes the failures along a why-location-fails axis, distinct from the nociceptive/neuropathic/nociplastic taxonomy (Kosek et al., 2016), and sets out the study that would test the one prediction still open.

q-bio.NC↗

Tensionless limit of M5-brane and conformal symmetry of its bosonic body

We obtain a tensionless limit of the M-theory 5-brane action and discuss its properties. The consistency of this limit requires that the field strength $H_3$ of the 2-form gauge field on the M5 brane worldvolume is non-vanishing and non-degenerate, so that one can also regard it as a strong $H_3$-field limit. The induced $6d$ worldvolume metric of this theory is non-degenerate in contrast to conventional tensionless (null) $p$-branes. We find that the tensionless M5-brane action is invariant under a super-Weyl symmetry acting on the bulk supergeometry, while in the purely bosonic case the action is invariant under an 11D target-space conformal symmetry, a property which is characteristic for the tensionless bosonic branes. Upon imposing a static gauge, which fixes worldvolume diffeomorphisms, this 11D conformal symmetry becomes a deformed 6d worldvolume conformal symmetry whose action on the worldvolume coordinates involves field dependent terms. When the transversal fluctuations of the 5-brane in the bulk are zero, the model reduces to the conformal non-linear chiral 2-form electrodynamics considered earlier by Gibbons and West, and by Townsend.

hep-th↗

Gaokerena: A Small Persian Medical Language Model Family

The integration of artificial intelligence into medical question-answering systems has advanced rapidly; however, research remains predominantly focused on English, leaving low-resource languages like Persian significantly underserved. To address this gap, this paper introduces Gaokerena, a novel family of compact Persian medical language models optimized for deployment on consumer-grade hardware. As a foundational step toward localized digital healthcare, we first present Gaokerena-V, developed by training a baseline model on a strategically selected subset of a newly curated 90-million-token Persian medical corpus (approximately 54 million tokens) together with 20,000 expert-vetted physician Q&A pairs (approximately 3 million tokens), for a total of 57 million new tokens. This training improved performance on a translated medical MMLU benchmark from 46.64% to 49.31%. Second, recognizing the critical demands of clinical reasoning, we developed Gaokerena-R by integrating a Chain-of-Thought approach with two novel Reinforcement Learning with AI Feedback (RLAIF) frameworks to optimize preference-based reasoning. Despite utilizing the same baseline architecture and a smaller dataset than Gaokerena-V, Gaokerena-R achieved a superior benchmark score of 52.98%. Furthermore, both models are equipped with custom-developed uncertainty heads that predict the models confidence in its responses based solely on internal hidden states. While these results demonstrate significant progress in Persian medical language modeling and proactive safety estimation, current performance levels remain insufficient for direct clinical application, highlighting the necessity for further research into robust knowledge acquisition and rigorous safety verification prior to real-world deployment.

cs.CL↗

Money Burning Mechanism Design: From Welfare to Surplus

We settle the worst-case approximability of consumer-surplus maximization in general multidimensional mechanism-design environments. We do so through two black-box reductions from welfare maximization to the agents' total utility. Our first reduction turns exact welfare maximization into a prior-free, universally truthful and ex-post individually rational mechanism that preserves at least a $1/H_n$ fraction of optimal welfare as expected consumer surplus. The guarantee holds for $n$ agents with arbitrary nonnegative valuations over a finite outcome space, where $H_n$ is the $n$-th harmonic number. The factor $H_n$ is worst-case optimal, including its constant, even for a single-item auction with a known i.i.d. prior and Bayesian incentive compatibility. Our second reduction allows existing truthful welfare approximation mechanisms to be reused for surplus maximization. For valuation classes closed under scaling, it converts any ex-post individually rational, truthful $α$-approximation for welfare with nonnegative payments into an $O(α\log(n))$-approximation for surplus. Our sharp guarantee resolves the welfare-approximation aspect of the open question of Hartline and Roughgarden [2008] on the power of money burning beyond $k$-unit auctions, and the question of Ezra et al. [2025] concerning optimal surplus guarantees for broader valuation classes. It also replaces the outcome-dependent $O(\log|\mathcal{O}|)$ guarantee of Fotakis et al. [2015] with the tight agent-dependent factor $H_n$. These results yield polynomial-time mechanisms with the exact $H_n$ guarantee for gross-substitutes. They also give prior-free, universally truthful approximations of $O(H_n\log^2\log m)$ for XOS valuations and $O(H_n\log^3\log m)$ for subadditive valuations using demand and value queries, where $m$ is the number of items.

cs.GT↗

PACE-QAOA: Physics-Constrained Quantum Optimization for Qubit-Efficient Power System Islanding

Increasing renewable-energy penetration heightens power-system variability and complicates disturbance containment. Controlled islanding mitigates cascading failures by partitioning a stressed network to limit disrupted power transfer while preserving each island's operational integrity, but this constrained partitioning problem is NP-hard. Although QAOA offers a complementary search strategy, limited near-term qubit capacity restricts conventional formulations. This paper presents a qubit-efficient hybrid quantum framework combining a physics-informed compact encoding with Lagrangian constraint handling and classical feasibility refinement. The encoding exploits grid structure while formally preserving the original feasible solution space and objective. For a fixed island count on sparse working graphs, the formulation reduces phase-separator and per-layer gate complexity from quadratic to linear scaling with system size. Tests on eight IEEE systems ranging from 9 to 89 buses and multiple quantum-provider backends produce feasible, high-quality islanding solutions under practical circuit and sampling budgets. Factorial ablation attributes resource and runtime improvements to the complementary effects of compact encoding and qubit-efficient constraint handling. Noise analysis shows stable solution quality under tested device noise, while landscape diagnostics reveal smoother, more consistently scaled QAOA cost surfaces and improved parameter-optimization behavior. These results offer a transferable approach for scaling constrained quantum optimization toward larger real-world applications on near-term hardware.

quant-ph↗

Wiring Beats Blending: Structure-Aware Compensation for Transformer Downscaling

Model families are trained size by size. Can a pretrained large model instead be converted into a smaller sibling? We study the 1.4B->410M conversion in Pythia end to end. Representations align strongly across sizes (ridge R^2=0.84); parameters align weakly. Dense weight projection is destructive; a bit-exact control places the fault in basis mixing, which breaks rotary, per-head, GELU, and LayerNorm structure. Residuals after the best-fit linear operator carry no learnable or transferable signal under shuffle controls, so conversion value lives in initialization. Matched-budget continued pre-training separates two independent levers: least-squares compensation (function lever, best zero-shot) and variance-preserving rescale (dynamics lever, best endpoints). Placement follows the architecture: compensation is well-posed exactly where no normalization sits between cut and read; norm-fronted paths take rescale. Compensation is a low-budget, token-efficiency win, not a universal one. At 30M tokens it beats the best subcloning variant on a width-reduced pair (84.0+-1.8 vs. 89.7+-3.7, 3/3 seeds) and a held-out depth-reduced pair (109.3 vs. 117.9, 3/3 seeds). Selection given the same activation statistics recovers under half of that gap (3/3 seeds): the gain is the re-fit, not the information. At 33x the budget the two reach parity (40.3+-0.3 vs. 40.3+-0.5, 3 seeds), both far ahead of from-scratch, which transfer always beats (up to 18x at low budget, narrowing at convergence and at the largest scale). At ~5x the donor scale (6.9B->1.4B) stacking both levers over-corrects, consistent with an ill-conditioned compensation solve at large width, pointing to dimension-aware regularization as a fix. The init also beats structured pruning plus distillation, the standard pipeline, at matched budget, and improves further combined with it. Code, checkpoints, and the frozen eval corpus are released.

cs.LG↗

Hulls, linear equivalence, and weighted superelliptic codes

The containment of the code of the meet $G\wedge A$ in the hull and the identity $G\vee A-D=K-G\wedge A$ exchanging meet and join are known; imposing that $G\wedge A$ be principal constructs algebraic geometry codes with one-dimensional hull. We turn that construction into a measurement. For arbitrary divisors $G$ and $A$ we compute $C_L(D,G)\cap C_L(D,A)$ exactly: it is the code of the meet together with an excess $\varepsilon(G,A)$, canonically their quotient and a subquotient of $H^1(\mathcal O(G\wedge A))$. So $\varepsilon$ vanishes exactly when the meet is non-special; otherwise it certifies that $K-G\wedge A$ is linearly equivalent to an effective divisor, at degree zero the vanishing of a single class in the Picard group: the hull detects a linear equivalence rather than being built from one. For superelliptic curves $y^n=f(x)$ both sides can be computed: their weighted plane models in $\mathbb P^2_{(1,n/c,d/c)}$, $c=\gcd(n,d)$, identify codes $C_s$ of weighted forms of degree $s$ with those of $sD_\infty$ and turn hulls into lattice counts. The range on which $\varepsilon$ is blind is an explicit interval of degrees, where $\dim\operatorname{Hull}(C_s)=cμ(s)-nδ+1-g_X$, $μ(s)=\min\{s,M-s\}$, depends only on its affine-point count. Outside it the meet and join are invariant under $s\mapsto M-s$ while $\varepsilon$ is not, so every asymmetry of the hull profile is excess and the threshold in $s$ refines the divisor class: two totally split curves of genus two, over $\mathbb F_7$ and over $\mathbb F_{11}$, present the same class at the same pair of degrees and are separated by the profile alone. If $0\leq°(G\wedge A)\leq2g_X-2$ the hull is at most $g_X+1$, so it is large only where it is blind, and over a prime field, under an explicit inequality on $(n,d,q)$, its maximum over the family is $\ell(\lfloor M/2\rfloor D_\infty)$, attained exactly on the totally split locus.

math.AG↗

Self-Evolving Coding Agents

Large language models are increasingly embedded in software engineering workflows as coding agents that can inspect repositories, invoke tools, execute tests, debug failures, and generate patches. Yet most existing coding agents remain largely static after deployment, even though software development is a dynamic, feedback-rich process in which repositories evolve, dependencies change, tests fail, and repair attempts leave reusable experience. This tension has motivated a growing body of work on self-evolving coding agents, where the agent improves its future behavior by persistently updating its framework, memory, skills and tools, components, workflow and topology, or environment and context from prior coding interactions. In this survey, we provide a structured synthesis of this emerging area. We first define the concept of self-evolving coding agents and distinguish it from conventional coding agents and general self-evolving agents. We then develop a taxonomy centered on the targets of evolution, complemented by two orthogonal perspectives: when evolution occurs and which code-specific signals drive it. We further examine the benchmarks used to measure the effect of evolution and related coding products. Across the literature, we find that executable feedback, repository-level context, and coding trajectories make software engineering a natural domain for agent self-evolution, but also introduce challenges in feedback reliability, benchmark overfitting, reversibility, system complexity, safety, cost, and generalization. By organizing existing work around these dimensions, this survey aims to clarify the conceptual boundaries of self-evolving coding agents and provide a foundation for designing more adaptive, reliable, and software-aware agentic systems.

cs.SE↗

TARL: Transaction-Aware Reliable Ledgers for Executable Memory Management in Long-Term Agents

Persistent memory helps long-term agents retain knowledge, yet a single update error can repeatedly distort future retrieval and reasoning. Most existing systems reduce memory updating to a binary Write/Hold decision, which cannot distinguish whether new information should be added, ignored, used to revise an outdated belief, rejected as unreliable, or deferred for verification. These choices may share the same binary label while producing fundamentally different memory states. We introduce TARL, a memory state update framework that maps each statement to one of five executable actions. TARL identifies the affected memory, resolves its temporal scope, compares source reliability, and updates accepted, pending, and rejected ledgers. It is further trained by comparing the memory states produced by alternative update operations, encouraging the model to select the operation that leads to the correct result. We also introduce TARL-Mem, a benchmark with fine-grained action labels and next-state targets. Across in-domain, cross-source, temporal, counterfactual, and sequential evaluations, TARL improves action prediction and state recovery, reduces memory pollution, preserves conflicting evidence, and limits cumulative corruption.

cs.AI↗

Q-CueGraph: Query-Conditioned Visual Evidence Graphs for Multimodal Reasoning

Multimodal large language models (MLLMs) can miss fine details in a full image that they recognize in a closer view. Recovering this evidence requires deciding where to look and how much surrounding context to retain. We present Q-CueGraph, a query-conditioned evidence acquisition method for frozen MLLMs. For text-rich images, it builds a reusable graph of OCR lines and layout relations. Each question activates anchors, expands them into contextual regions, and selects candidates for a single observation window. Query-conditioned object detections support natural-image search through the same region-selection and composition interface. A lightweight candidate scorer further learns which observations support correct answers from frozen-reader feedback and training answers, without evidence-box supervision. Across six benchmarks, we examine the roles of query conditioning, evidence composition, and learned answerability. With Qwen2.5-VL-7B, Q-CueGraph raises V*Bench accuracy from 0.696 to 0.832 using 19.1% of source-image area, and retains 92% of full-image ANLS on InfographicVQA using about half the image area. The analyses show that useful evidence depends on both its relevance to the question and the context available to the reader. Q-CueGraph makes these choices explicit before answer generation.

cs.CV↗

Finite-coefficient Gersten injectivity fails in ramified mixed characteristic

Let $V$ be a complete discrete valuation ring of mixed characteristic $(0,3)$ in which $3$ is a uniformizer, and put $A=V[[x,y]]/(3+x^2-y^3)$. We construct a nonzero class $a\in K_2(A;\mathbf Z/3)$ whose restriction to the fraction field of $A$ is zero. Thus Gersten injectivity for algebraic $K$-theory with $\mathbf Z/3$-coefficients fails for a two-dimensional ramified regular local ring. We also indicate the expected analogous construction for every odd prime. This counterexample does not contradict the integral Gersten conjecture but it rules out a naive reduction to finite coefficients.

math.KT↗

Multiscale Reward Hedging from Correct Demonstrations

Learning from correct demonstrations is harder than supervised learning when many answers are correct: after predicting, the learner sees one valid answer but not whether its own answer was valid, nor any reward. Existing reward-hedging guarantees consequently assume a finite reward class. We give the first horizon-free guarantee for continuous classes. The key is to hedge in one shared vote over tolerant optimality tests at every accuracy scale. A target reward has one surviving proxy per scale, and a prediction with gap above that scale doubles the proxy. This yields the simultaneous tail bound $|\{t:\ell_t>2^{-j}\}|\leq \log_2\mathcal N(\mathcal G,2^{-j-1})+j$, where $\mathcal G$ is the class of optimality-gap functions. Integrating the tails gives cumulative hidden gap bounded by a metric-entropy integral, independently of the number of rounds. Polynomial entropy $(A/ε)^d$ gives $O(d\log A)$ total gap and a fast $O(d/m)$ statistical rate. For bounded linear contextual recommendation, the result is $O(d)$ regret for arbitrary compact menus. This is the first polynomial finite bound without structural restrictions on the menus, at the price of improper prediction. Although the general vote can be expensive, it is exactly polynomial-time for one-dimensional Lipschitz parameter curves. Fixed-radius rank-two recommendation takes $O(KT^2)$ time for menus of size $K$. We also prove an $Ω(d)$ lower bound, low-rank and bounded ReLU-network corollaries, and a robust theorem that adds only the demonstrator's cumulative suboptimality. A reproducible adaptive stress test illustrates the predicted scale adaptation. After factorization, an exact MovieLens audit runs in 1.7 CPU seconds across ten users and improves mean latent gap over both a demonstrated-rating policy and a proper online baseline. The learner uses only action demonstrations and never observes a reward or a loss.

cs.LG↗

Aftab: A Progressive Design Study of Visual Encoders and Value Estimation for Replay-Free Parallelized Q-Learning

Replay-free parallelized Q-learning removes the large experience replay buffers and target networks used by conventional deep Q-learning, but the role of network architecture in this training regime remains comparatively underexplored. We investigate this question through a progressive three-phase study within the Parallelized Q-Network (PQN) framework. First, we compare eight convolutional encoder topologies on Atari-57 under a common training protocol while jointly considering performance and computational complexity. Second, we integrate Hadamax-style multiplicative feature interactions and explicit pooling into the selected encoder hierarchy. Third, with the visual representation fixed, we compare complete categorical-dueling, ensemble-dueling, and categorical ensemble-dueling value-estimation configurations. The resulting architecture, Aftab, achieves an interquartile mean human-normalized score of $6.592$ on Atari-57, compared with $2.715$ for our independently rerun PQN reference, with a game-level Probability of Improvement of $0.86$. After completing all architecture selection on Atari-57, we evaluate Aftab on Procgen Hard. Aftab achieves a terminal IQM normalized score of $0.418$ compared with $0.382$ for PQN and increases the normalized area under the learning curve from $0.216$ to $0.541$, although terminal performance remains heterogeneous across environments. These results show that visual topology, multiplicative representation, and downstream value-estimation design can substantially affect replay-free Q-learning, and that their benefits should be evaluated jointly with computational complexity. The complete Aftab framework, including model definitions, training configurations, reproducibility settings, and raw experimental logs, is open-sourced at https://github.com/tahashieenavaz/aftab

cs.LG↗