arXiv Science⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 1,567 records · Page 87Linked to original sources

Valid Stopping in Adaptive Generator-Verifier Loops

Numerous agentic workflows are based on a generator-verifier loop: a generator proposes candidates, a cheap verifier scores them, and the workflow terminates when a proposal is verified as good enough. The verifier typically proxies a more costly ground-truth oracle, and as the generator searches adaptively against it, false acceptances may accumulate. Proposals can pass the proxy but fail under the costlier ground-truth check. We study when to stop these loops while controlling the false discovery rate of the accepted proposals. Our construction introduces tools of independent interest in distribution-free statistical testing and conformal risk control, including analysis of $e$-values constructed through index betting and a novel conformal risk control procedure for non-monotone losses. We validate the approach in synthetic settings and on a protein-design benchmark.

stat.ML↗

Lyric: Wave-Domain Computing for Efficient Spoken-Digit Recognition

Low-power speech recognition requires reducing audio acquisition and feature-processing costs as well as neural inference. We present Lyric, a speech-recognition front end that uses wave-domain computing to extract features before digitization. Our proposed system uses passive acoustic resonators to separate speech by frequency and supplies their slowly varying envelopes to a compact temporal neural network. This moves spectral filtering into the acoustic structure, reducing the data and processing needed for recognition. We build a prototype and evaluate spoken-digit recognition on a speakerdisjoint AudioMNIST split across nine classifier families. The front end acquires 32 times fewer scalar samples than a 16-kHz waveform. In the EdgeSpeechNet-A comparison on a Raspberry Pi 4, we reduce preprocessing-plus-inference latency and estimated processor energy per inference by 98.6%, with a 3.58-percentage-point decrease in accuracy to 95.70%.

eess.AS↗

KESurv: A Kernel Ensemble Method for Patient-Specific Survival Prediction

Predicting patient-specific survival functions is crucial for clinicians in making informed decisions about patient care and treatment strategies. Among the various models available, the Survival Forest has demonstrated significant effectiveness in numerous scenarios. In this work, we propose an ensemble method that leverages the strengths of the Survival Forest as the master model, complemented by several base models. This ensemble incorporates the Beran estimator, a type of kernel estimator, to enhance predictions of patient-specific survival curves. We evaluated the performance of our proposed model using four distinct healthcare datasets. The results highlight the superiority of our ensemble method over baseline models in both calibration and ranking across most datasets. The findings suggest that our approach offers a more accurate and reliable estimation of patient-specific survival functions, providing a valuable tool for clinical decision-making.

stat.ML↗

Descartes' rule of signs for arbitrary fewnomial systems

We consider systems of $n$ real polynomial equations in $n$ variables with $n+k+1$ monomials. By the Gale duality of Bihan and Sottile, their positive solutions correspond to the solutions of a system of $k$ equations $\prod_ip_i^{B_{ij}}=1$ in a polyhedron $Δ\subset\mathbb{R}^k$, where the $p_i$ are affine functions, and a Khovanskii--Rolle argument bounds their number by the number of common zeros in $Δ$ of iterated Jacobians $Γ_k,\dots,Γ_1$ plus the number of noncompact branches of certain curves. We bound the first term by the Bézout number minus the numbers of zeros in the other chambers of the arrangement $\{p_i=0\}$, which we bound from below by boundary degrees given by a facet-count formula. The branches of the curves end at zeros of the Jacobians on faces of $Δ$, which we count on the flats of the arrangement through explicit reduced systems. The resulting recursion over all flats and chambers starts on lines with Descartes' rule of signs for circuits. We obtain upper bounds for the number of positive solutions which only depend on the oriented matroid of the coefficient matrix and on the oriented matroids of the exponent matrix and of its liftings.

math.CO↗

From Sparse AFM Observations to Probabilistic Macroscale Mechanics: A Generative-Physics Framework for Plant Cell Walls

Nanoscale imaging of developing plant cell walls is expensive, and a limited number of atomic force microscopy (AFM) scans cannot capture the full structural variability of the wall. We present a generative-physics workflow that links sparse AFM observations of cotton (Gossypium hirsutum) fiber cell walls at 8 days post-anthesis (DPA) to distributions of effective elastic properties and a larger-scale mechanical response. We adapt a Stable Diffusion model with Low-Rank Adaptation (LoRA) to expand the experimental scans into an ensemble of AFM-like microstructures and assess the synthetic structures using microfibril crossover count and crossover angle. Each microstructure is mapped to spatially varying Young's modulus and Poisson's ratio fields and analyzed using strain-controlled finite element homogenization. Repeating this process across the image ensemble and a prescribed sweep of matrix-to-fibril stiffness ratios produces distributions of effective Young's modulus and Poisson's ratio. A single constitutive pair cannot capture this variation. The resulting distributions depend strongly on the matrix-to-fibril stiffness ratio and, in several cases, exhibit apparent multimodality. We further propagate the paired effective properties into 20 stochastic realizations of a tensile simulation of a larger specimen, yielding a distribution of macroscale stress response. The framework provides a probabilistic link between sparse nanoscale observations and continuum-scale mechanics while retaining variability across scales. The present results establish the computational workflow, while calibration of the intensity-to-property mapping against nanomechanical measurements remains necessary for quantitative prediction at the fiber scale.

cs.CE↗

HeuFouFT: Task-Guided Metaheuristic Coordinate Search for Fourier Fine-Tuning

We introduce Heuristic-Guided Fourier Fine-Tuning (HeuFouFT), a task-guided framework for selecting trainable frequency coordinates in Fourier fine-tuning. Existing uniform and Gaussian band-pass schemes allocate a limited spectral budget through fixed, task-agnostic rules. HeuFouFT instead searches for coordinates using downstream performance. A coarse intensity map from lightweight block-level probes initializes three metaheuristic optimizers: Genetic Algorithm with Simulated Annealing (GA-SA), Particle Swarm Optimization (PSO), and Cuckoo Search (CS). During search, a Random Forest filters each population so that only the top 30% of candidates proceed to proxy fine-tuning. On E2E with GPT-2-Medium, all three variants outperform random-uniform FourierFT, Gaussian band-pass FourierFT, and LoRA across five metrics. PSO further outperforms LoCA, the best-performing baseline, on four metrics while using 37.6% fewer trainable spectral coefficients. Once coordinates are selected, HeuFouFT requires only 15--18% FLOPs of Full FT. These results show that task-guided search allocates limited spectral capacity more effectively than fixed sampling. Our code is publicly available.

cs.LG↗

Simpson filtrations and closure relations of Hodge moduli spaces

In this paper, we study closure relations of Simpson strata in Hodge moduli spaces over a smooth complex projective curve of genus $g\geq2$ and prove in every rank that the Heinloth weight strictly increases whenever a specialization changes the fixed component. For stable full chains, we refine this numerical condition by comparing the degrees of the corresponding filtration steps and construct counterexamples to Simpson's nestedness conjecture for the Dolbeault and de Rham moduli spaces in every rank $n\geq4$. Using the maximum property of this weight, we also discuss Simpson filtrations from the viewpoint of $Θ$ stratifications.

math.AG↗

Better Call Reward: Reward Hacking as Strategic Abstention in Legal Reasoning Models

What happens when a legal AI model learns to look like a lawyer instead of reasoning like one? We fine tune Qwen3-8B with Group Relative Policy Optimisation (GRPO) against a proxy built from three surface features: citation count, legalese density, and response length. The model does not learn to reason more effectively. It learns to withhold commitment. Across 16 yes or no legal reasoning tasks from LegalBench (N=320), overall accuracy collapses from 0.500 (chance) to 0.072 (McNemar p < 10^-36), driven entirely by the rate of properly formatted answers falling from 0.900 to 0.109. The model stops committing to answers. Yet when it does commit, accuracy rises from 0.556 to 0.657, showing that the collapse is not a failure of capability but a strategic response: the model has learned that verbose responses packed with citations but empty of a direct answer score higher than terse correct ones. We term this the Saul Goodman effect, a policy that becomes maximally lawyerly while becoming maximally noncommittal, and prove formally that it is the optimal response to any surface feature proxy that attaches no penalty to abstention. We further show that 89.3% of citations produced after training are structurally implausible hallucinations, many of them subtly corrupted names of real landmark cases, constructed in effect to survive a casual read and fail under scrutiny. To detect this failure mode before deployment, we introduce three diagnostic tools: the Confidence Theater Score (CTS), the Citation Plausibility Rate (CPR), and the Regret Gap (RG). In a domain where a confidently wrong answer can constitute malpractice, the broader lesson is direct: a reward function that measures how legal a response looks will produce a model that is maximally photogenic and minimally useful.

cs.LG↗

Genus-3 Lefschetz pencils on torus bundles

We use the breeding technique to construct genus-3 Lefschetz pencils for which the total spaces have $c_1=0$. We compute the fundamental groups of our Lefschetz pencils and identify them with the fundamental groups of certain torus bundles over tori. Using work of Friedl and Vidussi, the total spaces of our pencils are homeomorphic to torus bundles over tori. These bundles do not admit sections.

math.GT↗

Regularity for $H$-systems of bounded distortion

We consider weak solutions $u \in W^{1,n}(U,\mathbf{R}^{n+1})$ of the $H$-system $- \operatorname{div} ( |\mathrm{D} u|^{n-2} \mathrm{D} u ) = (H \circ u) \, \mathbf{J} u$, where $n \ge 2$, $U$ is the open unit ball in $\mathbf{R}^n$, and $H$ is bounded. We show that $u$ is locally bounded if its distortion $|\mathrm{D} u|^n / |\mathbf{J} u|$ is bounded on the set where $|u|$ is large. If moreover $H$ is Hölder continuous, then $u$ is locally Hölder continuous. The proof pushes forward $|\mathrm{D} u|^{n-2} \mathrm{D} u \circ \mathrm{D} u^*$ by $u$ to a linear functional on continuous maps of $\mathbf{R}^{n+1}$ into $\operatorname{Hom}(\mathbf{R}^{n+1},\mathbf{R}^{n+1})$ with compact support, whose values on derivatives are controlled by the equation. When the distortion is bounded, this functional satisfies the hypotheses of an inequality of Michael--Simon type for measures satisfying a first-order partial differential equation, due to De Philippis, Gennaioli, Pigati, and Rindler. The inequality yields lower bounds for the mass ratios of the push-forward of $|\mathrm{D} u|^n \mathscr{L}^n$ by $u$, and these imply boundedness.

math.AP↗

Torus Surgery Constraints and Intermediate Covers of the Cartwright--Steger Surface

We study torus surgery on the Cartwright--Steger surface and the degree-two cover construction proposed in \cite[Remark~1]{ASY}. Every torus in this surface is integrally nullhomologous, and surgery on a fixed disjoint collection of such tori cannot decrease the first Betti number. We then determine the nodal-curve subgroup and complete the fundamental-group calculation in the degree-two construction. The resulting minimal symplectic manifold has signature zero and fundamental group $\Z/2$; for compatible symplectic gluings, its universal cover is an exotic symplectic $31\CP^2\#31\overline{\CP}^{\,2}$. The construction on the original Cartwright--Steger surface gives a minimal symplectic manifold with signature $-1$ and fundamental group $(\Z/2)^2$.

math.SG↗

ARO: Aligned Representation learning for multi-Omics data

The high cost of functional molecular assays, and prevalence of missing modalities and unmatched samples in computational biology, create significant barriers to comprehensive multi-omic profiling, essential for capturing and reasoning over molecules, cells, tissues, and organisms. This work proposes a model that learns meaningful representations from multi-omics cancer data supporting the reconstruction of missing and unpaired modalities. Contrary to increasingly complex, larger models, e.g. Foundation Models (FMs), ARO prioritizes practical applicability in limited or incomplete data settings. ARO optimally reconstructs missing modalities (MSE of $0.15$ on the validation and test data in the Unmasked settings), with its learned latent embeddings enabling a downstream cancer classification task. Our findings indicate that analyzing diverse molecular layers as a single integrated system offers a reliable and cost-efficient approach, reducing dependence on large-scale experimental testing, while still supporting multi-omic exploration in limited data settings.

cs.LG↗

Choosing an energy-efficient software architecture for building system diagnostic support

Around 30\% of global energy expenditure can be attributed to the building sector, where a large portion of energy-consumption could be avoided by repairing existing faults. Fault detection and diagnosis (FDD) software addresses this issue; however, its creation and operation also have an environmental impact. The magnitude of this impact is influenced by the diagnosis architecture, as different architectures and methods have different energy demands. Yet, simply considering the energy consumed by the software itself is not sufficient to assess its overall environmental impact, since the diagnostic performance, e.g., number of detected faults or number of faults missed, also contributes to its ecological footprint. In this paper, we propose an energy-consumption model that considers FDD performance and energy spend directly by the diagnosis software. In an initial experiment, we compare several FDD architecture families, i.e., rule-based, model-based, classical machine learning, and large-language-model-based, in simulation using performance and energy-consumption values collected from prior literature. The results show that considering the computational energy and accuracy of FDD can change the relative benefit of the different approaches. Computationally efficient machine learning methods, such as random forest, provide the largest net savings on smaller buildings, whereas more resource-intensive approaches, such as fine-tuned large language models, become advantageous as building size increases. Our findings suggest that overall energy efficiency depends not only on the computational demand of the FDD software, but also on its diagnostic performance and the scale of the building.

cs.SE↗

Counting zero-sum subspaces for the multiplicative inverse function

Let $q=2^n$ with $n\ge6$, and let $\mathbb{F}_q$ be the finite field with $q$ elements. For a subspace $E$ of $\mathbb{F}_q$, define \[ S(E)=\sum_{x\in E\setminus\{0\}}x^{-1}.\] Let $N_{n,k}$ denote the number of $k$-dimensional $\mathbb{F}_2$-subspaces $E$ for which $S(E)=0$. We prove that, if $k\ge3$ and $n\ge2k+1$, then \[ \left|\frac{2^nN_{n,k}}{\genfrac{[}{]}{0pt}{}{n}{k}_{2}}-1\right|<2^{r^2+r}2^{n(1-r/2)},\] where $r=\lfloor\frac{k-1}{2}\rfloor$, and $\genfrac{[}{]}{0pt}{}{n}{k}_{2}$ is the Gaussian binomial coefficient. In particular, \[ N_{n,k}\sim2^{-n}\genfrac{[}{]}{0pt}{}{n}{k}_{2}\quad (n\to\infty)\] uniformly for $n/3<k\le(n-1)/2$. Combined with the known low-dimensional cases and the symmetry between dimensions $k$ and $n-k$, and the elementary middle-dimensional subfield construction, this proves a conjecture of Carlet: for every $3\le k\le n-3$, there exists a $k$-dimensional $\mathbb{F}_2$-subspace $E$ such that $S(E)=0$.

math.NT↗

The Assistance Dilemma: Learning to Teach via Multi-Turn Reinforcement Learning

Large language models (LLMs) trained to answer questions are natively poor at teaching. Reinforcement Learning (RL) against a simulated student is a promising approach to improve their pedagogy, but existing RL-trained tutors reward the student's success on the tutored problem with the tutor's words still in context. The reward is then easiest to raise by telling the student the answer, and a tuned penalty is needed to reduce telling. Drawing on learning sciences, we introduce a masked near-transfer post-test: the student is tested on an unseen variant of the tutored problem with the tutor's utterances masked, so the reward can rise only through what the student wrote in its own turns. This discourages cognitive offloading by the student and allows the continuous penalty to be replaced by two binary reward gates (factual correctness of tutor response, no solution handover). A leave-one-out ablation shows that the learning-gain reward on its own does not separate teaching from telling: the gates reduce solution handover while the near-transfer post-test improves out-of-domain transfer. Using these reward designs we develop Eduardo, a multi-turn RL recipe for training LLM tutors, and use it to train 4B, 9B, 14B and 27B models from two distinct LLM architectures. Our post-trained Eduardo-27B model matches Gemini-3.1-Pro on MathTutorBench and Claude Opus 4.8 on TutorMoments at 2.4-6.2x fewer thinking tokens than frontier models, which matters for interactive tutoring. Without being named in the reward, the model more than doubles its use of the push-for-justification teacher move while support fading (e.g., assigning independent work), whose payoff lies beyond a single-problem dialog episode, is trained out. We open-source our training environment, an 8,671-problem near-transfer dataset, and trained models for further development.

cs.CL↗

Time-series Foundation Models for Predictive Control: The Role of Excitation

Deploying model predictive control (MPC) requires constructing or identifying a predictive model for each target system. Time-series foundation models (TSFMs) offer an attractive option thanks to strong zero-shot forecasting capabilities across systems. However, low forecast error does not guarantee that a TSFM captures the system's response to the alternative actions considered by the controller. We study this gap using residential heat-pump control as a test bed, measuring the agreement between predicted and ground-truth effects of control interventions. Importantly, we find that TSFMs can recover the system's input-response relationship when the context contains sufficient independent control excitation. Common fine-tuning pipelines and feature smoothing reduce, but do not eliminate, the need for in-context excitation. Our results indicate that current TSFMs used for predictive control require sufficiently informative control variation in the inference context. Initial closed-loop results show promise for shorter context windows.

cs.LG↗

A Practical Introduction to VQE: Methods, Challenges, and Applications in Physics and Chemistry

The Variational Quantum Eigensolver (VQE) has become a central hybrid quantum-classical method for addressing electronic structure and materials science problems on near-term quantum hardware. As researchers across physics, chemistry, and quantum information engage with VQE, they encounter a wide range of algorithmic choices, including problem encoding, ansatz construction, classical optimization, and measurement strategies to choose from. This article provides a structured overview of these components and discusses practical factors that influence VQE performance in real settings, such as noise characteristics, available error-mitigation techniques, and hardware-dependent circuit considerations. Drawing partly on our early experiences, we highlight aspects that may be especially relevant for newcomers and outline common challenges observed in practical workflows. Our aim is to offer an accessible entry point for researchers approaching VQE-based quantum simulation and to support informed decision-making as the method and underlying hardware continue to evolve, which we support by discussion of examples from physics and quantum chemistry.

quant-ph↗

Exact vortex decomposition of the shift current

Does any integer survive in the shift current, and does the Chern number control its sign? We show that the shift vector is a singular field on the Brillouin torus. Its curl has two sources: the interband Berry curvature and quantised point charges at the zeros of the transition dipole. Inverting this relation gives an exact decomposition of the shift conductivity, valid at every frequency, into one sector linear in the integer charges and three continuous sectors. As a testing ground, we take a graphene-Haldane bilayer: a model that lets us sweep the interband Chern number through five values at fixed gap to isolate its role. We also test the formalism against the time-reversal-symmetric biased AB bilayer. Whenever a charge sits at the band edge, and trigonal warping is present, the integer sector alone accounts for the band-edge peak. The relevant integer is not the Chern number but the winding of that single charge. Its sign, combined with a valley-dependent profile, sets the sign of the photocurrent. In these bilayers, sign reversals across topological transitions follow this winding; the invariant changes with it only when the same gap closing does both. A band-edge charge needs no invariant. It only requires that the two band-edge sublattices not be connected by the nearest-neighbour hopping, and trigonal warping makes it visible.

cond-mat.mes-hall↗