arXiv Science⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 847 records · Page 47Linked to original sources

Not Every Token Is Worth Distilling: Selective Supervision for Direct-OPD

Direct On-Policy Distillation (Direct-OPD) transfers reinforcement-learning-induced policy improvements from a small model to a larger student by using the token-level log-ratio between post-RL and pre-RL checkpoints as dense supervision on the student's own rollouts. This transfer rewards the policy shift at every state, yet the log-ratio measures only relative change: it can stay fixed even as the probability mass that both checkpoints assign to the student's candidate tokens vanishes. Through an exact construction, we show that the Direct-OPD reward and its update can remain unchanged while the Jensen-Shannon divergence (JSD) and both KL directions between the checkpoints vanish with this mass, and we note that a small JSD bounds how much the teacher's behavior changed. Motivated by this analysis, we propose Selective Supervision for Direct-OPD (S$^2$D-OPD), which ranks student-sampled states by their teacher-reference JSD and masks Direct-OPD supervision at low-divergence states, retaining only the top 10% of states per response. Across two teacher pairs and four student models ranging from 1.7B to 8B parameters, S$^2$D-OPD improves held-out accuracy over dense Direct-OPD on AIME and HMMT benchmarks in seven of eight settings and matches it in the eighth, without extra forward passes. Our code is available at https://anonymous.4open.science/r/S2D-OPD-8868.

cs.LG↗

AI-Moderated Interviews for Market Research and Digital Twins Calibration

AI-moderated interviews are emerging as a scalable market-research method for generating consumer insights and building consumer "digital twins." Yet it remains unclear whether they match human-moderated interviews or improve on simpler, static data collection methods. In a pre-registered, between-subjects study (N = 317) with three industry partners, we compare AI-moderated (N = 139), human-moderated (N = 24), and static interviews (N = 154). AI moderation matches human moderation in depth, covers more themes, and, holding budget constant, recovers significantly more customer needs than human moderation or static interviews. However, participants sound more emotionally engaged when speaking to a live human. We then create digital twins using interview data and evaluate each twin against the participant's own held-out responses to six real-world marketing stimuli. We find that digital twins created from AI-moderated interviews predict consumer responses better than demographics-only personas. However, the additional richness from AI moderation does not translate into better quantitative predictions compared to static interviews. By analyzing open-ended thoughts generated from humans versus their twins, we find that prediction errors are connected both to differences in (self-reported) thinking styles between twins and humans, and to gaps between training and validation data (i.e., asking questions that are too far out of distribution).

cs.CY↗

Scope Before You Persist: Preventing Cross-Family Interference in Agent Memory

Persistent memory lets language-model agents improve prompts and skills without updating model weights. We show that matching retrieval scope to certification scope enables these edits to support reliable repeated adaptation across recurring task families. We study frozen-model agents on ProcStream-RSI, a 12-round code-repair stream, using Orthogonal Regression Control (ORC), an execution-grounded gate for persistent skill edits. In an intervention that holds proposals and gate decisions fixed, retrieving each accepted skill only for its originating family raises mean hidden trajectory utility from 0.713 under global memory to 0.816 and changes harmful deployments from six of eight to none. In 27 paired randomized-order streams, Scoped-ORC improves mean trajectory utility by 0.063 [0.037, 0.094] over Global-ORC, accepts 63 rather than 12 updates, and produces multiple accepted updates in 19/27 streams, with 0/63 harmful acceptances. The global control reaches 0.713, below the static agent's 0.775, because locally valid edits can interfere with unrelated families. These results establish scope matching as a complementary control for persistent agent memory: certification determines whether an edit is supported, while retrieval scope determines where that evidence authorizes its use.

cs.AI↗

Claim-Gated Source-Risk Auditing for Generative Search

A generative search answer can cite a supported passage yet omit a source relationship that changes its interpretation. We specify a claim-gated audit of the query-source-answer tuple. An omission is resolved only when relationship evidence, answer adoption, materiality, and disclosure are all observed; incomplete evidence remains unresolved rather than being treated as independence. The specification separates this endpoint from citation support and review priority, and binds decisions to versioned evidence spans. A reference checker makes the record contract executable. On an exhaustive synthetic suite, it reproduces all 81 three-state predicate combinations and rejects 192 deliberately malformed records. Common-guard baselines and predicate ablations isolate endpoint logic from missing-evidence handling, while controlled transitions check support separation and evidence removal. These are finite contract-conformance results, not detector accuracy or evidence of improved user outcomes. We define the independent annotation, held-out evaluation, and paired utility tests still required to establish semantic validity and deployment benefit.

cs.AI↗

BanglaKontho: Closing the Long-Form Gap in Bangla Text-to-Speech

Bangla, the seventh most spoken language in the world, remains under-resourced for neural text-to-speech. Public Bangla speech corpora are dominated by short read-prompt utterances collected for speech recognition, leaving long-form prosody and consistent single-speaker narration uncovered. We present BanglaKontho, a single-speaker Bangla TTS corpus of 20 hours derived from professional audiobook recordings: 7,050 segmented utterances with verified transcripts at 24 kHz. We also release a reusable Bangla text normalizer covering Bangladeshi-style digit grouping, currency and date expressions, Danda punctuation and Unicode normalization, together with the full preprocessing pipeline. An MB-iSTFT-VITS baseline trained from scratch reaches 9.5% WER and 4.46 naturalness MOS, against 16.0% and 3.16 for the same architecture retrained on the 12-hour IndicTTS-Bn corpus. The corpus is released openly under CC BY-NC 4.0.

cs.CL↗

Global well-posedness of defocusing cubic NLS in $M^{\infty,1}(\mathbb{R})$

We prove global well-posedness of the one-dimensional defocusing cubic nonlinear Schrödinger equation in the modulation space $M^{\infty,1}(\mathbb{R})$. This space imposes no spatial decay and contains $C_b^2(\mathbb{R})$ as well as all absolutely convergent sums of plane waves. The result applies to arbitrary data in this space, including large smooth quasiperiodic profiles and their localized perturbations. The proof constructs a nonnegative density satisfying a local conservation law from forward Weyl ratios, which are defined through half-line square-integrable solutions of the associated spectral problem. A suitable nonlinear combination of localized integrals of this density controls the modulation norm. Finally, choosing the spatial localization scale and NLS scaling in a coordinated way makes the accumulated boundary flux small enough to continue every mild solution globally.

math.AP↗

Exact Traces for Vector-Valued Morrey Spaces

Trace spaces determine exactly which initial values are compatible with an evolution class. We identify the exact trace generated by vector-valued Morrey control in time and show that it differs essentially from the classical $L^p$ theory. For a Banach couple $X_1\hookrightarrow X_0$, the trace of the natural Morrey evolution class is the weak real-interpolation space $(X_0,X_1)_{θ,\infty}$, where $θ=1-(1-λ)/p$. Every element of this space occurs as a trace through a bounded extension operator. The result is sharp in both interpolation parameters: in general the smoothness exponent cannot be increased and the fine index $\infty$ cannot be replaced by any finite index. Thus Morrey control changes the exact trace mechanism rather than merely strengthening an integrability estimate. At the limiting endpoint, bounded mean oscillation (BMO) control yields finite-index interpolation traces together with a logarithmic modulus of continuity. The result isolates the precise initial-data space naturally associated with local, scale-sensitive time regularity and provides a trace framework suited to evolution equations with Morrey-type maximal regularity. \keywords{Morrey spaces \and trace spaces \and real interpolation \and evolution equations \and bounded mean oscillation}

math.AP↗

Replica Thresholds for Stripeless Erasure Coding Based on Symmetric Block Designs

This paper investigates a fundamental question in stripeless erasure coding based on symmetric balanced incomplete block designs (SBIBDs): how many replicas per object are precisely required to guarantee recovery from any set of at most $p$ node failures? We refine the known sufficient recovery guarantee for the generalized SBIBD $(v,k,λ)$ construction. One replica is necessary and sufficient for $p=1$, and $q=λ(p-1)+2$ replicas ($q\le k$) guarantee recovery for $p\ge2$. Recovery takes one round for $p=2$ and at most two rounds in general. To study the tightness of this count, we define the universal replica threshold $q^\ast(A,p)$ for a fixed zero-diagonal SBIBD representative~$A$. We determine $q^\ast(A,p)$ for $p=1$ and $p=2$. For $p\ge3$, we identify pairwise separated sets and private target permutations as the structures determining whether the sufficient count is tight or can be further reduced. These conditions give exact thresholds for all but the case where $λ>1$ and $A$ contains no pairwise separated set of size~$p$. We improve the bounds for this remaining case and leave its exact threshold open. Finally, we give sufficient conditions for pairwise separated sets and show how affinity relabeling can realize private target permutations.

cs.IT↗

A Particle-Swarm-Assisted Gradient Meta-Learning Algorithm for Joint Transmit Precoding and STAR-RIS Coefficient Optimization

This paper investigates the joint optimization of the transmit precoder and the transmission/reflection coefficients of a simultaneously transmitting and reflecting reconfigurable intelligent surface (STAR-RIS) to maximize the weighted sum rate (WSR) in a multi-user downlink. We propose a particle-swarm-assisted gradient meta-learning (PSA-GML) algorithm for this non-convex problem. The original problem is first equivalently transformed via an amplitude-split parameterization and a collapsed precoder representation, which automatically satisfy the energy-conservation constraint and reduce the search dimension. Particle swarm optimization (PSO) then performs a global search over the STAR-RIS coefficients to yield a high-quality, initialization-robust warm start, with the transmit precoder obtained in closed form. Departing from conventional alternating optimization (AO), a coordinate-wise long short-term memory (LSTM) meta-optimizer trained by first-order gradient meta-learning further refines the coefficients and precoder jointly, learning per-coordinate adaptive update rules from data. The meta-optimizer is trained offline and applied to unseen channels without further adaptation. Numerical results show that PSA-GML attains an 11.06 bits/s/Hz WSR at 10 dB with N=32 elements and K=4 users, exceeding AO by 13.1% (and by 6.2% even with multiple random restarts) and the random-phase scheme by 35.1%. In the interference-limited regime it reaches 83.9% of the hand-designed Adam refinement without manual hyper-parameter tuning, and it transfers zero-shot across regimes, indicating that the learned update rule captures the intrinsic WSR landscape structure.

cs.LG↗

Recoverable Geographic Location Information in Earth-Observation Embeddings

Earth-observation (EO) foundation models provide reusable embeddings, yet downstream task accuracy does not reveal whether these representations encode geographic information, which may be beneficial for location-aware applications but potentially detrimental when representations invariant to geographic location are desired. We therefore evaluate the geographic coordinate robustness of Tessera v1, Tessera v1.1, and AlphaEarth by testing whether coordinates can be predicted from the embedding representations using 284 quality-verified European solar farms from 2024. We assessed geographic information content information through the association between cosine and geodesic distances and through prediction of projected coordinates in EPSG:3035. Embeddings from all three EO foundation models contain recoverable geographic information. All prediction models significantly outperform training-range uniform random sampling baselines, with AlphaEarth exhibiting the strongest distance association and lowest mean geodesic error. Both Tessera variants also yielded higher geographic distance correlations than the Sentinel-2 controls. These findings motivate geographic information content as an additional criterion for auditing EO foundation models.

cs.CV↗

Compact Embeddings of Vector-Valued Morrey Spaces

We develop compactness and defect-of-compactness results for Banach-valued evolution classes with Morrey control in time. For \[ \begin{aligned} \mathbb W_M^{p,λ}(0,T;E_0,E_1) =\{u\in\mathcal M^{p,λ}(0,T;E_0):\;& u'\in\mathcal M^{p,λ}(0,T;E_1)\},\\[-1mm] &0<λ<1. \end{aligned} \] the exact trace exponent \[ θ=1-\frac{1-λ}{p} \] governs both continuity and compactness. If $E_0\hookrightarrow\!\hookrightarrow E\hookrightarrow E_1$, bounded sets are compact in the \emph{same} Morrey space $\mathcal M^{p,λ}(0,T;E)$, not merely in $L^p(0,T;E)$. In contrast, compactness in $C([0,T];E)$ holds if and only if the exact trace space $(E_1,E_0)_{θ,\infty}$ embeds compactly into $E$. We also obtain compact lower-order Hölder embeddings and a sharp Hilbert-triple threshold. On unbounded domains we prove a tightness criterion for global Morrey compactness. In the Hilbert-valued case we establish, by a direct time-averaging argument, cocompactness modulo spatial translations and a translation profile decomposition whose remainder vanishes in every strictly subcritical time-Morrey--Sobolev target; a uniform little-Morrey condition removes the loss in the time exponent. At the doubly critical endpoint, heat and whole-space Stokes dynamics force balanced parabolic profiles, and the remainder vanishes in $\mathcal M_t^{2,λ}L_x^{2^*}$. Finally, in three-dimensional Navier--Stokes we show that the classical nonlinear profile decomposition has a scale-sensitive Morrey refinement: the remainder is small in $\mathcal M_t^{p,1-p/4}L_x^6$ for every $2\le p<4$, and orthogonal profiles have vanishing interaction in the corresponding Morrey forcing space.

math.AP↗

On degeneration of tetrahedra under longest-edge n-section refinement

For every integer $n \ge 2$, we construct explicit tetrahedral counterexamples showing that repeated longest-edge (LE) n-section refinement does not, in general, preserve shape regularity, contrary to long-standing conjectural expectations. For $n = 2$, the construction disproves the finite-family conjecture put forward by Adler in 1983 and the non-degeneracy conjecture for tetrahedral longest-edge bisection formulated by Rivara and Levin in 1992. It also shows that the sharper higher-dimensional diameter decay suggested by Stynes in 1983,following his 1980 planar finite-similarity result, cannot hold for arbitrary tetrahedra. For $n = 3$, it disproves the published non-degeneracy conjecture for tetrahedral longest-edge trisection formulated in 2011 by Suárez, Abellón, Abad and Plaza on the basis of numerical experiments. The resulting infinite descendant sequences violate both the minimum and maximum angle conditions. For $n \ge 3$, the selected edge is always the unique longest edge at every refinement step. The two-step recurrence used in the construction exhibits two distinct degeneration behaviors: a flat-tetrahedron regime for $2 \le n \le 5$ and a skinny-tetrahedron regime for $n \ge 6$. The degenerating sequences can also be realized within conforming partitions generated by the conforming LE n-section algorithm. For both the classical and the conforming LE n-section algorithms, the maximal element diameter tends to zero as the number of refinement steps tends to infinity, regardless of which longest edge is chosen. Thus diameter convergence, and even conformity, do not prevent shape degeneration. These algorithms should therefore be used with care in applications requiring uniform shape regularity.

math.NA↗

A Wrong Turn Does Not Ruin the Journey: Deviation-Guided Skill Self-Evolution for LLM Agents

Large language model agents increasingly rely on natural-language skills to solve complex tool-use tasks. However, such tasks often admit multiple valid solution paths, making it inappropriate to improve skills by forcing failed trajectories to match a fixed successful trajectory. Moreover, failed trajectories are rarely entirely wrong: an agent may first collect useful evidence and make meaningful progress, but later deviate into an erroneous suffix. We therefore argue that skill self-evolution should identify where productive problem solving begins to break down, rather than reflect coarsely over the entire failure. Based on this insight, we propose SkillPivot, a deviation-point-guided framework for skill self-evolution. SkillPivot detects the transition from a useful prefix to an erroneous suffix using execution validity, goal progress, and action diversity. A stronger teacher then continues from the same prefix and produces a successful alternative under the same interaction history. By contrasting the student's failed suffix with the teacher's successful suffix, SkillPivot generates localized skill updates while preserving already effective guidance. Experiments on ToolQA, LogicBench, and WildClawBench show that SkillPivot consistently outperforms competing skill-evolution methods, improves multiple agent models, and produces compact, transferable skill updates.

cs.AI↗

ProteoEM: probabilistic protein abundance estimation from iterative affinity traces

Single-molecule affinity mapping enables molecular-level measurement of proteins and proteoforms, but imperfect and nonspecific probe binding makes individual affinity traces compatible with multiple molecular identities. Accurate abundance estimation therefore requires apportionment of ambiguous traces by weight rather than assignment to a single candidate. We developed ProteoEM, an expectation-maximization framework for weighted proteoform quantification, inspired by transcript abundance estimation methods for RNA sequencing and released as an open-source Python package. ProteoEM evaluates each molecule against every candidate using fixed, pre-calibrated probe-response rates held separate from the abundance estimate, while retaining the full likelihood of the observed affinity features. The framework estimates proteoform abundances, reports indistinguishable proteoforms as groups when measurements cannot separate them, and accounts for differential observation yields to distinguish the composition of observed molecules from that of the source sample. In simulations, ProteoEM accurately recovered the underlying molecular composition where approaches that reduce each trace to a hard yes/no call introduced substantial errors. ProteoEM's performance was insensitive to a moderate, uniform calibration error but was biased by informative missing data and by proteoforms absent from the reference. When observation yields were known, it also recovered source-sample composition from observed molecular counts. ProteoEM provides an open-source, reproducible framework for quantitative analysis of single-molecule affinity measurements, and these results motivate validation on experimental molecule-level data.

q-bio.QM↗

Med-AR: Autoregressive Vision-Language Pretraining for Long-Tailed Chest X-Ray Classification and Uncertainty-Aware Evaluation

Long-tailed chest X-ray classification requires visual representations that capture both common abnormalities and subtle, infrequent findings. We propose Med-AR-8B and Med-AR-2B, two radiology-native autoregressive vision-language models pretrained with structured reports, abnormality-focused text, and region annotations. We evaluate the transfer of their visual encoders to multi-label classification against contrastive, self-supervised, and supervised pretrained encoders, including Med-CLIP, CheXFound, EVA-Base, ARK, and BioViL-T, using a common ML-Decoder classification head. To assess fine-grained recognition, we also construct LLM-expanded, report-derived label sets for MIMIC-CXR and CheXpert. Across PadChest, MIMIC-CXR, and CheXpert, Med-AR-8B outperforms Med-CLIP in mean AUROC and AUPRC for head, medium, and tail findings. On MIMIC-CXR, it increases tail-label mean AUPRC from 0.1033 to 0.1441. Med-AR-2B achieves the strongest discrimination results on PadChest. Across the broader encoder comparison, a Med-AR variant achieves the highest mean AUROC and AUPRC in every reported prevalence group on each public dataset. Both Med-AR variants also achieve lower excess area under the risk-coverage curve than Med-CLIP on all three public datasets, indicating improved selective-prediction performance under the evaluated protocol. Internal results are metric-dependent, with Med-CLIP retaining advantages in overall and tail AUPRC and in selective prediction. These findings establish Med-AR as a strong pretraining recipe for long-tailed chest X-ray classification on the evaluated public benchmarks and demonstrate the value of assessing discrimination and selective prediction together.

cs.CV↗

OREN-X: Octree Residual Network for Real-Time Multi-Modal Mapping

To achieve general-purpose autonomy over long horizons, a robot needs to maintain spatial environment information that supports a variety of tasks: geometry for planning and control, radiance for rendering and relocalization, and vision-language features for open-vocabulary grounding. Existing methods represent and estimate each modality separately, multiplying memory and compute cost while forgoing potential synergy among the representations. We develop OREN-X, an online mapping method that uses an octree in 3D space as a shared data structure for indexing and storing a multi-modal field, capturing geometric, radiance, and vision-language information. OREN-X provides efficient unified storage and retrieval of these data in explicit/implicit and full/compressed form. Our unified representation yields cross-modality synergy: SDF estimates are sharpened by occupancy and radiance, while GPU-based ray-octree traversal and octree query enable real-time rendering. We also use online dictionary learning to compress the vision-language features, shrinking them 3.7x below full per-vertex storage while raising the query accuracy. On Replica, OREN-X maps in real time (80+ fps for SDF and 30+ fps for all four modalities), improves near-surface SDF accuracy by 33% over single-modality baselines, and improves mean open-vocabulary 3D mIoU by 71% and mean accuracy by 61% over the best prior method.

cs.RO↗

A Morrey-to-Lebesgue Equivalence for Convolution Calderón--Zygmund Operators

We prove that, for vector-valued convolution Calderón--Zygmund operators, boundedness on a single nontrivial Morrey space is equivalent to the corresponding global $L^p$ boundedness. Thus one Morrey scale already contains the full finite-$p$ boundedness information. The implication from Morrey to $L^p$ is obtained by a separated-copy amplification argument that reconstructs the global norm from a single scale-local estimate; the converse is proved in the same vector-valued framework by a local/far-field decomposition. As a consequence, boundedness of the vector-valued Hilbert transform on one nontrivial Morrey space is equivalent to the UMD property. The result shows that Morrey estimates do not bypass the classical Banach-space obstruction: they detect it exactly.

math.AP↗

Functional dynamic mode decomposition: Learning infinite-dimensional systems from data

Dynamic mode decomposition (DMD) is a data-driven method that computes the best linear approximation of the underlying dynamical system and decomposes the dynamics into a superposition of characteristic spatiotemporal patterns. Originally introduced by the fluid dynamics community, DMD and its extensions have found widespread use in many other research areas such as molecular dynamics, climate science, engineering, finance, and neuroscience. Applications include dimensionality reduction, forecasting, system identification, control, and spectral clustering. In order to apply DMD to partial differential equations, the spatial domain is typically first discretized using finite difference or finite element techniques, thus implicitly rendering the problem finite-dimensional. We extend projected and exact DMD to infinite-dimensional systems. Rather than estimating matrices from vector-valued observations, our DMD variants learn finite-rank operators from functional data such as observables, densities, or wavefunctions. We show that conventional DMD algorithms can be regarded as special cases of their functional DMD counterparts. All results will be illustrated with the aid of guiding examples. We focus in particular on Koopman, Perron-Frobenius, and Koopman-von Neumann operators associated with graphons, ordinary differential equations, and stochastic differential equations.

math.DS↗