arXiv Science⌕ Search

SEARCH · arXiv Science

Search arXiv Science

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 991 records · Page 55Linked to original sources

Interpreting and Enhancing Emotional Circuits in Large Vision-Language Models via Cross-Modal Information Flow

Large Vision-Language Models (LVLMs) represent a significant leap towards empathetic agents, demonstrating remarkable capabilities in emotion understanding. However, the internal mechanisms governing how LVLMs translate abstract visual stimuli into coherent emotional narratives remain largely unexplored, primarily due to the scarcity of visual counterfactuals and the diffuse nature of emotional expression. In this paper, we bridge this gap by introducing a steering-vector-based causal attribution framework tailored for descriptive emotional reasoning. To this end, we construct a specialized dataset to demystify the emotional circuits underlying the three-stage ``Adapt-Aggregate-Execute'' mechanism. Crucially, we discover a functional decoupling: visual emotional cues are aggregated in middle layers via sentiment-specific attention heads, but are subsequently translated into narrative generation in deep layers through emotion-general pathways. Guided by these insights, we regulate the emotional information routing to strengthen attention flow and amplify the semantic activation to consolidate expression. Extensive experiments on the comprehensive MER-UniBench demonstrate that our methods significantly improve performance via inference-time intervention, effectively mitigating emotional hallucinations and corroborating the causal fidelity of the discovered circuits.

cs.CV↗

Extrinsic characterizations of biconservative surfaces in the $4$-dimensional hyperbolic space

Biconservative submanifolds arise as a natural relaxation of the biharmonic condition and play an important role in the submanifold theory. In this paper, we study non-CMC biconservative surfaces with parallel normalized mean curvature vector field (PNMC surfaces) in the four-dimensional hyperbolic space $\mathbb{H}^4$, for which we consider the hyperboloid model. We provide a local extrinsic description of such surfaces, showing that they are generated by a directrix curve lying in a totally geodesic hypersurface $\mathbb{H}^3$ of $\mathbb{H}^4$, through a certain normal flow. This extrinsic classification of non-CMC, PNMC biconservative surfaces in $\mathbb{H}^4$ splits naturally into three cases according to the type of a certain vector field, which can be non-zero null, spacelike or timelike. We also prove that these surfaces are invariant under the action of a parabolic, elliptic, and hyperbolic one-parameter group of isometries of $\mathbb{H}^4$, respectively. Moreover, their full groups of ambient isometries preserving the surfaces are determined. Together with the previous results, the classification of non-CMC, PNMC surfaces in four-dimensional space forms is now complete, from both intrinsic and extrinsic points of view.

math.DG↗

Arboreal Galois Groups of a PCF Map with Strictly Pre-periodic Critical Points

We study the arithmetic and geometric iterated monodromy groups associated to the postcritically finite (PCF) quadratic rational function $f(x)=\frac{2}{(x-1)^2}$ defined over a number field $k$, whose critical points are both strictly pre-periodic. We give explicit recursive descriptions of the topological generators of the geometric iterated monodromy group of $f$ and show that the arithmetic iterated monodromy group has Hausdorff dimension zero. We describe an explicit criterion to determine the values $a \in k$ for which the associated arboreal Galois group achieves its maximum possible size. In particular, we show that maximality of the arboreal Galois group can already be verified at level four, which is computationally accessible. We also determine the normalizer of the geometric IMG of $f$ and show that the arithmetic IMG equals the normalizer if $ζ_8$ is not contained in $k$. Finally, we show that when $k=\mathbb{Q}$, the constant field contains $\mathbb{Q}(μ_{2^{\infty}})$, providing the first full study of a PCF quadratic map with non-abelian constant field.

math.NT↗

pANO-F12: An atomic natural orbital-inspired route to more compact basis sets for F12 explicitly correlated methods

Explicitly correlated methods such as MP2-F12 and CCSD(F12*) exhibit much faster basis set convergence (asymptotically $\propto L^{-7}$, with L the highest angular momentum) than orbital-only approaches. Yet it has been pointed out that cc-pVnZ-F12 basis sets themselves are substantially larger than the corresponding cc-pVnZ, and specifically that cc-pVDZ-F12 is the size of cc-pVTZ. One way to generate compact basis sets in an orbital-only context are Atomic Natural Orbital (ANO) basis sets [J. Almlöf and P. R. Taylor, JCP 86, 4070 (1987)]. However, obtaining the required first-order reduced density matrix while properly accounting for the F12 geminal is problematic. In this work, we show that an energy minimization-based contraction process under linear independence constraints yields `pseudo-ANO' (pANO) basis sets that are functionally equivalent in quality. Subsequently, we apply this recipe to obtain pANO-F12 basis sets from the same elements, then validate them for several thermochemical benchmarks and for the hypersensitive out-of-plane vibrations of benzene. We show that, unlike cc-pVnZ-F12, pANO-F12 exhibits the familiar shell structure seen in cc-pVnZ and ANO basis sets, and that pANO-F12 offers a route to more compact F12 basis sets more amenable to medium-sized systems, especially in conjunction with localized pair natural orbital approaches. Overall, the pANO approach is most beneficial for the smaller double-and triple-zeta basis sets, offering either superior performance to cc-pVnZ-F12 at same cost, or similar performance at lower cost.

physics.chem-ph↗

AFDM as a Software Upgrade of OFDM: One Firmware Patch, a New Frontier

In this white paper, we summarize for the benefit of the wider research community on wireless communications, the two key results that we shared with the attendees of the 2026 IEEE Communication Theory Workshop in Azores, Portugal, about affine frequency division multiplexing (AFDM). Firstly, we show that in contrast to the wide perception by most researchers, AFDM can be implemented at marginal costs by means of a simple software upgrade (firmware patch) of conventional orthogonal frequency division multiplexing (OFDM), indicating that its adoption can potentially be achieved across a wide range of OFDM-based wireless infrastructure and systems. The most crucial relevance of this finding is that such an upgrade would enable, under the specific conditions of the corresponding systems and their applications, exploiting various advantageous features of AFDM, including robustness to doubly dispersive channels (i.e., to support high-mobility use-cases in 6G), inherent integrated sensing and communications (ISAC) compatibility (i.e., to support sensing use-cases in 802.11bf), and the straightforward introduction of low-complexity physical-layer security at the waveform level (as needed in next-generation IoT systems). Secondly, we also show that the same mathematical principles underpinning the aforementioned finding, also imply an inherent capability of AFDM to reap the full uncoded diversity of static linear time-invariant (LTI) channels, demonstrating that this simple upgrade taps into previously undiscovered strengths of multicarrier waveforms.

eess.SP↗

Intertwined quantum phase transitions in the even-even $^{90-100}$Sr isotopes

The even-even $^{90-100}$Sr isotopes are identified as a region of intertwined quantum phase transitions (IQPTs). In this scenario, a quantum phase transition involving the crossing of normal and intruder configurations is accompanied by a shape evolution within the intruder configuration. Using the interacting boson model with configuration mixing (IBM-CM), its is shown that the strontium chain exhibits the IQPT scenario, where the intruder configuration evolves from a near-spherical structure in $^{90\text{--}96}$Sr to a deformed one in $^{98,100}$Sr, while the normal and intruder configurations cross between $^{96}$Sr and $^{98}$Sr. As a result, the ground state changes abruptly from a weakly collective normal configuration to a deformed intruder configuration. Evidence for this scenario is provided by a detailed comparison with experimental excitation energies, isotope shifts, and monopole $E0$ transition strengths, together with the configuration and $n_d$ decompositions of the calculated wave functions. The results place the strontium isotopes alongside the neighboring zirconium chain as a realization of IQPTs in the intricate $A\approx100$ region.

nucl-th↗

Every Component Is a Lookup: One Linear Graph for Interaction, Composition and Attribution

Interpretability methods for transformers are typically built around separate questions: which components interact, how information routes to the output, and which input tokens contribute. Because these methods rely on different assumptions, their answers are difficult to relate. We argue that two architecturally motivated assumptions suffice to address all three questions: attention and MLPs share a key-value form, $ϕ(S)\,U$, in which $ϕ(S)$ selects over values $U$, and components read from an additive residual stream, the sum of component outputs. Holding these selections at their forward-pass values turns the model into a computational graph, of which component interactions, composition paths, and token attribution are different readouts. We develop Unpack, a backward attribution procedure over this graph, and validate each readout against the corresponding established test: interaction scores predict ablation effects across models from 160M to 6.9B parameters, recovered routes reproduce established circuits down to the key, query, or value branch the circuit specifies, and token attribution passes the same faithfulness test as dedicated attribution methods. The results suggest that these two assumptions suffice for the interpretability questions above. On a task with a known circuit, we find that contribution and causal effect can differ, and that the difference has a recognisable signature: components that matter for the task change their contribution when the task is removed from the input, while components that act like a bias term do not. Code is available at https://github.com/Fun-Cry/unpacklm.

cs.LG↗

A superlinear improvement on line-free sets in $\mathbb{F}_p^3$

Building on an earlier result of the author together with Elsholtz, Führer, Füredi, Pach, Simon and Velich, we present an improved construction for a line-free set in $\mathbb{F}_p^3$, showing that $r_p(\mathbb{F}_p^3)\ge (p-1)^3+\frac18 p^{3/2} - O(p)$ as $p\to \infty$. This results in the first superlinear-term improvement over the standard hypercube construction $\{0,1,\ldots,p-2\}^3$. Via complementation, a line-free set in $\mathbb{F}_p^3$ corresponds to a $2$-blocking set in the affine geometry $AG(3,p)$, hence we also obtain an upper bound of $3p^2-\frac18p^{3/2}+O(p)$ on the smallest size of such a $2$-blocking set.

math.CO↗

Benchmarking the Limits of In-Context Reinforcement Learning for Ad-Hoc Teamwork

In-Context Reinforcement Learning (ICRL) has enabled foundation agents to adapt instantaneously to novel tasks, yet its efficacy in Ad-Hoc Teamwork (AHT)-where coordination with unknown partners is required-remains unexplored. To rigorously evaluate this, we introduce a large-scale benchmark ICRL4AHT, built upon a high-throughput JAX implementation of Overcooked-V2. Our benchmark includes a large, diverse teammate suite spanning both RL and heuristic policies, enabling controlled train-test shifts, and provides a reproducible end-to-end pipeline for teammate generation, learning-history collection, dataset construction, and online multi-episode evaluation. We evaluate representative history-conditioned ICRL algorithms, including Algorithm Distillation (AD) and Decision-Pretrained Transformer (DPT), across millions of transitions. Results reveal notable limitations: contrary to their success in single-agent domains, these baselines fail to exhibit robust test-time adaptation in multi-agent settings. Specifically, these methods frequently underperform random baselines across both unseen teammate and unseen layout tracks, with no clear in-context improvement over long horizons. These findings highlight the challenges of strategic inference under partial observability within the OvercookedV2 AHT protocol, establishing our benchmark as a critical testbed for next-generation coordination algorithms.

cs.AI↗

Proactive for Uncertainty: Cause-Aware Error Diagnosis and Interactive Clarification for Spoken Dialogue Systems

Cascaded Automatic Speech Recognition - Large Language Model (ASR-LLM) pipelines remain popular for industrial Spoken Dialogue Systems (SDS), primarily because their decoupled design ensures perceptual verifiability. However, cascaded systems suffer from error propagation, as transcription failures inevitably cascade to subsequent components, thereby degrading the final interaction quality. Although ASR confidence scores offer a simple filter for unreliable inputs, this approach is fundamentally limited because it typically fails to detect deletion errors or to distinguish between acoustic (inability to hear clearly) and linguistic (inability to understand) mismatches, both of which require targeted recovery strategies. In this paper, we propose a cause-aware error recovery paradigm that fundamentally rethinks robustness in SDS. Unlike traditional confidence filtering, we introduce a suite of small precision-focused detectors that exploit deep ASR latent representations to disentangle token-level errors into perception, comprehension, and deletion failures. This fine-grained diagnostic intelligence empowers the LLM to orchestrate targeted, multi-turn clarification strategies, effectively transforming ambiguous signals into seamless user interactions. Experimental results validate the precision of our approach, which more than doubles the recall on domain-shift errors (57.96% vs. 23.66%) compared to baselines. Crucially, this diagnostic precision yields up to a 31% reduction in WER and a 19% improvement on the downstream task across diverse accents, distortions, and domains.

cs.CL↗

Lubin-Tate representations over nontrivial finite Galois extensions of $\mathbb{Q}_{p}$ are not Aut-intrinsically Hodge-Tate

In the present paper, we show that, for an odd prime number $p$ and a nontrivial finite Galois extension $k$ of $\mathbb{Q}_{p}$, the $p$-adic representation of the absolute Galois group of $k$ determined by a Lubin-Tate formal group over the ring of integers of $k$ is not Aut-intrinsically Hodge-Tate [in the sense of Hoshi]. This settles the odd-degree cases left open in the previous works of Hoshi and the author and, together with the known even-degree case, completes the picture for finite Galois extensions of $\mathbb{Q}_{p}$ in the case where $p$ is odd. This exhibits a sharp contrast, from the viewpoint of anabelian geometry, between the $p$-adic cyclotomic character and other $p$-adic Lubin-Tate characters.

math.NT↗

Nonlocal problems for Laplace equation in Bochner spaces

We study the Laplace equation posed in the unbounded rectangular domain $Π= I \times (0,\infty)$ with $I= (0,2π)$, and subject to nonlocal boundary conditions on $\partial Π$ in the trace sense. The analysis is carried out in the Bochner-Sobolev space $W^2_{p,1}(Π;X)$, associated with the Bochner space $L^{p,1}(Π;X)$, with $ p \in (1,\infty)$ and $X$ is a suitable Banach space. To solve the problem, we employ a generalized spectral method. In particular, we introduce the notion of $\otimes$-basis generated by tensor products and extend the classical scheme known from the scalar case to the present setting. Moreover, we prove that the system of root functions of the corresponding nonlocal spectral problem forms a $\otimes$-basis in $L^p(I;X)$.

math.AP↗

A Multimodal 3D Foundation Model for Light Sheet Fluorescence Microscopy Enables Few-Shot Segmentation, Classification, and Deblurring

Light sheet fluorescence microscopy (LSM) enables high-resolution, three-dimensional (3D) imaging of biological specimens, providing rich volumetric data for studying cellular organization, pathology, and vascular networks. However, the size, dimensionality, and annotation burden of LSM data make supervised deep learning approaches costly and difficult to scale. Additionally, despite the abundance of unannotated LSM volumes, foundation models for this modality remain underexplored due to computational challenges and the complexity of volumetric representation learning. In this work, we introduce a 3D foundation model for LSM data, pretrained on a large curated collection of 3D images spanning multiple organisms, stains, and imaging protocols. We learn transferable volumetric representations by jointly optimizing for masked reconstruction and image-text alignment. The pretrained backbone drastically reduces the annotation burden, enabling efficient, few-shot adaptation for varied downstream tasks. We evaluate this approach on downstream segmentation, classification, and deblurring. Our results demonstrate consistent improvements over baselines, (1) when measured using standard evaluation metrics and (2) when rigorously assessed by domain experts. This highlights the potential of foundation model pretraining to reduce annotation requirements while improving performance across diverse LSM analysis tasks. Pretrained model weights and code for pretraining and finetuning are publicly available: https://github.com/AdinaScheinfeld/lsm_fm_public_repo.git.

cs.CV↗

Optimal Uncertainty Relations for a Single Observable

Uncertainty relations are usually formulated as trade-offs between two or more observables. Yet the uncertainty of a single observable can already contain an intrinsically quantum contribution arising from its noncommutativity with the state. In the sense of Luo, such a contribution can be quantified by a quantum uncertainty $Q_ρ(A)$, with prominent examples including the Wigner--Yanase and Wigner--Yanase--Dyson skew informations and the quantum Fisher information. Here we ask how strong such single-observable uncertainty relations can be when the quantum state is fixed. We solve this fixed-state optimization problem by determining the largest coefficient $c_{\rm opt}(ρ)$ for which $V_ρ(A)\ge c_{\rm opt}(ρ)Q_ρ(A)$ holds for every observable $A$, and show that it depends only on the smallest and largest eigenvalues of $ρ$. The same optimization principle extends to a new power-commutator family $\frac{1}{2}|[ρ^s,A]|^2$ $(s\ge 1/2)$, providing a direct measure of state--observable noncommutativity beyond the standard quantum-uncertainty framework. More importantly, the optimal noncommutative contribution can be combined with the maximal classical contribution compatible with the state, yielding a sharpened variance decomposition without weakening the optimal bound. For states with exactly two distinct eigenvalues, this relation becomes an identity for every observable; in particular, qubit variances decompose exactly into classical and noncommutative parts. Finally, the corresponding variance-product bounds are no weaker than the Robertson relation for qubit spin observables and can be strictly stronger, showing that a substantial part of the uncertainty usually expressed as a trade-off between distinct observables is already encoded in their individual noncommutativity with the state.

quant-ph↗

Spectra and Ionization Efficiencies of Charged Decay Particles in Kilonova Ejecta

Ionization by radioactive decay products including $α$-particles, $β$-decay electrons, and fission fragments plays a central role in determining the nebular-phase ionization state and spectra of kilonovae. In this work, ionization cross sections, stopping powers, thermalization histories, and particle degradation spectra are calculated self-consistently for $α$-particles and fission fragments propagating through expanding kilonova ejecta. The treatment includes interactions with bound and free electrons, charge evolution of fission fragments, adiabatic energy losses, and particle spectra obtained from the continuous slowing-down approximation and Spencer--Fano formalism. The work per ion pair is evaluated for a range of ejecta compositions and ionization states. Despite the large differences in mass, charge, and injection energy between $β$-electrons, $α$-particles, and fission fragments, the resulting work per ion pair is found to be remarkably similar across all decay channels and target species considered. In particular, heavy-element ions exhibit nearly identical normalized ionization efficiencies for all decay products. This robustness arises because the ionization cross sections and stopping powers are governed by the same underlying collision physics, causing the ratio between ionization and energy loss to remain approximately constant over the relevant energy range. The results imply that the ionization state of late-time kilonova ejecta depends only weakly on the dominant radioactive decay channel, even in ejecta where $α$-decay and fission dominate the heating budget.

astro-ph.HE↗

Problem-Specific Basis Quantum State Readout via Proper Orthogonal Decomposition

Quantum computing offers a promising approach to accelerating partial differential equation (PDE) solvers for large-scale, real-world problems. However, reconstructing the classical representation of a solution from a quantum state remains a significant computational bottleneck. We propose a problem-specific method, termed proper orthogonal decomposition-based readout (PODR), to improve this efficiency. This method comprises offline and online stages. In the offline stage, a set of basis functions capturing the dominant features of the target problem is constructed using classical computations. In the online stage, the quantum state is projected onto this reduced basis, and only a small number of coefficients is extracted to reconstruct the solution. PODR is particularly advantageous for simulations involving varying parameters, common in computational fluid dynamics (CFD), because the POD basis functions are constructed only once in the offline stage and reused during the online stage. Applying PODR to benchmark CFD problems demonstrates a significant reduction in online computational cost compared with conventional readout methods.

quant-ph↗

Persona Prompting in Multimodal Urban Perception: Descriptive Convergence and Interpretive Variation

This study examines how persona prompting shapes language generated by two multimodal large language models in urban perception, a setting for examining subjective interpretations of shared visual evidence. We organize outputs into three functional layers: descriptive grounding (captions), intermediate semantic layer (perception tags), and interpretive framing (justifications). Using approximately 60,000 persona-conditioned annotations from each of two MLLMs, Qwen3-VL and Gemma4, we find that captions converge strongly across persona profiles and show only small attribute-associated differences. Justifications vary substantially more: economic status produces the largest difference in both models, with political orientation and personality also prominent. Paired image-level comparisons confirm larger justification than caption differences for these three attributes. For perception tags, personas sharing the same attribute level produce more similar tag sets than personas with different attribute levels, with the largest separation observed for economic status. Exploratory topic analysis further suggests persona-specific evaluative emphasis. Across models, profile-pair similarity patterns are strongly correlated for all three output types, although agreement is lowest for justifications. Overall, persona prompting affects interpretive framing more strongly than descriptive grounding.

cs.CL↗

A finite victory over de Bruijn-Erdős in interval discrepancy

We study a finite form of the classical interval discrepancy problem. Starting from the unit interval, one repeatedly splits an existing interval into two until $n$ intervals have been produced. The discrepancy of such a process is the maximum, over all intermediate stages, of the ratio between the longest interval and the shortest interval. A theorem of de Bruijn and Erdős from 1949 shows that this ratio must approach $2$ as $n\to\infty$, and they give a sharp construction achieving this bound. For fixed $n$, their construction gives the upper bound $\operatorname{disc}(n)\leq 2-\frac{3}{2n}+O\bigl(\frac 1{n^2}\bigr)$. In this paper, we prove that $\operatorname{disc}(n)=2^{1-1/\lceil n/2\rceil}=2-\frac{4\ln 2}{n}+O\bigl(\frac 1{n^2}\bigr)$ for every $n$.

math.CO↗