arXiv ScienceSearch

arXiv subjects

Xuan Luo

Publications and source records attributed to Xuan Luo.

At least 19 recordsLinked to original sources

Bioinfoysis Technical Report

Large language model agents have shown promise in bioinformatics, but most existing systems focus primarily on producing final answers, treating planning, tool use, and code execution as transient interactions. This design is poorly suited to long-horizon bioinformatics tasks, where conclusions must remain connected to the data, computations, and intermediate evidence that support them. We introduce \textbf{Bioinfoysis}, a multi-agent harness that represents each request as a persistent, artifact-grounded analysis run. Bioinfoysis combines global planning with step-wise, evidence-driven replanning: the planner maintains an executable checklist and revises pending steps using structured handoffs returned after each worker execution. These handoffs bind intermediate results to their responsible agent, checklist step, and plan generation, preventing stale evidence from being silently reused after replanning. A controlled runtime validates generated scripts, tables, and figures before they are used in downstream analysis or reporting, while role-specific context, persistent memory, and governed bioinformatics skills support reliable execution over long analysis trajectories. We evaluate Bioinfoysis on BixBench and two question-answering tracks of LAB-Bench 2. On BixBench, Bioinfoysis achieves state-of-the-art accuracy of 82.4\%. Across four underlying language models, Bioinfoysis increases average accuracy from 27.81\% to 64.13\% on SeqQA2 and from 3.13\% to 31.25\% on DbQA2. These results demonstrate that reliable bioinformatics automation depends not only on model capability, but also on the harness that governs planning, execution, memory, and evidence flow. We hope that the emergence of Bioinfoysis will play a driving and leading role in the development of the bioinformatics community. Our demo website can be seen in https://report.bioinfoysis.com/.

cs.AI

SAGE: From Direct Answering to Evidence-Grounded Inference for Chinese Ancient Document Understanding

Chinese ancient document understanding demands complex visual, linguistic, and historical reasoning. Current Large Vision-Language Models (LVLMs) typically rely on an opaque, single-pass generation paradigm, often producing overconfident and weakly grounded responses. To address this, we propose SAGE, an evidence-grounded multi-agent framework that reformulates Chinese ancient document understanding as evidence-grounded inference rather than direct answer generation. SAGE coordinates specialized agents for task-aware planning, tool-mediated evidence acquisition, claim-level verification, and bounded replanning under a constrained shared-state runtime. This design supports bounded evidence seeking, answer revision, and abstention when grounding is insufficient. Experiments on the AncientDoc benchmark show that SAGE consistently outperforms matched direct-answering baselines across three LVLM backbones. Remarkably, SAGE with Qwen3.5-9B surpasses much larger monolithic LVLMs on most evaluated metrics, highlighting the importance of structured, evidence-grounded inference beyond model scaling.

cs.CL

The role of triangle singularity in the $B^0 \to D^- \pi^+ a_0(980)(\pi^0 \eta)$ decay

The triangle singularity interpretation of the \replyhe{$a_1(1420)$} observed by the COMPASS Collaboration has been widely accepted. In this work, we investigate the triangle mechanism in the decay $B^0 \to D^- \pi^+ a_0(980)(\pi^0 \eta)$, where the $a_0(980)$ is treated as a dynamically generated state. The $\bar{K}^{*0}$-$K^+$-$K^-$ triangle loop originates from the decay $B^0 \to D^- K^+ \bar{K}^{*0}$, which has been observed by the Belle Collaboration, followed by the subsequent decay $\bar{K}^{*0} \to K^- \pi^+$. The triangle amplitude develops a pronounced peak around 1420 MeV, which is reflected in the invariant mass spectrum of the $\pi a_0(980)$ system. The differential decay width is calculated and exhibits a narrow peak around $980~\mathrm{MeV}$ in the $\pi^0 \eta$ invariant mass distribution. Furthermore, the invariant mass distribution of the $\pi^+ \pi^0 \eta$ system shows a clear peak around $1420~\mathrm{MeV}$, which further \replyhe{confirms} the triangle singularity explanation of the \replyhe{$a_1(1420)$}. We expect that the proposed $B^0$ decay mode could provide a potential platform for further exploring the triangle-singularity nature of the \replyhe{$a_1(1420)$} and could be tested in future experiments such as LHCb, BESIII, and Belle~II.

hep-ph

Unified spectroscopy of $P$-wave flavor-sextet heavy baryons from QCD sum rules

We investigate the $P$-wave flavor-sextet charmed and bottom baryons within heavy quark effective theory. A key motivation is provided by the recent evidence for the $\Xi_c(2882)^0$, which completes a sequence of four narrow $\Xi_c$ structures together with the $\Xi_c(2923)^0$, $\Xi_c(2939)^0$, and $\Xi_c(2965)^0$. Their characteristic mass-splitting pattern closely parallels that of the $\Omega_c(3000)^0$, $\Omega_c(3050)^0$, $\Omega_c(3066)^0$, and $\Omega_c(3090)^0$ states, providing strong constraints on their spectroscopic assignments. We classify the seven $P$-wave states in each of the $\Sigma_Q$, $\Xi_Q^\prime$, and $\Omega_Q$ sectors ($Q=c,b$), calculate their masses and intra-doublet mass splittings using QCD sum rules, and study their strong decays using light-cone sum rules. Configuration mixing between states with the same quantum numbers is also investigated. The combined analysis favors a common interpretation of the four narrow $\Xi_c$ and the four lowest narrow $\Omega_c$ structures in terms of the corresponding $\lambda$-mode excitations. Extending the same framework to the bottom sector, we obtain a coherent picture of the observed $\Sigma_b$, $\Xi_b$, and $\Omega_b$ structures. In particular, the $\Sigma_b(6097)$, $\Xi_b(6227)$, and $\Omega_b(6350)$ structures may each contain an unresolved pair of nearby $P$-wave states. We also predict two additional $\Sigma_b$ states, two additional $\Xi_b^\prime$ states, and one additional $\Omega_b$ state, all of which are expected to be relatively narrow and remain to be identified experimentally.

hep-ph

Simple-OPD: Demystifying Warm-up for On-policy Distillation

On-policy distillation (OPD) trains a student on its own rollouts with token-level supervision from teacher models, but its effectiveness can depend strongly on the warm-up stage before OPD. In this paper, we demystify warm-up for OPD from both data and training perspectives. For data, we find that effective warm-up relies on teacher-compatible chain-of-thought supervision, and that even incorrect teacher rollouts can provide comparable benefits to correct ones. This suggests that warm-up primarily transfers a teacher-compatible thinking pattern rather than merely correct answers. For training, we show that low-rank adaptation (LoRA) with a near-saturation training duration better balances in-domain adaptation and out-of-distribution generalization than full-parameter SFT. Based on these findings, we propose Simple-OPD, a plug-and-play initialization method that warms up the student on teacher-generated CoT with LoRA before OPD. Experiments across diverse settings demonstrate the effectiveness and robustness of Simple-OPD.

cs.CL

Quantum numbers of excited $\Xi_c^\prime$ and $\Omega_c$ baryons and the $P$-wave $\Sigma_c$ spectrum

Recent precision measurements of excited heavy baryons, combined with systematic theoretical studies, make it possible to resolve the fine structure of their spectra. We calculate the masses and strong-decay properties of the $P$-wave charmed baryons using QCD sum rules and light-cone sum rules within heavy quark effective theory (HQET). Although seven states are allowed in each flavor sector, we find that only four $\Sigma_c$, four $\Xi_c^\prime$, and five $\Omega_c$ states are expected to be experimentally resolvable. The similar mass-splitting patterns of $\Xi_c(2882)$, $\Xi_c(2923)$, $\Xi_c(2939)$, and $\Xi_c(2965)$ and of $\Omega_c(3000)$, $\Omega_c(3050)$, $\Omega_c(3066)$, and $\Omega_c(3090)$, together with our theoretical results, lead to the successive quantum-number assignments $J^P=1/2^-,3/2^-,3/2^-$, and $5/2^-$. We tentatively interpret $\Omega_c(3119)$ as a predominantly $\rho$-mode excitation with $J^P=3/2^-$. We also predict the masses, widths, and dominant decay modes of four resolvable $P$-wave $\Sigma_c$ states, which may overlap within the observed $\Sigma_c(2800)$ and $\Sigma_c(2900)$ structures. Precision spectroscopy of the narrow $\Xi_c$ and $\Omega_c$ states thus provides a route to resolving the $P$-wave $\Sigma_c$ spectrum.

hep-ph

Online Neural Space Time Memory for Dynamic Novel View Synthesis

Online novel view synthesis from multi-view streaming videos faces a fundamental trade-off: maintaining a persistent, long-horizon memory to reconstruct temporarily occluded regions while operating under strict real-time constraints. While Test-Time Training (TTT) offers a powerful memory mechanism, standard models mandate gradient-based memory updates at every frame to adapt to the changing motion in dynamic scenes. The computational cost of heavy memory updates precludes real-time application and can lead to instability over long contexts. Given that memory updates are more demanding than memory application and video content is largely redundant, we propose to decouple the frequencies of these two processes. Our approach performs periodic memory updates while applying the memory on a per-frame basis, using cross-view attention to manage deformations between the prior memory state and the current frame. To lock in the historical context, we introduce two critical mechanisms: an auxiliary Memory Loss that forces persistent internalization of the scene, and a Memory Caching strategy that regularizes active weights against catastrophic drift. Our method demonstrates real-time, state-of-the-art performance on scenes with dynamic human motion as well as minute-scale online memorization.

cs.CV

RaMark: Radioactive Watermarking for Generated Tabular Data

Recent advances in generative modeling have made generated tabular data a practical solution for privacy-sensitive data sharing, where watermarking enables ownership verification. However, existing watermarking methods fundamentally fail under retraining attacks, in which an adversary retrains a generative model on a watermarked dataset and regenerates high-utility data that no longer carries the watermark. We address this challenge by introducing radioactivity, the property that a watermark remains detectable after generative model retraining, and propose RaMark, a radioactive watermarking method that embeds a sinusoidal dependency as an intrinsic component of the data distribution. By coupling the watermark with the underlying distribution, RaMark ensures that any generative model preserving data utility also has to preserve the watermark. We theoretically show that with high probability removing watermark degrades utility and alters data distribution. Extensive experiments on two real-world tabular datasets, under a large-scale ownership verification setting with $10^5$ independent data owners, demonstrate that RaMark achieves substantially stronger radioactivity than seven state-of-the-art methods and consistently outperforms them against both retraining and data modification attacks.

cs.CR

Study of exotic hadron states in the $DD^{*}$ system via the complex momentum representation and Green's function method

In this paper, we propose a novel approach to investigate exotic hadronic states. For the $DD^{*}$ system, we employ the projection operator method to derive the momentum-space interaction potential. Subsequently, the complex momentum representation (CMR) method is adopted to realize a unified description of bound states, resonant states, and the continuum. By combining the Green's function and the CMR, the scattering phase shifts and cross sections are determined. This integrated approach provides a comprehensive framework for analyzing the scattering dynamics of the $DD^{*}$ system. In the hadronic molecular state framework, the $X(3872)$, $T_{cc}^+$, and $Z_c(3900)$ states can be consistently explained as bound states, while the $G(3900)$ can be interpreted as a $P$-wave resonant state. The decomposition of the scattering phase shifts and cross sections facilitates understanding the roles of resonant and continuum spectrum.

hep-ph

Spin-orbit-enabled Fermi-surface splitting in noncollinear antiferromagnetic SmBi

Spin-split electronic structures in compensated antiferromagnets are commonly sought in the nonrelativistic limit, where magnetic order lifts spin degeneracy without spin-orbit coupling (SOC). Whether SOC can instead be the indispensable symmetry-breaking ingredient remains largely unexplored. Here we combine quantum oscillations detected by ultrahigh-sensitivity ac magnetostriction, magnetic-symmetry analysis and first-principles calculations to resolve the bulk Fermi-surface evolution of SmBi across two successive antiferromagnetic (AFM) transitions. New oscillation branches emerge below TN and undergo a further reconstruction below T*, whereas isostructural SmSb shows no comparable change. For the candidate noncollinear orders of SmBi, breaking global parity-time symmetry is insufficient in the nonrelativistic limit because residual spin-space symmetries protect twofold band degeneracy; conversely, SOC alone cannot lift the degeneracy of the centrosymmetric paramagnetic (PM) phase. Only the coexistence of noncollinear order and SOC locks spin to the lattice and removes the residual protection. SmBi therefore realizes a cooperative, relativistic route to spin-split Fermi surfaces, broadening unconventional magnetism beyond systems whose splitting is already present in the nonrelativistic limit.

cond-mat.mtrl-sci

Dual Dimensionality for Local and Global Attention

Decoder-only Transformers compute attention over the KV cache of preceding tokens. Keys (and Values) are typically represented with the same dimensionality, regardless of its distance from the prediction target. In natural language, however, the next word is most strongly influenced by the immediately preceding tokens. We hypothesize that local and distant tokens impose asymmetric demands on representational capacity: local tokens are more critical for predicting immediate outputs and thus require richer representations, whereas distant tokens primarily serve as long-range memory, for which lower-dimensional representations may suffice. We formalize this idea as Distance-Adaptive Representation (DAR), implemented in a controlled setting that preserves full-dimensional representations within a local context window while assigning reduced-dimensional representations (e.g. 1/4 of the original dimensionality) to tokens beyond that window. Across multiple pretraining scales (70M to 410M parameters), as well as continued supervised fine-tuning on a 1B-scale model, this approach closely matches the performance of full-dimensional baselines. In contrast, uniformly reducing dimensionality across all token positions leads to worse performance. These results challenge the common assumption that key and value dimensionality should be uniform across token positions. Our findings suggest a new direction for designing attention architectures that adaptively allocate representational capacity across sequences, enabling further reductions in KV cache during inference.

cs.CL

BAIT: Boundary-Guided Disclosure Escalation via Self-Conditioned Reasoning

In this work, we propose BAIT (Boundary-Aware Iterative Trap), a three-step jailbreak framework that approaches malicious goals through internal disclosure. BAIT first asks the model to identify the protection boundary, then requires it to refine that boundary, and finally requests a detailed example. By expanding each step upon the model's previous responses, BAIT turns the model's own reasoning and consistency tendency into a disclosure pathway. Experiments on AdvBench, JailbreakBench, AIR-Bench, and SORRY-Bench demonstrate that BAIT consistently achieves strong attack success rates across top-tier large language models, significantly advancing conventional jailbreak baselines. Further analysis reveals that: 1) prevention-oriented framing significantly outperforms direct knowledge request; 2) the refinement step plays a critical role in disclosure escalation; and 3) the first two steps have a certain chance of eliciting harmful content while triggering little filtering.

cs.CR

Forecasting Japanese elections: A nonlinear machine-learning approach

Despite Japan being one of the world's largest advanced democracies, the development of election forecasting models for its national elections remains limited. This study introduces nonlinear machine-learning forecasting models, based on decision tree and ensemble learning methods, for predicting the outcomes of Japanese lower-house elections. To assess the methodological benefits of our approach, we replicated the theoretical framework and dataset of Lewis-Beck and Tien's (LBT) foundational statistical forecasting model for Japanese elections. Our models demonstrated moderately but consistently improved predictive accuracy compared to LBT's model in both in-sample and out-of-sample evaluations, suggesting that nonlinear algorithms offer an alternative approach to classical linear methods in capturing complex electoral dynamics. This study represents one of the earlier applications of nonlinear machine-learning techniques to single-country election forecasting. It offers a replicable framework that, when combined with the country-specific electoral theories of other nations, may enhance the predictive performance of forecasting models in broader national contexts.

physics.soc-ph

Anomalies in the thermal conductivity of honeycomb antiferromagnet MnPS$_{3}$

Intrinsic two-dimensional magnets serve as a good platform to explore collective, charge-neutral and low-energy excitations. Distinguishing the crucial role of them in experimental aspect remains a challenge for decades. Here, we study the thermal transport in honeycomb antiferromagnet MnPS$_{3}$ with $T_N$=78 K down to very low temperatures (<0.01$T_N$). At high temperatures (>0.1$T_N$), the field dependence of the thermal Hall conductivity exhibits a linear phonon Hall effect and a peak associated with the spin-flop transition due to a strong spin-lattice coupling, well reproducing the previous report (Phys. Rev. B 110, 165147 (2024)). Notably, below 2 K, we find that the field dependence of the thermal Hall conductivity exhibits sign reversals within the spin-flop phase, at which the field dependence of the longitudinal thermal conductivity also shows multiple valleys. We suggest that these anomalies are caused by the redistribution of Berry curvature in magnon bands, demonstrating the superior performance of the thermal Hall measurements to detect the Berry curvature distributions in magnetic insulators.

cond-mat.mtrl-sci

WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors

Commercial video generation systems such as Seedance2.0 and Veo3.1 have rapidly improved, strengthening the view that video generators may be evolving into "world simulators." Yet the community still lacks a benchmark that directly tests whether a model can reason about how an observed world should evolve over time. We introduce WorldReasonBench, which reframes video generation evaluation as world-state prediction: given an initial state and an action, can a model generate a future video whose state evolution remains physically, socially, logically, and informationally consistent? WorldReasonBench contains 436 curated test cases with structured ground-truth QA annotations spanning four reasoning dimensions and 22 subcategories. We evaluate generated videos with a human-aligned two-part methodology: Process-aware Reasoning Verification uses structured QA and reasoning-phase diagnostics to detect temporal and causal failures, while Multi-dimensional Quality Assessment scores reasoning quality, temporal consistency, and visual aesthetics for ranking and reward modeling. We further introduce WorldRewardBench, a preference benchmark with approximately 6K expert-annotated pairs over 1.4K videos, supporting pair-wise and point-wise reward-model evaluation. Across modern video generators, our results expose a persistent gap between visual plausibility and world reasoning: videos can look convincing while failing dynamics, causality, or information preservation. We will release our benchmarks and evaluation toolkit to support community research on genuinely world-aware video generation at https://github.com/UniX-AI-Lab/WorldReasonBench/.

cs.CV

Discovery of a hybridization-wave electronic order in a van der Waals Kondo lattice

Kondo lattice systems, in which localized magnetic moments coherently hybridize with itinerant electrons, exhibit a rich landscape of emergent quantum phenomena. Within this framework, the hybridization strength itself has been theoretically proposed as a spatially modulated order parameter, giving rise to a so-called hybridization wave. However, direct experimental evidence of this quantum state has remained an outstanding challenge. Here, we report the direct observation of a hybridization wave in the layered transition metal dichalcogenide 6R-TaS2, a naturally occurring heterostructure composed of alternating 1T- and 1H-TaS2 layers. Using scanning tunneling microscopy and spectroscopy (STM/STS), we identify the hybridization gap in 1T layer, demonstrating the establishment of a coherent Kondo lattice. Notably, we discover that the hybridization gap present a uniaxial unit-cell doubling modulation, which breaks the both translational and rotational symmetries of the underlying Star-of-David superlattice. Such unit-cell doubling is not caused by structural topography, and therefore, constitutes the real-space visualization of the hybridization-wave order. Furthermore, the hybridization wave correlates with an energy-dependent nematic order that shares the same periodicity and orientation, revealing intertwined electronic instabilities. Our findings not only validate a long-standing prediction but also establish layer-engineered van der Waals materials as a versatile platform for exploring and controlling hybridization-driven quantum phases.

cond-mat.str-el

Speech-Omni-Lite: Portable Speech Interfaces for Vision-Language Models

While large-scale omni-models have demonstrated impressive capabilities across various modalities, their strong performance heavily relies on massive multimodal data and incurs substantial computational costs. This work introduces Speech-Omni-Lite, a cost-efficient framework for extending pre-trained Visual-Language (VL) backbones with speech understanding and generation capabilities, while fully preserving the backbones' vision-language performance. Specifically, the VL backbone is equipped with two lightweight, trainable plug-and-play modules, a speech projector and a speech token generator, while keeping the VL backbone fully frozen. To mitigate the scarcity of spoken QA corpora, a low-cost data construction strategy is proposed to generate Question-Text Answer-Text-Speech (QTATS) data from existing ASR speech-text pairs, facilitating effective speech generation training. Experimental results show that, even with only thousands of hours of speech training data, Speech-Omni-Lite achieves excellent spoken QA performance, which is comparable to omni-models trained on millions of hours of speech data. Furthermore, the learned speech modules exhibit strong transferability across VL backbones.

eess.AS

Evaluating Proactive Risk Awareness of Large Language Models

As large language models (LLMs) are increasingly embedded in everyday decision-making, their safety responsibilities extend beyond reacting to explicit harmful intent toward anticipating unintended but consequential risks. In this work, we introduce a proactive risk awareness evaluation framework that measures whether LLMs can anticipate potential harms and provide warnings before damage occurs. We construct the Butterfly dataset to instantiate this framework in the environmental and ecological domain. It contains 1,094 queries that simulate ordinary solution-seeking activities whose responses may induce latent ecological impact. Through experiments across five widely used LLMs, we analyze the effects of response length, languages, and modality. Experimental results reveal consistent, significant declines in proactive awareness under length-restricted responses, cross-lingual similarities, and persistent blind spots in (multimodal) species protection. These findings highlight a critical gap between current safety alignment and the requirements of real-world ecological responsibility, underscoring the need for proactive safeguards in LLM deployment.

cs.CL