arXiv ScienceSearch

arXiv subjects

Yong Xu

Publications and source records attributed to Yong Xu.

At least 19 recordsLinked to original sources

One Prompt Does Not Fit All: Self-Meta-Evolve for Personalized Information Extraction

Large language models (LLMs) are increasingly deployed for enterprise information extraction (IE), where the same document must be reorganized differently for each user. Existing prompt optimization methods, however, rely on a single prompt optimized against a global objective, which is misaligned with the inherent user heterogeneity of real workplaces. We formulate enterprise IE as per-user prompt adaptation under interaction feedback and propose Self-Meta-Evolve, a hierarchical framework that maintains a dedicated prompt for each user and continuously refines it through a dual-loop process: an inner loop that edits structured prompts based on persona-conditioned feedback, and an outer loop that evolves the meta-prompt itself by distilling successful editing patterns. To enable scalable training and evaluation, we release a persona-driven IE benchmark of 292 simulated enterprise users, paired with a reproducible persona-generation pipeline grounded in O*NET occupational taxonomies. On this benchmark, Self-Meta-Evolve achieves a 74.58% success rate, outperforming the strongest prompt-optimization baseline by 13.56 absolute points, and reaches 52.54\% within only two iterations. A double-blind human study with twenty real professionals further confirms that prompts adapted by our framework win against static baselines in 71% of pairwise comparisons.

cs.AI

Time-Efficient Iterative Learning Planning for Safety-Critical Dynamic Obstacle Avoidance

Autonomous mobile robots require timeefficient planning and safety-critical dynamic obstacle avoidance under constrained onboard computation. While Iterative Learning Planning (ILP) offers lightweight and efficient traversal planning, it lacks explicit mechanisms for dynamic obstacle perception and avoidance. This article extends ILP to safety-critical navigation in dynamic environments by integrating an anticipatory risk-blended control barrier function (ARB-CBF). The extended ILP learns traversal-speed and steering-bias profiles via a fractionalpower update based on local obstacle risk, generating nominal control commands that ARB-CBF modifies at runtime for real-time safety guarantees. Algorithmic analysis demonstrates that the ILP replanning stage scales at O(kN) for k iterations and N waypoints, while ARB-CBF executes with linear complexity. Comprehensive simulations and real-world experiments validate the framework, demonstrating superior temporal efficiency and safety with lower computational overhead compared to optimizationbased baselines, making it highly suitable for resourceconstrained platforms.

cs.RO

Fractional Chern Insulators in Twisted Bilayer Optical Lattices

Twisted bilayer materials provide a versatile platform for realizing correlated topological states. Motivated by the recent experimental realization of atomic Bose-Einstein condensates in twisted bilayer optical lattices, we investigate the emergence of fractional topological phases in such systems. For experimentally realistic parameters, the single-particle spectrum hosts a pair of quasi-degenerate nearly flat moiré bands. We show using self-consistent Hartree-Fock calculations that atomic interactions spontaneously break time-reversal symmetry, lift the quasi-degeneracy, and generate an isolated nearly flat band with a nonzero Chern number. At fractional filling, exact diagonalization of the Hamiltonian projected to this band reveals a threefold degenerate ground-state manifold and a characteristic particle entanglement spectrum, providing strong evidence for a fractional Chern insulator. The resulting many-body gap is of order \(10^{-3}E_R\), corresponding to a nanokelvin energy scale accessible to state-of-the-art cold-atom experiments. Our scheme requires neither spin-orbit coupling, as in transition metal dichalcogenides, nor a finite magnetic field, as in existing twisted-bilayer-graphene experiments, providing a highly tunable route to strongly correlated topological phases in twisted bilayer optical lattices.

cond-mat.quant-gas

SSS: Semi-Supervised SAM-2 with Efficient Prompting for Medical Imaging Segmentation

In the era of information explosion, efficiently leveraging large-scale unlabeled data while minimizing the reliance on high-quality pixel-level annotations remains a critical challenge in the field of medical imaging. Semi-supervised learning (SSL) enhances the utilization of unlabeled data by facilitating knowledge transfer, significantly improving the performance of fully supervised models and emerging as a highly promising research direction in medical image analysis. Inspired by the ability of Vision Foundation Models (e.g., SAM-2) to provide rich prior knowledge, we propose SSS (Semi-Supervised SAM-2), a novel approach that leverages SAM-2's robust feature extraction capabilities to uncover latent knowledge in unlabeled medical images, thus effectively enhancing feature support for fully supervised medical image segmentation. Specifically, building upon the single-stream "weak-to-strong" consistency regularization framework, this paper introduces a Discriminative Feature Enhancement (DFE) mechanism to further explore the feature discrepancies introduced by various data augmentation strategies across multiple views. By leveraging feature similarity and dissimilarity across multi-scale augmentation techniques, the method reconstructs and models the features, thereby effectively optimizing the salient regions. Furthermore, a prompt generator is developed that integrates Physical Constraints with a Sliding Window (PCSW) mechanism to generate input prompts for unlabeled data, fulfilling SAM-2's requirement for additional prompts. Extensive experiments demonstrate the superiority of the proposed method for semi-supervised medical image segmentation on two multi-label datasets, i.e., ACDC and BHSD. Notably, SSS achieves an average Dice score of 53.15 on BHSD, surpassing the previous state-of-the-art method by +3.65 Dice. Code will be available at https://github.com/AIGeeksGroup/SSS.

cs.CV

A Self-Adaptive First-Principles Approach for Magnetic Excited States

The profound impact of excited magnetic states on the intricate interplay between electron and lattice behaviors in magnetic materials is a topic of great interest. Unfortunately, despite the significant strides that have been made in first-principles methods, accurately tracking these phenomena remains a challenging and elusive task. The crux of the challenge that lies before us is centered on the intricate task of characterizing the magnetic configuration of an excited state, utilizing a first-principle approach that is firmly rooted in the ground state of the system. We propose a versatile self-adaptive spin-constrained density functional theory formalism. By iteratively optimizing the constraining field alongside the electron wave function during energy minimization, we are able to obtain an accurate potential energy surface that captures the longitudinal and transverse variations of magnetization in itinerant ferromagnetic Fe. Moreover, this technique allows us to identify the subtle coupling between magnetic moments and other degrees of freedom by tracking energy variation, providing new insights into the intricate interplay between magnetic interactions, electronic band structure, and phonon dispersion curves in single-layered CrI$_3$. This new methodology represents a significant breakthrough in our ability to probe the complex and multifaceted properties of magnetic systems.

cond-mat.mtrl-sci

Ab initio-based Deep-Learning Prediction of Carrier Mobility in Strongly Anharmonic Materials

Predicting charge transport in strongly anharmonic materials, particularly ultralow thermal conductors, remains a major challenge for first-principles methods. In such systems, perturbative treatments of electron-phonon interactions and the harmonic phonon picture often break down, necessitating non-perturbative approaches. The ab initio Kubo-Greenwood(aiKG) formalism provides a rigorous framework for evaluating temperature-dependent carrier transport beyond the harmonic approximation. Nevertheless, its practical application is computationally demanding because it requires large supercells, extensive statistical sampling, and extrapolation to the zero-frequency limit. In this work, we introduce an artificial-intelligence(AI)-assisted aiKG framework that incorporates the deep-learning Hamiltonian model. By predicting the Kohn-Sham Hamiltonian with sub-meV accuracy for supercells of up to 250 atoms, the model bypasses the costly iterative self-consistent field calculations while retaining first-principles reliability within the scope of effects captured by the training data. Using a strongly anharmonic thermal insulator, potassium iodide(KI) as a benchmark system, we demonstrate that the proposed approach enables efficient simulations of electronic structure and transport properties from a large supercell. The framework reproduces temperature-dependent carrier mobilities, spectral functions, and effective masses in close agreement with the underlying density functional theory while reducing computational cost to 10%. These results suggest that the AI-assisted aiKG framework can make non-perturbative transport calculations tractable for strongly anharmonic materials, opening a scalable route towards realistic simulations and accelerated discovery of new functional materials.

cond-mat.mtrl-sci

Inverse Problems for Partial Differential Equations with Jump Discontinuities in Coefficients via Two-Stage Physics-Informed Deep Learning and Statistical Mixture Models

This work proposes a two-stage physics-informed deep learning framework that combines neural-network-based sampling with statistical inference and constrained parameter refinement. In the first stage, a dual-network physics-informed architecture is used, where a main network approximates the PDE solution and an auxiliary coefficient sub network provides a relaxed continuous surrogate of the true discontinuous coefficient field. A gradient-adaptive weighting strategy is incorporated into the physics residual to improve residual training and enhance sampling reliability near possible discontinuity regions. The sampled coefficient values are then analyzed using Bayesian learning for Gaussian mixture models and birth-death Markov chain model selection, which estimate the number of coefficient regimes and provide heuristic search intervals for coefficient values and candidate transition regions. In the second stage, the inverse problem is reformulated as a constrained physics-informed estimator, in which the coefficient is represented explicitly as a hard piecewise-constant function over the spatiotemporal domain. Numerical experiments on different PDE types with jump-discontinuous coefficients demonstrate that the proposed framework achieves accurate parameter estimation with acceptable computational costs compared to existing methods. This work provides an effective integrated workflow for inverse problems governed by PDEs with discontinuous parameter structures, particularly in nonstationary and heterogeneous systems.

stat.ML

Identifying parameter couplings and uncertainties of mixed-noise stochastic systems via full-covariance Gaussian mixture network

Parameter identification of stochastic dynamical systems driven by mixed noises is challenging due to intractable likelihood functions. We propose PENN-GMD, a parameter estimation neural network that maps partially observed trajectories to a Gaussian mixture distribution (GMD) over the system parameters. Unlike conventional uncertainty estimates, the GMD employs full covariance matrices to explicitly reveal parameter couplings and multi-modal likelihood structures. The network is trained by minimizing the negative log-likelihood via a surjective parameterization that hard-encodes all GMD constraints, thereby approximating the true likelihood. We validate the method on five numerical examples with increasing complexity, including systems driven by fractional Gaussian and Lévy noises, oscillators with colored noise, coupled neurons under different observability, and an aeroelastic airfoil with unidentifiable stochastic disturbances. Results demonstrate that PENN-GMD accurately recovers likelihood distributions, captures parameter couplings, and naturally diagnoses non-identifiability through variance broadening or mode splitting. These capabilities establish PENN-GMD as a practical tool for uncertainty-aware parameter identification in complex stochastic systems where conventional likelihood-based methods are infeasible.

stat.ML

StageWAM: Joint-Embedding Stage Prediction for World-Action Models in Robot Manipulation

Generalist robot policies aim to map multimodal observations and linguistic task instructions to actions across diverse tasks. However, existing methods typically represent the future as a fixed, short video-action chunk. This short-term future captures local scene evolution for action execution, but it does not explicitly describe the stage-level future that specifies how a task should progress from its current stage to the next. We therefore distinguish two complementary futures for robot manipulation: a short-term physical future to capture local scene evolution and a stage-level semantic future to represent task progress. We introduce StageWAM, which augments a Motus-based World Action Model (WAM) with Stage-JEPA, a goal-conditioned Joint-Embedding Predictive Architecture (JEPA) predictor. Given the current observation and task instruction, Stage-JEPA uses a frozen V-JEPA2 encoder to extract the current-state representation and predicts the latent target of the next inferred stage. Across 50 RoboTwin 2.0 tasks in clean and randomized environments, StageWAM achieves 90.25% overall success and reduces the mean number of execution steps in successful rollouts by 5.97% relative to the strongest baseline.

cs.RO

Spin-Chirality-Driven Bulk Photovoltaic Effect in van der Waals Magnet CrSBr

The bulk photovoltaic effect (BPVE) can be greatly enriched in magnetic materials. Here, we establish vector spin chirality as a tunable knob for generating an unconventional time-reversal-even magnetic BPVE, comprising the chiral shift current (CSC) and chiral injection current (CIC). Using bilayer antiferromagnetic (AFM) CrSBr as a prototype, we theoretically demonstrate the emergence of CSC and CIC. Compared with conventional photovoltaic currents arising from noncentrosymmetric crystal structures or collinear magnetic orderings, CSC and CIC not only possess comparable magnitudes but also exhibit exceptional tunability. Specifically, they can be switched on and off by magnetic-field-induced spin canting, reversed in direction upon canting-direction reversal, and continuously modulated in intensity via canting-angle variation. Furthermore, we reveal an unusual optical transition channel governing both currents in CrSBr. Our work establishes an unconventional magnetic BPVE with remarkable controllability, paving the way for applications in optoelectronics and magnetic sensing in noncollinear magnets.

cond-mat.mtrl-sci

NCGR: Noise-Conditional Gated Rectification for Camera Extrinsic Perturbations in BEV 3D Object Detection

Camera-based bird's-eye-view (BEV) 3D detection typically assumes accurate and fixed camera extrinsics. In detectors using spatial cross-attention (SCA), extrinsic perturbations displace the image-plane projections of BEV reference points, causing queries to sample features from incorrect regions and degrading detection performance. To address this failure mode, Noise-Conditional Gated Rectification (NCGR) is proposed to compensate for projection errors without explicitly estimating a full six-degree-of-freedom extrinsic correction. For each query-camera pair, a 2D rectification offset is predicted and modulated by a camera-level gate to rectify the base projection before native deformable sampling. During training, the perturbation-derived quantities used to construct the condition and gate are gradually replaced through scheduled interpolation by counterparts generated from an auxiliary scalar predicted from camera features. This transition enables blind inference without perturbation metadata. During training, a weight-shared clean-teacher/perturbed-student pair is used, and the rectification module is supervised by a BEV-consistency objective between the two branches. NCGR is evaluated on nuScenes with simulated dynamic and static extrinsic perturbations. In a five-camera dynamic stress test, NCGR achieves 39.69% NDS, compared with 28.00% for BEVFormer and 33.23% for CAPE. Under clean extrinsics, NCGR maintains performance comparable to that of BEVFormer.

cs.CV

Disorder-Induced Entanglement Phase Transitions in Non-Hermitian Systems with Skin Effects

Non-Hermitian dynamics is ubiquitous in various physical systems. While recent study shows that such a dynamics leads to an area-law scaling of the entanglement entropy due to the non-Hermitian skin effects, it remains unclear how disorder changes the behavior of the entanglement entropy in a non-Hermitian system with skin effects. Here we study the dynamics of a many-body state of free fermions in the paradigmatic Hatano-Nelson model with open boundaries, and find that the area-law behavior of the entanglement entropy in the pristine Hatano-Nelson model develops into a logarithmic scaling for small disorder strength. As we further increase the disorder strength, the system reenters an area-law regime through an entanglement phase transition. At the critical point, the entanglement entropy exhibits a universal algebraic scaling. We further demonstrate the absence of a conformal invariance in the log-law regime by examining the subsystem entanglement entropy, the connected correlation function and the mutual information. Finally, we show the existence of disorder induced entanglement phase transitions in the Hatano-Nelson model with periodic boundaries.

quant-ph

FreeShadow: Training-Free Shadow Removal via Illumination Transfer and Selective Content Preservation in Diffusion Models

Existing supervised and unsupervised shadow removal methods often suffer from limited generalization due to the insufficient diversity of available training datasets, while zero-shot methods tend to produce artifacts and require time-consuming test-time optimization. To address these issues, we propose FreeShadow, a training-free shadow removal method built upon pretrained diffusion models, which exploits diffusion priors for shadow removal without any training or optimization. For illumination recovery, we propose an illumination transfer attention (ITA), which re-weights the self-attention maps in diffusion model to transfer illumination cues from non-shadow to shadow regions. For content preservation, we analyze the effects of illumination variations on self-attention maps and latent high-frequency features in diffusion model, and selectively preserve illumination-invariant components to maintain content fidelity while suppressing residual shadows. We further propose local texture-preserving relighting (LTPR) to mitigate local texture misalignment caused by VAE compression. Extensive experiments demonstrate that our method achieves strong generalization and produces realistic shadow-free images.

cs.CV

Topological Codes from Space Groups: A Route beyond Translation Invariance

Translation invariance underlies all algebraic constructions of topological codes with geometrical locality. It has remained an open question whether codes that generically break this invariance can still be topological and simultaneously possess geometrical locality. Resolving this question is important both fundamentally---deepening our understanding of topological phases---and practically, as relaxing translation invariance could vastly expand the design space and potentially reduce resource overhead in fault-tolerant architectures. Here we introduce space-group codes, in which crystallographic point-group operations enter the bulk stabilizer algebra; bivariate bicycle (BB) codes arise as the translation-only limit. The key insight is that the point-group orbit resolves topology and locality together: it yields a computable algebraic criterion for topological order and a folded geometry in which point-group operations become local. We identify space-group codes whose code parameters exceed the reported same-blocklength, same-check-weight BB benchmarks. In five parameter-matched neutral-atom comparisons, reflection codes reduce the optimized movement cost in every case, by up to $60\%$, while folded placements also enable lower-overhead multilayer superconducting layouts. Treating spatial operations as a code-design variable therefore opens a route to topological codes jointly optimized for information protection and hardware geometry.

quant-ph

GMoT: Gated Motion-Aware Tokenization for Fine-Grained Micro-Gesture Video Reasoning with Multimodal LLMs

Micro-gesture recognition demands the detection of fleeting, spatially localized movements that are frequently overwhelmed by dominant static appearances and background noise. While Multimodal Large Language Models (MLLMs) excel at general video understanding, they inherently struggle with subtle kinematics and often rely on static posture priors. To this end, we propose GMoT, a Gated Motion-Aware Tokenization module that explicitly distills sparse kinematic evidence into a compact sequence prior to temporal modeling. GMoT dynamically spotlights action-relevant regions via spatially weighted pooling, extracts adjacent-frame temporal differencing to capture precise motion energy, and adaptively fuses these cues into the visual stream using a conservatively initialized semantic gate. To transition from simple classification to evidence-grounded reasoning, we further introduce a progressive reward-guided policy refinement paradigm, supported by a semi-supervised annotation pipeline that generates anatomically focused captions. Beyond achieving the best Top-1 accuracy among the compared methods on iMiGUE (67.32\%) and SMG (73.11\%), improving the Qwen3-VL-8B baseline by +6.80 and +3.11 points, our framework introduces Body-Region Grounding (BRG) Recall as an anatomical-grounding proxy conditioned on correct predictions, together with an overlapping-label cross-domain transfer protocol between iMiGUE and SMG. Extensive evaluations demonstrate that our GMoT-augmented model improves in-domain accuracy, retains clear gains under label-preserving corruptions, and improves accuracy-oriented cross-domain transfer under explicit small-split caveats while maintaining high anatomical grounding in its generated rationales.

cs.CV

Large and moderate deviation for rough slow-fast system with level 3 geometric rough path

This work is to give the large deviation principle (LDP) and moderate deviation principle (MDP) for a slow-fast system driven by the mixed fractional Brownian motion (FBM) (1/4,1/3) via the level-3 geometric rough path (RP). Firstly, we provide a different way to lift the translation of mixed FBM in the Cameron-Martin direction to the geometric RP. The LDP turns to the weak convergence of the controlled system by averaging the fast one. Lacking the invariant measure for the controlled fast one, a "replaced process" whose limiting measure could be decoupled from the control term and owes the exponential ergodicity is constructed. Here, the stability under level-3 RPs is established under more elaborate estimates. Besides, different from LDP we give the MDP, that requires more intricate bounded estimates of the deviation component. The MDP for level-2 RP system is recovered as a special case.

math.PR

Gravitational Wave from Graviton Bremsstrahlung during Reheating

We revisit graviton production via Bremsstrahlung from the decay of the inflaton during inflationary reheating. Using two complementary computational techniques, we first show that such 3-body differential decay rates differ from previously reported results in the literature. We then compute the stochastic gravitational wave (GW) background that forms during the period of reheating, when the inflaton perturbatively decays with the radiative emission of gravitons. By computing the number of relativistic degrees of freedom in terms of $ΔN_\text{eff}$, we constrain the resulting GW energy density from BBN and CMB. Finally, we project current and future GW detector sensitivities in probing such a stochastic GW background, which typically peaks in the GHz to THz ballpark, opening up the opportunity to be detected with microwave cavities and space-based GW detectors.

hep-ph

Central limit theorem for slow-fast system under mixed fractional Brownian motion

This work considers a type of slow-fast system, where the slow component is driven by fractional Brownian motion (FBM) with \(H > 1/2\) and the fast component is a Markovian stationary process. Our solution mapping is defined based on the Young-Wiener sense, which is constructed via the stochastic sewing lemma. Then, we aim to show the fluctuation from the averaging limit by applying the Poisson PDE method. Unlike the case of standard Brownian motion, the Poisson PDE method must be developed to a non-Markovian fractional setting. The first part addresses the central limit theorem problem for a slow-fast system under small FBM, which is fully coupled with the fast varying one. Firstly, a Wiener-Young-Ito formula is constructed for the Wiener-Young-Ito integral. The behavior of the deviation component is related to solutions of Poisson PDEs associated with the Laplace operator of the fast process. The fast one is assumed to satisfy "stricter" Holder conditions related to the time scale parameter. The tightness is then derived through the Holder semi-norm of the fast one, the properties of the Poisson PDE solution, and a Gronwall-type result for a linear Young differential equation. The weak limit shown includes an extra Gaussian process. The second part is dedicated to the problem under a general FBM. In contrast to the former one, here the FBM integral term will be more difficult to bound due to the coupling between the regularity of the Poisson PDE solution and dependence on the unbounded Holder seminorm of the fast one.

math.PR