arXiv ScienceSearch

arXiv subjects

Gang Xu

Publications and source records attributed to Gang Xu.

At least 19 recordsLinked to original sources

UniE2F: A Unified Diffusion Framework for Event-to-Frame Reconstruction with Video Foundation Models

Event cameras excel at high-speed, low-power, and high-dynamic-range scene perception. However, as they fundamentally record only relative intensity changes rather than absolute intensity, the resulting data streams suffer from a significant loss of spatial information and static texture details. In this paper, we address this limitation by leveraging the generative prior of a pre-trained video diffusion model to reconstruct high-fidelity video frames from sparse event data. Specifically, we first establish a baseline model by directly applying event data as a condition to synthesize videos. Then, based on the physical correlation between the event stream and video frames, we further introduce the event-based inter-frame residual guidance to enhance the accuracy of video frame reconstruction. Furthermore, we extend our method to video frame interpolation and prediction in a zero-shot manner by modulating the reverse diffusion sampling process, thereby creating a unified event-to-frame reconstruction framework. Experimental results on real-world and synthetic datasets demonstrate that our method outperforms previous reconstruction approaches on most quantitative metrics and achieves superior qualitative results, while achieving competitive zero-shot performance compared with dedicated methods explicitly trained for interpolation and prediction tasks. We also refer the reviewers to the video demo contained in the supplementary material for video results. The code will be publicly available at https://github.com/CS-GangXu/UniE2F.

cs.CV

Multiple Majorana zero modes realization based on superconducting topological crystalline metal ZrRuAs

The symmetry-protected multiple Majorana zero modes (MZMs) can be manipulated under external fields and have emerged as a promising pathway toward realizing topological quantum computing. While the suitable materials hosting multiple MZMs are still scarce, we propose a feasible candidate platform named superconducting topological crystalline metals (STCMs) that simultaneously possess symmetry-protected topological bands and intrinsic superconductivity. Model analyses demonstrate that the interplay among s-wave superconductivity, mirror symmetry-protected multiple surface Dirac cones, and the introduced spin splitting leads to high BdG Chern numbers of $\mathcal{N} = \pm C_M$, where $C_M$ is mirror Chern number of the STCM. First-principles calculations identify the experimentally synthesized superconductor ZrRuAs as a promising candidate with $C_M=2$, hosting two symmetry-protected surface Dirac cones. When integrated into a heterostructure with the ferromagnetic insulator (FMI) such as GdI$_{2}$, a topological superconducting phase with $\mathcal{N} = -2$ can be realized, giving rise to two branches of MZMs. This new scheme offers advantages of structural simplicity and tunability, making the FMI/STCM heterostructure an ideal platform for investigating multipole MZMs and novel topological qubit.

cond-mat.mes-hall

Momentum-Space-Engineered Spatial Photonic Ising Machine for Long-Range Interactions

Long-range Ising models (LRIMs) with dense nonlocal and competing interactions are central to statistical physics, quantum simulation, and complex networks. Although the spatial photonic Ising machine (SPIM) exploits intrinsic optical parallelism for Ising computation, its ability to faithfully encode dense long-range couplings and capture the resulting thermodynamic signatures remains underexplored. Here, we present a momentum-space-engineered SPIM framework that maps prescribed long-range coupling kernels onto momentum-space masks for parallel Hamiltonian evaluation. Based on a high-fidelity optical field propagation model, the annealing dynamics of LRIMs with power-law and Ruderman-Kittel-Kasuya-Yosida (RKKY) interactions are systematically investigated. For the power-law model, we investigate the modulation of the estimated critical temperature by the decay exponent σ and coupling cutoff radius R. For the RKKY model, we reproduce diverse ordered states induced by complex competing long-range interactions. A proof-of-principle experiment demonstrates the physical feasibility of our approach. This framework broadens the class of many-body systems accessible to the SPIM platform.

quant-ph

CIR-DDG: backbone-agnostic residual correction of antibody-antigen affinity changes with explicit cross-chain geometry

Motivation: Accurate prediction of mutation-induced protein--protein binding free-energy changes is important for antibody affinity maturation, yet scarce labels and complex interface geometry limit generalization. Heterogeneous predictors may process three-dimensional complexes without preserving the cross-chain signals most relevant to a mutation in their final scalar output. Results: We introduce CIR-DDG, a lightweight residual adapter that combines a fixed base prediction with 22 interpretable descriptors of cross-chain distance, contact density and site--partner context. In complex-level five-fold evaluation on SKEMPI 2.0 measurements from 343 complexes, CIR-DDG improved all six tested backbones on antibody--antigen interface mutations: Spearman correlation increased by 0.0346--0.1296, while RMSE decreased by 0.0074--0.0408\kcalmol. Cross-validated probing, equal-capacity controls and feature ablations support the complementarity of explicit geometry. On an independent SARS-CoV-2 RBD--ACE2 deep-mutational-scan benchmark of 3669 substitutions, the fold-specific adapters transferred without any retraining: the absolute interface Spearman correlation increased by 0.026--0.081 for all four evaluable backbones, showing that the learned geometric correction generalizes beyond SKEMPI thermodynamic measurements. Availability and implementation: CIR-DDG is available at https://github.com/ecnuabmlab/CIR-ddG.

q-bio.QM

WildShadowRemover: In-the-Wild Video Shadow Removal via Detail-Preserving Video Diffusion Models

Video shadow removal in the wild remains challenging due to complex illumination, diverse shadow appearances, and limited training data. Despite its importance to numerous vision and graphics applications, it remains largely unexplored in unconstrained real-world scenarios. To address this gap, we present WildShadowRemover, a framework that adapts a pretrained video diffusion model for robust video shadow removal via LoRA fine-tuning. To preserve fine image details while retaining the model's powerful generative prior, we augment the frozen VAE decoder with a detail injection module and introduce a shadow-mask-guided frequency-decomposed modulation module to selectively restore high-frequency textures while suppressing shadow artifacts. Monocular depth priors from Depth Anything 3 further provide geometry-aware guidance under challenging lighting conditions. We also construct WildShadow, a large-scale paired video shadow removal dataset and benchmark, covering diverse synthetic scenes. Extensive experiments demonstrate that our method outperforms existing approaches in shadow removal quality and temporal consistency, producing temporally coherent shadow-free videos with superior visual quality and strong generalization across challenging in-the-wild scenarios.

cs.CV

Robust Global Structure-from-Motion via View Graph Pruning

Structure-from-Motion (SfM) aims to estimate camera poses and reconstruct 3D structures from a collection of unordered images. Compared with incremental SfM, global SfM achieves better scalability by jointly estimating camera poses based on a view graph constructed from pairwise correspondences. However, its performance is highly sensitive to erroneous edges caused by visually ambiguous matches, which may lead to incorrect camera registration and reconstruction artifacts. In this work, we propose a subgraph-guided view graph pruning framework for robust global SfM. Our key idea is to exploit the internal consistency of reliable subgraphs to identify and remove unreliable connections. Specifically, we first partition the view graph into locally consistent subgraphs and perform global SfM within each subgraph to obtain reliable camera poses. We then apply RANSAC-based edge pruning across subgraphs to remove inconsistent edges, and finally perform global SfM on the refined view graph. Extensive experiments on ambiguous, sequential, and unordered image datasets demonstrate that our method improves the robustness of global SfM under challenging conditions. Further evaluation with neural rendering shows that the improved camera estimation leads to higher-quality novel view synthesis results.

cs.CV

Booster-based beam recycling for swap-out injection at the High Energy Photon Source

Fourth-generation synchrotron light sources employ ultralow-emittance storage rings with stringent injection requirements. On-axis swap-out injection alleviates the dependence on storage-ring dynamic aperture, but high-charge operation requires an efficient injector architecture capable of producing high-charge replacement bunches. This paper presents the accelerator physics design and performance analysis of a booster-based beam-recycling swap-out injection scheme implemented at the High Energy Photon Source (HEPS). In this approach, the full-energy booster serves as both an injector and a high-energy accumulator. An extracted storage-ring bunch is returned to the booster, merged with a low-charge bunch previously injected from the linac and accelerated to full energy. Following high-energy damping, the merged bunch is reinjected into the original storage-ring bucket. The scheme avoids the need for a dedicated accumulator ring while enabling high-charge bunch replacement. The recycling scheme was commissioned through staged machine studies. Full recycling-chain simulations, commissioning studies, and measured performance analysis are presented. The measured results characterize the recycling operation and quantify the transmission efficiency and performance limitations of the complete recycling loop. These results demonstrate the feasibility of the booster-based beam-recycling architecture and establish its operational basis for high-charge swap-out injection in future fourth-generation synchrotron light sources.

physics.acc-ph

Prediction-Enhanced Monte Carlo: A Machine Learning View on Control Variate

For many complex simulation tasks spanning areas such as healthcare, engineering, and finance, Monte Carlo (MC) methods are invaluable due to their unbiased estimates and precise error quantification. Nevertheless, Monte Carlo simulations often become computationally prohibitive, especially for nested, multi-level, or path-dependent evaluations lacking effective variance reduction techniques. While machine learning (ML) surrogates appear as natural alternatives, naive replacements typically introduce unquantifiable biases. We address this challenge by introducing Prediction-Enhanced Monte Carlo (PEMC), a framework that leverages modern ML models as learned predictors, using cheap and parallelizable simulation as features, to output unbiased evaluation with reduced variance and runtime. PEMC can also be viewed as a "modernized" view of control variates, where we consider the overall computation-cost-aware variance reduction instead of per-replication reduction, while bypassing the closed-form mean function requirement and maintaining the advantageous unbiasedness and uncertainty quantifiability of Monte Carlo. We illustrate PEMC's broader efficacy and versatility through three examples: first, equity derivatives such as variance swaps under stochastic local volatility models; second, interest rate derivatives such as swaption pricing under the Heath-Jarrow-Morton (HJM) interest-rate model. Finally, we showcase PEMC in a socially significant context - ambulance dispatch and hospital load balancing - where accurate mortality rate estimates are key for ethically sensitive decision-making. Across these diverse scenarios, PEMC consistently reduces variance while preserving unbiasedness, highlighting its potential as a powerful enhancement to standard Monte Carlo baselines.

stat.ML

Robust Activation Map Rectification for Weakly Supervised Volumetric Segmentation: Temporal Coherence as a Free Lunch

Weakly supervised segmentation relies heavily on class activation maps (CAMs) to initially localize target regions. However, CAMs are often noisy and prone to catastrophic failures. Existing remedies typically introduce additional training stages or prototype learning, increasing computational cost and reducing robustness. In this paper, we propose a training-free prototype-free framework that rectifies unreliable CAMs by exploiting temporal and structural coherence in volumetric data as a free lunch. Our approach is built on two key components. First, we introduce Variance-Reduced Activation Aggregation (VRAA) which suppresses noise and amplify coherent semantic signals. We provide a theoretical justification by modeling CAMs as high-dimensional random vectors and show that aggregation yields provable variance reduction. Second, we design a Bidirectional Extremity Rectification (BER) mechanism that detects and rectifies implausible activations through bidirectional extremity checks, effectively mitigating extreme-value failures without learning additional parameters. Our method is model-agnostic and can be seamlessly integrated with existing pipelines. Extensive experiments on multiple public benchmarks demonstrate substantial improvements over state-of-the-art weakly supervised methods, achieving up to 20% Dice and 40% mIoU gains while reducing inference time by more than 5 times. These results indicate that leveraging coherence as an implicit inductive bias yields a principled and efficient approach to stabilizing weakly supervised volumetric segmentation. Our code will be available.

cs.CV

Multiband transport hierarchy and large Nernst effect in EuAuBi: Establishing a Nernst scaling for asymmetric multiband systems

In correlated materials, coexisting pockets of vastly different carrier densities raise two fundamental questions: which pocket governs the various transport coefficients, and does the conventional Nernst scaling $ν/T \propto μ/E_F$, originally derived for single-band systems, still hold? We address both questions in the polar semimetal EuAuBi, where a dilute electron pocket ($n_e \sim 10^{16}~\mathrm{cm}^{-3}$) coexists with a dense hole pocket ($n_h \sim 10^{21}~\mathrm{cm}^{-3}$). We find a clear hierarchy: the hole pocket dominates the longitudinal resistivity; the Hall effect crosses from electron- to hole-dominance with increasing field; the Seebeck coefficient is dominated by the electron pocket at low temperature and by both pockets at high temperature. Remarkably, the Nernst effect is governed entirely by the ultrahigh-mobility electron pocket, yielding a large low-field signal of $\sim 5~μ\mathrm{V/K}$ near 1~T at 202~K, comparable to anomalous Nernst signals in magnetic Weyl semimetals. By analyzing the two-band thermoelectric conductivity, we show that the Nernst coefficient follows a scaling $ν/T \propto μ_e/{E_{F, tot}}$. This scaling originates from a compensation between the electron-to-hole conductivity ratio and the Fermi-energy ratio, establishing that the large Nernst effect is a semiclassical multiband phenomenon rather than a topological Berry-curvature contribution. This understanding advances the thermoelectric transport physics of multiband electronic systems and offers a guiding principle for low-field transverse thermoelectric design.

cond-mat.str-el

Giant anomalous Hall and Nernst effects in a heavy fermion ferromagnet

The anomalous Hall and Nernst effects refer to the perpendicular voltage drop generated by a magnetic material's magnetization in response to an applied current and temperature gradient. These effects can be harnessed to determine the Berry curvature and hold potential for future applications in electronic devices and thermoelectric energy conversion. We investigate the anomalous Hall and Nernst effects in the heavy-fermion ferromagnet CeCrGe$_3$ and its non-4f analog ferromagnet LaCrGe$_3$. We find that CeCrGe$_3$ exhibits a giant anomalous Hall angle and an anomalous Nernst coefficient, reaching values as high as 33% and ~ 10 $\mathrm{μV\ K}^{-1}$, respectively, among the largest reported for topological magnets. Based on electronic band-structure calculations, we identify a series of topological flat bands carrying strong Berry curvature with a pronounced Ce 4f orbital character in CeCrGe$_3$, which are absent in LaCrGe$_3$, highlighting the crucial role of Kondo flat bands in generating large anomalous transport responses. Furthermore, we identify a breakdown of the anomalous Hall scaling relation and the nonlinear anomalous Mott relation, which we attribute to the break of the topological Kondo flat bands at finite temperatures.

cond-mat.str-el

ORACLE: Anticipating Scams from Partial Trajectories in Streaming App Usage

Smartphone scams are increasingly prevalent and typically manifest as multi-stage, cross-application processes with gradually emerging intent. Effective intervention thus requires anticipating scams before the intent becomes explicit. This is inherently challenging, as decisions must rely on partial trajectories with temporally distributed evidence. In this paper, we propose \textbf{ORACLE} Online Reasoning for Anticipating Cross-temporal Latent thrEats, the first agentic framework for early scam anticipation from \textit{streaming app-usage} trajectories. To support this setting, we curate a real-world long-horizon benchmark of streaming app-usage trajectories, covering 12 scam types, spanning extended periods (15 days on average), involving diverse applications (95 apps), and interleaving normal and scam behaviors. To address fragmented evidence, we introduce a self-evolving context manager that adaptively consolidates entity-centric interactions over time, enabling more effective reconstruction of cross-temporal evidence from partial observations. To enhance sensitivity to latent early-stage signals, we propose an on-policy self-distillation scheme in which a teacher model, conditioned on summarized anti-scam reflections and clues by skills, supervises a student model without access to such reflections. This scheme thereby distills evidence-informed knowledge and improves recognition of emerging fraud patterns from partial trajectories. Experiments show that \method{} consistently improves early scam anticipation, yielding timely warnings while reducing false alerts in realistic streaming scenarios.

cs.LG

Risk Awareness Injection: Calibrating Vision-Language Models for Safety without Compromising Utility

Vision language models (VLMs) extend the reasoning capabilities of large language models (LLMs) to cross-modal settings, yet remain highly vulnerable to multimodal jailbreak attacks. Existing defenses predominantly rely on safety fine-tuning or aggressive token manipulations, incurring substantial training costs or significantly degrading utility. Recent research shows that LLMs inherently recognize unsafe content in text, and the incorporation of visual inputs in VLMs frequently dilutes risk-related signals. Motivated by this, we propose Risk Awareness Injection (RAI), a lightweight and training-free framework for safety calibration that restores LLM-like risk recognition by amplifying unsafe signals in VLMs. Specifically, RAI constructs an Unsafe Prototype Subspace from language embeddings and performs targeted modulation on selected high-risk visual tokens, explicitly activating safety-critical signals within the cross-modal feature space. This modulation restores the model's LLM-like ability to detect unsafe content from visual inputs, while preserving the semantic integrity of original tokens for cross-modal reasoning. Extensive experiments across multiple jailbreak and utility benchmarks demonstrate that RAI substantially reduces attack success rate without compromising task performance.

cs.AI

RS-WorldModel: a Unified Model for Remote Sensing Understanding and Future Sense Forecasting

Remote sensing world models aim to both explain observed changes and forecast plausible futures, two tasks that share spatiotemporal priors. Existing methods, however, typically address them separately, limiting cross-task transfer. We present RS-WorldModel, a unified world model for remote sensing that jointly handles spatiotemporal change understanding and text-guided future scene forecasting, and we build RSWBench-1.1M, a 1.1 million sample dataset with rich language annotations covering both tasks. RS-WorldModel is trained in three stages: (1) Geo-Aware Generative Pre-training (GAGP) conditions forecasting on geographic and acquisition metadata; (2) synergistic instruction tuning (SIT) jointly trains understanding and forecasting; (3) verifiable reinforcement optimization (VRO) refines outputs with verifiable, task-specific rewards. With only 2B parameters, RS-WorldModel surpasses open-source models up to 120$ \times $ larger on most spatiotemporal change question-answering metrics. It achieves an FID of 43.13 on text-guided future scene forecasting, outperforming all open-source baselines as well as the closed-source Gemini-2.5-Flash Image (Nano Banana).

cs.AI

Yo'City: Personalized and Boundless 3D Realistic City Scene Generation via Self-Critic Expansion

Realistic 3D city generation is fundamental to a wide range of applications, including virtual reality and digital twins. However, most existing methods rely on training a single diffusion model, which limits their ability to generate personalized and boundless city-scale scenes. In this paper, we present Yo'City, a novel agentic framework that enables user-customized and infinitely expandable 3D city generation by leveraging the reasoning and compositional capabilities of off-the-shelf large models. Specifically, Yo'City first conceptualizes the city through a top-down planning strategy that defines a hierarchical "City-District-Grid" structure. The Global Planner determines the overall layout and potential functional districts, while the Local Designer further refines each district with detailed grid-level descriptions. Subsequently, the grid-level 3D generation is achieved through a "produce-refine-evaluate" isometric image synthesis loop, followed by image-to-3D generation. To simulate continuous city evolution, Yo'City further introduces a user-interactive, relationship-guided expansion mechanism, which performs scene graph-based distance- and semantics-aware layout optimization, ensuring spatially coherent city growth. To comprehensively evaluate our method, we construct a diverse benchmark dataset and design six multi-dimensional metrics that assess generation quality from the perspectives of semantics, geometry, texture, and layout. Extensive experiments demonstrate that Yo'City consistently outperforms existing state-of-the-art methods across all evaluation aspects.

cs.CV

BioChemInsight: An Online Platform for Automated Extraction of Chemical Structures and Activity Data from Patents

The automated extraction of chemical structures and their corresponding bioactivity data is essential for accelerating drug discovery and enabling data-driven research. Current optical chemical structure recognition tools lack the capability to autonomously link molecular structures with their bioactivity profiles, posing a significant bottleneck in structure-activity relationship analysis. To address this, we present BioChemInsight, an open-source pipeline that integrates DECIMER Segmentation with MolNexTR for chemical structure recognition, GLM-4.5V for compound identifier association, and PaddleOCR combined with GLM-4.6 for bioactivity extraction and unit normalization. We evaluated BioChemInsight on 181 patents covering 15 therapeutic targets. The system achieved an average extraction accuracy of above 90% across three key tasks: chemical structure recognition, bioactivity data extraction, and compound identifier association. Our analysis indicates that the chemical space covered by patents is largely complementary to that contained in established public database ChEMBL. Consequently, by enabling systematic patent mining, BioChemInsight provides access to chemical information underrepresented in ChEMBL. This capability expands the landscape of explorable compound-target interactions, enriches the data foundation for quantitative structure-activity relationship modeling and targeted screening, and reduces data preprocessing time from weeks to hours. BioChemInsight is available at https://github.com/dahuilangda/BioChemInsight.

q-bio.QM

Symmetry-Broken Cavity Solitons and Collective Polarization Conformity in Fabry-Perot Kerr Resonators

We report on the experimental generation of polarization symmetry-broken cavity solitons (CSs) in a passive, fiber-based, coherently-driven, Fabry-Perot (FP) Kerr resonator. Polarization resolved measurements reveal the spontaneous transition of initially symmetric CSs into asymmetrical vectorial states, triggered by a cross-phase modulation-induced polarization bifurcation. Most notably, due to counter-propagation of light occurring in FP resonators, we unveil a collective polarization conformity effect, whereby multiple CSs circulating in the cavity converge to the same asymmetric polarization state once their number exceeds a certain threshold. These results demonstrate that Fabry-Perot resonators support novel collective soliton dynamics that are absent in ring architectures.

physics.optics

Uniaxial stress enhanced anisotropic magnetoresistance and superconductivity in the kagome superconductor LaRu$_{3}$Si$_{2}$

Elucidating the role of the kagome electronic structure in determining the various quantum ground states is of fundamental importance. In this work, we employ in-plane uniaxial stress as a tuning parameter to probe the electronic structure and its impact on the superconducting and normal-state properties of the kagome superconductor LaRu$_{3}$Si$_{2}$, combining magnetotransport measurements with first-principles calculations. We identify a pronounced anisotropy in both the upper critical field and the normal-state magnetoresistance, indicating strong electronic anisotropy despite the three-dimensional crystal structure. Furthermore, we find that the superconducting transition temperature $T_{\rm c}$ increases under in-plane stress applied within the kagome plane, although the enhancement is modest, reaching approximately 0.3 K at 0.6 GPa. Furthermore, the absolute magnetoresistance exhibits a pronounced increase from about 22${\%}$ at zero stress to 35${\%}$ at 0.6 GPa, indicating a substantial modification of the normal state above $T_{\rm c}$. Previous studies have reported time-reversal-symmetry (TRS) breaking below a temperature scale that coincides with the onset of magnetoresistance. The simultaneous enhancement of both $T_{\rm c}$ and magnetoresistance under stress therefore suggests a positive correlation between superconductivity and normal-state electronic and magnetic properties in LaRu$_{3}$Si$_{2}$. Detailed calculations demonstrate that stress-induced changes in $T_{\rm c}$ arise from the joint evolution of the total density of states and the flat band, whereas the large magnetoresistance enhancement is dominated by the stress-driven downward shift of the Ru $dz^{2}$ kagome flat band.

cond-mat.supr-con