arXiv ScienceSearch

arXiv subjects

Yunfei Bai

Publications and source records attributed to Yunfei Bai.

At least 19 recordsLinked to original sources

Accelerating dynamic simulations of photoexcited materials and their evolution by electron-informed machine learning

Nonadiabatic coupled electron-nuclear dynamics upon electronic excitation underpin the microscopic mechanism and rational modulation of diverse photoinduced functional phenomena in materials, yet their direct first-principles simulations remain computationally demanding. Here we develop a framework for nonadiabatic excited-state machine-learning molecular dynamics (EMLMD) simulations, where the nonequilibrium electronic information upon photoexcitation such as electron temperature is rigorously calibrated from high-precision real-time time-dependent density functional theory (rt-TDDFT) benchmark simulations, enabling accurate reconstruction of excited-state potential energy surfaces (PES). This framework natively incorporates the excited-state electron-phonon couplings and intrinsically captures photoinduced phonon anharmonicity, both of which are missing in standard machine learning molecular dynamics, thus delivering first-principles-level accuracy for excited-state atomic evolutions. Large-scale EMLMD simulations resolve time- and momentum-resolved phonon dynamics in photoexcited materials, directly uncovering the competition between photogenerated coherent phonons and thermal phonons during photoinduced phase transition of bismuth. It also simultaneously resolves elusive atomic-scale microscopic dynamics and global structural rearrangement for selenium photoamorphization. Balancing high accuracy and efficiency, EMLMD offers a versatile paradigm to tackle key challenges in the study of complex excited-state molecular dynamics.

cond-mat.mtrl-sci

Gated Multimodal Learning for Interpretable Property Energy Performance Prediction and Retrofit Scenario Analysis

Achieving resilient and sustainable cities requires scalable approaches to decarbonising residential buildings, which account for about 20% of UK greenhouse gas emissions and 25% of energy-related emissions in the European Union. Energy Performance Certificates (EPCs) support regulation and retrofit planning, but their reliance on on-site inspections limits timely city-scale assessment. This study introduces a gated multimodal model to predict Standard Assessment Procedure (SAP) energy efficiency and Environmental Impact (EI) scores by integrating EPC tabular variables, assessor-written free text, and Geographic Information System (GIS)-derived spatial features describing footprint geometry, height, area, and orientation. Sample-wise gating learns property-specific modality weights, while an auxiliary band classification head stabilises training. In a Westminster, London case study, the model predicts SAP and EI scores with MAEs of 4.03 and 4.76 points and R2 values of 0.757 and 0.748, respectively, achieving a mean MAE of 4.39. Ablation results show that full multimodal fusion outperforms unimodal and bimodal baselines for both score prediction and band-level classification. Interpretability analyses provide decision-relevant evidence: gating weights indicate strong reliance on assessor text; SHAP highlights main fuel, built form, and construction age band; text occlusion prioritises roof and wall fields; and spatial attribution is dominated by height and footprint area, with sensitivity to footprint shape. The validated framework is further applied to retrofit scenarios for wall insulation, roof insulation, and window glazing upgrades, indicating projected improvements in SAP, EI, annual energy cost, and equivalent CO2 emissions. Overall, the framework provides scalable property-level evidence for retrofit screening, intervention prioritisation, and net-zero housing transitions.

cs.LG

Chart-RL: Policy Optimization Reinforcement Learning for Enhanced Visual Reasoning in Chart Question Answering with Vision Language Models

The recent advancements in Vision Language Models (VLMs) have demonstrated progress toward true intelligence requiring robust reasoning capabilities. Beyond pattern recognition, linguistic reasoning must integrate with visual comprehension, particularly for Chart Question Answering (CQA) tasks involving complex data visualizations. Current VLMs face significant limitations in CQA, including imprecise numerical extraction, difficulty interpreting implicit visual relationships, and inadequate attention mechanisms for capturing spatial relationships in charts. In this work, we address these challenges by presenting Chart-RL, a novel reinforcement learning framework that enhances VLMs chart understanding through feedback-driven policy optimization of visual perception and logical inference. Our key innovation includes a comprehensive framework integrating Reinforcement Learning (RL) from Policy Optimization techniques along with adaptive reward functions, that demonstrates superior performance compared to baseline foundation models and competitive results against larger state-of-the-art architectures. We also integrated Parameter-Efficient Fine-Tuning through Low-Rank Adaptation (LoRA) in the RL framework that only requires single GPU configurations while preserving performance integrity. We conducted extensive benchmarking across open-source, proprietary, and state-of-the-art closed-source models utilizing the ChartQAPro dataset. The RL fine-tuned Qwen3-VL-4B-Instruct model achieved an answer accuracy of 0.634, surpassing the 0.580 accuracy of the Qwen3-VL-8B-Instruct foundation model despite utilizing half the parameter count, while simultaneously reducing inference latency from 31 seconds to 9 seconds.

cs.AI

When Models Learn to Ask Why: Adaptive Causal Reasoning for Trustworthy Medical Vision-Language Models

Vision-Language Models (VLMs) have enabled interpretable medical diagnosis by integrating visual perception with linguistic reasoning. Yet, existing medical chain-of-thought (CoT) models lack explicit mechanisms to represent and enforce causal reasoning, leaving them vulnerable to spurious correlations and limiting their clinical reliability. We pinpoint three core challenges in medical CoT reasoning: how to adaptively trigger causal correction, construct high-quality causal-spurious contrastive samples, and maintain causal consistency across reasoning trajectories. To address these challenges, we propose MedCausalX, an end-to-end framework explicitly models causal reasoning chains in medical VLMs. We first introduce the CRMed dataset providing fine-grained anatomical annotations, structured causal reasoning chains, and counterfactual variants that guide the learning of causal relationships beyond superficial correlations. Building upon CRMed, MedCausalX employs a two-stage adaptive reflection architecture equipped with $\langle$causal$\rangle$ and $\langle$verify$\rangle$ tokens, enabling the model to autonomously determine when and how to perform causal analysis and verification. Finally, a trajectory-level causal correction objective optimized through error-attributed reinforcement learning refines the reasoning chain, allowing the model to distinguish genuine causal dependencies from shortcut associations. Extensive experiments on multiple benchmarks show that MedCausalX consistently outperforms state-of-the-art methods, improving diagnostic consistency by +5.4 points, reducing hallucination by over 10 points, and attaining top spatial grounding IoU, thereby setting a new standard for causally grounded medical reasoning. The code and dataset are available at https://github.com/zhcz328/MedCausalX.

cs.AI

SQL-Trail: Multi-Turn Reinforcement Learning with Interleaved Feedback for Text-to-SQL

While large language models (LLMs) have substantially improved Text-to-SQL generation, a pronounced gap remains between AI systems and human experts on challenging benchmarks such as BIRD-SQL. We argue this gap stems largely from the prevailing single-pass paradigm, which lacks the iterative reasoning, schema exploration, and error-correction behaviors that humans naturally employ. To address this limitation, we introduce SQL-Trail, a multi-turn reinforcement learning (RL) agentic framework for Text-to-SQL. Rather than producing a query in one shot, SQL-Trail interacts with the database environment and uses execution feedback to iteratively refine its predictions. Our approach centers on two key ideas: (i) an adaptive turn-budget allocation mechanism that scales the agent's interaction depth to match question difficulty, and (ii) a composite reward panel that jointly incentivizes SQL correctness and efficient exploration. Across benchmarks, SQL-Trail sets a new state of the art and delivers strong data efficiency--up to 18x higher than prior single-pass RL state-of-the-art methods. Notably, our 7B and 14B models outperform substantially larger proprietary systems by 5% on average, underscoring the effectiveness of interactive, agentic workflows for robust Text-to-SQL generation.

cs.AI

Quasi-steady electron-excitonic complexes coupling in a two-dimensional semiconductor

Excitons and their complexes govern optical-related behaviors in semiconductors. Here, using angle-resolved photoemission spectroscopy (ARPES), we have elucidated the light-matter interaction mediated by quasi-steady excitonic complexes within a monolayer of the prototypical two-dimensional (2D) semiconductor WSe2. Under continuous incident light, we have observed the generation of quasi-steady excitons and their complexes, encompassing ground and excited state excitons, trions, as well as their intricate interplay. We further show spectral evidence of electronic excitation states within the background of quasi-steady excitonic complexes, characterized by valence band (VB) effective mass renormalization, the enhanced spin-orbit coupling (SOC), the formation of an excitonic gap near the Fermi level (EF ) of the conduction band (CB), and intervalley excitonic band folding. Our findings not only unveil a quasi-steady excitonic complex background for the creation of diverse electronic excitations in 2D semiconductors but also offer new insights into the role of excitons in the charge density wave (CDW) formation mechanism and facilitate the advancement of correlated electronic state engineering based on the coupling between electrons and excitonic complexes in a quasi-equilibrium state.

cond-mat.str-el

An Explainable Natural Language Framework for Identifying and Notifying Target Audiences In Enterprise Communication

In large-scale maintenance organizations, identifying subject matter experts and managing communications across complex entities relationships poses significant challenges -- including information overload and longer response times -- that traditional communication approaches fail to address effectively. We propose a novel framework that combines RDF graph databases with LLMs to process natural language queries for precise audience targeting, while providing transparent reasoning through a planning-orchestration architecture. Our solution enables communication owners to formulate intuitive queries combining concepts such as equipment, manufacturers, maintenance engineers, and facilities, delivering explainable results that maintain trust in the system while improving communication efficiency across the organization.

cs.AI

Observation of quasi-steady dark excitons and gap phase in a doped semiconductor

Exciton plays an important role in optics and optics-related behaviors and leads to novel correlated phases like charge order, exciton insulator, and exciton-polariton condensation. Dark exciton shows distinct properties from bright one. However, it cannot be directly detected by conventional optic measurements. The electronic modulation effect of dark excitons in quasi-equilibrium distribution, critical for electronic devices in working status, is still elusive. Here, using angle-resolved photoemission spectroscopy, we report creating, detecting, and controlling dark excitons in the quasi-equilibrium distribution in a doped semiconductor SnSe2. Surprisingly, we observe an excitonic gap phase, with a conduction band opening an anisotropic gap. Our results broaden the scope of dark excitons, extending their studies from the picosecond timescale in the ultrafast photoemission process to conditions occurring under quasi-equilibrium. We reveal the light-matter interaction in the engineering of electronic structures and provide a new way to realize the excitonic gap phase in semiconductors with large band gaps.

cond-mat.mtrl-sci

Scalable Canonical and Isothermal-Isobaric Sampling of Coupled Spin-Lattice Systems with Machine-Learning Potentials

Magnetic machine-learning potentials (MLPs) now reach near-first-principles accuracy on the spin-lattice potential energy surface, but the dynamics and sampling frameworks that convert this accuracy into quantitative finite-temperature thermodynamics have lagged behind. Landau-Lifshitz-Gilbert spin-lattice dynamics fixes the local moment magnitude and incurs $O(N)$ MLP evaluations per integration step, while hybrid molecular-dynamics/Monte-Carlo lacks rigorous isothermal-isobaric sampling and remains expensive. We introduce TSPIN, which promotes the spin to a canonical pair $(\mathbf{S}_i,\boldsymbol{\pi}_i)$ alongside the lattice $(\mathbf{R}_i,\mathbf{p}_i)$ within a Nos\'e-Hoover-chain / Martyna-Tobias-Klein construction, delivering rigorous canonical and isothermal-isobaric sampling, native access to longitudinal spin fluctuations through an unconstrained spin amplitude, and one MLP evaluation per integration step. Applied to itinerant Co and localized multiferroic BiFeO$_3$, TSPIN matches the MD/MC reference thermodynamics of Co at substantially lower cost and reproduces the Curie and N\'eel temperatures within $\sim 7\%$ and $\sim 2\%$ of experiment, respectively. The same unconstrained-amplitude dynamics resolves contrasting spin-amplitude behavior: pronounced spin-modulus softening in Co, but a nearly temperature-independent high-spin Fe$^{3+}$ moment in BiFeO$_3$. TSPIN thereby promotes magnetic MLPs from accurate energy models to predictive finite-temperature simulation engines.

cond-mat.mtrl-sci

Intrinsic exciton transport and recombination in single-crystal lead bromide perovskite

Photogenerated carrier transport and recombination in metal halide perovskites are critical to device performance. Despite considerable efforts, sample quality issues and measurement techniques have limited the access to their intrinsic physics. Here, by utilizing high-purity CsPbBr3 single crystals and contact-free transient grating spectroscopy, we directly monitor exciton diffusive transport from 26 to 300 K. As the temperature (T) increases, the carrier mobility ({\mu}) decreases rapidly below 100 K wtih a {\mu}~T^{-3.0} scaling, and then follows a more gradual {\mu}~T^{-1.7} trend at higher temperatures. First-principles calculations perfectly reproduce this experimental trend and reveal that optical phonon scattering governs carrier mobility shifts over the entire temperature range, with a single longitudinal optical mode dominating room-temperature transport. Time-resolved photoluminescence further identifies a substantial increase in exciton radiative lifetime with temperature, attributed to increased exciton population in momentum-dark states caused by phonon scattering. Our findings unambiguously resolve previous theory-experiment discrepancies, providing benchmarks for future optoelectronic design.

cond-mat.mtrl-sci

Yu-Shiba-Rusinov states in the s-wave superconducting kagome Hubbard model: Self-consistent Bogoliubov-de Gennes calculations

Significant research has recently been conducted into the Yu-Shiba-Rusinov (YSR) states in kagome superconductors through theoretical modeling and experimental investigations. However, additional efforts are still needed to further understand the local superconductivity near magnetic impurities in the kagome lattice and clarify how relevant quantities depend on the interaction strength $J$ between such impurities and electrons. In this study, we explore a self-consistent numerical solution of the Bogoliubov-de Gennes equations for an $s$-wave superconducting kagome model with a single classical magnetic impurity. Our study reveals that with increasing $J$, the local pair potential is systematically depressed in the vicinity of the impurity, similar to previous results obtained for the square and triangular lattices. Moreover, when further increasing $J$, the system undergoes a first-order phase transition with the appearance of stable and metastable states, reflecting the presence of the hysteresis loop in the pertinent quantities. As a consequence of this transition, the minimal energy of the stable YSR state is nonzero at any $J$, contrary to the expectations based on the assumption of a constant pair potential. A distinctive feature of the kagome lattice is that characteristics of the first-order transition are very sensitive to the position of the chemical potential within the kagome energy spectrum.

cond-mat.supr-con

Ab initio self-trapped excitons

We propose a new formalism and an effective computational framework to study self-trapped excitons (STE) in insulators and semiconductors from first principles. Using the many-body Bethe-Salpeter equation in combination with perturbation theory, we are able to obtain the mode- and momentum-resolved exciton-phonon coupling matrix element in a perturbative scheme, and explicitly solve the real space localization of the electron (hole), as well as the lattice distortion. Further, this method allows to compute the STE potential energy surface and evaluate the STE formation energy and Stokes shift. We demonstrate our approach using two-dimensional magnetic semiconductor chromium trihalides and a wide-gap insulator BeO, the latter of which features dark excitons, and make predictions of their Stokes shift and coherent phonon generation which we hope to spark future experiments such as photoluminescence and transient absorption studies.

cond-mat.mtrl-sci

Tailoring of the interference-induced surface superconductivity by an applied electric field

Nucleation of the pair condensate near surfaces above the upper critical magnetic field and the pair-condensate enhancement/suppression induced by changes in the electron-phonon interaction at interfaces are the most known examples of the surface superconductivity. Recently, another example has been reported, when the surface enhancement of the critical superconducting temperature occurs due to quantum interference. In this case the pair states spread over the entire volume of the system while exhibiting the constructive interference near the surface. In the present work we investigate how an applied electric field impacts the interference-induced surface superconductivity. The study is based on a numerical solution of the self-consistent Bogoliubov-de Gennes equations for a one-dimensional attractive Hubbard model. Our results demonstrate that the surface superconducting characteristics, especially the surface critical temperature, are sensitive to the applied electric field and can be tailored by changing its magnitude.

cond-mat.supr-con

Surface superconductor-insulator transition induced by an electric field

It is well-known that the electric field can induce phase transitions between superconducting, metallic and insulating states in thin-film materials due to its control of the charge carrier density. Since a similar effect on the charge carriers can also be expected for surfaces of bulk samples, here we investigate the transformation of the surface states in a superconductor under an applied screened electric field. Our study is performed by numerically solving the self-consistent Bogoliubov-de Gennes equations for the one-dimensional attractive Hubbard model. It is found that the surface insulating regime occurs at sufficiently large (but still experimentally accessible) electric fields. Our calculations yield the phase diagram of the surface superconducting, metallic, and insulating states for a wide range of temperatures and applied fields. Our results are in qualitative agreement with the phase diagram obtained by the transport measurements for (Li, Fe)OHFeSe thin flakes [Sci. Bull. 64, 653 (2019); ACS Nano 14, 7513 (2020)].

cond-mat.supr-con

Deep RL at Scale: Sorting Waste in Office Buildings with a Fleet of Mobile Manipulators

We describe a system for deep reinforcement learning of robotic manipulation skills applied to a large-scale real-world task: sorting recyclables and trash in office buildings. Real-world deployment of deep RL policies requires not only effective training algorithms, but the ability to bootstrap real-world training and enable broad generalization. To this end, our system combines scalable deep RL from real-world data with bootstrapping from training in simulation, and incorporates auxiliary inputs from existing computer vision systems as a way to boost generalization to novel objects, while retaining the benefits of end-to-end training. We analyze the tradeoffs of different design decisions in our system, and present a large-scale empirical validation that includes training on real-world data gathered over the course of 24 months of experimentation, across a fleet of 23 robots in three office buildings, with a total training set of 9527 hours of robotic experience. Our final validation also consists of 4800 evaluation trials across 240 waste station configurations, in order to evaluate in detail the impact of the design decisions in our system, the scaling effects of including more real-world data, and the performance of the method on novel objects. The projects website and videos can be found at \href{http://rl-at-scale.github.io}{rl-at-scale.github.io}.

cs.RO

On Designing a Learning Robot: Improving Morphology for Enhanced Task Performance and Learning

As robots become more prevalent, optimizing their design for better performance and efficiency is becoming increasingly important. However, current robot design practices overlook the impact of perception and design choices on a robot's learning capabilities. To address this gap, we propose a comprehensive methodology that accounts for the interplay between the robot's perception, hardware characteristics, and task requirements. Our approach optimizes the robot's morphology holistically, leading to improved learning and task execution proficiency. To achieve this, we introduce a Morphology-AGnostIc Controller (MAGIC), which helps with the rapid assessment of different robot designs. The MAGIC policy is efficiently trained through a novel PRIvileged Single-stage learning via latent alignMent (PRISM) framework, which also encourages behaviors that are typical of robot onboard observation. Our simulation-based results demonstrate that morphologies optimized holistically improve the robot performance by 15-20% on various manipulation tasks, and require 25x less data to match human-expert made morphology performance. In summary, our work contributes to the growing trend of learning-based approaches in robotics and emphasizes the potential in designing robots that facilitate better learning.

cs.RO

Interference-induced surface superconductivity:Enhancement by tuning the Debye energy

In the usual perception, surface superconductivity is associated with the surface nucleation of a superconducting condensate above the upper critical field in type-II superconductors or with a rearrangement of phonon properties and the electron-phonon coupling near surfaces/interfaces. Recently, it has been found that there is another example when the surface superconducting temperature is increased up to 20-25% as compared to the bulk one due to constructive interference of superconducting pair states. In the present work, we demonstrate that in fact, such an interferenceinduced enhancement can be much more pronounced, up to nearly 70%. Furthermore, here it is shown that such an interference enhancement persists over a wide range of microscopic parameters.

cond-mat.supr-con

Practical Imitation Learning in the Real World via Task Consistency Loss

Recent work in visual end-to-end learning for robotics has shown the promise of imitation learning across a variety of tasks. Such approaches are expensive both because they require large amounts of real world training demonstrations and because identifying the best model to deploy in the real world requires time-consuming real-world evaluations. These challenges can be mitigated by simulation: by supplementing real world data with simulated demonstrations and using simulated evaluations to identify high performing policies. However, this introduces the well-known "reality gap" problem, where simulator inaccuracies decorrelate performance in simulation from that of reality. In this paper, we build on top of prior work in GAN-based domain adaptation and introduce the notion of a Task Consistency Loss (TCL), a self-supervised loss that encourages sim and real alignment both at the feature and action-prediction levels. We demonstrate the effectiveness of our approach by teaching a mobile manipulator to autonomously approach a door, turn the handle to open the door, and enter the room. The policy performs control from RGB and depth images and generalizes to doors not encountered in training data. We achieve 72% success across sixteen seen and unseen scenes using only ~16.2 hours of teleoperated demonstrations in sim and real. To the best of our knowledge, this is the first work to tackle latched door opening from a purely end-to-end learning approach, where the task of navigation and manipulation are jointly modeled by a single neural network.

cs.RO