arXiv ScienceSearch

arXiv subjects

Zhen Fan

Publications and source records attributed to Zhen Fan.

At least 19 recordsLinked to original sources

Resolving support-mismatch by local basis rotation in variational Monte Carlo

Real-time dynamics after a local quench by a charged operator encodes the response functions measured in spectroscopic experiments, yet they have long posed a challenge for variational Monte Carlo calculations. The obstacle is a support mismatch: the projective action by a charged local operator forces an exponentially large number of configurations to vanish, but these configurations may still contribute to the dynamics, biasing the estimators and freezing the evolution at the very first step. This difficulty is an artifact of the chosen sampling basis, and the support mismatch generated by a charged local operator is itself local. We demonstrate that the missing support can be restored by a local rotation of the sampling basis, without changing the underlying variational dynamics. We propose a local basis-rotation sampling scheme that resolves the support-mismatch problem and can be readily incorporated into existing variational Monte Carlo algorithms. Benchmarks show that rotation sampling accurately captures long-time quantum dynamics, enabling variational Monte Carlo calculations of dynamical structure factors in one dimension and unbiased local-operator quench dynamics in two dimensions. We also show that this resolution of the support-mismatch problem extends beyond real-time dynamics, and may also be helpful for ground state variational Monte Carlo calculations.

cond-mat.str-el

Hessian-informed machine learning interatomic potential towards bridging theory and experiments

Local curvature of potential energy surfaces is critical for predicting certain experimental observables of molecules and materials from first principles, yet it remains far beyond reach for complex systems. In this work, we introduce a Hessian-informed Machine Learning Interatomic Potential (Hi-MLIP) that captures such curvature reliably, thereby enabling accurate analysis of associated thermodynamic and kinetic phenomena. To make Hessian supervision practically viable, we develop a highly efficient training protocol, termed Hessian INformed Training (HINT), achieving two to four orders of magnitude reduction for the requirement of expensive Hessian labels. HINT integrates critical techniques, including Hessian pre-training, configuration sampling, curriculum learning and stochastic projection Hessian loss. Enabled by HINT, Hi-MLIP significantly improves transition-state search and brings Gibbs free-energy predictions close to chemical accuracy especially in data-scarce regimes. Our framework also enables accurate treatment of strongly anharmonic hydrides, reproducing phonon renormalization and superconducting critical temperatures in close agreement with experiment while bypassing the computational bottleneck of anharmonic calculations. These results establish a practical route to enhancing curvature awareness of machine learning interatomic potentials, bridging simulation and experimental observables across a wide range of systems.

cs.LG

Egocentric Visibility-Aware Human Pose Estimation

Egocentric human pose estimation (HPE) using a head-mounted device is crucial for various VR and AR applications, but it faces significant challenges due to keypoint invisibility. Nevertheless, none of the existing egocentric HPE datasets provide keypoint visibility annotations, and the existing methods often overlook the invisibility problem, treating visible and invisible keypoints indiscriminately during estimation. As a result, their capacity to accurately predict visible keypoints is compromised. In this paper, we first present Eva-3M, a large-scale egocentric visibility-aware HPE dataset comprising over 3.0M frames, with 435K of them annotated with keypoint visibility labels. Additionally, we augment the existing EMHI dataset with keypoint visibility annotations to further facilitate the research in this direction. Furthermore, we propose EvaPose, a novel egocentric visibility-aware HPE method that explicitly incorporates visibility information to enhance pose estimation accuracy. Extensive experiments validate the significant value of ground-truth visibility labels in egocentric HPE settings, and demonstrate that our EvaPose achieves state-of-the-art performance in both Eva-3M and EMHI datasets.

cs.CV

2D ferroelectric narrow-bandgap semiconductor Wurtzite' type alpha-In2Se3 and its silicon-compatible growth

2D van der Waals ferroelectrics, particularly alpha-In2Se3, have emerged as an attractive building block for next-generation information storage technologies due to their moderate band gap and robust ferroelectricity stabilized by dipole locking. alpha-In2Se3 can adopt either the distorted zincblende or wurtzite structures; however, the wurtzite phase has yet to be experimental-ly validated, and its large-scale synthesis poses significant challenges. Here, we report an in-situ transport growth of centimeter-scale wurtzite type alpha-In2Se3 films directly on SiO2 substrates using a process combining pulsed laser deposition and chemical vapor deposition. We demonstrate that it is a narrow bandgap ferroelectric semiconductor, featuring a Curie tem-perature exceeding 620 K, a tunable bandgap (0.8-1.6 eV) modulated by charged domain walls, and a large optical absorption coefficient of 1.3 times 10 powers 6 per centemeter. Moreover, light absorption promotes the dynamic conductance range, linearity, and symmetry of the synapse devices, leading to a high recognition accuracy of 92.3 percent in a supervised pattern classification task for neuromorphic computing. Our findings demonstrate a ferroelectric polymorphism of In2Se3, highlighting its potential in ferroelectric synapses for neuromorphic computing.

cond-mat.mtrl-sci

Tree tensor network impurity solver based on Cayley-tree mapping

We introduce a tree tensor network (TTN) impurity solver that enables highly efficient and accurate real-time simulations of quantum impurity models. By decomposing a noninteracting bath Hamiltonian into a Cayley tree, the method provides a tensor network representation that naturally captures the multiscale entanglement structure intrinsic to impurity-bath systems. This geometry differs from conventional chain-based mappings and yields a substantial reduction of entanglement, allowing accurate ground-state properties and long-time dynamics to be captured at significantly lower bond dimensions. Benchmark calculations for the single-impurity Anderson model demonstrate that the TTN solver achieves markedly enhanced resolution of real-frequency spectral functions, without invoking analytic continuation. This impurity solver provides a balanced, scale-uniform description of impurity physics and offers a versatile approach for real-time dynamical mean-field theory and related applications involving quantum impurity models.

cond-mat.str-el

Thermodynamics of the Hubbard Model on the Bethe Lattice

We investigate the thermodynamic properties of the Hubbard model on the Bethe lattice with a coordination number of 3 using the thermal canonical tree tensor network method. Our findings reveal two distinct thermodynamic phases: a low-temperature antiferromagnetic phase, where spin SU(2) symmetry is broken, and a high-temperature paramagnetic phase. A key feature of the system is the separation of energy scales for charge and spin excitations, which is reflected in the temperature dependence of thermodynamic quantities and the disparity between spin and charge gaps extracted from their respective susceptibilities. At the critical point, both spin and charge susceptibilities exhibit singularities, suggesting that charge excitations are not fully decoupled from their spin counterparts. Additionally, the double occupancy number exhibits a non-monotonic temperature dependence, indicative of an entropy-driven Pomeranchuk effect. These results demonstrate that the loopless Bethe lattice effectively captures the essential physics of the Hubbard model while providing a computationally efficient framework for studying strongly correlated electronic systems.

cond-mat.str-el

EMHI: A Multimodal Egocentric Human Motion Dataset with HMD and Body-Worn IMUs

Egocentric human pose estimation (HPE) using wearable sensors is essential for VR/AR applications. Most methods rely solely on either egocentric-view images or sparse Inertial Measurement Unit (IMU) signals, leading to inaccuracies due to self-occlusion in images or the sparseness and drift of inertial sensors. Most importantly, the lack of real-world datasets containing both modalities is a major obstacle to progress in this field. To overcome the barrier, we propose EMHI, a multimodal \textbf{E}gocentric human \textbf{M}otion dataset with \textbf{H}ead-Mounted Display (HMD) and body-worn \textbf{I}MUs, with all data collected under the real VR product suite. Specifically, EMHI provides synchronized stereo images from downward-sloping cameras on the headset and IMU data from body-worn sensors, along with pose annotations in SMPL format. This dataset consists of 885 sequences captured by 58 subjects performing 39 actions, totaling about 28.5 hours of recording. We evaluate the annotations by comparing them with optical marker-based SMPL fitting results. To substantiate the reliability of our dataset, we introduce MEPoser, a new baseline method for multimodal egocentric HPE, which employs a multimodal fusion encoder, temporal feature encoder, and MLP-based regression heads. The experiments on EMHI show that MEPoser outperforms existing single-modal methods and demonstrates the value of our dataset in solving the problem of egocentric HPE. We believe the release of EMHI and the method could advance the research of egocentric HPE and expedite the practical implementation of this technology in VR/AR products.

cs.CV

HumanSplat: Generalizable Single-Image Human Gaussian Splatting with Structure Priors

Despite recent advancements in high-fidelity human reconstruction techniques, the requirements for densely captured images or time-consuming per-instance optimization significantly hinder their applications in broader scenarios. To tackle these issues, we present HumanSplat which predicts the 3D Gaussian Splatting properties of any human from a single input image in a generalizable manner. In particular, HumanSplat comprises a 2D multi-view diffusion model and a latent reconstruction transformer with human structure priors that adeptly integrate geometric priors and semantic features within a unified framework. A hierarchical loss that incorporates human semantic information is further designed to achieve high-fidelity texture modeling and better constrain the estimated multiple views. Comprehensive experiments on standard benchmarks and in-the-wild images demonstrate that HumanSplat surpasses existing state-of-the-art methods in achieving photorealistic novel-view synthesis.

cs.CV

NAFRSSR: a Lightweight Recursive Network for Efficient Stereo Image Super-Resolution

Stereo image super-resolution (SR) refers to the reconstruction of a high-resolution (HR) image from a pair of low-resolution (LR) images as typically captured by a dual-camera device. To enhance the quality of SR images, most previous studies focused on increasing the number and size of feature maps and introducing complex and computationally intensive structures, resulting in models with high computational complexity. Here, we propose a simple yet efficient stereo image SR model called NAFRSSR, which is modified from the previous state-of-the-art model NAFSSR by introducing recursive connections and lightweighting the constituent modules. Our NAFRSSR model is composed of nonlinear activation free and group convolution-based blocks (NAFGCBlocks) and depth-separated stereo cross attention modules (DSSCAMs). The NAFGCBlock improves feature extraction and reduces number of parameters by removing the simple channel attention mechanism from NAFBlock and using group convolution. The DSSCAM enhances feature fusion and reduces number of parameters by replacing 1x1 pointwise convolution in SCAM with weight-shared 3x3 depthwise convolution. Besides, we propose to incorporate trainable edge detection operator into NAFRSSR to further improve the model performance. Four variants of NAFRSSR with different sizes, namely, NAFRSSR-Mobile (NAFRSSR-M), NAFRSSR-Tiny (NAFRSSR-T), NAFRSSR-Super (NAFRSSR-S) and NAFRSSR-Base (NAFRSSR-B) are designed, and they all exhibit fewer parameters, higher PSNR/SSIM, and faster speed than the previous state-of-the-art models. In particular, to the best of our knowledge, NAFRSSR-M is the lightest (0.28M parameters) and fastest (50 ms inference time) model achieving an average PSNR/SSIM as high as 24.657 dB/0.7622 on the benchmark datasets. Codes and models will be released at https://github.com/JNUChenYiHong/NAFRSSR.

eess.IV

HMD-Poser: On-Device Real-time Human Motion Tracking from Scalable Sparse Observations

It is especially challenging to achieve real-time human motion tracking on a standalone VR Head-Mounted Display (HMD) such as Meta Quest and PICO. In this paper, we propose HMD-Poser, the first unified approach to recover full-body motions using scalable sparse observations from HMD and body-worn IMUs. In particular, it can support a variety of input scenarios, such as HMD, HMD+2IMUs, HMD+3IMUs, etc. The scalability of inputs may accommodate users' choices for both high tracking accuracy and easy-to-wear. A lightweight temporal-spatial feature learning network is proposed in HMD-Poser to guarantee that the model runs in real-time on HMDs. Furthermore, HMD-Poser presents online body shape estimation to improve the position accuracy of body joints. Extensive experimental results on the challenging AMASS dataset show that HMD-Poser achieves new state-of-the-art results in both accuracy and real-time performance. We also build a new free-dancing motion dataset to evaluate HMD-Poser's on-device performance and investigate the performance gap between synthetic data and real-captured sensor data. Finally, we demonstrate our HMD-Poser with a real-time Avatar-driving application on a commercial HMD. Our code and free-dancing motion dataset are available https://pico-ai-team.github.io/hmd-poser

cs.CV

Building Retrieval Systems for the ClueWeb22-B Corpus

The ClueWeb22 dataset containing nearly 10 billion documents was released in 2022 to support academic and industry research. The goal of this project was to build retrieval baselines for the English section of the "super head" part (category B) of this dataset. These baselines can then be used by the research community to compare their systems and also to generate data to train/evaluate new retrieval and ranking algorithms. The report covers sparse and dense first stage retrievals as well as neural rerankers that were implemented for this dataset. These systems are available as a service on a Carnegie Mellon University cluster.

cs.IR

Combating Adversarial Attacks with Multi-Agent Debate

While state-of-the-art language models have achieved impressive results, they remain susceptible to inference-time adversarial attacks, such as adversarial prompts generated by red teams arXiv:2209.07858. One approach proposed to improve the general quality of language model generations is multi-agent debate, where language models self-evaluate through discussion and feedback arXiv:2305.14325. We implement multi-agent debate between current state-of-the-art language models and evaluate models' susceptibility to red team attacks in both single- and multi-agent settings. We find that multi-agent debate can reduce model toxicity when jailbroken or less capable models are forced to debate with non-jailbroken or more capable models. We also find marginal improvements through the general usage of multi-agent interactions. We further perform adversarial prompt content classification via embedding clustering, and analyze the susceptibility of different models to different types of attack topics.

cs.CL

Superconductivity in nickelate and cuprate superconductors with strong bilayer coupling

The discovery of superconductivity at 80 K under high pressure in La$_3$Ni$_2$O$_7$ presents the groundbreaking confirmation that high-$T_c$ superconductivity is a property of strongly correlated materials beyond cuprates. We use density functional theory (DFT) calculations of the band structure of La$_3$Ni$_2$O$_7$ under pressure to verify that the low-energy bands are composed almost exclusively of Ni 3$d_{x^2-y^2}$ and O 2$p$ orbitals. We deduce that the Ni 3$d_{z^2}$ orbitals are essentially decoupled by the geometry of the high-pressure structure and by the effect of the Ni Hund coupling being strongly suppressed, which results from the enhanced interlayer antiferromagnetic interaction between $d_{z^2}$ orbitals and the strong intralayer hybridization of the $d_{x^2-y^2}$ orbitals with O 2$p$. By introducing a tight-binding model for the Fermi surfaces and low-energy dispersions, we arrive at a bilayer $t$-$t_\perp$-$J$ model with strong interlayer hopping, which we show is a framework unifying La$_3$Ni$_2$O$_7$ with cuprate materials possessing similar band structures, particularly the compounds La$_2$CaCu$_2$O$_6$, Pb$_2$Sr$_2$YCu$_3$O$_8$, and EuSr$_2$Cu$_2$NbO$_8$. We use a renormalized mean-field theory to show that these systems should have ($d$+$is$)-wave superconductivity, with a dominant $d$-wave component and the high $T_c$ driven by the near-optimally doped $\beta$ band, while the $\alpha$ band adds an $s$-wave component that should lead to clear experimental signatures.

cond-mat.supr-con

A Self-enhancement Approach for Domain-specific Chatbot Training via Knowledge Mining and Digest

Large Language Models (LLMs), despite their great power in language generation, often encounter challenges when dealing with intricate and knowledge-demanding queries in specific domains. This paper introduces a novel approach to enhance LLMs by effectively extracting the relevant knowledge from domain-specific textual sources, and the adaptive training of a chatbot with domain-specific inquiries. Our two-step approach starts from training a knowledge miner, namely LLMiner, which autonomously extracts Question-Answer pairs from relevant documents through a chain-of-thought reasoning process. Subsequently, we blend the mined QA pairs with a conversational dataset to fine-tune the LLM as a chatbot, thereby enriching its domain-specific expertise and conversational capabilities. We also developed a new evaluation benchmark which comprises four domain-specific text corpora and associated human-crafted QA pairs for testing. Our model shows remarkable performance improvement over generally aligned LLM and surpasses domain-adapted models directly fine-tuned on domain corpus. In particular, LLMiner achieves this with minimal human intervention, requiring only 600 seed instances, thereby providing a pathway towards self-improvement of LLMs through model-synthesized training data.

cs.CL

Tripling energy storage density through order-disorder transition induced polar nanoregions in PbZrO3 thin films by ion implantation

Dielectric capacitors are widely used in pulsed power electronic devices due to their ultrahigh power densities and extremely fast charge/discharge speed. To achieve enhanced energy storage density, both maximum polarization (Pmax) and breakdown strength (Eb) need to be improved simultaneously. However, these two key parameters are inversely correlated. In this study, order-disorder transition induced polar nanoregions (PNRs) have been achieved in PbZrO3 thin films by making use of the low-energy ion implantation, enabling us overcome the trade-off between high polarizability and breakdown strength, which leads to the tripling of the energy storage density from 20.5 J/cm3 to 62.3 J/cm3 as well as the great enhancement of breakdown strength. This approach could be extended to other dielectric oxides to improve the energy storage performance, providing a new pathway for tailoring the oxide functionalities.

cond-mat.mtrl-sci

Topologically Protected Ferroelectric Domain Wall Memory with Large Readout Current

The discovery and precise manipulation of atomic-size conductive ferroelectric domain defects, such as geometrically confined walls, offer new opportunities for a wide range of prospective electronic devices, and the so-called walltronics is emerging consequently. Here we demonstrate the highly stable and fatigue-resistant nonvolatile ferroelectric memory device based on deterministic creation and erasure of conductive domain wall geometrically confined inside a topological domain structure. By introducing a pair of delicately designed co-axial electrodes onto the epitaxial BiFeO3 film, one can easily create quadrant center topological polar domain structure. More importantly, a reversible switching of such center topological domain structure between the convergent state with highly conductive confined wall and the divergent state with insulating confined wall can be realized, hence resulting in an apparent resistance change with a large On/Off ratio > 104 and a technically preferred readout current (up to 40 nA). Owing to the topological robustness of the center domain structure, the device exhibits the excellent restoration repeatability over 106 cycles and a long retention over 12 days (> 106 s). This work demonstrates a good example for implementing the exotic polar topologies in high-performance nanoscale devices, and would spur more interest in exploring the rich emerging applications of these exotic topological states.

physics.app-ph

Localize, Group, and Select: Boosting Text-VQA by Scene Text Modeling

As an important task in multimodal context understanding, Text-VQA (Visual Question Answering) aims at question answering through reading text information in images. It differentiates from the original VQA task as Text-VQA requires large amounts of scene-text relationship understanding, in addition to the cross-modal grounding capability. In this paper, we propose Localize, Group, and Select (LOGOS), a novel model which attempts to tackle this problem from multiple aspects. LOGOS leverages two grounding tasks to better localize the key information of the image, utilizes scene text clustering to group individual OCR tokens, and learns to select the best answer from different sources of OCR (Optical Character Recognition) texts. Experiments show that LOGOS outperforms previous state-of-the-art methods on two Text-VQA benchmarks without using additional OCR annotation data. Ablation studies and analysis demonstrate the capability of LOGOS to bridge different modalities and better understand scene text.

cs.CV

Nanoscale Non-Destructive Ferroelectric Characterization with Non-Contact Heterodyne Electrostrain Force Microscopy

Perceiving nanoscale ferroelectric phenomena from real space is of great importance for elucidating underlying ferroelectric physics. During the past decades, nanoscale ferroelectric characterization has mainly relied on the Piezoresponse Force Microscopy (PFM), however, the fundamental limitations of PFM have made the nanoscale ferroelectric studies encounter significant bottlenecks. In this study, a high-resolution non-contact ferroelectric measurement, named Non-Contact Heterodyne Electrostrain Force Microscopy (NC-HEsFM), has been introduced firstly. It has been unambiguously demonstrated that NC-HEsFM can operate on multiple eigenmodes to perform ideal high-resolution ferroelectric domain mapping, standard ferroelectric hysteresis loop measurement and controllable domain manipulation. With using quartz tuning fork (QTF) sensor and heterodyne detection, NC-HEsFM shows an unprecedented capability in achieving real non-contact yet non-destructive ferroelectric characterization with negligible electrostatic force effect. It is believed that NC-HEsFM can be extensively used in various ferroelectric or piezoelectric studies with providing substantially improved characterization performance. Meanwhile, the QTF-based force detection makes NC-HEsFM highly compatible for high-vacuum and low-temperature environments, providing ideal conditions for achieving an ultra-high spatial resolution to investigate the most intrinsic ferroelectric phenomena.

cond-mat.mtrl-sci