arXiv ScienceSearch

arXiv subjects

Wei Cui

Publications and source records attributed to Wei Cui.

At least 19 recordsLinked to original sources

SOLO: Stable Omni-terrain Long-Horizon Perceptive Humanoid Locomotion

Humans traverse complex terrain over long distances without losing balance, whereas perceptive humanoid policies become fragile as perception and control errors accumulate. We present SOLO, a unified framework addressing two compounding causes of this long-horizon fragility: dense terrain reconstruction smooths action-critical details, and pointwise imitation lacks temporal credit assignment. Its Query Reconstructor (QR) uses Fourier-encoded cell queries to retrieve spatially specific evidence from depth-proprioception tokens, preserving sharp terrain boundaries. Trajectory-Aware MSE (TA-MSE) Distillation adds next-state teacher-student disagreement to the PPO reward, enabling Generalized Advantage Estimation to propagate future disagreement penalties to preceding actions. In simulation, QR reduces height-map L1 error by factors of 3.3-4.0, while TA-MSE surpasses PPO and MSE+PPO in curriculum progression. On stress-test terrains, SOLO achieves 97.5% mean traversal success and 96% stepping-stone success, versus 75.0-75.6% and 0-3% for dense-reconstructor variants. Deployed zero-shot with only a chest-mounted depth camera and proprioception, SOLO completes a continuous 1.5-km outdoor route and an indoor mixed-terrain course. Project page: https://sunpihai-up.github.io/solo/

cs.RO

Machine Learning Topological Order from Defect Partition Functions

We introduce a machine learning framework for extracting Ising topological order from defect partition functions of the two-dimensional Ising model on a torus. Restricted Boltzmann Machines (RBMs) are trained on Ising model data sampled at criticality across topological sectors. We take a component-wise square-root map of the learned distributions which naturally produces candidate wavefunctions for the (2+1)-dimensional Ising TQFT. As a nontrivial consistency check, we extract the modular S-matrix from overlaps of the resulting states and recover the expected Ising modular data. Our results demonstrate that neural network representations can capture both critical fluctuations and emergent topological structure, providing a data-driven route from lattice statistical mechanics to topological quantum field theory.

cond-mat.dis-nn

DIffuse X-ray Explorer (DIXE): Sky Survey Strategy and Collimator Response Demodulation

DIffuse X-ray Explorer (DIXE) is a proposed high-resolution X-ray spectroscopic surveyor aimed at studying large structures of hot gas in the Milky Way. Its payload is designed to have a field of view (FoV) of $10^\circ$ (half-power diameter) and an energy resolution of better than 6 eV, covering an energy range of 0.1-10 keV. It will be mounted on the China Space Station (CSS) and follow the CSS orbit to conduct the survey with fixed zenith pointing in order to optimize the coverage of key science targets. The payload will avoid the Sun passively via an operable sunshade, where a minimum $25^\circ$ angular separation between the pointing axis and the direction of the Sun is required. Two Sun-avoidance strategies are considered: one focusing on minimizing mechanical risk and the other on maximizing exposure time. The one-year exposure maps indicate that DIXE will cover approximately $72.5\%$ of the sky, with typical exposure times of 26 ks and 68 ks for the two strategies, respectively. Although mechanically collimated, the imaging performance of the payload can be enhanced with a demodulation method based on Markov Chain Monte Carlo sampling using the collimator response. Through simulation, we found that the method could achieve a localization accuracy of $1^\circ$ for point-like sources and a spatial resolution of $3^\circ$ for the extended sources of complex surface brightness distribution, both of which are significantly smaller than the FoV.

astro-ph.IM

Detector Development for HUBS I: Initial Testing of Small-Area TES Microcalorimeters

We report progress on the ongoing development of microcalorimeter detector technology for the Hot Universe Baryon Surveyor (HUBS) mission. We show the results from testing and characterizing selected pixels in a 10$\times$10 microcalorimeter array. The microcalorimeter is based on a Mo/Cu transition-edge sensor (TES) coupled to an Au absorber. To better understand the properties of the devices, we have first measured the energy resolution of a selected pixel in a TES array of the same design with a pulsed laser system that produces 3 eV photons, and found that individual photon peaks are easily resolved with the TES, indicating good performance. We have then exposed the microcalorimeter array to radiation from a $^{55}$Fe source, and found that the pixels tested show energy resolutions as good as 3.7$\pm$0.1 eV at 5.9 keV. The energy resolution is found to vary monotonically with the bias point for all the devices, showing little evidence for the presence of the so-called excess noise. This is consistent with the results from modeling the measured noise spectrum. The effects of thermal crosstalk are evident, leading to the degradation of energy resolution.

astro-ph.IM

CanLegalRAGBench: Evaluating Retrieval-Augmented Generation on Canadian Case Law

RAG-based legal assistants have been growing in popularity, but LLM hallucinations remain a key issue and potentially undermines justice. While benchmarks have been developed to evaluate progress, many rely on synthetic queries rather than realistic legal scenarios. Moreover, Canadian law remains underrepresented in existing evaluations. To address this gap, we introduce CanLegalRAGBench, a Canadian legal QA benchmark based on realistic queries and expert-annotated answers grounded in case law. Our evaluation shows that retrieval performance is sensitive to design choices and that open-source embedding models are competitive with closed source models. However, it also reveals the limitation of automatic evaluations that penalize systems for retrieving alternative relevant documents. We also find that generated answers often diverge from gold responses, either with hallucinations or by producing overly detailed or irrelevant content, with 8-29% of claims not being supported by the retrieved documents. We hope this benchmark will help drive continued progress in addressing limitations of legal RAG systems.

cs.CL

Conf-Gen: Conformal Uncertainty Quantification for Generative Models

Conformal prediction (CP) and its extension, conformal risk control (CRC), are established frameworks for quantifying uncertainty in supervised machine learning through formal guarantees. However, recent breakthroughs in artificial intelligence (AI) have been driven by unsupervised generative models, such as large language models (LLMs) and image generators, which are not directly compatible with CP or CRC. In this work we introduce conformal generation (Conf-Gen), a general framework adapting CRC to generative tasks while relaxing its theoretical assumptions. Conf-Gen unifies and generalizes previous attempts to apply CP to LLMs, and extends conformal methodology to entirely new domains. We demonstrate the flexibility of Conf-Gen through some novel applications, including obtaining conformal guarantees on: image generators producing non-memorized images, conversational AI systems having asked enough clarifying questions, and the output of AI agents being correct.

cs.LG

Room-temperature THz photon detection via nonlinear upconversion with 2% full-system efficiency

Sensitive detection of terahertz (THz) radiation is fundamental to progress in spectroscopy, advanced wireless communication, and the realization of emerging quantum technologies. However, the intrinsically low photon energies in the THz range combined with thermal background radiation tend to constrain detector performance when operating at ambient temperatures. Here, we demonstrate efficient room-temperature THz detection based on nonlinear upconversion in the organic crystal N-benzyl-2-methyl-4-nitroaniline (BNA) to resolve frequencies from 1 to 7.5 THz. The system encompassing spectral filters and a single-photon counter achieves an overall detection efficiency of 2% for sum-frequency generated photons. This enables the detection of a train of 50 000 terahertz pulses carrying, on average, fewer than 0.04 photons per pulse, with a signal-to-noise ratio of unity. At a higher flux, when ~60 photons per pulse impinge on the BNA crystal, the per-pulse detection probability reaches 50%. After accounting for loss mechanisms in the setup, the nonlinear THz-to-near-infrared conversion efficiency in BNA exceeds 75%. These results demonstrate the feasibility of quantum experiments relying on single-photon-level THz detection via upconversion in nonlinear crystals in ambient conditions.

physics.optics

Half-Spacetime Gauging of 2-Group Symmetry in 3d

We construct a class of non-invertible duality defects, in (2+1)d quantum field theories, arising from half-spacetime gauging of a 2-group symmetry. Starting from a parent theory with two discrete and Abelian 0-form symmetries and a prescribed mixed anomaly, we show that gauging one factor produces a theory with a 2-group symmetry, while gauging the other yields a theory with a non-invertible 0-form symmetry, whose fusion rules we derive explicitly. When the parent theory possesses three such symmetries with a cyclic anomaly structure, gauging different factors can produce mutually dual theories and the half-spacetime gauging of the 2-group is implemented by a non-invertible duality defect, whose fusion rules we obtain. We illustrate the construction with explicit examples, including a $U(1)\times U(1)\times U(1)$ gauge theory and a general class of product theories. We also include a self-contained pedagogical introduction to the cohomological tools employed throughout the article.

hep-th

One-shot learning for the complex dynamical behaviors of weakly nonlinear forced oscillators

Extrapolative prediction of complex nonlinear dynamics remains a central challenge in engineering. This study proposes a one-shot learning method to identify global frequency-response curves from a single excitation time history by learning governing equations. We introduce MEv-SINDy (Multi-frequency Evolutionary Sparse Identification of Nonlinear Dynamics) to infer the governing equations of non-autonomous and multi-frequency systems. The methodology leverages the Generalized Harmonic Balance (GHB) method to decompose complex forced responses into a set of slow-varying evolution equations. We validated the capabilities of MEv-SINDy on two critical Micro-Electro-Mechanical Systems (MEMS). These applications include a nonlinear beam resonator and a MEMS micromirror. Our results show that the model trained on a single point accurately predicts softening/hardening effects and jump phenomena across a wide range of excitation levels. This approach significantly reduces the data acquisition burden for the characterization and design of nonlinear microsystems.

cs.LG

RobotPan: A 360$^\circ$ Surround-View Robotic Vision System for Embodied Perception

Surround-view perception is increasingly important for robotic navigation and loco-manipulation, especially in human-in-the-loop settings such as teleoperation, data collection, and emergency takeover. However, current robotic visual interfaces are often limited to narrow forward-facing views, or, when multiple on-board cameras are available, require cumbersome manual switching that interrupts the operator's workflow. Both configurations suffer from motion-induced jitter that causes simulator sickness in head-mounted displays. We introduce a surround-view robotic vision system that combines six cameras with LiDAR to provide full 360$^\circ$ visual coverage, while meeting the geometric and real-time constraints of embodied deployment. We further present \textsc{RobotPan}, a feed-forward framework that predicts \emph{metric-scaled} and \emph{compact} 3D Gaussians from calibrated sparse-view inputs for real-time rendering, reconstruction, and streaming. \textsc{RobotPan} lifts multi-view features into a unified spherical coordinate representation and decodes Gaussians using hierarchical spherical voxel priors, allocating fine resolution near the robot and coarser resolution at larger radii to reduce computational redundancy without sacrificing fidelity. To support long sequences, our online fusion updates dynamic content while preventing unbounded growth in static regions by selectively updating appearance. Finally, we release a multi-sensor dataset tailored to 360$^\circ$ novel view synthesis and metric 3D reconstruction for robotics, covering navigation, manipulation, and locomotion on real platforms. Experiments show that \textsc{RobotPan} achieves competitive quality against prior feed-forward reconstruction and view-synthesis methods while producing substantially fewer Gaussians, enabling practical real-time embodied deployment.

cs.RO

Simulation of non X-ray background for the DIffuse X-ray Explorer (DIXE) mission

DIffuse X-ray Explorer (DIXE) is a proposed high-resolution spectroscopic survey mission onboard the China Space Station. Equipped with microcalorimeters based on the Transition-edge sensor technology, it aims to survey the hot gas in the Milky Way. The performance of DIXE depends on the understanding of non X-ray background (NXB), which can strongly affect observations of diffuse X-ray emission. In this work, we simulated the NXB of DIXE in a low-earth orbit (LEO) using \textsc{Geant4}. A detailed mass model of the payload was constructed, and the major sources of NXB were identified, including cosmic rays, albedo neutrons and albedo photons. These components were implemented in \textsc{Geant4} with realistic angular and spectral distributions. We simulated the relevant physical processes of space radiation interacting with the instrument and calculated the resulting NXB. We also evaluated the delayed background from trapped protons in the South Atlantic Anomaly (SAA). Our simulations show that, at the geomagnetic equator and under solar minimum conditions, the NXB is on average $4.46 \times 10^{-2} ~\mathrm{counts~s^{-1}~cm^{-2}~keV^{-1}}$ in 0.1--10 keV energy band, with dominant contributions from the induced particles generated by primary cosmic protons. The NXB increases toward higher geomagnetic latitudes, reaching a maximum of $1.55 \times 10^{-1} ~\mathrm{counts~s^{-1}~cm^{-2}~keV^{-1}}$. The delayed background induced by the SAA decays rapidly after exiting the anomaly and becomes negligible within approximately 5 minutes. The simulated NXB is consistent with that of similar X-ray observatories in LEOs.

astro-ph.IM

MeshMimic: Geometry-Aware Humanoid Motion Learning through 3D Scene Reconstruction

Humanoid motion control has witnessed significant breakthroughs in recent years, with deep reinforcement learning (RL) emerging as a primary catalyst for achieving complex, human-like behaviors. However, the high dimensionality and intricate dynamics of humanoid robots make manual motion design impractical, leading to a heavy reliance on expensive motion capture (MoCap) data. These datasets are not only costly to acquire but also frequently lack the necessary geometric context of the surrounding physical environment. Consequently, existing motion synthesis frameworks often suffer from a decoupling of motion and scene, resulting in physical inconsistencies such as contact slippage or mesh penetration during terrain-aware tasks. In this work, we present MeshMimic, an innovative framework that bridges 3D scene reconstruction and embodied intelligence to enable humanoid robots to learn coupled "motion-terrain" interactions directly from video. By leveraging state-of-the-art 3D vision models, our framework precisely segments and reconstructs both human trajectories and the underlying 3D geometry of terrains and objects. We introduce an optimization algorithm based on kinematic consistency to extract high-quality motion data from noisy visual reconstructions, alongside a contact-invariant retargeting method that transfers human-environment interaction features to the humanoid agent. Experimental results demonstrate that MeshMimic achieves robust, highly dynamic performance across diverse and challenging terrains. Our approach proves that a low-cost pipeline utilizing only consumer-grade monocular sensors can facilitate the training of complex physical interactions, offering a scalable path toward the autonomous evolution of humanoid robots in unstructured environments.

cs.RO

ZipFlow: a Compiler-based Framework to Unleash Compressed Data Movement for Modern GPUs

In GPU-accelerated data analytics, the overhead of data transfer from CPU to GPU becomes a performance bottleneck when the data scales beyond GPU memory capacity due to the limited PCIe bandwidth. Data compression has come to rescue for reducing the amount of data transfer while taking advantage of the powerful GPU computation for decompression. To optimize the end-to-end query performance, however, the workflow of data compression, transfer, and decompression must be holistically designed based on the compression strategies and hardware characteristics to balance the I/O latency and computational overhead. In this work, we present ZipFlow, a compiler-based framework for optimizing compressed data transfer in GPU-accelerated data analytics. ZipFlow classifies compression algorithms into three distinct patterns based on their inherent parallelism. For each pattern, ZipFlow employs generalized scheduling strategies to effectively exploit the computational power of GPUs across diverse architectures. Building on these patterns, ZipFlow delivers flexible, high-performance, and holistic optimization, which substantially advances end-to-end data transfer capabilities. We evaluate the effectiveness of ZipFlow on industry-standard benchmark, TPC-H. Overall, ZipFlow achieves an average improvement of 2.08 times over the state-of-the-art GPU compression library (nvCOMP) and 3.14 times speedup against CPU-based query processing engines (e.g., DuckDB).

cs.DB

Broadband THz spectroscopy system beyond 25 THz using BNA crystals and a tunable single-ring-fiber compressor

We present a terahertz time-domain spectroscopy (THz-TDS) system which accesses a broadband spectrum, efficiently covering the so-called "new THz gap" between 5 and 15 THz and extending beyond 25 THz. The system exploits nonlinear interactions within the organic crystal BNA (N-benzyl-2-methyl-4-nitroaniline) to generate and detect THz radiation upon excitation by a near-infrared (NIR) pulse centered at 1.03 $\mu$m. To enable broadband THz spectral monitoring, the NIR pulse from a Yb-based solid-state laser undergoes spectral broadening in a gas-filled single-ring hollow-core photonic crystal fiber, followed by a pulse compression to achieve durations as short as 31 fs. This approach paves the way for broadband spectroscopy in hard-to-access THz regions using widely available near-infrared ultrafast sources.

physics.optics

Central Charges and Vacuum Moduli of 2d $\mathcal{N}=(0,4)$ Theories from Class $\mathcal{S}$

We investigate 2d $\mathcal{N}=(0,4)$ supersymmetric theories obtained from a topologically-twisted reduction of 4d $\mathcal{N}=2$ class $\mathcal{S}$ theories on a Riemann surface. This study addresses subtle aspects of central charges, unbroken gauge groups, and emergent superconformal R-symmetries of these theories. Focusing on infrared vacuum structures, we propose conjectural formulas for the central charges. For theories with the gauge group $SU(2)$, we use a Lagrangian description to analyze the vacuum moduli spaces. In particular, we examine two distinct branches -- the special Higgs branch and the twisted Higgs branch -- by computing their Hilbert series, and find agreement with the proposed central charge formulas.

hep-th

SARMAE: Masked Autoencoder for SAR Representation Learning

Synthetic Aperture Radar (SAR) imagery plays a critical role in all-weather, day-and-night remote sensing applications. However, existing SAR-oriented deep learning is constrained by data scarcity, while the physically grounded speckle noise in SAR imagery further hampers fine-grained semantic representation learning. To address these challenges, we propose SARMAE, a Noise-Aware Masked Autoencoder for self-supervised SAR representation learning. Specifically, we construct SAR-1M, the first million-scale SAR dataset, with additional paired optical images, to enable large-scale pre-training. Building upon this, we design Speckle-Aware Representation Enhancement (SARE), which injects SAR-specific speckle noise into masked autoencoders to facilitate noise-aware and robust representation learning. Furthermore, we introduce Semantic Anchor Representation Constraint (SARC), which leverages paired optical priors to align SAR features and ensure semantic consistency. Extensive experiments across multiple SAR datasets demonstrate that SARMAE achieves state-of-the-art performance on classification, detection, and segmentation tasks. Code and models will be available at https://github.com/MiliLab/SARMAE.

cs.CV

Universal Quantum Interconnects via Phase-Coherent Four-Wave Mixing

Quantum transduction, which enables the coherent conversion of quantum information between disparate physical platforms, is a cornerstone for realizing scalable and interoperable quantum networks. Among various approaches, parametric frequency mixing processes such as four-wave mixing (FWM) offer a promising pathway toward efficient and low-noise transduction. In this work, we demonstrate the feasibility of coherent quantum state transfer by indirectly verifying high-fidelity wavefunction's phase mapping (>99%) from the input field to the generated output field wave. Using a gas-filled hollow-core capillary fiber, we systematically investigate spectral phase evolution across a broad range, including infrared (IR) to ultraviolet (UV) transitions, as well as conversions from telecom-band (1550 nm) to visible (516 nm) and deep-UV (308 nm) wavelengths. Our results reveal that strong phase coherence can be maintained throughout these diverse conversion regimes. Because quantum properties such as coherence and entanglement are intrinsically encoded in both the amplitude and phase of a photonic wavefunction, preserving spectral phase is essential for faithful quantum information transfer. We further show that efficient and phase-preserving transduction can be achieved by tuning system parameters, offering valuable insights into nonlinear coupling dynamics. These findings establish a promising foundation for advancing FWM-based quantum transduction schemes and open new avenues for integrating heterogeneous quantum systems across wide spectral domains within future quantum communication networks.

physics.optics

An XMM-Newton View of the ANdromeda Galaxy as Explored in a Legacy Survey (New-ANGELS) II: Luminosity Function of X-ray Sources

As part of the New-ANGELS program, we systematically investigate the X-ray luminosity functions (XLFs) of 4506 X-ray sources projected within a radius of 2.5 deg centering on M31. We construct XLFs for different regions in the disk and halo of M31, accounting for the incompleteness with an effective sensitivity map. Assuming that the halo regions contain (mostly) foreground stars and background active galactic nuclei, they are taken as "background" for deriving the XLFs of the sources in the disk. Through modeling XLFs, we decompose the X-ray sources into distinct populations for each region. We find that low-mass X-ray binaries are the dominant X-ray population throughout the disk of M31. The XLFs of M31 reveal a consistently lower integrated LMXB luminosity per stellar mass ($\alpha_\mathrm{LMXB}$) compared to other galaxies, likely due to M31's prolonged period of quiescent star formation. Variations in the XLF shape and $\alpha_\mathrm{LMXB}$ across different regions of M31 suggest that the relationship between integrated luminosity and stellar mass may vary within the galaxy. Additionally, the relatively low integrated luminosity observed in the inner-arm region provides crucial evidence for a rapid fading of M31's LMXBs around 1 Gyr, a finding consistent with recent observations of other nearby galaxies.

astro-ph.GA