arXiv ScienceSearch

arXiv subjects

Yintao Ma

Publications and source records attributed to Yintao Ma.

11 recordsLinked to original sources

Do World Action Models Generalize Better than VLAs? A Robustness Study

Robot action planning in the real world is challenging as it requires not only understanding the current state of the environment but also predicting how it will evolve in response to actions. Vision-language-action (VLA), which repurpose large-scale vision-language models for robot action generation using action experts, have achieved notable success across a variety of robotic tasks. Nevertheless, their performance remains constrained by the scope of their training data, exhibiting limited generalization to unseen scenarios and vulnerability to diverse contextual perturbations. More recently, world models have been revisited as an alternative to VLAs. These models, referred to as world action models (WAMs), are built upon world models that are trained on large corpora of video data to predict future states. With minor adaptations, their latent representation can be decoded into robot actions. It has been suggested that their explicit dynamic prediction capacity, combined with spatiotemporal priors acquired from web-scale video pretraining, enables WAMs to generalize more effectively than VLAs. In this paper, we conduct a comparative study of prominent state-of-the-art VLA policies and recently released WAMs. We evaluate their performance on the LIBERO-Plus and RoboTwin 2.0-Plus benchmarks under various visual and language perturbations. Our results show that WAMs achieve strong robustness, with LingBot-VA reaching 74.2% success rate on RoboTwin 2.0-Plus and Cosmos-Policy achieving 82.2% on LIBERO-Plus. While VLAs such as $\pi_{0.5}$ can achieve comparable robustness on certain tasks, they typically require extensive training with diverse robotic datasets and varied learning objectives. The evaluation code for the RoboTwin2.0-Plus benchmark is available at: https://robot-robustness.github.io/RoboTwin2.0-Plus/.

cs.RO

WALDO: Where Unseen Model-based 6D Pose Estimation Meets Occlusion

Accurate 6D object pose estimation is vital for robotics, augmented reality, and scene understanding. For seen objects, high accuracy is often attainable via per-object fine-tuning but generalizing to unseen objects remains a challenge. To address this problem, past arts assume access to CAD models at test time and typically follow a multi-stage pipeline to estimate poses: detect and segment the object, propose an initial pose, and then refine it. Under occlusion, however, the early-stage of such pipelines are prone to errors, which can propagate through the sequential processing, and consequently degrade the performance. To remedy this shortcoming, we propose four novel extensions to model-based 6D pose estimation methods: (i) a dynamic non-uniform dense sampling strategy that focuses computation on visible regions, reducing occlusion-induced errors; (ii) a multi-hypothesis inference mechanism that retains several confidence-ranked pose candidates, mitigating brittle single-path failures; (iii) iterative refinement to progressively improve pose accuracy; and (iv) series of occlusion-focused training augmentations that strengthen robustness and generalization. Furthermore, we propose a new weighted by visibility metric for evaluation under occlusion to minimize the bias in the existing protocols. Via extensive empirical evaluations, we show that our proposed approach achieves more than 5% improvement in accuracy on ICBIN and more than 2% on BOP dataset benchmarks, while achieving approximately 3 times faster inference.

cs.CV

AnyBox: Efficient Zero-Shot 9DoF Pose Estimation of Boxes for Robotic Manipulation

Recovering the 9D pose of objects, both their 6D pose and 3D dimensions, under clutter and occlusion is a core requirement for warehouse automation, logistics, and manufacturing. Model-based methods are accurate but assume an instance-specific CAD model for every object, which is costly to maintain as inventories change. Model-free and category-level methods relax this assumption, yet they remain vulnerable to the symmetry, weak texture, and heavy occlusion that characterize stacked storage boxes, and they ignore the strong structural priors such scenes provide. We present \textbf{AnyBox}, an efficient zero-shot framework that exploits the geometric regularity of boxes to jointly recover pose and dimensions from a single RGB-D observation. Starting from a canonical category template, AnyBox alternates between pose and scale estimation, using the discrepancy between the reprojected template and the observed mask to drive a binary search over box dimensions. Two lightweight components make this practical: a depth-consistency filter that rejects the implausible hypotheses induced by box symmetry, and an early-stopping rule that replaces the remaining search with a single closed-form update. On public benchmarks and an in-house warehouse dataset, AnyBox improves detection AP by up to 36 points, more than doubling the previous best, and approaches instance-level pipelines that have access to ground-truth CAD models. These gains transfer downstream, raising success by 28\% on a cluttered robotic box-shelving task.

cs.CV

MEMS Vapor Cells-based Rydberg-atom Electrometry Toward Miniaturization and High Sensitivity

Rydberg-atom electrometry, as an emerging cutting-edge technology, features high sensitivity, broad bandwidth, calibration-free operation, and beyond. However, until now the key atomic vapor cells used for confining electric field-sensitive Rydberg atoms nearly made with traditional glass-blown techniques, hindering the miniaturization, integration, and batch manufacturing. Here, we present the wafer-level MEMS atomic vapor cells with glass-silicon-glass sandwiched structure that are batch-manufactured for both frequency stability and electric field measurement. We use specially customized ultra-thick silicon wafers with a resistivity exceeding 10,000 cm, three orders of magnitude higher than that of typical silicon, and a thickness of 6 mm, providing a 4-fold improvement in optical interrogation length. With the as-developed MEMS atomic vapor cell, we configured a high-sensitivity Rydberg-atom electrometry with the minimal detectable microwave field to be 2.8 mV/cm. This combination of miniaturization and sensitivity represents a significant advance in the state-of-the-art field of Rydberg-atom electrometry, paving the way for chip-scale Rydberg-atom electrometry and potentially opening up new applications in a wider variety of fields.

physics.atom-ph

COMS-Integrated Atomic Vapor Cells with Ultra-long Optical Access for Highly Sensitive and Scalable Quantum Sensors

The most appealing features of chip-scale quantum sensors are their capability to maintain extreme sensitivity while enabling large-scale batch manufacturing. This necessitates high-level integration and wafer-level fabrication of atomic vapor cells. In this paper, we describe a micromachining paradigm for wafer-level atomic vapor cells functionalized by CMOS-compatible non-magnetic heaters and temperature sensors and demonstrate several innovative applications. Leveraging standard micro-nanofabrication technology, the integrated vapor cells achieved an ultra-long optical access of 5 mm, nearly four time that of previously microfabricated vapor cells. The feasibility of the integrated atomic vapor cells fabrication process was verified by a consecutive 30-day aging test in a harsh environment (operating temperature of 473 K and vacuum of approximately 1 Pa). Benefiting from the ultra-long optical path, we observed several typical quantum effects, including the saturation absorption and spin fluctuations, a regime previously inaccessible with conventional micromachined vapor cells. Finally, a zero-field quantum magnetometry with an ultra-high magnetic sensitivity of 12 fT/Hz1/2 was also demonstrated. Our achievements broaden the potential applications of microfabricated atomic vapor cells and pave the way for scalable manufacturing of ultrasensitive, chip-scale quantum sensors.

quant-ph

Theoretical study on rotation measurement with a quantum vibration oscillator based on Penning trapped ions

In traditional mechanics, harmonic oscillators can be used to measure force, acceleration, or rotation. Herein, we describe a quantum harmonic oscillator based on a penning trapped calcium ion crystal. Similar to traditional oscillators, the Coriolis force induced axial oscillation amplitude is precisely measured to determine the input velocity. We show that the magnetron motion can be controlled through the rotating wall driving and treated as the driving oscillator. The Coriolis force couples with the magnetron motion and induces vibration in the axial direction or the $z$ direction. The center of mass motion of the ion crystal in the axial direction could be precisely detected by the entanglement between the spins of the ions and the harmonic motion through lasers. The frequency of the magnetron motion needs to meet that of the axial motion under certain conditions and thus the axial motion could be tuned to the resonance peak for maximum detection signal. We gave the parameter spaces for the meeting of the magnetron frequencies to that of the axial frequencies. The measurement sensitivity was calculated in details and results show that rotation angular velocity of $3.0\times10^{-9}rad/s/\sqrt{Hz}$ could achieve with 10000 ions. Amplitude sensing could reach sensitivity of 0.4$pm/\sqrt{Hz}$. With spin squeezing, the sensitivity could be further improved.

quant-ph

Cavity-enhanced detection of spin polarization in a microfabricated atomic vapor cell

We demonstrate continuous Pound-Drever-Hall (PDH) nondestructive monitoring of the electron spin polarization of an atomic vapor in a microfabricated vapor cell within an optical resonator. The two-chamber silicon and glass cell contains $^{87}$Rb and 1.3 amagat of N$_{2}$ buffer gas, and is placed within a planar optical resonator formed by two mirrors with dichroic dielectric coatings to resonantly enhance the coupling to phase-modulated probe light near the D$_2$ line at 780 nm. We describe the theory of signal generation in this system, including the spin-dependent complex refractive index, cavity optical transfer functions, and PDH signal response to spin polarization. We observe cavity transmission and PDH signals across $\approx 200$ GHz of detuning around the atomic resonance line. By resonant optical pumping on the 795 nm D$_1$ line, we observe spin-dependent cavity line shifts, in good agreement with theory. We use the saturation of the line shift vs. optical pumping power to calibrate the number density and efficiency of the optical pumping. In the unresolved sideband regime, we observe quantum-noise-limited PDH readout of the spin polarization density, with a flat noise floor of $9 \times 10^9$ spins cm$^{-3}$ Hz$^{-1/2}$ for frequencies above 700 Hz. We note possible extensions of the technique.

physics.atom-ph

Nuclear spin self compensation system for moving MEG sensing with optical pumped atomic spin co-magnetometer

Recording the moving MEGs of a person in which a person's head could move freely as we record the brain's magnetic field is a hot topic in recent years. Traditionally, atomic magnetometers are utilized for moving MEGs recording and a large compensation coil system is utilized for background magnetic field compensation. Here we described a new potential candidate: an optically pumped atomic co-magnetometer(OPACM) for moving MEGs recording. In the OPACM, hyper-polarized nuclear spins could produce a magnetic field which will shield the background fluctuation low frequency magnetic field noise while the the fast changing MEGs signal could be recorded. The nuclear spins look like an automatic magnetic field shields and dynamically compensate the fluctuated background magnetic field noise. In this article, the magnetic field compensation is studied theoretically and we find that the compensation is closely related to several parameters such as the electron spin magnetic field, the nuclear spin magnetic field and the holding magnetic field. Based on the model, the magnetic field compensation could be optimized. We also experimentally studied the magnetic field compensation and the responses of the OPACM to different frequencies of magnetic field are measured. We show that the OPACM owns a clear suppression of low frequency magnetic field under 1Hz and response to magnetic field's frequencies around the band of the MEGs. Magnetic field sensitivity of $3fT/Hz^{1/2}$ has been achieved. Finally, we do a simulation for the OPACM as it is utilized for moving MEGs recording. For comparison, the traditional compensation system for moving MEGs recording is based on a coil which is around 2m in dimension while our compensation system is only 2mm in dimension. Moreover, our compensation system could work in situ and will not affect each other.

physics.app-ph

Quadrupolar interaction induced frequency shift of 131Xe nuclear spins on the surface of silicon

The combination of micro-machined technology with the Atomic Spin Gyroscope(ASG) devices could fabricated Chip Scale Atomic Spin Gyroscope(CASG). The core of the gyroscope is a micro-machined vapor cell which contains alkali metal and isotope enriched noble gases such as 129Xe and 131Xe. The quadrupolar frequency shift of 131Xe is key parameters which could affect the drift of the ASG and is related to the material of the cell in which they are contained. In micro machined technology, the typical utilized material is silicon. In this article, we studied the electric quadrupolar frequency shift of 131Xe atoms with the silicon wall of the micro-machined vapor cell. A cylinder micro-machined vapor cell is utilized in the experiment and a large part of the inner cell surface is composed of silicon material. We studied the temperature dependence of the 129Xe spin relaxation and 131Xe frequency shifts to evaluate the interaction of the nuclear spin with container wall and the alkali metal atoms. The results show that the average twisted angle of the 131Xe nuclear spins as they collide with the silicon wall is measured to be 29 *10^-6 rad. The desorption energy for the 131Xe nuclear spin to escape from the silicon surface is Esi = 0.009eV . This study could help to improve the bias stability of the CASG which is a key parameter for the gyroscope as well as may developes a method to study the surface property of various material.

physics.ins-det

A single beam Cs-Ne SERF magnetometer with differential laser power noise suppression method

We describe a single beam compact Spin Exchange Relaxation Free(SERF) magnetometer whose configuration is compatible with the silicon-glass bonding micro-machining method. A cylindrical vapor cell with 3mm diameter and 3mm in length is utilized in the magnetometer. In order to reduce the wall relaxation which could not be neglected in micro-machined SERF magnetometer, 3 Amagats(1Amagat=2.69$\times$ 10$^{19}$/cm$^3$) neon buffer gas is filled in the vapor cell and this is the first demonstration of a Cs-Ne SERF magnetometer. We also did a simulation to show that neon is a better buffer gas than nitrogen and helium which is typical utilized in vapor cells. In order to reduce the laser amplitude noise and the large background detection offset which is reported to be the main noise source of a single beam absorption SERF magnetometer, we developed a laser power differential method and a factor of 2 improvement of the power noise suppression has been demonstrated. Finally, we did an optimization of the magnetometer and sensitivity of 40$fT/Hz^{1/2}$@30Hz has been achieved.

physics.ins-det

A Spontaneous Driver Emotion Facial Expression (DEFE) Dataset for Intelligent Vehicles

In this paper, we introduce a new dataset, the driver emotion facial expression (DEFE) dataset, for driver spontaneous emotions analysis. The dataset includes facial expression recordings from 60 participants during driving. After watching a selected video-audio clip to elicit a specific emotion, each participant completed the driving tasks in the same driving scenario and rated their emotional responses during the driving processes from the aspects of dimensional emotion and discrete emotion. We also conducted classification experiments to recognize the scales of arousal, valence, dominance, as well as the emotion category and intensity to establish baseline results for the proposed dataset. Besides, this paper compared and discussed the differences in facial expressions between driving and non-driving scenarios. The results show that there were significant differences in AUs (Action Units) presence of facial expressions between driving and non-driving scenarios, indicating that human emotional expressions in driving scenarios were different from other life scenarios. Therefore, publishing a human emotion dataset specifically for the driver is necessary for traffic safety improvement. The proposed dataset will be publicly available so that researchers worldwide can use it to develop and examine their driver emotion analysis methods. To the best of our knowledge, this is currently the only public driver facial expression dataset.

cs.CV