arXiv ScienceSearch

arXiv subjects

Haiyi Liu

Publications and source records attributed to Haiyi Liu.

6 recordsLinked to original sources

UMI-Bridge: Action-Anchored Latent Alignment across Human and Robot Manipulation Data

Real-robot demonstrations are limited, motivating the use of human manipulation data collected without robots, including egocentric videos and handheld Universal Manipulation Interface (UMI) demonstrations. However, differences in viewpoint, embodiment, and available action supervision make it difficult to align representations across these sources according to manipulation motion rather than visual appearance. We introduce UMI-Bridge, which uses UMI as an intermediate domain to align representations according to action equivalence rather than pixel similarity. UMI action supervision anchors the latent representation to end-effector motion and gripper behavior, while synchronized head-wrist observations and paired ego-UMI clips support alignment across views and domains. We train a dual-view latent action model (LAM) on human manipulation data without robot demonstrations, then freeze its wrist teacher and dynamics model to regularize vision-language-action (VLA) post-training on UMI and robot data. The shared wrist interface enables this training-time supervision across both domains while preserving the policy's standard inference architecture. Across three real-robot tasks, UMI-Bridge achieves 91.7% mean success versus 73.3% for Naive Co-training with matched UMI and robot data. On two data-efficiency tasks, it surpasses a full-data Robot-only baseline using 25% of the robot demonstrations together with UMI data. It also achieves 85% and 90% success on two additional tasks learned from UMI demonstrations without task-specific robot demonstrations. These results support action-anchored latent alignment for data-efficient robot learning and UMI-to-robot task transfer.

cs.RO

SEA-Nav: Efficient Policy Learning for Safe and Agile Quadruped Navigation in Cluttered Environments

Efficiently learning safe and agile quadruped navigation in densely cluttered environments remains difficult: existing methods often lack safety and agility, or become conservative in complex scenes and require long training schedules. We propose SEA-Nav (Safe, Efficient, and Agile Navigation), a safe reinforcement learning framework for quadruped navigation in cluttered environments. A differentiable control barrier function (CBF) shield constrains the policy to produce safe velocity commands. An adaptive collision-state initialization mechanism increases the probability of learning from safety-critical near-collision experience. An action regularization term further suppresses infeasible commands for physical deployment. The policy converges after about one hour of training on a single RTX 4090 and transfers zero-shot to real-world cluttered scenes.

cs.RO

SigLoMa: Learning Open-World Quadrupedal Loco-Manipulation from Ego-Centric Vision

Designing an open-world quadrupedal loco-manipulation system is highly challenging. Traditional reinforcement learning frameworks utilizing exteroception often suffer from extreme sample inefficiency and massive sim-to-real gaps. Furthermore, the inherent latency of visual tracking fundamentally conflicts with the high-frequency demands of precise floating-base control. Consequently, existing systems lean heavily on expensive external motion capture and off-board computation. To eliminate these dependencies, we present SigLoMa, a fully onboard, ego-centric vision-based pick-and-place framework. At the core of SigLoMa is the introduction of Sigma Points, a lightweight geometric representation for exteroception that guarantees high scalability and native sim-to-real alignment. To bridge the frequency divide between slow perception and fast control, we design an ego-centric Kalman Filter to provide robust, high-rate state estimation. On the learning front, we alleviate sample inefficiency via an Active Sampling Curriculum guided by Hint Poses, and tackle the robot's structural visual blind spots using temporal encoding coupled with simulated random-walk drift. Real-world experiments validate that, relying solely on a 5Hz (200 ms latency) open-vocabulary detector, SigLoMa successfully executes dynamic loco-manipulation across multiple tasks, achieving performance comparable to expert human teleoperation.

cs.RO

OCC-VO: Dense Mapping via 3D Occupancy-Based Visual Odometry for Autonomous Driving

Visual Odometry (VO) plays a pivotal role in autonomous systems, with a principal challenge being the lack of depth information in camera images. This paper introduces OCC-VO, a novel framework that capitalizes on recent advances in deep learning to transform 2D camera images into 3D semantic occupancy, thereby circumventing the traditional need for concurrent estimation of ego poses and landmark locations. Within this framework, we utilize the TPV-Former to convert surround view cameras' images into 3D semantic occupancy. Addressing the challenges presented by this transformation, we have specifically tailored a pose estimation and mapping algorithm that incorporates Semantic Label Filter, Dynamic Object Filter, and finally, utilizes Voxel PFilter for maintaining a consistent global semantic map. Evaluations on the Occ3D-nuScenes not only showcase a 20.6% improvement in Success Ratio and a 29.6% enhancement in trajectory accuracy against ORB-SLAM3, but also emphasize our ability to construct a comprehensive map. Our implementation is open-sourced and available at: https://github.com/USTCLH/OCC-VO.

cs.RO

Femtosecond Laser Filamentation in Atmospheric Turbulence

The effects of turbulence intensity and turbulence region on the distribution of femtosecond laser filaments are experimentally elaborated. Through the ultrasonic signals emitted by the filaments, and it is observed that increasing turbulence intensity and expanding turbulence active region cause an increase in the start position of the filament, and a decrease in filament length, which can be well explained by the theoretical calculation. It is also observed that the random perturbation of the air refractive index caused by atmospheric turbulence expanded the spot size of the filament. Additionally, when turbulence intensity reaches , multiple filaments are formed. Furthermore, the standard deviation of the transverse displacement of filament is found to be proportional to the square root of turbulent structure constant under the experimental turbulence parameters in this paper. These results contribute to the study of femtosecond laser propagation mechanisms in complex atmospheric turbulence conditions

physics.optics

Filament based Ionizing Radiation Sensing Technology

Accidental exposure to overdose ionizing radiation will inevitably lead to severe biological damage, thus detecting and localizing radiation is essential. Traditional measurement techniques are generally restricted to the limited detection range of few centimeters, posing a great risk to operators. The potential in remote sensing makes femtosecond laser filament technology great candidates for constructively address this challenge. Here we propose a novel filament-based ionizing radiation sensing technology (FIRST), and clarify the interaction mechanism between filaments and ionizing radiation. Specifically, it is demonstrated that the energetic electrons and ions produced by α radiation in air can be effectively accelerated within the filament, serving as seed electrons, which will markedly enhance nitrogen fluorescence. The extended nitrogen fluorescence lifetime of ~1 ns is also observed. These findings provide insights into the intricate interaction among ultra-strong light filed, plasma and energetic particle beam, and pave the way for the remote sensing of ionizing radiation.

physics.plasm-ph