arXiv ScienceSearch

arXiv subjects

Yan Chang

Publications and source records attributed to Yan Chang.

At least 19 recordsLinked to original sources

Hydra-0: Action Flow for Generalist World Modeling and Control

We introduce Hydra-0, a generalist world model conditioned on action flow, which represents robot actions as pixel motion. This shared visual interface enables generalist world modeling and control by learning action consequences across embodiments, tasks, environments, and video-generation backbones. Our best configuration achieves 90.4% lower robot-motion error and 60.2% lower object-motion error than our action-conditioned baseline, while supporting zero-shot composition and data-efficient adaptation. On the RoboLab benchmark, Hydra-0 achieves a Pearson correlation of r=0.96 between replayed and reference success rates. Finally, we uncover an emergent inverse mode of this interface: a world action model that predicts compatible robot motion from desired object flow transferred from a human demonstration. A trained action head maps the resulting latent features to executable actions without requiring task-specific expert robot demonstrations. Together, these results demonstrate the potential of action flow as a shared control interface connecting heterogeneous training data, open-loop policy evaluation, and robot control.

cs.RO

Learning Dexterous Manipulation Using Contact Wrench Guidance From Human Demonstration

Dexterous robot manipulation can benefit from the abundance of human demonstrations, but transferring such demonstrations to robot policies remains challenging. We present Contact Wrench Guidance from Human Demonstration in Robotic Dexterous Manipulation (CHORD), a framework for long-horizon manipulation of rigid and articulated objects with reinforcement learning. The key idea is object-centric contact wrench space guidance: we represent human and robot motions by the forces and torques they can induce on the object, enabling similarity to be measured by the induced instantaneous motions. This guidance makes reinforcement learning more scalable for contact-rich dexterous manipulation. We further introduce a large-scale simulation benchmark with 4,739 bimanual dexterous manipulation tasks, constructed from motion-capture datasets and reconstructed in-house videos. Evaluated on 1,831 benchmark tasks, CHORD achieves an average success rate of 82.12%, demonstrating strong scalability. CHORD also generalizes to whole-body manipulation from hand-only and third-person demonstrations, achieving a 90.77% success rate, and the learned policies transfer to the real world in both open-loop and closed-loop settings.

cs.RO

Cosmos 3: Omnimodal World Models for Physical AI

We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-transformers architecture. By supporting highly flexible input-output configurations, Cosmos 3 seamlessly unifies critical modalities for Physical AI -- effectively subsuming vision-language models, video generators, world simulators, and world-action models into a single framework. Our evaluation demonstrates that Cosmos 3 establishes a new state-of-the-art across a diverse suite of understanding and generation tasks, demonstrating omnimodal world models as scalable, general-purpose backbones for embodied agents. Our post-trained Cosmos 3 models were ranked as the best open-source Text-to-Image and Image-to-Video models by Artificial Analysis, and the best policy model by RoboArena at the time the technical report was written. To accelerate open research and deployment in Physical AI, we make our code, model checkpoints, curated synthetic datasets, and evaluation benchmark available under the Linux Foundation's OpenMDW-1.1 License at https://github.com/nvidia/cosmos and https://huggingface.co/collections/nvidia/cosmos3. The project website is available at https://research.nvidia.com/labs/cosmos-lab/cosmos3.

cs.CV

Hidden Quantum Advantage near the Decoding Threshold of Decoded Quantum Interferometry

Where is the true boundary of the quantum advantage region of decoded quantum interferometry (DQI)? The best existing answer is provided by Theorem 7.1 in the Supplementary Material of Jordan et al. (2025), yet we show that this answer systematically underestimates the extent of quantum advantage. On the standard partial-win LDPC benchmark instance, there exist 26 consecutive parameter points ($\ell \in [642, 667]$) at which Jordan's analysis declares no quantum advantage ($\langle s\rangle/m < 0.5$), while quantum advantage is in fact present with an approximation ratio reaching $0.66$. The root cause is that Jordan's bound penalizes the entire system with the worst-case Hamming-layer decoding failure rate $\varepsilon = \max_k \varepsilon_k$, discarding the spectral structure of the DQI tridiagonal matrix. Exploiting the concentration of the Perron eigenvector, we replace the uniform penalty with the eigenvector-weighted average $\bar\varepsilon = \sum_k \varepsilon_k w_k^2$ and establish a unified lower bound (Master Theorem) valid over arbitrary finite fields $\mathbb{F}_q$, proving that it strictly improves upon the relaxed form of Jordan's bound by replacing the operator-norm penalty $2\varepsilon(q-1)(m+1)$ with a tighter Rayleigh-quotient penalty $2\bar\varepsilon\lambda_{\max}$.

quant-ph

AGILE: A Comprehensive Workflow for Humanoid Loco-Manipulation Learning

Recent advances in reinforcement learning (RL) have enabled impressive humanoid behaviors in simulation, yet transferring these results to new robots remains challenging. In many real deployments, the primary bottleneck is no longer simulation throughput or algorithm design, but the absence of systematic infrastructure that links environment verification, training, evaluation, and deployment in a coherent loop. To address this gap, we present AGILE, an end-to-end workflow for humanoid RL that standardizes the policy-development lifecycle to mitigate common sim-to-real failure modes. AGILE comprises four stages: (1) interactive environment verification, (2) reproducible training, (3) unified evaluation, and (4) descriptor-driven deployment via robot/task configuration descriptors. For evaluation stage, AGILE supports both scenario-based tests and randomized rollouts under a shared suite of motion-quality diagnostics, enabling automated regression testing and principled robustness assessment. AGILE also incorporates a set of training stabilizations and algorithmic enhancements in training stage to improve optimization stability and sim-to-real transfer. With this pipeline in place, we validate AGILE across five representative humanoid skills spanning locomotion, recovery, motion imitation, and loco-manipulation on two hardware platforms (Unitree G1 and Booster T1), achieving consistent sim-to-real transfer. Overall, AGILE shows that a standardized, end-to-end workflow can substantially improve the reliability and reproducibility of humanoid RL development.

cs.RO

The Development of a Preclinical Alpha Irradiation Platform with Versatile Control of Dose, Dose Rate, and Spatiotemporal Irradiation Patterns

Objectives. This study develops and validates a vacuum-based alpha irradiation platform to support preclinical radiobiology. We aim to demonstrate precise, independent control over incident energy, fluence rate, and spatiotemporal patterns, which are critical to the mechanisms underlying targeted alpha therapies and low-dose risk assessments. Approach. A vacuum-based system with a radioactive alpha source was designed and fabricated. The platform provides independent modulation of: (i) temporal patterns via a programmable gate valve; (ii) fluence rate across two orders of magnitude by varying source-to-aperture distance (57 to 381 mm); (iii) incident energy (0 to 4.6 MeV) using adjustable absorption layers; and (iv) spatial distributions via a 3D motion stage. Temporal precision was assessed via synchronized audio-electronic recordings. Fluence rates and energies were validated using CR-39 detectors and Monte Carlo (MC) simulations. Spatial precision was verified through programmed continuous and discrete trajectories. Main results. Validation experiments demonstrated high system fidelity. Measured irradiation durations deviated from programmed values by less than 0.3 s. Measured and computed fluence rates agreed within 3%. For energy validation, CR-39 track diameters matched MC model predictions within one standard deviation. Recorded spatial patterns and dimensions aligned well with programmed trajectories. Significance. We successfully validated a versatile vacuum-based platform that overcomes energy-degradation constraints of gas-filled systems. By providing multi-parametric control over alpha-particle delivery, this system enables systematic investigation into how energy, dose rate, and spatiotemporal patterns influence radiobiological responses. This platform is poised to optimize targeted alpha therapies and refine radiation protection frameworks.

physics.med-ph

A Deep Learning-Enhanced Fourier Method for the Multi-Frequency Inverse Source Problem with Sparse Far-Field Data

This paper introduces a hybrid computational framework for the multi-frequency inverse source problem governed by the Helmholtz equation. By integrating a classical Fourier method with a deep convolutional neural network, we address the challenges inherent in sparse and noisy far-field data. The Fourier method provides a physics-informed, low-frequency approximation of the source, which serves as the input to a U-Net. The network is trained to map this coarse approximation to a high-fidelity source reconstruction, effectively suppressing truncation artifacts and recovering fine-scale geometric details. To enhance computational efficiency and robustness, we propose a high-to-low noise transfer learning strategy: a model pre-trained on high-noise regimes captures global topological features, offering a robust initialization for fine-tuning on lower-noise data. Numerical experiments demonstrate that the framework achieves accurate reconstructions with noise levels up to 100%, significantly outperforms traditional spectral methods under sparse measurement constraints, and generalizes well to unseen source geometries.

math.AP

CARI4D: Category Agnostic 4D Reconstruction of Human-Object Interaction

Accurate capture of human-object interaction from ubiquitous sensors like RGB cameras is important for applications in human understanding, gaming, and robot learning. However, inferring 4D interactions from a single RGB view is highly challenging due to the unknown object and human information, depth ambiguity, occlusion, and complex motion, which hinder consistent 3D and temporal reconstruction. Previous methods simplify the setup by assuming ground truth object template or constraining to a limited set of object categories. We present CARI4D, the first category-agnostic method that reconstructs spatially and temporarily consistent 4D human-object interaction at metric scale from monocular RGB videos. To this end, we propose a pose hypothesis selection algorithm that robustly integrates the individual predictions from foundation models, jointly refine them through a learned render-and-compare paradigm to ensure spatial, temporal and pixel alignment, and finally reasoning about intricate contacts for further refinement satisfying physical constraints. Experiments show that our method outperforms prior art by 38% on in-distribution dataset and 36% on unseen dataset in terms of reconstruction error. Our model generalizes beyond the training categories and thus can be applied zero-shot to in-the-wild internet videos. Our code and pretrained models will be publicly released.

cs.CV

SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control

Despite the rise of billion-parameter foundation models trained across thousands of graphical processing units (GPUs), similar scaling gains have not been shown for humanoid control. Current neural controllers for humanoids remain modest in size, target a limited set of behaviors, and are trained on a handful of GPUs. We show that scaling model capacity, data, and compute yields a generalist humanoid controller capable of natural, robust whole-body movements. We position motion tracking as a scalable task for humanoid control, leveraging dense supervision from diverse motion-capture data to acquire human motion priors without manual reward engineering. We build a foundation model for motion tracking by scaling along three axes: network size (1.2M to 42M parameters), dataset volume (100M+ frames from 700 hours of motion capture), and compute (21k GPU hours). Beyond demonstrating the benefits of scale, we further show downstream utility through a real-time kinematic planner that bridges motion tracking to tasks such as navigation, enabling natural and interactive control, as well as a unified token space that supports virtual reality (VR) teleoperation and vision-language-action (VLA) models with a single policy. Through this interface, we demonstrate autonomous VLA-driven whole-body loco-manipulation requiring coordinated hand and foot placement. Scaling motion tracking exhibits favorable properties: performance improves steadily with compute and data diversity, and learned policies generalize to unseen motions, establishing motion tracking at scale as a practical foundation for humanoid control.

cs.RO

Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning

We present Isaac Lab, the natural successor to Isaac Gym, which extends the paradigm of GPU-native robotics simulation into the era of large-scale multi-modal learning. Isaac Lab combines high-fidelity GPU parallel physics, photorealistic rendering, and a modular, composable architecture for designing environments and training robot policies. Beyond physics and rendering, the framework integrates actuator models, multi-frequency sensor simulation, data collection pipelines, and domain randomization tools, unifying best practices for reinforcement and imitation learning at scale within a single extensible platform. We highlight its application to a diverse set of challenges, including whole-body control, cross-embodiment mobility, contact-rich and dexterous manipulation, and the integration of human demonstrations for skill acquisition. Finally, we discuss upcoming integration with the differentiable, GPU-accelerated Newton physics engine, which promises new opportunities for scalable, data-efficient, and gradient-based approaches to robot learning. We believe Isaac Lab's combination of advanced simulation capabilities, rich sensing, and data-center scale execution will help unlock the next generation of breakthroughs in robotics research.

cs.RO

Efficient hybrid variational quantum algorithm for solving graph coloring problem

In the era of Noisy Intermediate Scale Quantum (NISQ) computing, available quantum resources are limited. Many NP-hard problems can be efficiently addressed using hybrid classical and quantum computational methods. This paper proposes a hybrid variational quantum algorithm designed to solve the $k$-coloring problem of graph vertices. The hybrid classical and quantum algorithms primarily partition the graph into multiple subgraphs through hierarchical techniques. The Quantum Approximate Optimization Algorithm (QAOA) is employed to determine the coloring within the subgraphs, while a classical greedy algorithm is utilized to find the coloring of the interaction graph. Fixed coloring is applied to the interaction graph, and feedback is provided to correct any conflicting colorings within the subgraphs. The merging process into the original graph is iteratively optimized to resolve any arising conflicts. We employ a hierarchical framework that integrates feedback correction and conflict resolution to achieve $k$-coloring of arbitrary graph vertices. Through experimental analysis, we demonstrate the effectiveness of the algorithm, highlighting the rapid convergence of conflict evolution and the fact that iterative optimization allows the classical algorithm to approximate the number of colorings. Finally, we apply the proposed algorithm to optimize the scheduling of a subway transportation network, demonstrating a high degree of fairness.

quant-ph

COMPASS: Cross-embodiment Mobility Policy via Residual RL and Skill Synthesis

As robots are increasingly deployed in diverse application domains, enabling robust mobility across different embodiments has become a critical challenge. Classical mobility stacks, though effective on specific platforms, require extensive per-robot tuning and do not scale easily to new embodiments. Learning-based approaches, such as imitation learning (IL), offer alternatives, but face significant limitations on the need for high-quality demonstrations for each embodiment. To address these challenges, we introduce COMPASS, a unified framework that enables scalable cross-embodiment mobility using expert demonstrations from only a single embodiment. We first pre-train a mobility policy on a single robot using IL, combining a world model with a policy model. We then apply residual reinforcement learning (RL) to efficiently adapt this policy to diverse embodiments through corrective refinements. Finally, we distill specialist policies into a single generalist policy conditioned on an embodiment embedding vector. This design significantly reduces the burden of collecting data while enabling robust generalization across a wide range of robot designs. Our experiments demonstrate that COMPASS scales effectively across diverse robot platforms while maintaining adaptability to various environment configurations, achieving a generalist policy with a success rate approximately 5X higher than the pre-trained IL policy on unseen embodiments, and further demonstrates zero-shot sim-to-real transfer.

cs.RO

A Pioneering Neural Network Method for Efficient and Robust Fluid Simulation

Fluid simulation is an important research topic in computer graphics (CG) and animation in video games. Traditional methods based on Navier-Stokes equations are computationally expensive. In this paper, we treat fluid motion as point cloud transformation and propose the first neural network method specifically designed for efficient and robust fluid simulation in complex environments. This model is also the deep learning model that is the first to be capable of stably modeling fluid particle dynamics in such complex scenarios. Our triangle feature fusion design achieves an optimal balance among fluid dynamics modeling, momentum conservation constraints, and global stability control. We conducted comprehensive experiments on datasets. Compared to existing neural network-based fluid simulation algorithms, we significantly enhanced accuracy while maintaining high computational speed. Compared to traditional SPH methods, our speed improved approximately 10 times. Furthermore, compared to traditional fluid simulation software such as Flow3D, our computation speed increased by more than 300 times.

cs.CV

Identifying an acoustic source in a two-layered medium from multi-frequency phased or phaseless far-field patterns

This paper presents a method for reconstructing an acoustic source located in a two-layered medium from multi-frequency phased or phaseless far-field patterns measured on the upper hemisphere. The interface between the two media is assumed to be flat and infinite, while the source is buried in the lower half-space. In the phased case, a Fourier method is proposed to identify the source based on far-field measurements. This method assumes that the source is compactly supported and can be represented by a sum of Fourier basis functions. By utilizing the far-field patterns at different frequencies, the Fourier coefficients of the source can be determined, allowing for its reconstruction. For the case where phase information is unavailable, a phase retrieval formula is developed to retrieve the phase information. This formula exploits the fact that the far-field patterns are related to the source through a linear operator that preserves phase information. By developing a suitable phase retrieval algorithm, the phase information can be recovered. Once the phase is retrieved, the Fourier method can be adopted to recover the source function. Numerical experiments in two and three dimensions are conducted to validate the performance of the proposed methods.

math.NA

Recovering the sources in the stochastic wave equations from multi-frequency far-field patterns

This paper concerns the inverse source scattering problems of recovering random sources for acoustic and elastic waves. The underlying sources are assumed to be random functions driven by an additive white noise. The inversion process aims to find the essential statistical characteristics of the mean and variance from the radiated random wave field at multiple frequencies. To this end, we propose a non-iterative algorithm by approximating the mean and variance via the truncated Fourier series. Then, the Fourier coefficients can be explicitly evaluated by sparse far-field measurements, resulting in an easy-to-implement and efficient approach for the reconstruction. Demonstrations with extensive numerical results are presented to corroborate the feasibility and robustness of the proposed method.

math.NA

X-MOBILITY: End-To-End Generalizable Navigation via World Modeling

General-purpose navigation in challenging environments remains a significant problem in robotics, with current state-of-the-art approaches facing myriad limitations. Classical approaches struggle with cluttered settings and require extensive tuning, while learning-based methods face difficulties generalizing to out-of-distribution environments. This paper introduces X-Mobility, an end-to-end generalizable navigation model that overcomes existing challenges by leveraging three key ideas. First, X-Mobility employs an auto-regressive world modeling architecture with a latent state space to capture world dynamics. Second, a diverse set of multi-head decoders enables the model to learn a rich state representation that correlates strongly with effective navigation skills. Third, by decoupling world modeling from action policy, our architecture can train effectively on a variety of data sources, both with and without expert policies: off-policy data allows the model to learn world dynamics, while on-policy data with supervisory control enables optimal action policy learning. Through extensive experiments, we demonstrate that X-Mobility not only generalizes effectively but also surpasses current state-of-the-art navigation approaches. Additionally, X-Mobility also achieves zero-shot Sim2Real transferability and shows strong potential for cross-embodiment generalization.

cs.RO

ReMEmbR: Building and Reasoning Over Long-Horizon Spatio-Temporal Memory for Robot Navigation

Navigating and understanding complex environments over extended periods of time is a significant challenge for robots. People interacting with the robot may want to ask questions like where something happened, when it occurred, or how long ago it took place, which would require the robot to reason over a long history of their deployment. To address this problem, we introduce a Retrieval-augmented Memory for Embodied Robots, or ReMEmbR, a system designed for long-horizon video question answering for robot navigation. To evaluate ReMEmbR, we introduce the NaVQA dataset where we annotate spatial, temporal, and descriptive questions to long-horizon robot navigation videos. ReMEmbR employs a structured approach involving a memory building and a querying phase, leveraging temporal information, spatial information, and images to efficiently handle continuously growing robot histories. Our experiments demonstrate that ReMEmbR outperforms LLM and VLM baselines, allowing ReMEmbR to achieve effective long-horizon reasoning with low latency. Additionally, we deploy ReMEmbR on a robot and show that our approach can handle diverse queries. The dataset, code, videos, and other material can be found at the following link: https://nvidia-ai-iot.github.io/remembr

cs.RO

The inverse obstacle scattering with incident tapered waves

This paper is concerned with the reconstruction of the shape of an acoustic obstacle. Based on the use of the tapered waves with very narrow widths illuminating the obstacle, the boundary of the obstacle is reconstructed by a direct imaging algorithm. The stability of the imaging scheme is mathematically analyzed. We emphasize that different from the incident plane waves or point sources, the tapered waves with narrow widths bring several benefits in the inverse scattering: 1. local property. A tapered wave can illuminate only on a local part of the boundary of the obstacle, which generates the scattered field; 2. high resolution. We need only reconstruct the boundary near the beam, which improves the quality of some well-known algorithms; 3. fast and easy to implement. Numerical examples are included to demonstrate the effectiveness of the tapered waves.

math.NA