arXiv Science⌕ Search

arXiv · 2609.34085

AD-E2E-JEPA: A Joint-Embedding Predictive Architecture For End-to-End Autonomous Driving

Abstract

Autonomous driving requires \textit{world models} that can understand the physical world, reason and plan, and operate safely. In this paper, we first systematically evaluate existing action-conditioned joint-embedding predictive architecture (JEPA) world models, including LeWM, DINO-WM, and JEPA-WM for end-to-end autonomous driving (E2EAD). To isolate world-model quality from policy learning, we employ a goal-conditioned zero-shot planning setting that evaluates these models using ground-truth future observations as goals, without training any driving policy. We find that existing JEPA-based world models are either accurate for driving but computationally expensive, or computationally efficient but insufficient for planning. To address this trade-off, we propose \textbf{AD-E2E-JEPA}, which introduces a SIGReg-regularized learnable projector applied to projected patch embeddings. The projector reduces the number of planning patches by $16\times$ and the embedding dimension by $4\times$, achieving a $100\times$ inference speedup while retaining planning performance, with a 0.8-second runtime for an 8-frame rollout over 256 candidate trajectories. \textit{Without} training any driving policy, the world model itself reaches the goals located 20 meters away on average within the displacement of respectively 4.0/2.8 meters, using world-model rollouts over trajectory vocabularies of respectively 256/8,192 candidates. On the NAVSIMv2 benchmark, it achieves 67.3/72.9 EPDMS with multiplicative safety metrics and 84.1/86.5 EPDMS$^{\dagger}$ without them in goal-conditioned zero-shot planning. Experiments further show that the self-supervised pretrained projector improves downstream imitation learning performance from 80.2 to 85.4 EPDMS. The source code is available at https://github.com/HaoranZhuExplorer/AD-E2E-JEPA

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Haoran Zhu, Wancong Zhang, Yann LeCun, Anna Choromanska. 2026-09-28. AD-E2E-JEPA: A Joint-Embedding Predictive Architecture For End-to-End Autonomous Driving. https://arxiv.org/abs/2609.34085

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

iTeach: In the Wild Interactive Teaching for Failure-Driven Adaptation of Robot Perception

We present iTeach, a deployable system that lets any co-located human fix a robot's perception failures on the spot without expertise, a workstation, or offline retraining. The operator wears a mixed reality (MR) headset, sees the robot's segmentation predictions overlaid on the real scene, and corrects failures hands-free: rearranging objects (HumanPlay), annotating via gaze and voice, and triggering SAM2 backward mask propagation. Each ~20 s interaction yields 150-300 densely labeled training frames; the system fine-tunes the perception model onboard, keeps the better model, and redeploys, all without leaving the deployment site. The full loop requires only an RGB-D camera, onboard GPU, and an MR headset: any mobile robot, any environment. Starting from 26.1 on cluttered real-world scenes, 45 teaching interactions (13K frames) raise segmentation to 80.7 with no catastrophic forgetting; on three standard benchmarks the model never trained on, performance improves as well. Downstream pick-and-place on SceneReplica reaches 72/100, surpassing a model-based pipeline requiring CAD models. A 12-participant user study confirms non-experts match experts on annotation accuracy (~95% box IoU), speed, and task load (NASA-TLX 21/100). The framework is architecture-agnostic: any fine-tunable perception model can serve as backbone.

cs.RO↗

Hyper Yoshimura: How a slight tweak on a classical folding pattern unleashes meta-stability for deployable robots

Deployable structures inspired by origami have provided lightweight, compact, and reconfigurable solutions for various robotic and architectural applications. However, creating an integrated structural system that can effectively balance the competing requirements of high packing efficiency, simple deployment, and precise morphing into multiple load-bearing configurations remains a significant challenge. This study introduces a new class of hyper-Yoshimura origami, which exhibits a wide range of kinematically admissible and locally metastable states, including newly discovered symmetric "self-packing" and asymmetric "pop-out" states. This metastability is achieved by breaking a design rule of Yoshimura origami that has been in place for many decades. To this end, this study derives a new set of mathematically rigorous design rules and geometric formulations. Based on this, forward and inverse kinematic strategies are developed to stack hyper-Yoshimura modules into deployable booms that can approximate complex 3D shapes. Finally, this study showcases the potential of hyper-Yoshimura with a meter-scale pop-up cellphone charging station deployed at our university's bus transit station, along with a 3D-printed, scaled prototype of a space crane that can function as an object manipulator, solar tracking device, or high-load-bearing structure. These results establish hyper-Yoshimura as a promising platform for deployable and adaptable robotic systems in both terrestrial and space environments.

cs.RO↗

Gondola: Grounded Vision Language Planning for Robotic Manipulation

Vision-language-action (VLA) models have shown promising progress in robotic manipulation. However, directly mapping visual observations and language instructions to low-level actions often results in limited interpretability and weak robustness in complex, long-horizon tasks. To address these challenges, we employ a modular manipulation framework that separates high-level planning from low-level control. At its core is Gondola, a grounded vision-language planning model that generates structured plans with explicit pixel-level object grounding before action execution. Given multi-view observations and planning history, Gondola predicts the next-step plan as interleaved textual instructions and multi-view segmentation masks corresponding to target objects and goal locations. To train Gondola, we construct synthetic datasets that provide explicit supervision for short-horizon grounded planning, multi-view referring expression, and long-horizon compositional reasoning. By coupling grounded plan generation with a 3D-based execution policy, our framework achieves state-of-the-art performance on the challenging GemBench benchmark. The system further demonstrates promising transfer to real robots. Ablation studies confirm that pixel-level grounding and the proposed planning-oriented supervision are critical for effective high-level reasoning. Project webpage: https://cshizhe.github.io/projects/robot_gondola.html

cs.RO↗