arXiv Science⌕ Search

arXiv · 2610.05765

Bilinear Flow Policy: Distributional Extrapolation for Goal-Conditioned Visuomotor Imitation

Abstract

Goal-conditioned imitation learning (GCIL) with flow matching is a promising framework that can represent multimodal behaviors while adapting to diverse, user-specified goals, yet often fails when goals lie outside the demonstration support. To extrapolate to such unseen goals without collapsing multimodality - a problem we call distributional extrapolation - we introduce Bilinear Flow Policy (BFP), a generative visuomotor policy that combines transductive retrieval with a bilinear conditional flow. Given an unseen observation-goal pair, BFP retrieves an "anchor" training example and transductively reformulates the unseen pair as this familiar anchor plus a residual term. For this decomposition to guide action prediction, the residual must compactly encode how the current observation-goal pair differs from the anchor, and the anchor must be chosen so that this difference is predictive of the corresponding action distribution. BFP achieves this with pretrained visual features and a novel learned anchor-selection algorithm. The novel bilinear flow then models how the anchor and the residual jointly determine the multimodal action distribution. We prove that, for bilinear flow under suitable assumptions, action distribution error at unseen goals is bounded by the in-distribution flow-matching error up to problem-dependent factors. Across five manipulation tasks in simulation, BFP achieves 2.63x the out-of distribution success rate of a GCIL policy and 1.36x that of the strongest extrapolation-targeted baseline. On two real-world tasks, BFP improves over GCIL by 32%. Finally, our theory yields practical, pre deployment diagnostics for predicting which trained policies will extrapolate well and to which unseen goal.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Wonsuhk Jung, Sundhar Vinodh Sangeetha, Chen Xu, Abhishek Gupta, Masha Itkina, Shreyas Kousik, Haruki Nishimura. 2026-10-05. Bilinear Flow Policy: Distributional Extrapolation for Goal-Conditioned Visuomotor Imitation. https://arxiv.org/abs/2610.05765

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Disambiguate Gripper State in Grasp-Based Tasks: Pseudo-Tactile as Feedback Enables Pure Simulation Learning

Grasp-based manipulation tasks are fundamental to robots interacting with their environments, yet gripper state ambiguity significantly reduces the robustness of imitation learning policies for these tasks. Data-driven solutions face the challenge of high real-world data costs, while simulation data, despite its low costs, is limited by the sim-to-real gap. We identify the root cause of gripper state ambiguity as the lack of tactile feedback. To address this, we propose a novel approach employing pseudo-tactile as feedback, inspired by the idea of using a force-controlled gripper as a tactile sensor. This method enhances policy robustness without additional data collection and hardware involvement, while providing a noise-free binary gripper state observation for the policy and thus facilitating pure simulation learning to unleash the power of simulation. Experimental results across three real-world grasp-based tasks demonstrate the necessity, effectiveness, and efficiency of our approach.

cs.RO↗

Video Replanning via Latent Embedding Refinement and Rejection-Based Sampling

Video planning has emerged as a flexible framework for robot manipulation, in which a generative model predicts a video of task completion, and a downstream module translates the predicted frames into actions. However, existing methods typically ignore information from past interactions, limiting their ability to adapt to latent physical properties that can only be revealed through trial and error, such as whether a door should be pushed or pulled, or how friction affects object dynamics. When a plan fails, these methods usually replan from scratch without leveraging the information revealed by the failure. To address this limitation, we introduce RELIC, REplanning with Latent embedding refInement and Candidate rejection, a video planning framework that adapts to hidden physical properties from test-time failures. RELIC optimizes a latent embedding that captures the environment's hidden physical properties from interaction videos and introduces a rejection-based sampling mechanism that filters out hypotheses inconsistent with prior failures. Across eight tasks in two simulation suites, RELIC consistently reduces the number of replanning steps required for success, and linear probes show that its embedding captures the hidden parameters from the interaction itself rather than from scene appearance. Across four challenging real-world robotic manipulation tasks involving hidden interaction modes, e.g., friction, center of mass, and object mass, RELIC raises the one-shot replanning success rate of a video planning baseline from 30.0% to 63.8% after a single physical interaction.

cs.RO↗

Least Restrictive Hyperplane Control Barrier Functions

Control Barrier Functions (CBFs) can provide provable safety guarantees for dynamic systems. However, finding a valid CBF for a system of interest is often non-trivial, especially for systems with low computational resources, higher-order dynamics, and moving close to obstacles of complex shape. A common solution to this problem is to use a purely distance-based CBF. In this paper, we study Hyperplane CBFs (H-CBFs), where a hyperplane separates the agent from the obstacle. First, we note that the common distance-based CBF is a special case of an H-CBF where the hyperplane is a supporting hyperplane of the obstacle that is orthogonal to a line between the agent and the closest point of the obstacle. We then show that a less conservative CBF can be found by optimising over the orientation of the supporting hyperplane, in order to find the Least Restrictive Hyperplane CBF. This enables us to maintain the safety guarantees while allowing controls that are closer to the desired ones, especially when moving fast and passing close to obstacles. We illustrate the approach on a double integrator dynamical system with acceleration constraints, moving through a group of arbitrarily shaped static and moving obstacles, and show that the proposed approach reduces the average gap between desired and safe controls by an order of magnitude.

cs.RO↗