arXiv · 2603.02935
Temporal Consistency Improves Generalization in Contextual Offline Meta Reinforcement Learning
Abstract
Offline meta-reinforcement learning seeks to learn a policy that generalizes to new related tasks online. Context-based methods infer a task representation from transition histories, yet learning an effective task representation without supervision remains challenging. Existing methods relying on contrastive learning learn discriminative task representations, but fail to identify task-specific dynamics, while relying on reconstruction can be insufficient to model long-horizon dependencies, limiting generalization to new tasks. We investigate the impact of temporal consistency in latent space on task representation learning, showing that enforcing multi-step predictions in latent space encourages task representations that are able to capture task-dependent dynamics while preventing representation collapse. We provide theoretical analysis characterizing sources of error in value estimation and show through extensive experiments on MuJoCo, Contextual DeepMind Control, and MetaWorld benchmarks that temporal consistency significantly improves both zero-shot and few-shot generalization.
Explore related subjects
Keep this discovery
Mohammadreza Nakheai, Aidan Scannell, Kevin Luck, Joni Pajarinen. 2026-03-03. Temporal Consistency Improves Generalization in Contextual Offline Meta Reinforcement Learning. https://arxiv.org/abs/2603.02935
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.