arXiv · 2610.10013
Scalable Patch-Level Self-Supervised Learning
Abstract
Self-supervised learning (SSL) at scale produces powerful visual representations. However, most scalable SSL methods rely on ad hoc combinations of multiple objectives and stabilization mechanisms. Taking a step back, we ask if we can design a high-performing, yet principled SSL algorithm. Starting from the multi-view assumption, stipulating that task-relevant content is captured by the information common to different views, we construct an information-theoretic objective decomposing into interpretable terms. This derivation yields JEM, a student-teacher method that learns by aligning corresponding patch representations across views, explicitly regularized by information and structure preservation losses. JEM trains stably from 300M to 7B parameters, and, to our knowledge, is the first latent-space patch-level method demonstrated at 7B scale. Across all scales, JEM reaches strong performance on both global and dense probing tasks, on segmentation benchmarks consistently surpassing the DINOv2 algorithm, an influential foundation for today's strongest visual SSL methods. Notably, at 7B parameters, it exceeds the performance of DINOv3 on panoptic segmentation, despite being trained on $12\times$ less data without refinement stages. These results demonstrate that we can indeed design an SSL algorithm that learns strong representations, is principled and stable.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Maximilian Seitzer, Gabriele Trivigno, Antonín Vobecký, Seungeun Yi, Maxime Oquab, Huy V. Vo, Oriane Siméoni, Piotr Bojanowski. 2026-10-07. Scalable Patch-Level Self-Supervised Learning. https://arxiv.org/abs/2610.10013
Cite the original work for its findings. Save a collection to share your selection of sources.