arXiv · 2512.11016
SoccerMaster: A Vision Foundation Model for Soccer Understanding
Abstract
Soccer understanding has recently garnered growing research interest due to its domain-specific complexity and unique challenges. Unlike prior works that typically rely on isolated, task-specific expert models, this work aims to propose a unified model to handle diverse soccer visual understanding tasks, ranging from fine-grained perception (e.g., athlete detection and identification) to high-level semantic reasoning (e.g., event classification). Concretely, our contributions are threefold: (i) we present SoccerMaster, the first soccer-specific vision foundation model that unifies diverse tasks within a single framework via supervised multi-task pretraining; (ii) we develop an automated data curation pipeline, SoccerFactory, to generate scalable spatial annotations, and integrate multiple existing soccer video datasets as a comprehensive pretraining data resource for multi-task pretraining; and (iii) we conduct extensive evaluations demonstrating that SoccerMaster consistently outperforms task-specific expert models across diverse downstream tasks, highlighting its breadth and superiority. The data, code, and model will be publicly available.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Haolin Yang, Jiayuan Rao, Haoning Wu, Weidi Xie. 2025-12-11. SoccerMaster: A Vision Foundation Model for Soccer Understanding. https://arxiv.org/abs/2512.11016
Cite the original work for its findings. Save a collection to share your selection of sources.