arXiv · 2408.15651
Online pre-training with long-form videos
Abstract
In this study, we investigate the impact of online pre-training with continuous video clips. We will examine three methods for pre-training (masked image modeling, contrastive learning, and knowledge distillation), and assess the performance on downstream action recognition tasks. As a result, online pre-training with contrast learning showed the highest performance in downstream tasks. Our findings suggest that learning from long-form videos can be helpful for action recognition with short videos.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Itsuki Kato, Kodai Kamiya, Toru Tamaki. 2024-08-28. Online pre-training with long-form videos. https://arxiv.org/abs/2408.15651
Cite the original work for its findings. Save a collection to share your selection of sources.