arXiv · 2212.10596
Open-Vocabulary Temporal Action Detection with Off-the-Shelf Image-Text Features
Abstract
Detecting actions in untrimmed videos should not be limited to a small, closed set of classes. We present a simple, yet effective strategy for open-vocabulary temporal action detection utilizing pretrained image-text co-embeddings. Despite being trained on static images rather than videos, we show that image-text co-embeddings enable openvocabulary performance competitive with fully-supervised models. We show that the performance can be further improved by ensembling the image-text features with features encoding local motion, like optical flow based features, or other modalities, like audio. In addition, we propose a more reasonable open-vocabulary evaluation setting for the ActivityNet data set, where the category splits are based on similarity rather than random assignment.
Explore related subjects
Keep this discovery
Vivek Rathod, Bryan Seybold, Sudheendra Vijayanarasimhan, Austin Myers, Xiuye Gu, Vighnesh Birodkar, David A. Ross. 2022-12-20. Open-Vocabulary Temporal Action Detection with Off-the-Shelf Image-Text Features. https://arxiv.org/abs/2212.10596
Cite the original work for its findings. Save a collection to share your selection of sources.