arXiv · 1509.01947
An end-to-end generative framework for video segmentation and recognition
Abstract
We describe an end-to-end generative approach for the segmentation and recognition of human activities. In this approach, a visual representation based on reduced Fisher Vectors is combined with a structured temporal model for recognition. We show that the statistical properties of Fisher Vectors make them an especially suitable front-end for generative models such as Gaussian mixtures. The system is evaluated for both the recognition of complex activities as well as their parsing into action units. Using a variety of video datasets ranging from human cooking activities to animal behaviors, our experiments demonstrate that the resulting architecture outperforms state-of-the-art approaches for larger datasets, i.e. when sufficient amount of data is available for training structured generative models.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Hilde Kuehne, Juergen Gall, Thomas Serre. 2016-03-17. An end-to-end generative framework for video segmentation and recognition. https://arxiv.org/abs/1509.01947
Cite the original work for its findings. Save a collection to share your selection of sources.