arXiv · 2004.05232
End-to-end Learning Improves Static Object Geo-localization in Monocular Video
Abstract
Accurately estimating the position of static objects, such as traffic lights, from the moving camera of a self-driving car is a challenging problem. In this work, we present a system that improves the localization of static objects by jointly-optimizing the components of the system via learning. Our system is comprised of networks that perform: 1) 5DoF object pose estimation from a single image, 2) association of objects between pairs of frames, and 3) multi-object tracking to produce the final geo-localization of the static objects within the scene. We evaluate our approach using a publicly-available data set, focusing on traffic lights due to data availability. For each component, we compare against contemporary alternatives and show significantly-improved performance. We also show that the end-to-end system performance is further improved via joint-training of the constituent models.
Explore related subjects
Keep this discovery
Mohamed Chaabane, Lionel Gueguen, Ameni Trabelsi, Ross Beveridge, Stephen O'Hara. 2020-04-10. End-to-end Learning Improves Static Object Geo-localization in Monocular Video. https://arxiv.org/abs/2004.05232
Cite the original work for its findings. Save a collection to share your selection of sources.