arXiv · 2609.24850
Revisiting Multi-View Stereo: A Sequence-to-Sequence Formulation
Abstract
Computing accurate geometry from multi-view images is a fundamental problem in computer vision. Recent feed-forward (FF) models jointly estimate 3D geometry and camera parameters, but they typically suffer from geometry distortion caused by reconstruction ambiguity, even when ground-truth camera parameters are supplied. In this paper, we study the multi-view stereo (MVS) problem with known camera parameters and propose a novel approach that bridges conventional MVS and FF methods. Rather than casting MVS as a sequence-to-one mapping that predicts depth only for a single reference view, we reformulate it as a sequence-to-sequence task, akin to FF models, that jointly predicts geometry for all input views. We introduce a global transformer-based architecture with two components that explicitly exploit camera-induced priors: ray-map embeddings that inject camera parameters into image patch tokens, making the transformer camera-aware, and a unified global cost volume that replaces conventional per-view cost volumes to jointly capture 3D structure across all views. Extensive experiments on multiple public benchmarks show our approach achieves state-of-the-art performance, surpassing both MVS and FF reconstruction baselines.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Aoxiang Fan, Corentin Dumery, Nicolas Talabot, Pascal Fua. 2026-09-21. Revisiting Multi-View Stereo: A Sequence-to-Sequence Formulation. https://arxiv.org/abs/2609.24850
Cite the original work for its findings. Save a collection to share your selection of sources.