arXiv · 2609.32027
Depth Any Seen: Which Surfaces and How Far?
Abstract
When several surfaces are visible along a ray, recovering visible 3D structure from one image requires jointly estimating their presence and metric depth. Depth Any Seen represents these surfaces as image-conditioned multi-Bernoulli depth sets, whose components each contribute one depth or remain absent. Its auxiliary-free Exact Multi-Bernoulli objective (ExactMB) learns depth and presence by marginalizing one-to-one assignments to complete, distinct targets. Our analysis shows that matching expected count can leave component-surface assignment unresolved. We extend real and synthetic layered-depth benchmarks to evaluate depth accuracy, recovered support, and overprediction. Compared to depth stacking, ExactMB reduces overprediction by a relative 88.2% on LD-Real and 80.5% on MD-3K while retaining most ordinal accuracy, with comparable conditional metric-depth error on LD-Syn. Further ablation studies show that ordered assignment improves depth-accurate recall and precision over marginalization, whereas the count-regularized configuration achieves higher deeper-rank precision than ordered assignment at lower recall. Our code will be publicly released.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Xiaohao Xu, Xiaonan Huang. 2026-09-25. Depth Any Seen: Which Surfaces and How Far?. https://arxiv.org/abs/2609.32027
Cite the original work for its findings. Save a collection to share your selection of sources.