arXiv · 2609.28083
ZoomDiff: A High-Fidelity Diffusion Model for Dual-Camera Smooth Zooming
Abstract
Digital zoom transitions between dual cameras often exhibit conspicuous discontinuities in geometric structure and chromatic consistency, degrading the user experience. While recent dual-camera smooth zoom (DCSZ) methods attempt to mitigate this by fine-tuning frame interpolation (FI) models on DCSZ data, they struggle with the large cross-view disparities and complex geometric transformations. Considering that the generative prior of diffusion models is suitable for addressing this problem, we explore their application to DCSZ. However, naively applying existing diffusion-based FI models still yields low-fidelity transitions due to insufficient conditional guidance, high-frequency information loss during VAE encoding, as well as inadequate temporal consistency. To address this, we propose ZoomDiff, a high-fidelity diffusion model that leverages dual-camera inputs in both latent and pixel spaces for photo-realistic transitions. Specifically, we first strengthen dual-image conditional guidance during the multi-step denoising process to improve geometric consistency. Then we inject flow-aligned multi-scale features from the VAE encoder into the VAE decoder to recover high-frequency details, where flow-guided temporal consistency supervision are introduced to produce more smooth transitions. Extensive experiments on both synthetic and real-world datasets demonstrate that ZoomDiff outperforms state-of-the-art methods quantitatively and qualitatively. Project page: https://jiayi-hit.github.io/ZoomDiff.github.io/.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jiayi Zhang, Renlong Wu, Yukang Ding, Sibin Deng, Wangmeng Zuo. 2026-09-23. ZoomDiff: A High-Fidelity Diffusion Model for Dual-Camera Smooth Zooming. https://arxiv.org/abs/2609.28083
Cite the original work for its findings. Save a collection to share your selection of sources.