arXiv · 2511.03272
Unified Long Video Inpainting and Outpainting via Overlapping High-Order Co-Denoising
Abstract
Diffusion-based text-to-video models are increasingly capable, but mask-based editing over hundreds of frames remains challenging: na\"ive long-video generation suffers from memory blow-up, window seams, and temporal drift, while existing editors often require specialized modules or heavy fine-tuning. We present Overlapping High-Order Co-Denoising, a lightweight framework that turns a single pre-trained text-to-video model into a unified inpainting-outpainting editor. We train only LoRA adapters using mixed interior and border masks together with a dual-region loss that improves synthesis inside the mask while explicitly preserving known content. At inference, we denoise long latent sequences using overlapping windows, apply second-order Heun sampling within each window, and fuse overlaps with Hamming-weighted blending to reduce boundary artifacts and improve temporal coherence. On InpaintBench (30 real-world videos, 81--300 frames), our method outperforms Wan 2.1 variants and VACE in background faithfulness (SSIM/LPIPS), temporal consistency (tLPIPS), and text alignment (CLIP), and scales to long horizons, demonstrated up to 800 frames, with memory bounded by the chosen window size.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Shuangquan Lyu, Jian Mao, Yue Ma. 2025-11-05. Unified Long Video Inpainting and Outpainting via Overlapping High-Order Co-Denoising. https://arxiv.org/abs/2511.03272
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.