arXiv · 2409.10476
SimInversion: A Simple Framework for Inversion-Based Text-to-Image Editing
Abstract
Diffusion models demonstrate impressive image generation performance with text guidance. Inspired by the learning process of diffusion, existing images can be edited according to text by DDIM inversion. However, the vanilla DDIM inversion is not optimized for classifier-free guidance and the accumulated error will result in the undesired performance. While many algorithms are developed to improve the framework of DDIM inversion for editing, in this work, we investigate the approximation error in DDIM inversion and propose to disentangle the guidance scale for the source and target branches to reduce the error while keeping the original framework. Moreover, a better guidance scale (i.e., 0.5) than default settings can be derived theoretically. Experiments on PIE-Bench show that our proposal can improve the performance of DDIM inversion dramatically without sacrificing efficiency.
Explore related subjects
Keep this discovery
Qi Qian, Haiyang Xu, Ming Yan, Juhua Hu. 2024-09-16. SimInversion: A Simple Framework for Inversion-Based Text-to-Image Editing. https://arxiv.org/abs/2409.10476
Cite the original work for its findings. Save a collection to share your selection of sources.