arXiv · 2409.09668
EditBoard: Towards a Comprehensive Evaluation Benchmark for Text-Based Video Editing Models
Abstract
The rapid development of diffusion models has significantly advanced AI-generated content (AIGC), particularly in Text-to-Image (T2I) and Text-to-Video (T2V) generation. Text-based video editing, leveraging these generative capabilities, has emerged as a promising field, enabling precise modifications to videos based on text prompts. Despite the proliferation of innovative video editing models, there is a conspicuous lack of comprehensive evaluation benchmarks that holistically assess these models' performance across various dimensions. Existing evaluations are limited and inconsistent, typically summarizing overall performance with a single score, which obscures models' effectiveness on individual editing tasks. To address this gap, we propose EditBoard, the first comprehensive evaluation benchmark for text-based video editing models. EditBoard encompasses nine automatic metrics across four dimensions, evaluating models on four task categories and introducing three new metrics to assess fidelity. This task-oriented benchmark facilitates objective evaluation by detailing model performance and providing insights into each model's strengths and weaknesses. By open-sourcing EditBoard, we aim to standardize evaluation and advance the development of robust video editing models.
Explore related subjects
Keep this discovery
Yupeng Chen, Penglin Chen, Xiaoyu Zhang, Yixian Huang, Qian Xie. 2024-09-15. EditBoard: Towards a Comprehensive Evaluation Benchmark for Text-Based Video Editing Models. https://arxiv.org/abs/2409.09668
Cite the original work for its findings. Save a collection to share your selection of sources.