arXiv · 2312.15247
Prompt-Propose-Verify: A Reliable Hand-Object-Interaction Data Generation Framework using Foundational Models
Abstract
Diffusion models when conditioned on text prompts, generate realistic-looking images with intricate details. But most of these pre-trained models fail to generate accurate images when it comes to human features like hands, teeth, etc. We hypothesize that this inability of diffusion models can be overcome through well-annotated good-quality data. In this paper, we look specifically into improving the hand-object-interaction image generation using diffusion models. We collect a well annotated hand-object interaction synthetic dataset curated using Prompt-Propose-Verify framework and finetune a stable diffusion model on it. We evaluate the image-text dataset on qualitative and quantitative metrics like CLIPScore, ImageReward, Fedility, and alignment and show considerably better performance over the current state-of-the-art benchmarks.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Gurusha Juneja, Sukrit Kumar. 2023-12-23. Prompt-Propose-Verify: A Reliable Hand-Object-Interaction Data Generation Framework using Foundational Models. https://arxiv.org/abs/2312.15247
Cite the original work for its findings. Save a collection to share your selection of sources.