arXiv · 2608.00829
GeminiPainter's sequence-formed pipeline comprised of perception, cognition, planning, and action stages
Abstract
We present an autonomous robotic portrait-generation system combining real-time face detection, AI-based sketch generation, and robotic drawing. The system captures video frames, extracts facial regions, converts them into minimalist single-line sketches using the Gemini Vision API, optimizes stroke order through graph-based path planning, and executes smooth trajectories on a 6-DoF collaborative manipulator. This perception-cognition-action pipeline integrates computer vision, neural artistic abstraction, motion optimization, and robot control. User ratings on a 5-point scale were high for sketch quality 4.33, perceived execution 4.53, and user experience 4.65, indicating recognizable, appealing, and engaging robotic portraits.
Explore related subjects
Keep this discovery
Miguel Altamirano Cabrera, Aleksey Fedoseev, Iana Zhura, Dzmitry Tsetserukou. 2026-08-01. GeminiPainter's sequence-formed pipeline comprised of perception, cognition, planning, and action stages. https://arxiv.org/abs/2608.00829
Cite the original work for its findings. Save a collection to share your selection of sources.