arXiv Science⌕ Search

arXiv · 2610.09382

ScribbleEdit: A Benchmark for Scribble-Only Image Editing

Abstract

Scribble-based interaction provides a lightweight and intuitive way for users to specify image editing intents in interactive editing tools. However, current image editing models based on VLMs or LLMs struggle to understand and execute edits based solely on scribble inputs. To systematically study this problem, we construct a new benchmark, ScribbleEdit, that evaluates the ability of image editing models to perform image editing conditioned on scribbles. This task requires both a deep understanding of the intention of the scribble and an accurate interpretation of its spatial information. In ScribbleEdit, we design an automated data construction pipeline and introduce a dedicated evaluation protocol that explicitly measures intention alignment. Our analysis reveals that existing VLM/LLM-based editing models fail to accurately capture scribble intentions. To guide future progress on scribble-only image editing, we propose a simple yet effective soft-token baseline, which enhances the model's understanding of scribble semantics and outperforms standard image editing models on our benchmark. Our evaluation and baseline together provide a concrete foundation for assessing and improving the scribble-driven image editing.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jie Ren, Hao Kang, Kai Guo, Yiding Yang, Bo Liu, Liming Jiang, Qing Yan, Zichuan Liu, Yizhi Song, Yue Xing, Hui Liu, Xin Lu. 2026-10-07. ScribbleEdit: A Benchmark for Scribble-Only Image Editing. https://arxiv.org/abs/2610.09382

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Event-T2M: Event-level Conditioning for Complex Text-to-Motion Synthesis

Text-to-motion generation has advanced with diffusion models, yet existing systems often collapse complex multi-action prompts into a single embedding, leading to omissions, reordering, or unnatural transitions. In this work, we shift perspective by introducing a principled definition of an event as the smallest semantically self-contained action or state change in a text prompt that can be temporally aligned with a motion segment. Building on this definition, we propose Event-T2M, a diffusion-based framework that decomposes prompts into events, encodes each with a motion-aware retrieval model, and integrates them through event-based cross-attention in Conformer blocks. Existing benchmarks mix simple and multi-event prompts, making it unclear whether models that succeed on single actions generalize to multi-action cases. To address this, we construct HumanML3D-E, the first benchmark stratified by event count. Experiments on HumanML3D, KIT-ML, and HumanML3D-E show that Event-T2M matches state-of-the-art baselines on standard tests while outperforming them as event complexity increases. Human studies validate the plausibility of our event definition, the reliability of HumanML3D-E, and the superiority of Event-T2M in generating multi-event motions that preserve order and naturalness close to ground-truth. These results establish event-level conditioning as a generalizable principle for advancing text-to-motion generation beyond single-action prompts.

cs.GR↗

SECA: Strict-Elastic Compact Activation for Differentiable Elastoplastic Deformation

Return mapping keeps an exact elastic domain but switches its response tangent at yield, while smoothing laws with a sub-yield tail admit plastic flow below the threshold. We present SECA, a Strict-Elastic Compact Activation method for differentiable elastoplastic simulation. SECA prescribes a bounded $C^2$ activation over a finite positive-overstress interval and integrates it into a Perzyna-type flow law, so the implicit update returns exactly zero plastic increment for sub-yield trials while smoothing the onset. For scalar and proportional isotropic $J_2$ updates we derive an exact decomposition of the plastic-increment discrepancy into finite-width and viscosity contributions, and an explicit algebraic budget for the transition width. The update is embedded in finite-deformation loading histories with contact and complete unloading, and differentiated to released equilibria. Controlled indentation results stay closely matched. In the tested coarse cube-indentation run, where the transition is crossed without intermediate increments, SECA reduces whole-run Newton iterations by 47% and backtracking halvings by 93% against return mapping. A four-method comparison evaluates released-shape error and computational cost across loading partitions, and on a contact scene the path derivative drives a bounded control task to its target in three iterations. Further examples show elastic recovery and residual deformation under compression, tension and bending.

cs.GR↗

emg2face: Expressive Facial Animation with High-Density Surface EMG

Facial movements convey subtle and important information that is critical for human social communication. Optical methods for face capture are difficult or impossible to use when the face is occluded by head-mounted devices (HMDs), such as VR headsets. Even with a clear line of sight, such methods raise privacy concerns and require head-mounted capture rigs that offset cameras and lighting from the face. We show that high-density surface electromyography (HD-sEMG) provides a viable non-optical alternative that addresses these challenges. We measured 64 EMG channels, using two textile EMG grids, with 32 from the forehead (typically occluded by an HMD) and 32 from the side of the face. EMG data were digitized at 2048 Hz and filtered. Facial movements were simultaneously recorded and used to estimate 478 3D facial landmarks using MediaPipe's Face Landmarker. A major challenge in such multimodal recordings is synchronizing EMG and video data, which have different sampling frequencies and independent clocks. We developed a novel synchronization method using analog audio bursts that is capable of sub-millisecond synchronization. We also developed a staged fitting method that fits a recent high-resolution parametric head model (GNM), with 253 identity blendshapes and 383 expression blendshapes, to the MediaPipe landmarks as participants performed different facial expressions. We trained a deep neural network comprising per-grid spatial encoders followed by a dilated temporal convolutional network (TCN) to predict blendshape parameters from HD-sEMG signals at 100 Hz. Once trained, the network can predict expression blendshapes solely from HD-sEMG recordings. The output can be rendered using standard real-time blendshape animation methods. We demonstrate the methods using recordings from 25 participants, and direct expression transfer to a variety of human faces and non-human characters.

cs.GR↗