arXiv · 2609.23067
LD-RSVIS: A Large-Scale and Diverse Benchmark for Referring Surgical Video Instrument Segmentation
Abstract
Referring surgical video instrument segmentation (RSVIS) aims at segmenting the instrument in a surgical video, given a textual description. Despite recent progress, current models are trained and assessed on relatively small-scale benchmarks, hindering the development of more general RSVIS. In addition, existing benchmarks support only the single-target expression that refers to one instrument in the video, while overlooking multi-target and no-target referring expressions, restricting the applicability of RSVIS in practical scenarios. Addressing these issues, we propose LD-RSVIS, a new benchmark aiming to facilitate more robust and general RSVIS. Specifically, LD-RSVIS consists of 3,536 surgical videos with 1.09 million frames and covers a broad set of 30 instrument classes from 25 various procedures. By including abundant videos and classes, LD-RSVIS could benefit large-scale training and evaluation of more general RSVIS methods. Besides, unlike existing datasets, LD-RSVIS offers diverse referring settings, including no-target, single-target, and multi-target expressions, which enables the development of more practical RSVIS models in real applications. In order to ensure high-quality annotations, all videos in LD-RSVIS are manually labeled with multiple rounds of inspection and refinement. To our knowledge, LD-RSVIS is the largest and most diverse benchmark for RSVIS. To analyze LD-RSVIS and to provide comparison for future research, we evaluate 12 representative methods, and the results reveal that more efforts are required for improvements. To encourage future research, we present a simple yet effective RSVIS method, dubbed Cascade-RSVIS, that first mines target-specific cues using the complementary multi-cue text information and then employs such cues and textual information for segmentation, achieving promising performance. Our benchmark and code will be released.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Zan Wang, Yunhe Feng, Dong Nie, Oluwatosin Oluwadare, Kewei Sha, Yan Huang, Heng Fan. 2026-09-19. LD-RSVIS: A Large-Scale and Diverse Benchmark for Referring Surgical Video Instrument Segmentation. https://arxiv.org/abs/2609.23067
Cite the original work for its findings. Save a collection to share your selection of sources.