arXiv · 2609.34863
Revisit to Segment: Working Memory Distillation for Reasoning Segmentation
Abstract
Multimodal large language models (MLLMs) have approached image segmentation by reasoning about visual content and predicting target locations. Their generated responses contain reasoning traces and localization proposals that can serve as working memory when revisiting the same image and query. Our exploration reveals that MLLMs benefit from using this self-generated working memory as context, leading to enhanced reasoning segmentation. Motivated by this finding, we seek to strengthen the backbone model's reasoning segmentation capabilities by distilling the guidance gained from revisiting prior attempts, enabling it to benefit with or without working memory at inference time. To this end, we propose Reasoning Segmenter with Working Memory (SWiM), a working-memory distillation framework for reasoning segmentation. Specifically, SWiM selects rollouts based on segmentation quality to construct working memory and uses the memory-conditioned model as a teacher. The teacher provides token-level distributional supervision along student-generated trajectories, while the student receives only the original image and query. Joint optimization of on-policy self-distillation and outcome-based reinforcement learning combines working-memory guidance with direct feedback on segmentation quality. Extensive experiments on reasoning segmentation benchmarks demonstrate that SWiM achieves state-of-the-art performance, validating the effectiveness of working-memory distillation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Cilin Yan, Yilun Qiu, Wanyang Zhang, Rui Zu, Xiaolong Jiang, Yao Hu, Jiayin Cai. 2026-09-28. Revisit to Segment: Working Memory Distillation for Reasoning Segmentation. https://arxiv.org/abs/2609.34863
Cite the original work for its findings. Save a collection to share your selection of sources.