Revisit to Segment: Working Memory Distillation for Reasoning Segmentation

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
The reasoning segmentation performance of multimodal large language models is significantly constrained in the absence of working memory. To address this limitation, this work proposes SWiM, a framework that introduces the first working memory distillation paradigm. Specifically, it supervises student trajectories using teacher model distributions, integrating online policy self-distillation, token-level distribution supervision, and outcome-based reinforcement learning to enable autonomous reasoning optimization without external memory. This approach achieves state-of-the-art performance across multiple reasoning segmentation benchmarks, validating the effectiveness of working memory distillation in enhancing model reasoning capabilities under memory-free input conditions.
📝 Abstract
Multimodal large language models (MLLMs) have approached image segmentation by reasoning about visual content and predicting target locations. Their generated responses contain reasoning traces and localization proposals that can serve as working memory when revisiting the same image and query. Our exploration reveals that MLLMs benefit from using this self-generated working memory as context, leading to enhanced reasoning segmentation. Motivated by this finding, we seek to strengthen the backbone model's reasoning segmentation capabilities by distilling the guidance gained from revisiting prior attempts, enabling it to benefit with or without working memory at inference time. To this end, we propose Reasoning Segmenter with Working Memory (SWiM), a working-memory distillation framework for reasoning segmentation. Specifically, SWiM selects rollouts based on segmentation quality to construct working memory and uses the memory-conditioned model as a teacher. The teacher provides token-level distributional supervision along student-generated trajectories, while the student receives only the original image and query. Joint optimization of on-policy self-distillation and outcome-based reinforcement learning combines working-memory guidance with direct feedback on segmentation quality. Extensive experiments on reasoning segmentation benchmarks demonstrate that SWiM achieves state-of-the-art performance, validating the effectiveness of working-memory distillation.
Problem

Research questions and friction points this paper is trying to address.

Reasoning Segmentation
Multimodal Large Language Models
Working Memory Distillation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Working Memory Distillation
Reasoning Segmentation
Multimodal Large Language Models
Self-Distillation
Reinforcement Learning