arXiv · 2305.05841
A Self-Training Framework Based on Multi-Scale Attention Fusion for Weakly Supervised Semantic Segmentation
Abstract
Weakly supervised semantic segmentation (WSSS) based on image-level labels is challenging since it is hard to obtain complete semantic regions. To address this issue, we propose a self-training method that utilizes fused multi-scale class-aware attention maps. Our observation is that attention maps of different scales contain rich complementary information, especially for large and small objects. Therefore, we collect information from attention maps of different scales and obtain multi-scale attention maps. We then apply denoising and reactivation strategies to enhance the potential regions and reduce noisy areas. Finally, we use the refined attention maps to retrain the network. Experiments showthat our method enables the model to extract rich semantic information from multi-scale images and achieves 72.4% mIou scores on both the PASCAL VOC 2012 validation and test sets. The code is available at https://bupt-ai-cz.github.io/SMAF.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Guoqing Yang, Chuang Zhu, Yu Zhang. 2023-05-10. A Self-Training Framework Based on Multi-Scale Attention Fusion for Weakly Supervised Semantic Segmentation. https://arxiv.org/abs/2305.05841
Cite the original work for its findings. Save a collection to share your selection of sources.