arXiv · 2610.07913
Multimodal Knowledge Distillation for Gastric Adenocarcinoma Classification from Whole-Slide Images
Abstract
Gastric adenocarcinoma (GA) is a leading cause of cancer-related mortality worldwide, and accurate histopathological subtype classification from whole-slide images (WSIs) is essential for effective treatment planning. While multimodal approaches that integrate pathology report text with WSIs can improve classification, existing methods often depend on computationally expensive transformer architectures and large language models. We propose a multimodal knowledge distillation (MKD) framework that combines a pretrained WSI image encoder and a clinical text encoder using Low-Rank Multimodal Fusion (LMF) to efficiently model cross-modal interactions during training. Each WSI is represented as a bag of patches paired with a slide-level diagnostic caption. The teacher model learns fused image-text representations for subtype classification, while the student model distills this knowledge to enable accurate image-only inference. We evaluate our method on the PatchGastric benchmark dataset and achieve at least 3.35% higher mean accuracy than state-of-the-art approaches, without relying on transformer-based fusion, multi-task learning, or large language models. The source code is available at https://github.com/helomelo1/MKD-LMF.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Shrihari Dumbre, Bikash Santra. 2026-10-06. Multimodal Knowledge Distillation for Gastric Adenocarcinoma Classification from Whole-Slide Images. https://arxiv.org/abs/2610.07913
Cite the original work for its findings. Save a collection to share your selection of sources.