arXiv Science⌕ Search

arXiv · 2610.08838

Deep Learning for Longitudinal Medical Imaging: A Scoping Review

Abstract

Longitudinal medical imaging analysis is a cornerstone of modern medical practice and patient care. Deep learning applied to longitudinal imaging offers wide potential to enhance diagnosis and track disease progression by capturing spatial changes over time. With major advances in single-timepoint deep learning for imaging, there has been growing interest in longitudinal image analysis, given its increased clinical relevance, though technical challenges remain. Several recent innovations may lead to a new era of multi-timepoint image evaluation, yet the scientific landscape, recent progress, and areas of need remain under-characterized. To address this gap, we conducted a scoping review of deep learning methodologies applied to longitudinal medical imaging, yielding 102 studies published between 2018 and 2025. Neurological disorders (48%) and ophthalmic conditions (12%) were the most common clinical applications, with MRI serving as the predominant imaging modality (67%). Sequential feature modeling approaches combining convolutional neural networks (CNNs) with temporal models (LSTM/RNN) were the most frequent methodology (40%), followed by direct feature aggregation across timepoints (23%). Most studies targeted classification tasks (56%), while external validation was performed in only 24% of studies. Our findings highlight that deep learning-based longitudinal imaging analysis remains a promising field, though newer temporal architectures and larger datasets may improve success and clinical adoption of these tools.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Francesca Mussa, Divyanshu Tak, Atlas H. Avval, Sarah Brueningk, Ray H. Mak, Hugo J. W. L. Aerts, Andreas M Rauschecker, Benjamin H. Kann. 2026-09-29. Deep Learning for Longitudinal Medical Imaging: A Scoping Review. https://arxiv.org/abs/2610.08838

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Research on Deep Learning-Based Semantic Segmentation Algorithms for Subcortical Brain Structures

Segmentation of subcortical brain structures is fundamental to computer-aided diagnosis and treatment in neurology and related clinical fields. To improve the accuracy of subcortical structure segmentation, this thesis develops two related deep learning approaches for brain MR images: DenseMedic and the alternate connected neural network (ACNN). The first part of the work develops DenseMedic. First, the OreoDown method accelerates receptive-field growth by introducing larger convolutional strides at earlier layers and restores network depth by interleaving size-preserving convolutional layers, thereby translating faster receptive-field growth into an effective increase in receptive field. Second, DenseMedic instantiates the OreoDown framework using the construction principle of DenseNet and obtains multiscale contextual information through densely connected feature-extraction operations. The second part develops ACNN. First, alternate connections between layers are proposed to instantiate OreoDown as a single-path ACNN. Second, the single path is divided centrally to form a multipath ACNN without changing the stated parameter count, providing a unified architecture for single- and multimodal segmentation. Experiments were conducted on the public IBSR and MRBrainS18 datasets for subcortical brain segmentation. Performance was evaluated using the Dice similarity coefficient (DSC), intersection over union (IoU), 95th-percentile Hausdorff surface distance (HSD95), and average surface distance (ASD). The experimental results showed greater regional overlap and closer boundary agreement between the predicted and reference structures, supporting the accuracy and robustness of both methods for the evaluated subcortical structures.

eess.IV↗

SAIPAS: Simulating aspect-angles-invariant physical adversarial attacks on SAR target recognition models

Synthetic aperture radar (SAR) enables versatile, all-time, all-weather remote sensing. Coupled with automatic target recognition (ATR) leveraging machine learning (ML), SAR is empowering a wide range of Earth observation and surveillance applications. However, the surge of attacks based on adversarial perturbations against the ML algorithms underpinning SAR ATR is prompting the need for systematic research into adversarial perturbation mechanisms. Research in this area began in the digital (image) domain and evolved into the (simulated) physical domain, resulting in physical adversarial attacks (PAAs) that strategically exploit corner reflectors as attack vectors to evade ML-based ATR. Existing PAAs assume that the attacker knows the SAR platform's aspect angles, restricting their applicability to idealized scenarios. We propose the Simulated Aspect-angle-Invariant Physical Adversarial SAR attack (SAIPAS), a framework that determines adversarially effective positions and orientations of any given set of reflectors, regardless of their number or size, even when the attacker lacks knowledge of the SAR platform's aspect angles. This is enabled by rigorous physics-based modeling of the reflected signal and the SAR imaging process. To facilitate mapping between image and scene coordinates, we additionally propose a method for generating bounding boxes in densely sampled azimuthal SAR images, allowing the target object to serve as a spatial reference. The resulting adversarial configurations offer a clear physical interpretation while maintaining high fooling rates across continuous aspect trajectories under more realistic operational assumptions (69.1% for AConvNet for a four-reflector white-box attack). This paper has supplementary material available, which demonstrates the SAIPAS.

eess.IV↗

Hybrid++: The Bridge between PDE Models and Deep Learning for Gamma Noise Removal

Multiplicative gamma noise is one of the dominant noise factors in Synthetic Aperture Radar (SAR) and medical ultrasound images. They are dependent on pixel level noises due to which they are highly varying across the image and harder to handle as compared to additive noise. The denoising methods to address this noise currently include classical partial differential equation (PDE) methods which are good for interpretation but lack the restoration ability as compared to the state-of-the-art models while deep convolutional networks like DnCNN achieve high performance but at the cost of transparency due to which practitioners are skeptical to use them in high-risk critical fields like medicine. This paper presents Hybrid++, a novel trainable nonlinear reaction-diffusion architecture that addresses both the concerns - staying interpretable while offering performance close to huge black-box models. It combines a fully learnable PDE initialization with a 3-stage reaction-diffusion network having 64-channel multiscale filter banks, 4-layer Squeeze-and-Excitation attention-based influence functions and a 64-dimensional noise level embedding. It uses a two-phase training strategy, stage-wise optimization followed by joint end-to-end refinement which enables co-adaptation of all learnable parameters. On the FoE benchmark, Hybrid++ substantially improves over classical PDE, BM3D and the original TNRD baselines. In the severe-noise setting L=1, it comes within 0.23 dB PSNR of a separately trained DnCNN while using only about 8% of its parameters. We therefore position Hybrid++ not as a universal state-of-the-art image restoration backbone, but as a compact, physically structured reaction-diffusion model for multiplicative gamma noise.

eess.IV↗