arXiv Science⌕ Search

arXiv · 2610.04258

FloVMos: Optical Flow-based Medical Video Mosaicking

Abstract

Biomedical imaging modalities often require a trade-off among resolution, field of view (FOV), and acquisition speed. Video mosaicking offers a strategy to overcome this limitation by computationally stitching sequential high-resolution frames into a wide-FOV composite. However, existing methods struggle with non-rigid deformations, and modality-specific artifacts arising in clinical and research imaging. Here, we present FloVMos, a generalizable, optical-flow-based deep learning framework for real-time video mosaicking across diverse biomedical imaging modalities. FloVMos achieves robust, pixel-level registration by fine-tuning an optical flow model on synthetic training data with ground-truth deformation fields. We introduce a pipeline for generating this training data, simulating realistic tissue motion and imaging distortions from existing mosaics or raw videos. Our automated synthetic data generation and optical flow model training based on this data allow users to adapt FloVMos to different imaging modalities. To demonstrate this, we applied FloVMos to seven diverse imaging modalities: reflection confocal microscopy, open-top light-sheet microscopy, fetoscopy, laparoscopy, dermoscopy, sparse spectral microscopy, and endoscopy. FloVMos outperforms conventional baselines in accuracy, robustness, and speed for all the tested modalities. This adaptable and training-efficient framework enables large-area visualization with real-time performance and may support broader use of video-based biomedical imaging in research and clinical workflows.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jinyang Liu, Sandesh Ghimire, Chaman Singh, Jennifer Dy, Milind Rajadhyaksha, Dana H. Brooks, Octavia Camps, Kivanc Kose. 2026-10-03. FloVMos: Optical Flow-based Medical Video Mosaicking. https://arxiv.org/abs/2610.04258

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

A Unified Deep Learning Framework for Motion Correction in Medical Imaging

Deep learning has shown significant value in medical image registration for motion correction; however, current techniques are either limited by the type and range of motion they can handle or require iterative inference and/or retraining for new imaging data. To address these limitations, we introduce UniMo, a Unified Motion Correction framework that uses deep neural networks to correct various types of motion in medical imaging. UniMo uses an alternating optimization scheme with a unified loss function to train an integrated model of 1) an equivariant neural network for global rigid motion correction and 2) an encoder-decoder network for local deformations. It features a geometric deformation augmenter that 1) enhances the robustness of global motion correction by addressing local deformations, whether caused by non-rigid motion or geometric distortions, and 2) generates augmented data to improve training. As a hybrid model that uses both image intensities and shapes, UniMo is robust to appearance variations and generalizes to various imaging modalities without retraining. We trained and tested UniMo for motion tracking in fetal magnetic resonance imaging, which is challenging due to 1) both large rigid and non-rigid motion and 2) large variations in image appearance. We then tested the trained model, without retraining, on three public datasets: MedMNIST, lung CT, and BraTS. UniMo surpassed existing motion correction methods in accuracy and, notably, enabled one-time training on a single modality while maintaining high stability and adaptability across multiple unseen imaging datasets. By offering a unified solution to motion correction, UniMo marks a significant advance in challenging applications with a mixture of bulk motion and local deformations. Code is available at https://github.com/IntelligentImaging/UNIMO

eess.IV↗

MedForj: An open, large-scale foundational generative prior for high-resolution 3D brain MRI

This work introduces MedForj, a suite of 3D foundational generative priors based on diffusion models. The MedForj models were trained on $72{,}659$ 1~mm isotropic 3D $T_1$-weighted MRI human brain image volumes from $38{,}174$ subjects, drawn from a curated corpus of $80{,}675$ volumes from $42{,}506$ subjects spanning $38$ publicly available datasets. These training images were manually inspected to exclude those with poor quality and excessive pathology, and otherwise were minimally processed. The models include six different diffusion training strategies: rectified flow, latent diffusion rectified flow, flow matching, velocity prediction, clean prediction, and noise prediction. Image samples produced by each of these models were compared to each other and against real, ground truth data under downstream segmentation distributions, FID, five inverse problems, and blind human inspection in an observer study. Flow matching was the strongest strategy overall, achieving the best inverse problem solving results at $28.80$~dB PSNR and $0.874$ SSIM averaged over the five forward problems, the highest rate of reconstructions judged real by blind human raters at $72.6\%$, and the closest per-structure match to real segmented anatomy in a permutation test. It was not best everywhere: rectified flow produced the most convincing unconditional samples in the observer study and the best FID, and the latent rectified-flow model achieved the smallest joint distributional distance to real anatomy. No other strategy, however, performed consistently well across all four evaluations. We therefore recommend flow matching as the default MedForj prior, while releasing every strategy so that the choice can be revisited per application. All model weights and corresponding code are publicly available at https://github.com/piksl-research/medforj.

eess.IV↗

Gaussian Surrogates for Poisson Imaging: Some Theoretical and Empirical Results

In imaging inverse problems with Poisson-distributed measurements, it is common to use objectives derived from the Poisson likelihood. But performance is often evaluated by mean squared error (MSE), which raises a practical question: how much does a Poisson objective matter for MSE, even at low dose? We analyze the MSE of Poisson and Gaussian surrogate reconstruction objectives under Poisson noise. In a stylized diagonal model, we show that the unregularized Poisson maximum-likelihood estimator can incur large MSE at low dose, while Poisson MAP mitigates this instability through regularization. We then study two Gaussian surrogate objectives: a heteroscedastic quadratic objective motivated by the normal approximation of Poisson data, and a homoscedastic quadratic objective that yields a simple linear estimator. We show that both surrogates can achieve MSE comparable to Poisson MAP in the low-dose regime, despite departing from the Poisson likelihood. Numerical computed tomography experiments indicate that these conclusions extend beyond the stylized setting of our theoretical analysis.

eess.IV↗