arXiv Science⌕ Search

arXiv · 2610.09098

On-Device Super-Resolution Imaging for a Low-Cost SPAD Array on a 640-KB-SRAM Microcontroller

Abstract

This work presents a super-resolution (SR) deep-learning (DL) architecture to simultaneously upscale low-resolution (LR) depth and intensity images acquired by a consumer-grade VL53L9CX time-of-flight (ToF) sensor, a compact single-photon avalanche diode (SPAD) ranging device, while targeting deployment on a resource-constrained microcontroller (MCU) with only 640 KB SRAM and 2 MB Flash. The VL53L9CX provides 54 $\times$ 42 depth and intensity measurements. The proposed framework performs $\times$4 spatial SR to reconstruct 216 $\times$ 168 depth and intensity images simultaneously. The network employs separate reconstruction branches for intensity and high-resolution (HR) depth. We train the network with efficient loss functions and investigate the performance by testing multiple compact, hardware-oriented backbones, including Spatially Adaptive Feature Modulation (SAFM), Swift Parameter-Free Attention Network (SPAN), and Residual Local Feature Network (RLFN). We evaluate the network on both synthetic and real measurements, with particular attention to constrained activation and memory storage. We export the selected network to ONNX, quantize it to INT8, generate C inference code for an STM32H563ZI Arm Cortex-M33 MCU, and integrate it with our simplified sensor's firmware. This work is the first to demonstrate a compact DL SR model on a low-cost bare-metal MCU for a consumer-grade LR SPAD sensor, integrated with customized MCU firmware to enable on-device SR inference.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Zhenya Zang, Istvan Gyongy, Mike Davies. 2026-10-06. On-Device Super-Resolution Imaging for a Low-Cost SPAD Array on a 640-KB-SRAM Microcontroller. https://arxiv.org/abs/2610.09098

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Research on Deep Learning-Based Semantic Segmentation Algorithms for Subcortical Brain Structures

Segmentation of subcortical brain structures is fundamental to computer-aided diagnosis and treatment in neurology and related clinical fields. To improve the accuracy of subcortical structure segmentation, this thesis develops two related deep learning approaches for brain MR images: DenseMedic and the alternate connected neural network (ACNN). The first part of the work develops DenseMedic. First, the OreoDown method accelerates receptive-field growth by introducing larger convolutional strides at earlier layers and restores network depth by interleaving size-preserving convolutional layers, thereby translating faster receptive-field growth into an effective increase in receptive field. Second, DenseMedic instantiates the OreoDown framework using the construction principle of DenseNet and obtains multiscale contextual information through densely connected feature-extraction operations. The second part develops ACNN. First, alternate connections between layers are proposed to instantiate OreoDown as a single-path ACNN. Second, the single path is divided centrally to form a multipath ACNN without changing the stated parameter count, providing a unified architecture for single- and multimodal segmentation. Experiments were conducted on the public IBSR and MRBrainS18 datasets for subcortical brain segmentation. Performance was evaluated using the Dice similarity coefficient (DSC), intersection over union (IoU), 95th-percentile Hausdorff surface distance (HSD95), and average surface distance (ASD). The experimental results showed greater regional overlap and closer boundary agreement between the predicted and reference structures, supporting the accuracy and robustness of both methods for the evaluated subcortical structures.

eess.IV↗

Nonparametric Distribution Matching for Self-Supervised Whole-Slide Image Condensation

Histological whole-slide images (WSIs) are central to computational pathology but pose severe computational challenges due to their extremely high resolution, often spanning several gigabytes per slide. To enable scalable learning, existing methods apply self-supervised data condensation to reduce computational cost, but typically rely on heuristic prototype learning and do not explicitly preserve learning-relevant feature distributions for downstream tasks. In response, we introduce a principled reformulation of WSI condensation as a distribution-matching problem under a fixed representational lens, and develop NICER, a tractable approximation framework based on a nonparametric prior with slide-adaptive capacity. Experiments on five histopathology datasets, together with clinical evaluation from a board-certified pathologist, show that NICER consistently outperforms prior methods, achieving an average accuracy improvement of 7.44% while offering improved efficiency-accuracy trade-offs, highlighting the benefits of principled, distribution-aware condensation for scalable histological representation learning. Source codes are available in https://github.com/nmduonggg/NICER.

eess.IV↗

Diagnosing Diversity Collapse and Validating Mask-Conditioned Diffusion for Labeled Microtubule Microscopy

Mask-conditioned diffusion models have become standard for generating labeled training data in biomedical imaging. But the tools used to evaluate them were built for a different problem. Standard metrics like FID and KID score images through ImageNet-trained features. Those features do not transfer to microscopy, leaving the metrics poorly calibrated for specialized domains. Even with a domain-appropriate feature space, a single blended number can still hide a real failure: a model can produce visually plausible images that still transfer poorly to downstream tasks. We identify a specific failure mode: Late in training, realism keeps improving while texture diversity collapses. The model settles onto a fixed color palette, converging to a high-similarity, low-diversity state that standard metrics do not penalize. Therefore, we propose a diversity-aware diagnostic in a domain-validated feature space. It combines a realism axis based on inter-similarity to real images with a diversity axis based on intra-similarity among generated samples. This diagnostic selects a checkpoint that domain experts cannot reliably distinguish from real recordings in a forced-choice study. It also yields useful downstream segmentation: a segmenter trained only on synthetically labeled data reaches a median skeletonised IoU comparable to one trained on real data, at lower cross-fold variance, and clearly ahead of the best available public alternative in this domain, a parametric renderer. We also reproduce a data-efficient hyperparameter-transfer experiment from prior work, tuning several foundation segmenters on a small labeled subset and evaluating on real images: transfer is stronger with our data. We release DiffuMT on HuggingFace, including the triplet dataset, code to reproduce the downstream-utility validation, and a standalone diagnostic tool for evaluating mask-conditioned diffusion models.

eess.IV↗