arXiv ScienceSearch

arXiv · 2507.11152

Latent Space Consistency for Sparse-View CT Reconstruction

Abstract

Computed Tomography (CT) is a widely utilized imaging modality in clinical settings. Using densely acquired rotational X-ray arrays, CT can capture 3D spatial features. However, it is confronted with challenged such as significant time consumption and high radiation exposure. CT reconstruction methods based on sparse-view X-ray images have garnered substantial attention from researchers as they present a means to mitigate costs and risks. In recent years, diffusion models, particularly the Latent Diffusion Model (LDM), have demonstrated promising potential in the domain of 3D CT reconstruction. Nonetheless, due to the substantial differences between the 2D latent representation of X-ray modalities and the 3D latent representation of CT modalities, the vanilla LDM is incapable of achieving effective alignment within the latent space. To address this issue, we propose the Consistent Latent Space Diffusion Model (CLS-DM), which incorporates cross-modal feature contrastive learning to efficiently extract latent 3D information from 2D X-ray images and achieve latent space alignment between modalities. Experimental results indicate that CLS-DM outperforms classical and state-of-the-art generative models in terms of standard voxel-level metrics (PSNR, SSIM) on the LIDC-IDRI and CTSpine1K datasets. This methodology not only aids in enhancing the effectiveness and economic viability of sparse X-ray reconstructed CT but can also be generalized to other cross-modal transformation tasks, such as text-to-image synthesis. We have made our code publicly available at https://anonymous.4open.science/r/CLS-DM-50D6/ to facilitate further research and applications in other domains.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Duoyou Chen, Yunqing Chen, Can Zhang, Zhou Wang, Cheng Chen, Ruoxiu Xiao. 2025-07-15. Latent Space Consistency for Sparse-View CT Reconstruction. https://arxiv.org/abs/2507.11152

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Hyperspectral Image Restoration and Super-resolution with Physics-Aware Deep Learning for Biomedical Applications

Hyperspectral imaging is a powerful bioimaging tool which can uncover novel insights, thanks to its sensitivity to the intrinsic properties of materials. However, this enhanced contrast comes at the cost of system complexity, constrained by an inherent trade-off between spatial, spectral, and temporal resolution. To overcome this limitation, we present a self-supervised deep learning-based approach that restores and enhances pixel resolution post-acquisition without requiring external training data beyond the images to be restored. Fine-tuned using metrics aligned with the imaging model, our physics-aware method achieves a 16$\times$ pixel super-resolution enhancement and a 12$\times$ imaging speedup without the need of additional training data for transfer learning. Applied to both synthetic and experimental data from five different sample types, including healthy and diseased tissues, we demonstrate that the model preserves biological integrity, as we did not detect systematic loss of biological features or biologically consequential hallucinations in tested datasets. We also concretely demonstrate the model's ability to reveal disease-associated metabolic changes that would otherwise remain undetectable. Furthermore, we provide physical insights into the model's inner workings, paving the way for future refinements that could potentially reveal novel high resolution features in an explainable manner. All methods are available as open-source software on GitHub.

eess.IV

OASIS: Online Adaptive Video Compression via Closed-loop Feedback Control

Computer vision systems are a key building block in an autonomous vehicle, responsible for a range of perception tasks. However, they incur massive data transmission over long communication links from multiple cameras, creating a critical bandwidth and energy bottleneck. Although conventional codecs such as H.264 can reduce data rates, they are ill-suited for real-time vision systems due to high processing latency and energy consumption, as well as their reliance on static user-defined compression settings. In light of these challenges, we propose OASIS, an adaptive video compression framework that integrates lightweight in-sensor compression with task-aware compression ratio control. Based on the real-time task performance, it dynamically updates the optimal compression ratio. Experimental results demonstrate that OASIS generalizes across multiple vision tasks, achieving on average a 6x data compression, 5.8x reduction in link power consumption, and 2.5x reduction in link latency, with at most 1.5% performance degradation.

eess.IV

Multisource Remote Sensing and Geospatial Analysis of Vineyard Wildfire Impacts and Resilience: The 2019 Kincade Fire

Working agricultural landscapes are often treated as background to wildfire disasters, even though they are managed fuel mosaics, productive assets, and parts of regional infrastructure systems. We examine vineyard wildfire resilience during the electrically initiated 2019 Kincade Fire in Sonoma County, California, using an open, event-anchored geospatial framework spanning 4,581 vineyard fields (8,813.2 ha), wildland vegetation, surveyed structures, roads, overhead smoke, and post-fire greenness. Sentinel-2, OpenET, gridMET, soils, terrain, NOAA smoke polygons, an ignition-date OpenStreetMap network, and three-dimensional data inventories were analyzed at native decision scales. Vineyard pixels showed substantially lower descriptive dNBR than wildland pixels inside the perimeter (means 0.130 and 0.337). This contrast did not identify a universal vineyard firebreak effect: a segment-clustered boundary model gave a small negative contrast at 100 m (tau = -0.0166) but changed across bandwidths, failed slope continuity, disappeared in a 100 m donut, and produced a wrong-signed placebo. A 250 m spatial GAM reversed the unconditional pattern: after conditioning on location, terrain, and water use, vineyard fraction was positively associated with dNBR, while residual Moran's I remained 0.519. Beyond spectral impact, all mapped vineyards intersected overhead smoke on at least one day (mean 7.78 potential smoke-days per field), 34.2% of road-network nodes were dead ends, and inside-perimeter vineyards showed a larger greenness deficit through 2021 (recovery ratios 0.815 inside and 0.854 outside). Lower immediate spectral impact therefore did not imply complete resilience. The study provides a reproducible urban-rural informatics template separating descriptive contrasts, conditional associations, exposure indicators, and recovery evidence for decisions in working landscapes.

eess.IV