arXiv Science⌕ Search

arXiv · 2610.05182

TIRMamba: A Thermal-Prior-Modulated State-Space Network for Sub-Million-Parameter Infrared Image Super-Resolution

Abstract

Infrared image super-resolution is currently led by Mamba-based networks with 26 to 37 million parameters, which are difficult to deploy on the airborne and handheld platforms where thermal imaging is most needed. This paper presents TIRMamba, a network with 896K to 910K parameters for single-channel thermal imagery. A Thermal Prior Highway computes gradient, local-contrast and spectral cues once at the input and, through one adapter per residual group, modulates a weight-tied bidirectional state-space trunk and gates its dual-scale detail branch; a tri-path reconstruction adds the learned residual to a bicubic radiometric baseline. Because the standard benchmark provides only 265 infrared training images and evaluates fusion products on full images, we train with a replay strategy: grayscale DIV2K pre-training followed by fine-tuning on 64-pixel patches drawn with equal probability from the infrared and natural corpora. At scale factor 4, TIRMamba matches the strongest protocol-trained methods on both official test sets with 29 to 40 times fewer parameters and 2.8 to 9.4 times lower latency; at scale factor 2 it gives the highest SSIM on both. A variant with prior-conditioned selectivity, TIRMamba-Rad, corrects a 3 dB raw-thermal failure of an intermediate size-invariant design and gives the best results at scale factor 4 on raw-thermal, unmanned-aerial-vehicle and independent-sensor test sets. Code and trained models will be released at https://github.com/julian135707/TIRMamba upon acceptance.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Chun-An Lin, Tsung-Jung Liu, Yen-Chieh Ouyang. 2026-10-04. TIRMamba: A Thermal-Prior-Modulated State-Space Network for Sub-Million-Parameter Infrared Image Super-Resolution. https://arxiv.org/abs/2610.05182

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Distribution-Aligned Representation Adaptation for DJSCC over Hybrid Wireless-Wired Networks

Deep joint source-channel coding (DJSCC) has emerged as a robust alternative to traditional separate coding for communications through wireless channels. Existing DJSCC approaches focus primarily on point-to-point wireless communication scenarios, while neglecting end-to-end communication efficiency in hybrid wireless-wired networks such as 5G and 6G communication systems. Considerable redundancy in DJSCC symbols against wireless channels becomes inefficient for long-distance wired transmission. Furthermore, DJSCC symbols must adapt to the varying transmission rate of the wired network to avoid congestion. In this paper, we propose a novel framework designed for efficient wired transmission of DJSCC symbols within hybrid wireless-wired networks, namely Rate-Controllable Wired Adaptor (RCWA). RCWA achieves redundancy-aware coding for DJSCC symbols to improve transmission efficiency, which removes considerable redundancy present in DJSCC symbols for wireless channels and encodes only source-relevant information into bits. Moreover, we leverage the Lagrangian multiplier method to achieve controllable and continuous variable-rate coding, which can encode given features into expected rates, thereby minimizing end-to-end distortion while satisfying given constraints. Extensive experiments on diverse datasets demonstrate the superior RD performance and robustness of RCWA compared to existing baselines, validating its potential for wired resource utilization in hybrid transmission scenarios. Specifically, our method can obtain peak signal-to-noise ratio gain of up to 0.7dB and 4dB compared to neural network-based methods and digital baselines on CIFAR-10 dataset, respectively.

eess.IV↗

Reliability assessment and multicenter clinical application of magnetic resonance methods for knee cartilage quantification

Background: This study evaluated interreader agreement and longitudinal performance of MRI methods for knee cartilage volume, thickness, and defect-area quantification. Methods: AI-presegmented masks from 1,189 phase III examinations underwent independent correction by two readers and adjudication. Cartilage volume, three-dimensional ray-tracing thickness (3D-RT), and ray-based defect area (3D-RBA), defined by a 1.5-mm thickness threshold, were calculated. Agreement was assessed using segmentation metrics, intraclass correlation coefficients (ICCs), repeated-measures Bland-Altman analysis, and minimal detectable change at 95% confidence (MDC95). The 3D-RBA framework was evaluated in 120 digital-phantom experiments from 40 participants. Longitudinal analyses included 374 participants, alternative-method comparisons included 65, and retrospective phase II analysis included 24 participants with four visits. Results: Overall AI-to-adjudicated-mask Dice was 0.964 +/- 0.029. Interreader ICCs for volume, thickness, and defect area were 0.956, 0.904, and 0.932; corresponding MDC95 values were 1,596.9 mm^3, 0.227 mm, and 147.4 mm^2. Geometric mean absolute percentage error for defect area was 5.62%, with spatial Dice of 0.961. In 374 participants, volume changes correlated positively with thickness changes (rho=0.431) and negatively with defect-area changes (rho=-0.221). Within-participant phase II correlations followed the same directions in both groups. Conclusions: The workflow demonstrated good interreader agreement. Controlled geometric results and longitudinal associations supported the feasibility of threshold-based defect-area estimation. Volume, thickness, and defect area provide complementary measures of cartilage structure.

eess.IV↗

Adapting Appearance-Based Gaze Estimation to Narrow-Range, Long-Duration Screen Viewing

Appearance-based gaze estimation offers a low-cost alternative to infrared eye tracking for screen-based behavioral and clinical applications, however existing models are typically developed for wide ranges of gaze angle and head pose. Prolonged screen viewing presents a distinct regime in which gaze remains near the screen center, head motion is limited, and calibration drift accumulates over time. In this work, we benchmark six published estimators along with a proposed differential-gaze model on 140 long-duration facial video recordings with synchronized eye tracking under subject-disjoint evaluation and a common budget of calibration frames. Differential estimation was the only static method that significantly improved upon a baseline predictor of each recording's mean gaze, and was further boosted by addition of temporal context. It also yielded the strongest agreement with reference fixations and saccades, demonstrating that low angular error alone could not validate eye movement reconstruction. These findings establish differential estimation as a promising foundation for reliable gaze tracking in long-duration, narrow-range settings.

eess.IV↗