arXiv Science⌕ Search

arXiv subjects

Mohammad Kazzazi

Publications and source records attributed to Mohammad Kazzazi.

3 recordsLinked to original sources

Distilling Vision-Language Models for On-Device Fire Understanding

Vision-language models (VLMs) offer a promising alternative to conventional fire detection systems by reasoning about the semantic context of a scene and thus reducing false alarms, yet their large model size makes deployment on embedded fire sensors impractical. In this paper, we study how domain-specialized VLMs can be compressed for fully on-device deployment without losing the safety-critical behavior required for fire detection. We develop a teacher-student knowledge distillation framework in which large VLMs fine-tuned for fire understanding can be distilled into lightweight students. Experiments across multiple VLM families and model scales show that compact students preserve most of their teachers' fire-understanding capability. We further deploy the distilled models on our commercial Detectium fire detection sensor and jointly evaluate reasoning accuracy, latency, and memory usage. The results show that compression and deployment affect not only accuracy but also model failure modes, with Qwen2.5-0.5B providing the strongest overall deployment trade-off. Our findings provide broader guidance for deploying domain-specialized VLMs in resource-constrained, safety-critical settings.

cs.AI↗

CLEAR: A Closed-Form Minimal-Sensor TDOA/FDOA Estimator for Moving-Source IoT Localization

This paper presents CLEAR -- a closed-form localization estimator with a reduced sensor network. The proposed method is a computationally efficient, two-stage estimator that fuses time-difference-of-arrival (TDOA) and frequency-difference-of-arrival (FDOA) measurements with a minimal number of sensors. CLEAR localizes a moving source in N-dimensional space using only N+1 sensors, achieving the theoretical minimum sensor count. The first stage introduces auxiliary range and range-rate parameters to construct a set of pseudo-linear equations, solved via weighted least squares. An algebraic elimination using Sylvester's resultant then reduces the problem to a quartic equation, yielding closed-form estimates for the nuisance variables. A second, lightweight linear refinement stage is applied to mitigate residual bias. Under mild Gaussian noise assumptions, the estimator's position and velocity estimates are statistically efficient, closely approaching the Cramer-Rao lower bound (CRLB). Extensive Monte Carlo simulations in 2-D and 3-D scenarios demonstrate CRLB-level accuracy and consistent performance gains over representative two-stage and iterative baselines, confirming the method's high suitability for power-constrained, distributed Internet of Things (IoT) applications such as UAV tracking and smart transportation.

eess.SP↗

Enhancing Skin Cancer Diagnosis (SCD) Using Late Discrete Wavelet Transform (DWT) and New Swarm-Based Optimizers

Skin cancer (SC) stands out as one of the most life-threatening forms of cancer, with its danger amplified if not diagnosed and treated promptly. Early intervention is critical, as it allows for more effective treatment approaches. In recent years, Deep Learning (DL) has emerged as a powerful tool in the early detection and skin cancer diagnosis (SCD). Although the DL seems promising for the diagnosis of skin cancer, still ample scope exists for improving model efficiency and accuracy. This paper proposes a novel approach to skin cancer detection, utilizing optimization techniques in conjunction with pre-trained networks and wavelet transformations. First, normalized images will undergo pre-trained networks such as Densenet-121, Inception, Xception, and MobileNet to extract hierarchical features from input images. After feature extraction, the feature maps are passed through a Discrete Wavelet Transform (DWT) layer to capture low and high-frequency components. Then the self-attention module is integrated to learn global dependencies between features and focus on the most relevant parts of the feature maps. The number of neurons and optimization of the weight vectors are performed using three new swarm-based optimization techniques, such as Modified Gorilla Troops Optimizer (MGTO), Improved Gray Wolf Optimization (IGWO), and Fox optimization algorithm. Evaluation results demonstrate that optimizing weight vectors using optimization algorithms can enhance diagnostic accuracy and make it a highly effective approach for SCD. The proposed method demonstrates substantial improvements in accuracy, achieving top rates of 98.11% with the MobileNet + Wavelet + FOX and DenseNet + Wavelet + Fox combination on the ISIC-2016 dataset and 97.95% with the Inception + Wavelet + MGTO combination on the ISIC-2017 dataset, which improves accuracy by at least 1% compared to other methods.

cs.CV↗