arXiv Science⌕ Search

arXiv subjects

Cap Dang Xuan Kiet

Publications and source records attributed to Cap Dang Xuan Kiet.

3 recordsLinked to original sources

Dynamic Time Step Prediction in Inverse Heat Dissipation for Blur-Like Image Restoration Tasks

When using diffusion models to target image restoration problems, diffusion inversion is typically employed to retain relevant image information from the degraded images. Instead of inverting back to the initial time step (i.e., T), many methods invert to a pre-determined intermediate time step, in order to better preserve information from degraded source images. However, a pre-determined time step for inversion is not ideal for reconstruction, as a severely degraded image requires an earlier starting time step than a mildly degraded one. In addition, DDIM-based models corrupt the original signal by adding Gaussian noise, which can be mismatched to the nature of blur-like degradations, such as blur, haze, and low-light. To address these problems, we propose two solutions: (1) we adopt an alternative diffusion process, called the Inverse Heat Dissipation Model, that diffuses the input image by gradually blurring a data point (2) we propose to implement a time predictor to estimate the starting time step for the inversion, with the model learning to adapt to the degradation severity. Extensive experiments on standard benchmarks show that our method achieves state-of-the-art performance in both quantitative and qualitative evaluations, with excellent generalization to many restoration tasks.

cs.CV↗

SIEVE: Selective attention-value Suppression for Vision-Language Models Unlearning

The ability of vision-language models (VLMs) to associate visual identities with biographical information creates a need for selective unlearning of personally identifiable information (PII) while preserving permitted knowledge about the same individual. This setting is challenging because both sensitive and retained information can share the same visual inputs and intermediate representations. We introduce SIEVE, a simple and effective framework for selective VLM unlearning. SIEVE directly regularizes attention-value representations while also controlling model outputs. SIEVE suppresses attention values for forget examples toward a constant zero, while preserving retain-example representations by matching them to a frozen reference model. These objectives are combined with sequence-level forget and retain supervision, enabling targeted forgetting without largely affecting retained knowledge. Extensive experiments show that SIEVE achieves state-of-the-art performance on unlearning with multiple model-modality settings, while maintaining competitive retained utility. Ablation studies further show that value suppression and negative cross-entropy contribute complementary forgetting signals, while reference-based value matching substantially reduces utility degradation. These results demonstrate that attention values provide an effective intervention point for selective multimodal unlearning when sensitive and retained knowledge are closely related.

cs.LG↗

Learning from Failure: Leveraging Unreliable Predictions in Semi-Supervised Real-World Adverse Weather Removal

Adverse weather image restoration aims to recover images degraded by rain, haze, snow, and other weather-induced artifacts, thereby improving the robustness of outdoor vision systems. Existing unified restoration models exhibit limited generalization to real-world scenes due to their reliance on synthetic supervision and insufficient semantic constraints. In this paper, we propose a novel student--teacher semi-supervised framework that addresses both challenges. Specifically, we introduce an unreliable database that preserves failed teacher predictions as informative negative samples for contrastive learning, while a reliable database stores high-quality teacher predictions as positive samples. By jointly exploiting reliable pseudo-ground truths and unreliable teacher outputs, the proposed framework learns to enhance desirable restoration characteristics while avoiding common failures. We further propose a phase spectrum-based semantic constraint that replaces computationally expensive text-based supervision with an efficient and naturally aligned semantic prior. An adaptive phase consistency loss is also designed to dynamically balance supervision between the degraded input and teacher pseudo-ground truths according to degradation severity. Extensive experiments on real-world benchmarks demonstrate that the proposed method consistently outperforms existing state-of-the-art approaches in restoration quality and perceptual fidelity while exhibiting stronger generalization to real-world adverse weather conditions.

cs.CV↗