arXiv · 2610.05825
Usefulness of Quantile-Aware Diffusion Modeling for Highly Imbalanced Tabular Data
Abstract
Classification problem in the context of highly imbalanced data is a major challenge in many real-world applications (e.g., FinTech, healthcare, etc.). In these cases, the vast majority of instances belong to a single class and a small fraction represent the minority class (often the most critical class). Recently, diffusion models have emerged as powerful approaches to reduce the degree of ``imbalanced-ness'' in the dataset; they work by generating synthetic data by capturing complex data distributions using iterative transformations. However, standard diffusion models are not inherently suited to highly skewed or heavy-tailed data, due to inbuilt quadratic error loss, which lacks the structural sensitivity to capture rare, extreme values, and minority-class nuances. We propose a novel approach, namely, Quantile-TabDDPM, based on a quantile-regularized denoising objective that combines the standard quadratic error loss with a quantile loss term to explicitly capture rare events while preserving the theoretical grounding of the original denoising objective. We extensively evaluated our approach on a real-world credit card transaction dataset characterized by extreme class imbalance. The results demonstrate that the integration of diffusion-based synthetic data generation with a quantile-regularized denoising objective provides a robust and effective framework for fraud detection in highly imbalanced datasets.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Abu Talha, Peng Liu, Souradyuti Paul. 2026-10-05. Usefulness of Quantile-Aware Diffusion Modeling for Highly Imbalanced Tabular Data. https://arxiv.org/abs/2610.05825
Cite the original work for its findings. Save a collection to share your selection of sources.