arXiv · 2609.20247
Unifying Image Quality Assessment Datasets: MOSAIQ-500K and MOSAIQ-Bench
Abstract
Image quality assessment (IQA) datasets use different subjective protocols and rating scales, so their scores are not directly comparable. The lack of a common perceptual scale hinders multi-dataset training and precludes direct inter-dataset evaluation. We address this by conducting a new subjective experiment and using its ratings as perceptual anchors to fit monotonic mappings that place the existing scores of 23 IQA datasets on a common quality scale while preserving within-dataset rankings. The resulting dataset, MOSAIQ-500K, contains over 500,000 images and is, to our knowledge, the largest IQA dataset with perceptually aligned subjective scores. We also propose MOSAIQ-Bench, an inter-dataset benchmark, and use it to evaluate 31 IQA methods, revealing substantial gaps between intra- and inter-dataset performance, particularly on authentic distortions. The aligned scores offer a key advantage: replacing specialized multi-dataset training mechanisms with standard regression losses yields comparable intra-dataset accuracy and better inter-dataset performance. MOSAIQ thus provides a practical foundation for combining heterogeneous IQA datasets to train more generalizable IQA models. Code and dataset will be made public at ivc.uwaterloo.ca/projects/unifying_iqa_datasets/ upon acceptance.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Wenbo Yang, Zhongling Wang, Jialu Xu, Jinghan Zhou, Zhou Wang. 2026-09-17. Unifying Image Quality Assessment Datasets: MOSAIQ-500K and MOSAIQ-Bench. https://arxiv.org/abs/2609.20247
Cite the original work for its findings. Save a collection to share your selection of sources.