Beyond point estimation: explaining the posterior in Bayesian cluster analysis
The Bayesian approach to clustering is often appreciated for its ability to provide uncertainty in the partition structure. However, summarising the posterior distribution over partitions remains challenging, due to its discrete, unordered and extremely high-dimensional nature. Existing approaches typically focus on a single representative clustering, which can obscure important features of the posterior, particularly in the presence of multimodality. To address this limitation, we propose a WASserstein Approximation for Bayesian Inference (WASABI), which represents the posterior distribution using a small number of weighted clustering solutions ("particles"), each capturing a distinct region of substantial posterior mass. The approximation is obtained by minimising the Wasserstein distance defined on the space of partitions, equipped with a suitable metric. This formulation naturally leads to a k-means-like algorithm on the partition space that divides posterior samples in representative groups, each represented by one of the clustering solutions. Through synthetic and real-data applications, we demonstrate that WASABI provides an interpretable and computationally efficient low-dimensional summary of the posterior, particularly in settings with overlapping clusters or model misspecification.