Confidence regions for SAFE AI compliance scores
Artificial Intelligence trustworthiness scores are moving from research dashboards into compliance claims and procurement decisions. A compliance claim is a statistical decision based on the comparison between an estimated compliance metric and a set threshold. The recently proposed integrated SAFE AI metrics are based on three basic rank-graduation measures of accuracy, explainability and robustness, expressed on a common footing. The metrics are integrated into a single compliance score, so far reported as a point value. In this paper we propose to add an uncertainty layer to the SAFE AI compliance score. All component metrics are estimated on a shared test sample, so their errors covary. We estimate the full covariance of the component vector with a paired bootstrap, and quantify departures from the source-wise independence baseline suggested by the integrated-metrics framework's variability decomposition. A closed-form identity splits such departure into an uncertainty-weighted effective dimension and an exposure-weighted data-driven correlation, with both depending on the aggregator and the estimated uncertainty of the component metrics. The application of our proposal to both real and simulated data shows that the assumption of independence can either narrow or widen confidence bands, while the direction and magnitude of the effect depend on the aggregator type, machine learning model, and perturbation family. The resulting confidence region supports compliance decisions, which can be documented in a formal uncertainty certificate.