ArXiv · 2026
In probabilistic classification, calibration error (CE) measures the average divergence of predicted probabilities f(X) from P(Y|f(X)), the true class distribution for that predicted probability. While being a useful diagnostic tool, it is hard to estimate: popular binning-based estimators are often inconsistent and scale poorly beyond two classes. Recent work rewrites the CE as the excess risk of a model compared to the best recalibration of its own predictions, measured with a proper loss. However, this only works for Bregman-divergence-based calibration errors like the squared error, excluding the more popular L₁-distance-based CE. We show that using prediction-dependent proper scores can alleviate this restriction, allowing us to estimate CEs with general convex divergences, including Lₚ distances with closed-form losses in the binary and multiclass settings. To estimate the excess risk, we introduce a more accurate recalibrator that fits a residual to temperature scaling with gradient boosting. The resulting variational estimator, V-ECE, needs no bins or clusters and lower-bounds the true calibration error in expectation. On a benchmark of semi-synthetic tasks built from real classifiers, with known true CE, V-ECE is among the most accurate binary estimators for every calibration error and significantly outperforms all multiclass estimators. Our results are accompanied by additional theory on Lₚ CE, estimator bias, and over- or under-confidence estimation.
Try inveni