ArXiv

Uncertainty quantification for trustworthy deep learning: Methods and measures

Authors
H. Martin Gillis, Thomas Trappenberg
Categories
stat.ML, cs.LG
arXiv
https://arxiv.org/abs/2607.28248v1
PDF
https://arxiv.org/pdf/2607.28248v1

Brief

Uncertainty quantification in deep learning is surveyed by H. Martin Gillis and Thomas Trappenberg (arXiv 2026-07-30), which reviews ensemble-based and approximate Bayesian UQ, organizing methods into five families and distinguishing predictive-distribution methods from uncertainty measures. The authors evaluate theoretical motivation, implementations, ensemble diversity and measures (entropy vs pairwise divergence), consolidate evaluation protocols, and highlight open research directions. Summary based on abstract-only.

Why it matters

Survey by H. Martin Gillis and Thomas Trappenberg (arXiv 2026-07-30) organizes UQ methods into five families: Bayesian neural networks, Monte Carlo Dropout, deep ensembles, efficient ensemble approximations, and last-layer/single-pass approaches.

Key details

  • The paper separates methods that produce predictive distributions from the measures that summarize uncertainty, contrasts entropy decomposition with pairwise divergence measures, and consolidates evaluation methodology while reviewing ensemble diversity theory.
  • It situates adjacent work (evidential/prior networks, conformal prediction, post-hoc calibration) and decision-time tasks (out-of-distribution detection, selective prediction), and lists open directions: efficient epistemic measures for classification, last-layer diversity, diversity and calibration under shift, and hybrid architectures.
Source evidence

Abstract

The deployment of deep neural networks in safety-critical domains demands reliable estimates of predictive confidence, yet conventional architectures lack principled uncertainty quantification. This survey provides a structured, critical review of methods for Uncertainty Quantification (UQ) in deep learning, scoped to ensemble-based and approximate Bayesian approaches and the measures used to summarize their outputs. Relative to existing UQ surveys, our contribution is depth on efficient ensemble approximations and single-pass methods, and a unified treatment that separates the method producing a predictive distribution from the measure that summarizes its uncertainty. We organize methods into five families: Bayesian neural networks, Monte Carlo Dropout, deep ensembles, efficient ensemble approximations, and last-layer or single-pass approaches. We situate adjacent work on evidential and prior networks, conformal prediction, and post-hoc calibration, together with the decision-time tasks of out-of-distribution detection and selective prediction. For each, we examine theoretical motivation, implementation, empirical performance, and limitations. We then review ensemble diversity theory and uncertainty measures and their decompositions, contrasting the entropy decomposition with pairwise divergence measures, and consolidate evaluation methodology so that our qualitative comparisons share a common basis. We close with a brief treatment of uncertainty in large language models and open research directions, including efficient epistemic measures for classification, last-layer diversity, diversity and calibration under shift, and hybrid architectures.