ArXiv

Provable Limits and Certified Deferral for Verbalized Uncertainty in Small Language Models

Authors
Jianru Shen
Categories
cs.CL, cs.AI, cs.LG
arXiv
https://arxiv.org/abs/2608.05064v1
PDF
https://arxiv.org/pdf/2608.05064v1

Brief

The paper studies when small (0.5B–14B) open‑weight language models should defer to humans by using verbalized confidence. It proves three calibration limits (including an infeasibility result for temperature scaling and a 200‑question Clopper–Pearson certification) and tests eleven instruction‑tuned models on ARC‑Challenge and TruthfulQA (25,168 predictions). Platt scaling achieves ECE≈0.02 but certified autonomy is scarce (3 pairs at 20% risk, 0 at 10%). Abstract only; full text not included here.

Why it matters

Three theoretical limits: (1) strictly monotone calibration preserves the risk–coverage frontier and error‑detection AUROC; (2) temperature scaling cannot calibrate models whose verbalized confidence stays >0.5 while accuracy <0.5; (3) a Clopper–Pearson procedure turns a 200‑question calibration set into a finite‑sample risk certificate under i.i.d. deployment.

Key details

  • Empirical evaluation across eleven instruction‑tuned models (0.5B–14B parameters) on ARC‑Challenge and TruthfulQA (25,168 local predictions): eight of 22 model×task pairs reached the temperature‑scaling infeasibility floor within 1 percentage point; Platt scaling reduced ECE to as low as 0.02; only three model×task pairs achieved certified autonomy at a 20% risk budget and none at 10%.
  • Additional contributions: identified and repaired an answer‑ordering artifact in TruthfulQA (multiple‑choice form). Paper accepted to MIWAI 2026 (Springer LNAI); summary based on the abstract (full text not provided here).
Source evidence

Abstract

Small open-weight language models increasingly run in private, offline, and cost-sensitive settings, where the key deployment question is not only what a model answers but when it should defer to a human. We study whether verbalized confidence can support risk-controlled deferral, evaluating eleven instruction-tuned models from three families, 0.5B to 14B parameters, on ARC-Challenge and TruthfulQA with 25,168 local predictions. Three theoretical results delimit what calibration can provide: strictly monotone calibration preserves the risk-coverage frontier and error-detection AUROC; temperature scaling cannot calibrate models whose confidence stays above one half while accuracy falls below it; and a Clopper-Pearson procedure converts a 200-question calibration set into a finite-sample risk certificate under an i.i.d. deployment assumption. Empirically, eight of 22 model-task pairs hit the temperature-scaling infeasibility floor within one percentage point of the predicted bound. Platt scaling reduces ECE to as low as 0.02, yet certified autonomy at a 20% risk budget is granted to only three model-task pairs and to none at 10%. We also identify and repair an answer-ordering artifact in the multiple-choice form of TruthfulQA. Calibration gives confidence semantics; certified deferral determines when small models are safe to use.

Comment: Accepted at MIWAI 2026 (The 19th International Conference on Multi-disciplinary Trends in Artificial Intelligence), to appear in Springer LNAI