ArXiv

Statistically Valid Hyperparameter Selection: From Tuning to Guarantees

Authors
Amirmohammad Farzaneh, Osvaldo Simeone
Categories
stat.ML, cs.IT, cs.LG, math.ST
arXiv
https://arxiv.org/abs/2606.25601v1
PDF
https://arxiv.org/pdf/2606.25601v1

Brief

The monograph by Amirmohammad Farzaneh and Osvaldo Simeone formalizes hyperparameter selection using a learn-then-test (LTT) multiple-hypothesis-testing approach that yields finite-sample error-probability guarantees for constraints such as average risk, quantile risk, and information-theoretic measures. The work supplies p-values, e-values and concentration inequalities; full text available on arXiv (2606.25601v1).

Why it matters

Monograph (Farzaneh & Simeone, 2026-06-24; arXiv:2606.25601v1) introduces a unified statistical framework that casts hyperparameter selection as multiple hypothesis testing under a learn-then-test (LTT) paradigm, providing explicit finite-sample control of error probabilities for selected hyperparameters.

Key details

  • Framework supports provable guarantees for application-specific constraints — including bounds on average risk, quantile risk, and information-theoretic constraints — and develops the supporting machinery (p-values, e-values, concentration inequalities) from first principles, contrasting with common heuristic methods like grid search or Bayesian optimization.
Source evidence

Abstract

Hyperparameter selection is a critical step in the deployment of modern artificial intelligence systems, given the need to tune degrees of freedom such as inference-time parameters, implementation-level settings, and thresholds driving decision rules. Despite its practical importance, hyperparameter selection is typically performed using best-effort empirical methods such as grid search or Bayesian optimization, which provide no formal statistical guarantees on reliability or safety. This monograph presents a unified statistical framework for reliable hyperparameter selection, centered on the learn-then-test (LTT) paradigm, which formulates the problem as multiple hypothesis testing over a candidate set of hyperparameters. The framework enables the selection of hyperparameters that provably satisfy application-specific reliability requirements -- such as bounds on average risk, quantile risk, or information-theoretic constraints -- with explicit, finite-sample control of error probabilities. The supporting statistical machinery, namely p-values, e-values, and concentration inequalities, is developed from first principles in a dedicated appendix.