ArXiv

AV-AIVAT: 74x Cheaper Agent Evaluation with Certified Anytime-Valid Stopping in Imperfect-Information Games

Authors
Boning Li, Yu Chen, Longbo Huang
Categories
cs.GT, cs.AI, cs.CL, cs.LG, cs.MA
arXiv
https://arxiv.org/abs/2608.06362v1
PDF
https://arxiv.org/pdf/2608.06362v1

Brief

AV-AIVAT combines AIVAT’s conditional mean-zero variance corrections with continuously monitored Confidence Sequences to stop agent-vs-agent evaluations the moment evidence suffices while preserving stated coverage. On 71,439 paired HUNL hands AIVAT gave a median 54× variance reduction; under the Asymptotic CS this translated to a median 74× reduction in required hands (95%, ±1 BB). EB-CS yields exact finite-sample certification when corrected-payoff bounds hold (authors derive such bounds for Leduc), enabling verifiable, early stopping.

Why it matters

At the nominal 95% level with target precision ±1 Big Blind, AV-AIVAT (AIVAT + continuously monitored Confidence Sequences) makes raw outcomes need a median 74× as many HUNL hands as AIVAT-corrected outcomes to stop under the Asymptotic CS; AIVAT alone gave a median 54× variance reduction across 15 LLM agent configurations over 71,439 paired HUNL hands.

Key details

  • Exact finite-sample certification uses an Empirical-Bernstein Confidence Sequence (EB-CS), which requires an independently justified bound on corrected payoffs; the authors structurally establish such a bound for Leduc Hold'em and report descriptive HUNL EB-CS runs with a median 1.37× stopping-time ratio. AV-AIVAT’s online value model learns only from past games so no game scores its own correction, enabling auditable early stopping.
Source evidence

Abstract

Deciding which of two agents is stronger means playing games until skill outweighs luck, and every game costs money, model inference, or expert time. Since the number of games needed is unknown, fixed-budget evaluations either keep paying after the result is settled or stop before the agents can be told apart, while naive optional stopping with an ordinary confidence interval invalidates the stated level. We make such an evaluation stop as soon as its evidence suffices, with the guarantee intact. The Action-Informed Value Assessment Tool (AIVAT) reduces variance in imperfect-information games through conditional mean-zero corrections, by a median $54\times$ across 15 LLM agent configurations spanning 71,439 paired Heads-Up No-Limit Hold'em (HUNL) hands, but does not say when to stop. We combine AIVAT with continuously monitored Confidence Sequences (CSs) into anytime-valid AIVAT (AV-AIVAT), whose online value model learns only from past games so that no game scores its own correction. At the nominal 95\% level and a target precision of $\pm1$ Big Blind, raw outcomes need a median $74\times$ as many hands as AIVAT-corrected outcomes to stop under the Asymptotic CS (AsympCS). Exact finite-sample certification uses the Empirical-Bernstein CS (EB-CS), which needs an independently justified bound on corrected payoffs. We establish such a bound structurally for Leduc hold'em and characterize a width floor set by the CS's bet cap and that bound, which governs how much of a variance gain becomes earlier stopping; the descriptive HUNL EB-CS runs show a median $1.37\times$ stopping-time ratio. AV-AIVAT turns variance reduction into efficient, auditable early stopping while separating asymptotic screening from exact certification, so an evaluation can stop the moment its evidence suffices and hand a third party everything needed to recheck the verdict at that very stopping time.

Comment: 34 pages, 5 figures