ArXiv

Hassoon & Dredze (published 2026-08-06) theoretically analyze innovation-residual…

Authors
Ahmed Hassoon, Mark Dredze
Categories
cs.AI, cs.LG, stat.ML
arXiv
https://arxiv.org/abs/2608.05490v1
PDF
https://arxiv.org/pdf/2608.05490v1

Brief

Auditing autonomous data-analysis agents, Hassoon and Dredze analyze innovation-residual methods that flag operations deviating from a model trained on sound analyses. They show scoring choice governs localization—single-step surprise misses propagated errors while longer reconstructions diffuse flags—derive false-flag control procedures under exchangeability, quantify robustness losses under misspecification/selection, and prove a slow-decaying detectability limit tied to representation dimension.

Why it matters

Hassoon & Dredze (published 2026-08-06) theoretically analyze innovation-residual audits and show that scoring an operation by surprise relative to only its immediate predecessor makes inherited errors indistinguishable from correct steps—one upstream mistake typically produces a single detectable flag.

Key details

  • Scores computed against longer reconstructions spread a single mistake across many operations; the paper quantifies that spread and gives guidance on choosing the reconstruction/comparison length when errors accumulate gradually.
  • The authors provide procedures that control the proportion of falsely flagged operations assuming sound analyses are exchangeable (not requiring the fitted model to be correct), quantify how guarantees degrade under model misspecification or content-dependent selection, and prove a detection limit: errors below a certain magnitude are un-attributable, and a 100× increase in sound analyses reduces that limit by under 2%, making representation dimension the bottleneck.
Source evidence

Abstract

Autonomous agents now carry out entire data analyses, selecting cohorts, joining tables, and fitting models with little step-by-step supervision. When such an analysis turns out to be wrong, someone must determine which operation caused it. A recent approach does this without any labelled mistakes, learning instead from analyses known to be sound and flagging operations that depart from what that model predicts; how reliable such audits are has not been studied. This paper supplies that analysis. The choice of score determines whether an error can be localized at all. If each operation is scored by how surprising it is given the operation immediately preceding it, then operations that merely inherit an earlier error are indistinguishable from correct ones, so one mistake produces one flag; scores computed against a longer reconstruction of the intended analysis instead spread a single mistake across many operations. We quantify how far they spread, and how to choose the comparison length when an error accumulates gradually rather than at once. We then give procedures that control the proportion of falsely flagged operations within a single audited analysis, requiring only that sound analyses be exchangeable rather than that the fitted model be correct, and we quantify how much the guarantees weaken when the model is imperfect or when the analysis was selected for review in a way that depends on its content. Finally we establish a limit on what any such audit can report: errors below a certain magnitude cannot be attributed at all, being indistinguishable from ordinary variation among sound analyses. This limit falls so slowly as more sound analyses are collected that at the representation sizes now in use a hundredfold increase reduces it by under two percent, so the dimension of the representation rather than the volume of training data is the binding constraint.