ArXiv

APEX-Accounting

Authors
Julien Benchek, Austin Bennett, Jasmin Kern...
Categories
cs.CL, cs.AI, cs.HC
arXiv
https://arxiv.org/abs/2607.27189v1
PDF
https://arxiv.org/pdf/2607.27189v1

Brief

APEX‑Accounting is a professional accounting benchmark (160 private tasks, 10 realistic worlds) assessing reconciliation, accruals, posting, and reporting using expert‑authored problems and rubrics. Evaluation across nine frontier LLMs finds Claude‑Fable‑5 (Max) highest on Mean Criteria@3 (56.4%) and mixed pass rates (highest Pass@8 21.5%). A token‑budget sweep ($1→$50) increases aggregate performance but shows a within‑harness Simpson’s paradox. The benchmark is closed; a public dev set is provided.

Why it matters

APEX-Accounting is a closed benchmark of 160 private evaluation tasks across 10 'worlds' (spreadsheets, PDFs, accounting systems); every task was authored and solved by professional accountants and includes expert grading rubrics (public dev set available on Hugging Face).

Key details

  • Across nine frontier models, Claude‑Fable‑5 (Max) leads with 56.4% Mean Criteria@3 and Muse‑Spark‑1.1 (xHigh) scores 52.6%; no model exceeds 2.6% on Pass^8 (GPT‑5.6‑Sol (Max+Pro) noted) while the highest reported Pass@8 is 21.5% (Muse‑Spark‑1.1 (xHigh)).
  • A token‑budget experiment (from $1 to $50) raised overall scores as budget increased but revealed Simpson’s paradox: within any fixed budget harness, tasks where the model spent more tokens had lower scores.
Source evidence

Abstract

We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models can do the real work of accountants. Tasks include reconciling accounts, accruing expenses, posting transactions, and producing reports. The private eval set comprises 160 tasks, split across 10 worlds. Each world contains an accounting system, as well as spreadsheets, PDFs, and other files. Every task was authored and solved by experts in accounting and bookkeeping, who also wrote grading rubrics. Across nine frontier models, Claude-Fable-5 (Max) leads with 56.4% Mean Criteria@3, ahead of Muse-Spark-1.1 (xHigh) at 52.6%. No model scores more than 2.6% Pass^8 (GPT-5.6-Sol (Max+Pro)) and the highest Pass@8 is 21.5% (Muse-Spark-1.1 (xHigh)). We experiment with increasing the token budget from $1 to $50 and observe an instance of Simpson's paradox: scores increase as the token budget increases but within a given budget-constrained harness, scores are lower on tasks where the model spends more tokens. As APEX-Accounting is a closed benchmark, leaderboard evals can be run for any frontier model on request.

Comment: Public dev set: https://huggingface.co/datasets/mercor/apex-accounting