APEX-Accounting

📅 2026-07-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study evaluates the practical capabilities of state-of-the-art large language models on real-world accounting tasks, including reconciliation, accruals, journal entry preparation, and financial statement generation. To this end, we introduce the first high-quality, closed-domain benchmark tailored to accounting practice, comprising 10 synthetic enterprises and 160 expert-designed and scored tasks that support multimodal financial documents such as PDFs and spreadsheets, all evaluated under a fixed token budget. Our analysis uncovers a Simpson’s paradox between model performance and token consumption. While Claude-Fable-5 (Max) achieves the highest Mean Criteria@3 score at 56.4%, no model exceeds 21.5% on Pass@8, revealing substantial limitations in current models’ ability to perform complex accounting reasoning.
📝 Abstract
We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models can do the real work of accountants. Tasks include reconciling accounts, accruing expenses, posting transactions, and producing reports. The private eval set comprises 160 tasks, split across 10 worlds. Each world contains an accounting system, as well as spreadsheets, PDFs, and other files. Every task was authored and solved by experts in accounting and bookkeeping, who also wrote grading rubrics. Across nine frontier models, Claude-Fable-5 (Max) leads with 56.4% Mean Criteria@3, ahead of Muse-Spark-1.1 (xHigh) at 52.6%. No model scores more than 2.6% Pass^8 (GPT-5.6-Sol (Max+Pro)) and the highest Pass@8 is 21.5% (Muse-Spark-1.1 (xHigh)). We experiment with increasing the token budget from $1 to $50 and observe an instance of Simpson's paradox: scores increase as the token budget increases but within a given budget-constrained harness, scores are lower on tasks where the model spends more tokens. As APEX-Accounting is a closed benchmark, leaderboard evals can be run for any frontier model on request.
Problem

Research questions and friction points this paper is trying to address.

accounting
benchmark
large language models
financial tasks
AI evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

accounting benchmark
frontier language models
real-world task evaluation
token budget analysis
expert-authored tasks
🔎 Similar Papers
No similar papers found.