leaderboard

Leaderboard for accounting task performance.

September 8 roster across current models and accounting agents.

Updated

table

Overall Ranking

Sort by score, reliability, latency, or cost index.

Overall ranking · pack v1.2 · harness 1.2.3 · runs Sep 2–6, 2026 · medians of three passes
Rank Participant Overall Categorization Context Reliability Latency Cost idx

Scores are medians of three seeded passes. Reliability is the share of requests that returned schema-valid output on the first attempt. Latency is median seconds per 32-row request. Cost index is list-price USD per 1,000 input rows at Sep 8 pricing, retries included.

run details

September 8 Run

What was run, how, and what moved since June 29.

Run window

Runs executedSep 2–6, 2026 (UTC)
Packv1.2 (512 rows, 490 scored)
Harnesseval.tax harness 1.2.3
Participants30 · 13 orgs · 1 agent
Requests16 per pass · 48 per participant

Protocol

Three seeded passes per participant, 32 statement rows per request, median track score reported. Temperature 0 where the API accepts sampling parameters; vendor default otherwise (Claude 4.6+ family, GPT-6 Astra). Up to two retries on schema-invalid output, each counted against reliability.

Access

Hosted models were run against public APIs at list price. Claude Mythos 5.1 was scored through a partner allocation and GPT-6 Astra through the limited-availability preview rollout; both are marked as preview or partner access in the roster notes. Open-weight rows (Kimi K3, Qwen3.8, DeepSeek V4) used the vendor-hosted API, not self-hosting.

Since June 29

Rows added21 (20 new, 1 re-labeled)
Rows retired21
Rows re-run9, on the same pack
Top model overall93.3 → 94.4
Models at 90.0+8 → 16
Largest pass spread0.6 (GPT-5.6 Luna)
tracks

Track Tables