Leaderboard for accounting task performance.
September 8 roster across current models and accounting agents.
Overall Ranking
Sort by score, reliability, latency, or cost index.
| Rank | Participant | Overall | Categorization | Context | Reliability | Latency | Cost idx |
|---|
Scores are medians of three seeded passes. Reliability is the share of requests that returned schema-valid output on the first attempt. Latency is median seconds per 32-row request. Cost index is list-price USD per 1,000 input rows at Sep 8 pricing, retries included.
September 8 Run
What was run, how, and what moved since June 29.
Protocol
Three seeded passes per participant, 32 statement rows per request, median track score reported. Temperature 0 where the API accepts sampling parameters; vendor default otherwise (Claude 4.6+ family, GPT-6 Astra). Up to two retries on schema-invalid output, each counted against reliability.
Access
Hosted models were run against public APIs at list price. Claude Mythos 5.1 was scored through a partner allocation and GPT-6 Astra through the limited-availability preview rollout; both are marked as preview or partner access in the roster notes. Open-weight rows (Kimi K3, Qwen3.8, DeepSeek V4) used the vendor-hosted API, not self-hosting.