changelog

Changelog

Release history for the pack, the harness, and the public roster.

Updated

· harness 1.2.3

  • Roster refresh: 21 rows added (20 new models, DeepSeek V4-Flash re-labeled from its preview name), 21 rows retired, 9 rows re-run on the same pack.
  • New rows include Claude Mythos 5.1, Claude Fable 5.1, Claude Opus 5, Claude Sonnet 5, GPT-6 Astra Preview, the GPT-5.6 Sol Pro / Sol / Terra / Luna tiers, Gemini 3.8 and 3.7 Flash, Grok 4.6 and 4.5, DeepSeek V4-Pro, Qwen3.8-Max and Qwen3.8-Flash, Kimi K3, Muse Spark, and Amazon Nova 2 Pro.
  • Category alias table extended from 214 to 231 entries; re-run rows moved +0.3 to +1.1. Pack and gold set unchanged.
  • Runs executed Sep 2–6, 2026 (UTC). Results archive published as JSON and CSV under /results/.
  • Per-track pages now list scored rows, the match rule, and the most common miss.

· pack v1.2

  • Three-pass protocol adopted: three seeded passes per participant, median track score reported.
  • Reliability, latency, and cost index columns added to the overall table.
  • Roster refresh: OpenAI GPT-5.6 Sol Preview and GPT-5.5 tiers, Anthropic Claude Mythos 5, Fable 5, Opus 4.8; Google, xAI, DeepSeek, Qwen, Mistral, Cohere, Meta, Amazon, Perplexity, and Moonshot rows refreshed.
  • Track pages split out into dedicated URLs.

· pack v1.1

  • Split detection and tax sensitivity tracks added (64 mixed-use rows, 96 tax-sensitive rows).
  • Gold set re-adjudicated by two reviewers; 31 rows changed label.
  • Client profile and 90-day vendor history attached to every run.

· pack v1.0

  • Initial public release: 512-row statement pack (Jan 31 – Apr 29, 2026), one target chart of accounts, four tracks.
  • Public leaderboard, methodology notes, and statement sample published.

Versioning: the pack version changes when rows, gold labels, or tracks change. The harness version changes when scoring, normalization, or the run protocol changes. Roster refreshes reuse the current pack and harness.