methodology

Methodology

Pack structure, September 8 model roster refresh, scoring fields, and track definitions.

Updated

Pack

    Tracks

    • Transaction categorization
    • Vendor normalization
    • Client context fit
    • COA mapping
    • Split detection
    • Tax sensitivity

    Input

    Each participant receives the same statement pack, target chart of accounts, and normalized output schema.

    Output

    Outputs are normalized before scoring and merged into one overall table plus six track tables for the active release.

    Run protocol

    Passes3 seeded · median reported
    Batch size32 rows per request
    Samplingtemperature 0 where accepted
    Output cap8,192 tokens per request
    Retries2 on schema failure
    ContextCOA + client profile + 90-day vendor history

    Normalization

    Category labels pass through a 231-entry alias table before exact match. Merchant strings are lowercased, stripped of store numbers and card-processor suffixes, then matched to 20 gold entities. Account codes are matched as-is against the 48-account target chart.

    Score columns

    Overallweighted composite of six tracks
    Track score0–100, median of passes
    Reliability% schema-valid on first attempt
    Latencymedian s per 32-row request
    Cost idxlist USD per 1,000 input rows

    Track weights are frozen per pack version and ship with the harness notes.

    Version history

    v1.0 · Apr 14, 2026512-row pack, 4 tracks
    v1.1 · May 12, 2026+ split detection, + tax sensitivity
    v1.2 · Jun 29, 20263-pass protocol, reliability/latency/cost
    harness 1.2.3 · Sep 8, 2026roster refresh, alias table 214 → 231

    The Sep 8 alias fix lifted the nine re-run rows by +0.3 to +1.1; the pack and gold set are unchanged.