Workload · v1.0.0
Pricing Intelligence Research
Given frozen pricing sources for three AI providers and a fixed monthly workload volume, calculate the true monthly cost of every plan, select the cheapest plan, and compute the landed cost per accepted result — with every claim cited to a frozen source.
Workload volume
- Requests / month
- 50,000
- Avg tokens in/out
- 800 / 400
- Human review
- 5 min × $75/hr
- Accepted outcomes
- 10 per run
Frozen sources
Provider Alpha — Pricing (captured 2026-07-30)
SHA-256
196516c063af3603b3a6935312c17dac111597dd12ef7b8d977a91fad13899b6
Provider Beta — Pricing (captured 2026-07-30)
SHA-256
151b508ad9deb5f3d26a6a000a45521704c9173846bbf63fb1cf15490b7e9eb2
Provider Gamma — Pricing (captured 2026-07-30)
SHA-256
d43121acac45ea4798a83c92f364e7eb09e4b98d7e4a8ffef746211815cb30c5
Acceptance criteria
- Minimum score
- 85/100
- Mandatory rules
- 7
- Tolerance
- ± $0.01
- Validation
- deterministic, no LLM judge
Deterministic rules
- json_schema — Output matches output_schema.json
- required_fields — All required fields present
- numeric_extraction — All plan numerics parse as floats
- plan_calculation — Each plan total matches the deterministic formula
- winner_selection — Winner is the cheapest plan
- landed_cost — Landed cost reconciles with the winner total and review cost
- source_mapping — All citations reference frozen sources
- unsupported_claims — Every plan in calculations has supporting evidence
- duplicate_evidence — No duplicate (sourceId, planId, field) citations
- uncertainty_disclosure — Uncertainties field present and well-formed
- total_reconciliation — Winner calculation total equals claimed winner plan total
- threshold_boundary — Score meets the acceptance threshold
- cost_ceiling — Run cost stayed within the configured ceiling
Currently fought by: The Flagship Fight. Future fights can reuse this workload without new code.
See results