PayingForAI Test Lab · By OverpayingForAI
Every flagship.
Same fight.
Real cost.
We test leading AI models on the same real workload and expose every result, failure, repair, and dollar spent.
Now holding · executed
The Flagship Fight
Every flagship. Same fight. Real cost.. One current flagship from every eligible major provider. The same workload, evidence, rules, and acceptance threshold.
10
Contestants
3
Runs each
30
Runs executed
The tale of the tape
Full results →Cheapest accepted result
xAI: Grok 4.5
6.2592 USD/accepted result
Most reliable
Google: Gemini 3.1 Pro (Preview)
1 eventual acceptance rate
Fastest accepted
xAI: Grok 4.5
12752 median runtime ms
Best first attempt
Google: Gemini 3.1 Pro (Preview)
1 first-attempt acceptance rate
Primary metric
Landed cost per
accepted result
API price lists tell you what a token costs. They don't tell you what a finished, verified answer costs — retries, repairs, failures, and human review included. We measure that.
Cheapest accepted result
$6.26
xAI: Grok 4.5
Largest overpayment gap
+$0.12
Google: Gemini 3.1 Pro (Preview) pays this much more per accepted result than the cheapest accepted contestant.
No cherry picking
Failures are
evidence too
Every failed run is preserved, hashed, and published. A model that fails three times in three tells you more than a leaderboard that hides it.
7 contestants · 0 accepted runs
OpenAI: GPT-5.1
openai/gpt-5.1
unreconciled_arithmetic, unreconciled_arithmetic, unreconciled_arithmetic, unreconciled_arithmetic, incorrect_winning_option, unreconciled_arithmetic, unreconciled_arithmetic
Anthropic: Claude Opus 4.5
anthropic/claude-opus-4.5
incorrect_winning_option, unreconciled_arithmetic
DeepSeek: DeepSeek V3.2
deepseek/deepseek-v3.2
incorrect_winning_option, unreconciled_arithmetic
Alibaba: Qwen3 Max
qwen/qwen3-max
unreconciled_arithmetic, unreconciled_arithmetic, unreconciled_arithmetic, unreconciled_arithmetic, unreconciled_arithmetic, incorrect_winning_option, unreconciled_arithmetic, unreconciled_arithmetic
Mistral: Mistral Large 2512
mistralai/mistral-large-2512
unreconciled_arithmetic, unreconciled_arithmetic, unreconciled_arithmetic, unreconciled_arithmetic, incorrect_winning_option, unreconciled_arithmetic, unreconciled_arithmetic
Meta: Llama 4 Maverick
meta-llama/llama-4-maverick
unreconciled_arithmetic, unreconciled_arithmetic, unreconciled_arithmetic, unreconciled_arithmetic, incorrect_winning_option, unreconciled_arithmetic, unreconciled_arithmetic
Cohere: Command A
cohere/command-a
malformed_required_output
How the fight works
-
01
Discover
Query the OpenRouter catalogue. One eligible flagship per provider — no minis, no aliases, no specialists.
-
02
Freeze
Exact canonical model IDs, pricing, and metadata are frozen into a roster that requires explicit human approval.
-
03
Fight
3 identical runs per model. Same prompt, same frozen evidence, same schema, temperature 0. No browsing, no tools.
-
04
Judge
Deterministic pure-Python validators check every number. No LLM judges arithmetic. Mandatory tripwires fail the run.
-
05
Publish
Medians, not best runs. Every failure preserved. Every artifact hashed.
Methodology
Judged by arithmetic,
not vibes
- Deterministic validators recompute every claim from frozen sources.
- Landed cost includes API spend, repairs, retries, and human review.
- Every cost is labelled by basis: provider-reported, calculated, estimated.
- Fixture data is labelled fixture. It is never benchmark evidence.
Next in the ring
Upcoming
The Efficient-Model Fight
The same workload, fought by the cheap seats. When is a small model enough?
Upcoming
Single Agent vs Multi-Agent
Does orchestration pay for itself on real work?
Upcoming
The Coding-Agent Fight
Cost per accepted pull request — not per token.
Your call
What should we
test next?
Name the fight. If it measures the cost of finished work, it belongs in this lab.
PayingForAI produces the evidence.
OverpayingForAI helps you decide whether the result is worth paying for.
By OverpayingForAI · payingforai.com