Skip to content
PayingForAI Test Lab By OverpayingForAI
MENU

PayingForAI Test Lab · By OverpayingForAI

Every flagship.
Same fight.
Real cost.

We test leading AI models on the same real workload and expose every result, failure, repair, and dollar spent.

Suited figures walking toward a monumental test arena
The contenders arrive. One workload. No excuses.

Now holding · executed

The Flagship Fight

Every flagship. Same fight. Real cost.. One current flagship from every eligible major provider. The same workload, evidence, rules, and acceptance threshold.

10

Contestants

3

Runs each

30

Runs executed

OpenAIGPT-5.1
AnthropicClaude Opus 4.5
GoogleGemini 3.1 Pro (Preview)
xAIGrok 4.5
DeepSeekDeepSeek V3.2
Alibaba QwenQwen3 Max
Mistral AIMistral Large 2512
MetaLlama 4 Maverick
Moonshot AIKimi K2.5
CohereCommand A
OpenAIGPT-5.1
AnthropicClaude Opus 4.5
GoogleGemini 3.1 Pro (Preview)
xAIGrok 4.5
DeepSeekDeepSeek V3.2
Alibaba QwenQwen3 Max
Mistral AIMistral Large 2512
MetaLlama 4 Maverick
Moonshot AIKimi K2.5
CohereCommand A

The tale of the tape

Full results →

Cheapest accepted result

xAI: Grok 4.5

6.2592 USD/accepted result

Most reliable

Google: Gemini 3.1 Pro (Preview)

1 eventual acceptance rate

Fastest accepted

xAI: Grok 4.5

12752 median runtime ms

Best first attempt

Google: Gemini 3.1 Pro (Preview)

1 first-attempt acceptance rate

Primary metric

Landed cost per
accepted result

API price lists tell you what a token costs. They don't tell you what a finished, verified answer costs — retries, repairs, failures, and human review included. We measure that.

Cheapest accepted result

$6.26

xAI: Grok 4.5

Largest overpayment gap

+$0.12

Google: Gemini 3.1 Pro (Preview) pays this much more per accepted result than the cheapest accepted contestant.

No cherry picking

Failures are
evidence too

Every failed run is preserved, hashed, and published. A model that fails three times in three tells you more than a leaderboard that hides it.

An official lowering a scorecard

7 contestants · 0 accepted runs

OpenAI: GPT-5.1

openai/gpt-5.1

0/3 accepted

unreconciled_arithmetic, unreconciled_arithmetic, unreconciled_arithmetic, unreconciled_arithmetic, incorrect_winning_option, unreconciled_arithmetic, unreconciled_arithmetic

Anthropic: Claude Opus 4.5

anthropic/claude-opus-4.5

0/3 accepted

incorrect_winning_option, unreconciled_arithmetic

DeepSeek: DeepSeek V3.2

deepseek/deepseek-v3.2

0/3 accepted

incorrect_winning_option, unreconciled_arithmetic

Alibaba: Qwen3 Max

qwen/qwen3-max

0/3 accepted

unreconciled_arithmetic, unreconciled_arithmetic, unreconciled_arithmetic, unreconciled_arithmetic, unreconciled_arithmetic, incorrect_winning_option, unreconciled_arithmetic, unreconciled_arithmetic

Mistral: Mistral Large 2512

mistralai/mistral-large-2512

0/3 accepted

unreconciled_arithmetic, unreconciled_arithmetic, unreconciled_arithmetic, unreconciled_arithmetic, incorrect_winning_option, unreconciled_arithmetic, unreconciled_arithmetic

Meta: Llama 4 Maverick

meta-llama/llama-4-maverick

0/3 accepted

unreconciled_arithmetic, unreconciled_arithmetic, unreconciled_arithmetic, unreconciled_arithmetic, incorrect_winning_option, unreconciled_arithmetic, unreconciled_arithmetic

Cohere: Command A

cohere/command-a

0/3 accepted

malformed_required_output

How the fight works

  1. 01

    Discover

    Query the OpenRouter catalogue. One eligible flagship per provider — no minis, no aliases, no specialists.

  2. 02

    Freeze

    Exact canonical model IDs, pricing, and metadata are frozen into a roster that requires explicit human approval.

  3. 03

    Fight

    3 identical runs per model. Same prompt, same frozen evidence, same schema, temperature 0. No browsing, no tools.

  4. 04

    Judge

    Deterministic pure-Python validators check every number. No LLM judges arithmetic. Mandatory tripwires fail the run.

  5. 05

    Publish

    Medians, not best runs. Every failure preserved. Every artifact hashed.

Inspectors examining documents and a calculator

Methodology

Judged by arithmetic,
not vibes

  • Deterministic validators recompute every claim from frozen sources.
  • Landed cost includes API spend, repairs, retries, and human review.
  • Every cost is labelled by basis: provider-reported, calculated, estimated.
  • Fixture data is labelled fixture. It is never benchmark evidence.

Next in the ring

Silhouetted contestants entering

Upcoming

The Efficient-Model Fight

The same workload, fought by the cheap seats. When is a small model enough?

Upcoming

Single Agent vs Multi-Agent

Does orchestration pay for itself on real work?

Upcoming

The Coding-Agent Fight

Cost per accepted pull request — not per token.

Your call

What should we
test next?

Name the fight. If it measures the cost of finished work, it belongs in this lab.

If the form is unavailable, email your suggestion to aniruddh@overpayingforai.com.

A restrained ceremonial acknowledgement

PayingForAI produces the evidence.
OverpayingForAI helps you decide whether the result is worth paying for.

By OverpayingForAI · payingforai.com