Skip to content
PayingForAI Test Lab By OverpayingForAI
MENU

Fight card · pricing-intelligence-research-001 v1.0.0

The Flagship Fight

Every flagship. Same fight. Real cost.. One current flagship from every eligible major provider. The same workload, evidence, rules, and acceptance threshold.

Vintage officials standing beneath the fight card like judges

10

Providers

10

Models

30/30

Runs executed

6

Excluded

Status

executed

Tested: 2026-07-30
Evidence: manifests + SHA-256 hashes per run
Primary metric: landed cost per accepted result
Current leader: xAI: Grok 4.5 — $6.26

The contestants

GPT-5.1

OpenAI · openai/gpt-5.1

Claude Opus 4.5

Anthropic · anthropic/claude-opus-4.5

Gemini 3.1 Pro (Preview)

Google · google/gemini-3.1-pro-preview

Grok 4.5

xAI · x-ai/grok-4.5

DeepSeek V3.2

DeepSeek · deepseek/deepseek-v3.2

Qwen3 Max

Alibaba Qwen · qwen/qwen3-max

Mistral Large 2512

Mistral AI · mistralai/mistral-large-2512

Llama 4 Maverick

Meta · meta-llama/llama-4-maverick

Kimi K2.5

Moonshot AI · moonshotai/kimi-k2.5

Command A

Cohere · cohere/command-a