LLM Value Rankings — Cost Per Benchmark Point

Which model gives the most intelligence per dollar? Rank LLMs by benchmark score relative to their cost for your specific input/output ratio.

Value Rankings
Settings
Blended price = input × 70% + output × 30%
Value score = benchmark_score / blended_price_per_M
Higher value = more quality per dollar
1
Gemini 2.5 Flash (Google)Score: 77.7 · $0.142/M
545
pts/$M
2
Mistral Small 3.2 (Mistral)Score: 69.0 · $0.160/M
431
pts/$M
3
Gemini 2.0 Flash (Google)Score: 74.3 · $0.190/M
391
pts/$M
4
GPT-5 mini (OpenAI)Score: 76.0 · $0.475/M
160
pts/$M
5
Claude Haiku 4.5 (Anthropic)Score: 73.0 · $1.76/M
41
pts/$M
6
Mistral Large 2411 (Mistral)Score: 76.3 · $3.20/M
24
pts/$M
7
Gemini 2.5 Pro (Google)Score: 87.0 · $3.88/M
22
pts/$M
8
GPT-5.4 (OpenAI)Score: 89.5 · $4.75/M
19
pts/$M
9
Claude Sonnet 5 (Anthropic)Score: 84.3 · $6.60/M
13
pts/$M
10
Claude Opus 5 (Anthropic)Score: 88.2 · $33.00/M
3
pts/$M

Benchmark scores are approximate community averages. Pricing verified 2026-08-20. Always run task-specific evals before committing to a model at scale.

How to Use

Select a benchmark

Choose MMLU (general reasoning), HumanEval (code), MATH, or the blended average across all three.

Set your token ratio

Adjust the input vs output ratio to match your workload. Long-form generation is output-heavy; RAG and classification are input-heavy.

Find the value leader

The top-ranked model has the highest score per dollar at your token ratio. Green bars indicate relative value efficiency.

Validate on your task

These are benchmark averages. Run your task-specific evals on the top 2–3 models before committing.

Why frontier models often lose on value

Frontier models (GPT-5.4, Claude Opus 5) score 5–15% higher on benchmarks but cost 10–60× more than mid-tier models. For most production use cases, a model scoring 82 MMLU at $0.25/M outperforms one scoring 90 MMLU at $15/M on cost-adjusted value — unless that extra 8 points meaningfully changes task success rate. Quantify your task's sensitivity to benchmark score before paying frontier prices.

FAQ

Have more questions? Contact us

Last verified: 2026-08-20 · methodology · data sources