LLM Value Rankings — Cost Per Benchmark Point
Which model gives the most intelligence per dollar? Rank LLMs by benchmark score relative to their cost for your specific input/output ratio.
input × 70% + output × 30%benchmark_score / blended_price_per_MBenchmark scores are approximate community averages. Pricing verified 2026-08-20. Always run task-specific evals before committing to a model at scale.
How to Use
Select a benchmark
Choose MMLU (general reasoning), HumanEval (code), MATH, or the blended average across all three.
Set your token ratio
Adjust the input vs output ratio to match your workload. Long-form generation is output-heavy; RAG and classification are input-heavy.
Find the value leader
The top-ranked model has the highest score per dollar at your token ratio. Green bars indicate relative value efficiency.
Validate on your task
These are benchmark averages. Run your task-specific evals on the top 2–3 models before committing.
Why frontier models often lose on value
Frontier models (GPT-5.4, Claude Opus 5) score 5–15% higher on benchmarks but cost 10–60× more than mid-tier models. For most production use cases, a model scoring 82 MMLU at $0.25/M outperforms one scoring 90 MMLU at $15/M on cost-adjusted value — unless that extra 8 points meaningfully changes task success rate. Quantify your task's sensitivity to benchmark score before paying frontier prices.
FAQ
Have more questions? Contact us
Last verified: 2026-08-20 · methodology · data sources