WiseCrewAI → Reports

LLM Price Index — July 2026

Published 2026-07-31 · WiseCrewAI LLM Price Index
Key finding

Gemini 2.5 Flash price cut and Llama 4 launch drive sub-$0.30 capable inference to mainstream availability

July 2026 saw two major events reshape the LLM pricing landscape. Google cut Gemini 2.5 Flash to $0.30/M input (from $0.40), making it the cheapest capable frontier model at launch. Simultaneously, Meta released Llama 4 Scout and Maverick — open-weight models available free to self-host, and at $0.18-0.27/M through inference providers. The combined effect: the cost floor for high-quality inference dropped roughly 25% month-over-month.

Biggest movers this month

Gemini 2.5 FlashPrice cut to $0.30/M input (-25%)
Llama 4 ScoutNEW: $0.18/M input via Together AI
Llama 4 MaverickNEW: $0.27/M input via Together AI
DeepSeek R1 DistillWider provider availability ($0.15-0.88/M)

Cheapest model per capability tier

TierCheapest modelPrice
Budget (<$0.20/M input)Llama 4 Scout$0.18/M
Mid-range ($0.20–$1/M input)Gemini 2.5 Flash$0.30/M
Premium (>$1/M input)Claude Sonnet 5$3.00/M

Analysis

The Llama 4 launch fundamentally changed the open-weight market. Maverick benchmarks near GPT-4o quality at $0.27/M — a 9× cost advantage. For inference providers, this creates significant pricing pressure on everything in the $1-3/M range. Models like Claude Haiku and GPT-4o mini now compete with open-weight alternatives at a fraction of their price.

Calculate your costs with current pricing
Prompt cost calculator →Model comparisonFull changelog

Frequently Asked Questions

How did Llama 4 affect API pricing?

Directly: inference providers can now offer near-GPT-4 quality at $0.18-0.27/M, creating a new price floor. Indirectly: this pressures proprietary model providers to cut prices or differentiate on quality/ecosystem. We expect Claude Haiku and GPT-4o mini to face price pressure in coming months.

WiseCrewAI LLM Price Index publishes monthly. See methodology · sources.