LLM Price Index — July 2026
Gemini 2.5 Flash price cut and Llama 4 launch drive sub-$0.30 capable inference to mainstream availability
July 2026 saw two major events reshape the LLM pricing landscape. Google cut Gemini 2.5 Flash to $0.30/M input (from $0.40), making it the cheapest capable frontier model at launch. Simultaneously, Meta released Llama 4 Scout and Maverick — open-weight models available free to self-host, and at $0.18-0.27/M through inference providers. The combined effect: the cost floor for high-quality inference dropped roughly 25% month-over-month.
Biggest movers this month
Cheapest model per capability tier
| Tier | Cheapest model | Price |
|---|---|---|
| Budget (<$0.20/M input) | Llama 4 Scout | $0.18/M |
| Mid-range ($0.20–$1/M input) | Gemini 2.5 Flash | $0.30/M |
| Premium (>$1/M input) | Claude Sonnet 5 | $3.00/M |
Analysis
The Llama 4 launch fundamentally changed the open-weight market. Maverick benchmarks near GPT-4o quality at $0.27/M — a 9× cost advantage. For inference providers, this creates significant pricing pressure on everything in the $1-3/M range. Models like Claude Haiku and GPT-4o mini now compete with open-weight alternatives at a fraction of their price.
Frequently Asked Questions
How did Llama 4 affect API pricing?
Directly: inference providers can now offer near-GPT-4 quality at $0.18-0.27/M, creating a new price floor. Indirectly: this pressures proprietary model providers to cut prices or differentiate on quality/ecosystem. We expect Claude Haiku and GPT-4o mini to face price pressure in coming months.
WiseCrewAI LLM Price Index publishes monthly. See methodology · sources.