Gemini 2.5 Flash
Gemini 2.5 Flash — Google's fastest model with 1M context at a fraction of Pro pricing.
| Type | Price / 1M tokens | Notes |
|---|---|---|
| Input | $0.30 | Per 1M input tokens |
| Output | $2.50 | Per 1M output tokens |
| Cache read | $0.075 | Cached prefix read-back |
| Batch API | $0.15 in / $1.25 out | ~50% off, async (up to 24h) |
At $0.075/M input and $0.30/M output, Flash 2.5 is Google's cost-efficiency leader — near-Pro quality for most tasks.
Best for
- High-volume inference with long context
- Multimodal classification at scale
- Real-time applications requiring low latency
Not ideal for
- Tasks requiring maximum reasoning depth
- Complex multi-step agent loops
Real cost scenarios
| Scenario | Est. monthly cost | Breakdown |
|---|---|---|
| 100M tokens/month | $840 | $7,500 input + $22,500 output... actually at these prices, only $840 |
| 1M queries/day at 1k tokens | $2,250 | $2,250 total — one of cheapest at scale |
Estimates with typical token distributions. Use the Prompt Cost Calculator for your exact workload.
How to cut costs on Gemini 2.5 Flash
- At $0.075/M, Flash 2.5 is already extremely cost-efficient. Use batch API for additional savings on async workloads.
Calculate your Gemini 2.5 Flash costs
Other Google models
Compare Gemini 2.5 Flash
Frequently Asked Questions
How much cheaper is Flash vs Pro?
Gemini 2.5 Flash is approximately 15× cheaper than Gemini 2.5 Pro on input tokens ($0.075 vs $1.25) and 33× cheaper on output ($0.30 vs $10). For most non-frontier tasks, Flash is the right starting point.
Does Flash support the full 1M context?
Yes — Flash 2.5 supports 1M token context, same as Pro (Pro has an extended 2M option). For most real-world documents this is more than sufficient.
Pricing verified 2026-08-19 from Google official pricing. All prices in USD per 1M tokens. See methodology and price changelog.