Model ComparisonGoogle

Gemini 2.5 Flash

Gemini 2.5 Flash — Google's fastest model with 1M context at a fraction of Pro pricing.

VisionTool usePrompt cachingReasoningFast
Pricing (verified 2026-08-19)
TypePrice / 1M tokensNotes
Input$0.30Per 1M input tokens
Output$2.50Per 1M output tokens
Cache read$0.075Cached prefix read-back
Batch API$0.15 in / $1.25 out~50% off, async (up to 24h)
Source: Google official pricing page · Last verified 2026-08-19
Context Window
1M tokens
750K words · 3 pages of text

At $0.075/M input and $0.30/M output, Flash 2.5 is Google's cost-efficiency leader — near-Pro quality for most tasks.

Best for

  • High-volume inference with long context
  • Multimodal classification at scale
  • Real-time applications requiring low latency

Not ideal for

  • Tasks requiring maximum reasoning depth
  • Complex multi-step agent loops

Real cost scenarios

ScenarioEst. monthly costBreakdown
100M tokens/month$840$7,500 input + $22,500 output... actually at these prices, only $840
1M queries/day at 1k tokens$2,250$2,250 total — one of cheapest at scale

Estimates with typical token distributions. Use the Prompt Cost Calculator for your exact workload.

How to cut costs on Gemini 2.5 Flash

  1. At $0.075/M, Flash 2.5 is already extremely cost-efficient. Use batch API for additional savings on async workloads.

Calculate your Gemini 2.5 Flash costs

Prompt Cost Calculator
Exact cost for your prompt
Agent Cost Calculator
Multi-step agent loops
Prompt Caching Calculator
Caching savings
Conversation Cost
Chat history strategies

Other Google models

Gemini 3 Flash$0.50/MGemini 2.5 Pro$1.25/MGemini 2.5 Flash-Lite$0.10/M

Compare Gemini 2.5 Flash

gemini 3 flash vs gpt 5 minigemini 3 flash vs claude haiku 4 5gemini 3 flash vs gemini 3 pro

Frequently Asked Questions

How much cheaper is Flash vs Pro?

Gemini 2.5 Flash is approximately 15× cheaper than Gemini 2.5 Pro on input tokens ($0.075 vs $1.25) and 33× cheaper on output ($0.30 vs $10). For most non-frontier tasks, Flash is the right starting point.

Does Flash support the full 1M context?

Yes — Flash 2.5 supports 1M token context, same as Pro (Pro has an extended 2M option). For most real-world documents this is more than sufficient.

Pricing verified 2026-08-19 from Google official pricing. All prices in USD per 1M tokens. See methodology and price changelog.