Gemini API Pricing
Google lists Gemini API prices on its official pricing page, dated 2026-10-01 when we checked it. Several models have a free tier on the Developer API.
| Model | Input | Cached input | Output |
|---|---|---|---|
| Gemini 3.8 Flash | $0.750 → $1.50 from 2027-01-01 | $0.075 → $0.150 from 2027-01-01 | $3.75 → $7.50 from 2027-01-01 |
| Gemini 3.7 Flash | $0.750 → $1.50 from 2027-01-01 | $0.075 → $0.150 from 2027-01-01 | $3.75 → $7.50 from 2027-01-01 |
| Gemini 3.6 Flash | $0.750 → $1.50 from 2027-01-01 | $0.075 → $0.150 from 2027-01-01 | $3.75 → $7.50 from 2027-01-01 |
| Gemini 3.5 Flash | $1.50 | $0.150 | $9.00 |
| Gemini 3.1 Pro (Preview) | $2.00 ($4.00 above 200k) | $0.200 ($0.400 above 200k) | $12.00 ($18.00 above 200k) |
| Gemini 2.5 Pro | $1.25 ($2.50 above 200k) | $0.125 ($0.250 above 200k) | $10.00 ($15.00 above 200k) |
| Gemini 2.5 Flash Input price shown is for text, image and video; audio input costs more. | $0.300 | $0.030 | $2.50 |
| Gemini 2.5 Flash-Lite Input price shown is for text, image and video; audio input costs more. | $0.100 | $0.010 | $0.400 |
What to watch
- Scheduled price change: Gemini 3.6, 3.7 and 3.8 Flash cost $0.75 input and $3.75 output per million tokens through 2026-12-31, then $1.50 and $7.50 from 2027-01-01. Budget for the later price if your project will run past that date.
- Long prompts: Pro models charge more for prompts above 200k tokens. The table shows both prices.
- Audio input on 2.5 Flash and Flash-Lite costs more than text, image or video input.
Example monthly costs
| Model | Chatbot | Document Q&A (RAG) | Bulk summarization |
|---|---|---|---|
| Gemini 3.8 Flash | $188 | $259 | $675 |
| Gemini 3.7 Flash | $188 | $259 | $675 |
| Gemini 3.6 Flash | $188 | $259 | $675 |
| Gemini 3.5 Flash | $420 | $555 | $1,440 |
| Gemini 3.1 Pro (Preview) | $560 | $740 | $1,920 |
| Gemini 2.5 Pro | $425 | $525 | $1,350 |
| Gemini 2.5 Flash | $105 | $129 | $330 |
| Gemini 2.5 Flash-Lite | $22.00 | $32.00 | $84.00 |
- Chatbot: 100,000 requests, 1,000 input and 300 output tokens each.
- Document Q&A (RAG): 50,000 requests, 8,000 input (half from cache) and 500 output tokens.
- Bulk summarization: 200,000 documents, 3,000 input and 300 output tokens, Batch where documented.
Costs use the prices valid on the day we checked. Compare with Claude and OpenAI.