LLM API Cost Calculator
Enter your expected traffic. The table ranks every model by estimated monthly cost.
| Model | Provider | $ / 1M in | $ / 1M out | Per request | Per month |
|---|
How the estimate works
- Cost per request = input tokens × input price + output tokens × output price, per million tokens.
- Cached input: the share you set is billed at the provider's cached-read price. Cache write costs are not included, so heavy cache writing makes real bills higher.
- Batch: only applied where the provider documents a discount (Claude: 50%). Other rows ignore the box.
- Tokenizer adjustment: Anthropic states that Claude 4.7 and later models produce about 30% more tokens for the same text. Tick the box to scale those models' token counts; the real change depends on your content.
- Long prompts: Gemini Pro models switch to a higher price when a single prompt exceeds 200k tokens, and the calculator applies it.
- Dated prices: Gemini 3.6, 3.7 and 3.8 Flash change price on 2027-01-01; the calculator uses today's date.
What it leaves out
Taxes, volume discounts, free tiers, tool and search fees, image and audio pricing, and the cost of retries. Models differ in quality and speed, so the cheapest row is not automatically the best choice. Check each provider's official pricing page before you budget.