AI Cost Calculator for LLM API Calls
Set the token size of a real request, pick a model, and see what it actually costs — per call, per month, per year.
Promotional rate through 31 Dec 2026
$0.00352 — the prompt, system message and tool schemas you send.
$0.00146 — billed at 5.0× the input rate on this model. Reasoning tokens count here.
Cache hits bill at $0.075 / 1M.
$598.50/yr assumes the introductory rate holds for 12 months. It does not. From 1 January 2027 this becomes Gemini 3.6 Flash (standard) at $1.5 / $7.5, or $1,197/yr — 100% more. Budget against that.
weighted = 4,700 + 5.0 × 390 = 6,650
→ 7 credits at 1,000 weighted tokens each
One credit backs $0.00071 of spend on this model, for any input/output mix.
The same call, every model
4,700 input and 390 output tokens priced against all 45 models. Cheapest to most expensive — a 540× spread for identical work.
| # | Model | Provider | Per call | Per year | vs cheapest |
|---|---|---|---|---|---|
| 1 | gpt-5-nano | OpenAI | $0.00039 | $46.92 | — |
| 2 | GLM-4.7-FlashX | Z.ai | $0.00049 | $58.20 | 1.2× |
| 3 | GLM-5.3-Flash | Z.ai | $0.0009 | $108.00 | 2.3× |
| 4 | GLM-4.5-Air | Z.ai | $0.00137 | $164.28 | 3.5× |
| 5 | gpt-5.6-luna | OpenAI | $0.00141 | $168.96 | 3.6× |
| 6 | gpt-5.4-nano | OpenAI | $0.00143 | $171.30 | 3.7× |
| 7 | Gemini 3.1 Flash-Lite | $0.00176 | $211.20 | 4.5× | |
| 8 | deepseek-flash (was V4 Flash) | DeepSeek | $0.00188 | $225.36 | 4.8× |
| 9 | gpt-5-mini | OpenAI | $0.00196 | $234.60 | 5.0× |
| 10 | Gemini 3.5 Flash-Lite | $0.00238 | $286.20 | 6.1× | |
| 11 | GLM-4.7 | Z.ai | $0.00368 | $441.36 | 9.4× |
| 12 | Gemini 3.6 Flash (promo) | $0.00499 | $598.50 | 12.8× | |
| 13 | Gemini 3.7 Flash (promo) | $0.00499 | $598.50 | 12.8× | |
| 14 | Gemini 3.8 Flash (promo) | $0.00499 | $598.50 | 12.8× | |
| 15 | gpt-5.4-mini | OpenAI | $0.00528 | $633.60 | 13.5× |
| 16 | GLM-5 | Z.ai | $0.00595 | $713.76 | 15.2× |
| 17 | Kimi K2.6 | Moonshot | $0.00602 | $723.00 | 15.4× |
| 18 | Claude Haiku 4.5 | Anthropic | $0.00665 | $798.00 | 17.0× |
| 19 | deepseek-v4-pro | DeepSeek | $0.00775 | $929.81 | 19.8× |
| 20 | GLM-5.2 | Z.ai | $0.0083 | $995.52 | 21.2× |
| 21 | GLM-5.3 | Z.ai | $0.0083 | $995.52 | 21.2× |
| 22 | gpt-5 | OpenAI | $0.00977 | $1,173 | 25.0× |
| 23 | gpt-5.1 | OpenAI | $0.00977 | $1,173 | 25.0× |
| 24 | Gemini 3.6 Flash (standard) | $0.00997 | $1,197 | 25.5× | |
| 25 | Gemini 3.7 Flash (standard) | $0.00997 | $1,197 | 25.5× | |
| 26 | Gemini 3.8 Flash (standard) | $0.00997 | $1,197 | 25.5× | |
| 27 | Gemini 3.5 Flash | $0.0106 | $1,267 | 27.0× | |
| 28 | Claude Sonnet 5 | Anthropic | $0.0133 | $1,596 | 34.0× |
| 29 | gpt-5.2 | OpenAI | $0.0137 | $1,642 | 35.0× |
| 30 | gpt-5.6-terra | OpenAI | $0.0141 | $1,690 | 36.0× |
| 31 | Gemini 3.1 Pro | $0.0141 | $1,690 | 36.0× | |
| 32 | gpt-5.4 | OpenAI | $0.0176 | $2,112 | 45.0× |
| 33 | Kimi K3 | Moonshot | $0.0199 | $2,394 | 51.0× |
| 34 | gpt-5.6-solpremium | OpenAI | $0.0266 | $3,192 | 68.0× |
| 35 | Claude Opus 5 | Anthropic | $0.0333 | $3,990 | 85.0× |
| 36 | gpt-5.5premium | OpenAI | $0.0352 | $4,224 | 90.0× |
| 37 | gpt-6-astrapremium | OpenAI | $0.0665 | $7,980 | 170.1× |
| 38 | Claude Fable 5premium | Anthropic | $0.0665 | $7,980 | 170.1× |
| 39 | Claude Fable 5.1premium | Anthropic | $0.0665 | $7,980 | 170.1× |
| 40 | Claude Mythos 5premium | Anthropic | $0.0665 | $7,980 | 170.1× |
| 41 | Claude Mythos 5.1premium | Anthropic | $0.0665 | $7,980 | 170.1× |
| 42 | gpt-5-propremium | OpenAI | $0.1173 | $14,076 | 300.0× |
| 43 | gpt-5.2-propremium | OpenAI | $0.1642 | $19,706 | 420.0× |
| 44 | gpt-5.4-propremium | OpenAI | $0.2112 | $25,344 | 540.2× |
| 45 | gpt-5.5-propremium | OpenAI | $0.2112 | $25,344 | 540.2× |
What changed since the last check
Every rate movement caught between 21 August 2026 and 13 September 2026, re-read from each provider's own documentation. The change column is the move in the cost of a 4,700 / 390 call, not in the headline rate.
| Model | Provider | Was | Now | Per call |
|---|---|---|---|---|
| Claude Sonnet 5 The scheduled 1 Sep increase to $3 / $15 was cancelled — the $2 / $10 introductory rate is now the standard price | Anthropic | $3.00 / $15.00 | $2.00 / $10.00 | -33% |
| Claude Fable 5.1 New model at the Fable 5 rate; cache hits $0.25 (0.025×) instead of $1.00 | Anthropic | — | $10.00 / $50.00 | new |
| Claude Mythos 5.1 New model, limited availability, same rates as Fable 5.1 | Anthropic | — | $10.00 / $50.00 | new |
| gpt-6-astra New premium model; cached input $1.00 | OpenAI | — | $10.00 / $50.00 | new |
| Gemini 3.8 Flash New model at the 3.6 / 3.7 Flash promotional rate through 31 Dec 2026, then $1.50 / $7.50 | — | $0.75 / $3.75 | new | |
| deepseek-flash (was deepseek-v4-flash) Renamed and cut; cache hits $0.006; peak hours now weekdays only | DeepSeek | $0.44 / $1.32 | $0.30 / $1.20 | -27% |
| GLM-5.3-Flash New budget model | Z.ai | — | $0.15 / $0.50 | new |
| Kimi K2.6 Newly tracked here — Moonshot's budget general model; K3 itself is unchanged | Moonshot | — | $0.95 / $4.00 | new |
| gpt-5.6-luna | OpenAI | $1.00 / $6.00 | $0.20 / $1.20 | -80% |
| gpt-5.6-terra | OpenAI | $2.50 / $15.00 | $2.00 / $12.00 | -20% |
| gpt-5.6-sol | OpenAI | $5.00 / $30.00 | $4.00 / $20.00 | -24% |
| Gemini 3.6 Flash Promotional rate through 31 Dec 2026, then back to $1.50 / $7.50 | $1.50 / $7.50 | $0.75 / $3.75 | -50% | |
| Gemini 3.7 Flash New model, introductory rate through 31 Dec 2026 | — | $0.75 / $3.75 | new | |
| deepseek-v4-flash Repriced into peak / off-peak tiers; peak shown | DeepSeek | $0.14 / $0.28 | $0.44 / $1.32 | +237% |
| deepseek-v4-pro Repriced into peak / off-peak tiers; peak shown | DeepSeek | $0.435 / $0.87 | $1.32 / $3.96 | +225% |
| GLM-5.3 New model at the GLM-5.2 rate | Z.ai | — | $1.40 / $4.40 | new |
How to read this. Rates are per million tokens, taken from each provider's own documentation on 13 September 2026. Published pricing changes often — check the source before you budget against it.
- The default 4,700 / 390 split is a planning or tool-calling request: a large prompt, a short structured answer. Chat workloads invert it, which changes the ranking.
- Cached input is only applied to models whose provider publishes a flat cache-hit rate. The rest ignore the slider rather than guess.
- Batch APIs on OpenAI, Google and Anthropic discount both directions by 50% — halve any figure here if the work is not interactive.
The reasoning behind the credit formula, the output-weight ratios, and where each rate came from is written up in what an AI feature actually costs per call.
Questions about AI pricing.
Rates come from each provider's own documentation — the full write-up is in the LLM cost article.