WP Webhooks/ Tools/ AI cost calculator

AI Cost Calculator for LLM API Calls

Set the token size of a real request, pick a model, and see what it actually costs — per call, per month, per year.

Rates verified / 45 models / 6 providers
cost = (input_tokens × input_rate + output_tokens × output_rate) ÷ 1,000,000

Promotional rate through 31 Dec 2026

4,700

$0.00352 — the prompt, system message and tool schemas you send.

390

$0.00146 — billed at 5.0× the input rate on this model. Reasoning tokens count here.

0%

Cache hits bill at $0.075 / 1M.

10,000
Input cost
$0.00352
Output cost
$0.00146
Per-call cost
$0.00499
Output is 7.7% of the tokens but 29% of the cost.
Per month
$49.88
Per year
$598.50
Rate expires 31 December 2026

$598.50/yr assumes the introductory rate holds for 12 months. It does not. From 1 January 2027 this becomes Gemini 3.6 Flash (standard) at $1.5 / $7.5, or $1,197/yr — 100% more. Budget against that.

If you resell this as credits

weighted = 4,700 + 5.0 × 390 = 6,650

7 credits at 1,000 weighted tokens each

One credit backs $0.00071 of spend on this model, for any input/output mix.

Cheapest for this call is gpt-5-nano at $0.00039 — 12.8× less than your pick. Over a year that is $551.58 of difference.

The same call, every model

4,700 input and 390 output tokens priced against all 45 models. Cheapest to most expensive — a 540× spread for identical work.

#ModelProviderPer callPer yearvs cheapest
1gpt-5-nanoOpenAI$0.00039$46.92
2GLM-4.7-FlashXZ.ai$0.00049$58.201.2×
3GLM-5.3-FlashZ.ai$0.0009$108.002.3×
4GLM-4.5-AirZ.ai$0.00137$164.283.5×
5gpt-5.6-lunaOpenAI$0.00141$168.963.6×
6gpt-5.4-nanoOpenAI$0.00143$171.303.7×
7Gemini 3.1 Flash-LiteGoogle$0.00176$211.204.5×
8deepseek-flash (was V4 Flash)DeepSeek$0.00188$225.364.8×
9gpt-5-miniOpenAI$0.00196$234.605.0×
10Gemini 3.5 Flash-LiteGoogle$0.00238$286.206.1×
11GLM-4.7Z.ai$0.00368$441.369.4×
12Gemini 3.6 Flash (promo)Google$0.00499$598.5012.8×
13Gemini 3.7 Flash (promo)Google$0.00499$598.5012.8×
14Gemini 3.8 Flash (promo)Google$0.00499$598.5012.8×
15gpt-5.4-miniOpenAI$0.00528$633.6013.5×
16GLM-5Z.ai$0.00595$713.7615.2×
17Kimi K2.6Moonshot$0.00602$723.0015.4×
18Claude Haiku 4.5Anthropic$0.00665$798.0017.0×
19deepseek-v4-proDeepSeek$0.00775$929.8119.8×
20GLM-5.2Z.ai$0.0083$995.5221.2×
21GLM-5.3Z.ai$0.0083$995.5221.2×
22gpt-5OpenAI$0.00977$1,17325.0×
23gpt-5.1OpenAI$0.00977$1,17325.0×
24Gemini 3.6 Flash (standard)Google$0.00997$1,19725.5×
25Gemini 3.7 Flash (standard)Google$0.00997$1,19725.5×
26Gemini 3.8 Flash (standard)Google$0.00997$1,19725.5×
27Gemini 3.5 FlashGoogle$0.0106$1,26727.0×
28Claude Sonnet 5Anthropic$0.0133$1,59634.0×
29gpt-5.2OpenAI$0.0137$1,64235.0×
30gpt-5.6-terraOpenAI$0.0141$1,69036.0×
31Gemini 3.1 ProGoogle$0.0141$1,69036.0×
32gpt-5.4OpenAI$0.0176$2,11245.0×
33Kimi K3Moonshot$0.0199$2,39451.0×
34gpt-5.6-solpremiumOpenAI$0.0266$3,19268.0×
35Claude Opus 5Anthropic$0.0333$3,99085.0×
36gpt-5.5premiumOpenAI$0.0352$4,22490.0×
37gpt-6-astrapremiumOpenAI$0.0665$7,980170.1×
38Claude Fable 5premiumAnthropic$0.0665$7,980170.1×
39Claude Fable 5.1premiumAnthropic$0.0665$7,980170.1×
40Claude Mythos 5premiumAnthropic$0.0665$7,980170.1×
41Claude Mythos 5.1premiumAnthropic$0.0665$7,980170.1×
42gpt-5-propremiumOpenAI$0.1173$14,076300.0×
43gpt-5.2-propremiumOpenAI$0.1642$19,706420.0×
44gpt-5.4-propremiumOpenAI$0.2112$25,344540.2×
45gpt-5.5-propremiumOpenAI$0.2112$25,344540.2×

What changed since the last check

Every rate movement caught between 21 August 2026 and 13 September 2026, re-read from each provider's own documentation. The change column is the move in the cost of a 4,700 / 390 call, not in the headline rate.

ModelProviderWasNowPer call
Claude Sonnet 5 The scheduled 1 Sep increase to $3 / $15 was cancelled — the $2 / $10 introductory rate is now the standard priceAnthropic$3.00 / $15.00$2.00 / $10.00-33%
Claude Fable 5.1 New model at the Fable 5 rate; cache hits $0.25 (0.025×) instead of $1.00Anthropic$10.00 / $50.00new
Claude Mythos 5.1 New model, limited availability, same rates as Fable 5.1Anthropic$10.00 / $50.00new
gpt-6-astra New premium model; cached input $1.00OpenAI$10.00 / $50.00new
Gemini 3.8 Flash New model at the 3.6 / 3.7 Flash promotional rate through 31 Dec 2026, then $1.50 / $7.50Google$0.75 / $3.75new
deepseek-flash (was deepseek-v4-flash) Renamed and cut; cache hits $0.006; peak hours now weekdays onlyDeepSeek$0.44 / $1.32$0.30 / $1.20-27%
GLM-5.3-Flash New budget modelZ.ai$0.15 / $0.50new
Kimi K2.6 Newly tracked here — Moonshot's budget general model; K3 itself is unchangedMoonshot$0.95 / $4.00new
gpt-5.6-luna OpenAI$1.00 / $6.00$0.20 / $1.20-80%
gpt-5.6-terra OpenAI$2.50 / $15.00$2.00 / $12.00-20%
gpt-5.6-sol OpenAI$5.00 / $30.00$4.00 / $20.00-24%
Gemini 3.6 Flash Promotional rate through 31 Dec 2026, then back to $1.50 / $7.50Google$1.50 / $7.50$0.75 / $3.75-50%
Gemini 3.7 Flash New model, introductory rate through 31 Dec 2026Google$0.75 / $3.75new
deepseek-v4-flash Repriced into peak / off-peak tiers; peak shownDeepSeek$0.14 / $0.28$0.44 / $1.32+237%
deepseek-v4-pro Repriced into peak / off-peak tiers; peak shownDeepSeek$0.435 / $0.87$1.32 / $3.96+225%
GLM-5.3 New model at the GLM-5.2 rateZ.ai$1.40 / $4.40new

How to read this. Rates are per million tokens, taken from each provider's own documentation on 13 September 2026. Published pricing changes often — check the source before you budget against it.

  • The default 4,700 / 390 split is a planning or tool-calling request: a large prompt, a short structured answer. Chat workloads invert it, which changes the ranking.
  • Cached input is only applied to models whose provider publishes a flat cache-hit rate. The rest ignore the slider rather than guess.
  • Batch APIs on OpenAI, Google and Anthropic discount both directions by 50% — halve any figure here if the work is not interactive.

The reasoning behind the credit formula, the output-weight ratios, and where each rate came from is written up in what an AI feature actually costs per call.

FAQ

Questions about AI pricing.

Rates come from each provider's own documentation — the full write-up is in the LLM cost article.

How do you calculate the cost of an LLM API call? +
Multiply input tokens by the model input rate, output tokens by the output rate, and divide by 1,000,000. On Gemini 3.6 Flash ($0.75 in / $3.75 out) a 4,700-input, 390-output call is (4,700 × 0.75 + 390 × 3.75) ÷ 1,000,000 = $0.00499. Every provider bills in exactly that shape, so the only hard part is knowing your real token size.
Why does the calculator default to 4,700 input and 390 output tokens? +
That is a measured production request for a planning or tool-calling task: a large system prompt plus tool schemas going in, a short structured decision coming back. It is deliberately not a chat workload, which inverts the ratio and changes which model is cheapest. Move the sliders to match your own traffic — and expect real calls to run 1.5 to 2 times larger than a thin test sample.
Why are output tokens more expensive than input tokens? +
Generating tokens is sequential, while input can be processed in parallel, and providers price accordingly. The ratio is typically 5× on Anthropic, Google Flash and Kimi K3, 6–8× across the OpenAI GPT-5 line, and 3× on DeepSeek V4. Because reasoning tokens are billed as output, a model that always thinks can cost more than its visible answer suggests.
Why is the cached-input slider disabled for some models? +
Because that provider does not publish a flat per-token cache-hit rate. OpenAI, Anthropic, Moonshot, DeepSeek and Z.ai all do, so the slider applies their real discount — typically 0.1× the base input rate, and as low as 0.03× on DeepSeek. Only Google is disabled here: its context-caching price carries a separate per-hour storage charge, so a single per-token rate would understate the real cost.
Which LLM API is the cheapest? +
At the default call size, gpt-5-nano ($0.05 / $0.40 per million) and GLM-4.7-FlashX ($0.07 / $0.40) are the cheapest paid options, followed by GLM-4.5-Air and gpt-5.6-luna at around $0.0014 a call. GLM-5.3-Flash ($0.15 / $0.50) is new in that group. The DeepSeek Flash model left it in August 2026 at $0.44 / $1.32 and is back on its edge since September as deepseek-flash at $0.30 / $1.20. Cheapest only matters if quality holds for your task, so A/B a candidate on real requests before switching, and note that naming is not a price guide: Kimi K3 costs half as much again per call as Claude Sonnet 5.
How should I charge credits if I resell AI usage? +
Charge from the token counts the provider returns, not a flat price per request, because a flat rate always undercharges your largest calls. Weight output tokens by the model output-to-input price ratio, divide by a fixed tokens-per-credit constant, and settle after the call succeeds. One credit then backs a fixed amount of provider cost for any input/output mix.
How current are these prices? +
Every rate was read from the provider's own documentation on 13 September 2026 — never from aggregator sites, which drift. Published pricing changes often, and introductory rates expire — Gemini 3.6, 3.7 and 3.8 Flash all double on 1 January 2027 — though not always: Claude Sonnet 5's planned move from $2/$10 to $3/$15 on 1 September 2026 was cancelled, and $2/$10 is now its standard rate. The change log on this page records every rate movement we have caught since 29 July 2026, including a 50% cut on Gemini 3.6 Flash and a roughly 3× rise on DeepSeek V4. Verify against the provider before you budget against any figure here.
Ready

Your next automation is
one sentence away.

$ wp plugin install flowsystems-webhook-actions --activate