AI API Cost Calculator

Compare API pricing across GPT, Claude, Gemini, Mistral, and DeepSeek — enter your usage and see monthly cost projections instantly. Free, no sign-up.

Not sure how many tokens your prompts use?

Get an exact token count with the Token Counter tool

Token Counter

Your Usage

Quick Presets

Models to Compare

OpenAI

Anthropic

Google

DeepSeek

Currently Using (optional)

See potential savings if you switch to the cheapest option.

Switch to Gemini 1.5 Flash and save

$123.67/month

$1,484/year compared to GPT-4o

Monthly Cost Comparison

Gemini 1.5 Flash
CHEAPEST
$3.82
GPT-4o mini
$7.65
DeepSeek V3
$13.95
GPT-4o
$127.50
Claude 3.5 Sonnet
$180.00
ModelPer RequestDailyMonthlyAnnual
Gemini 1.5 FlashCHEAPEST
$0.00013$0.1275$3.82$46.54
GPT-4o mini
$0.00025$0.2550$7.65$93.07
DeepSeek V3
$0.00047$0.4650$13.95$169.73
GPT-4oCURRENT
$0.00425$4.25$127.50$1,551
Claude 3.5 Sonnet
$0.00600$6.00$180.00$2,190

Pricing shown is a snapshot of published rates and does not include prompt caching, batch discounts, or fine-tuning costs. Always verify current pricing on the provider's official page before budgeting.

Like this tool? Share it 💰

Help other developers compare AI costs too.

Free AI API Cost Calculator — Compare 16 Models Side by Side

Choosing between GPT-4o, Claude, Gemini, Mistral, and DeepSeek isn't just about capability — cost per token varies by up to 800× between the cheapest and most expensive models. This calculator lets you enter your actual usage pattern (tokens per request, requests per day) and instantly compare projected monthly costs across every major provider, ranked from cheapest to most expensive.

Not sure how many tokens your prompts actually use? Get an exact count with the Token Counter tool, then bring those numbers back here to model your real costs.

Why Input and Output Pricing Differ

Every major provider charges more for output tokens than input tokens — typically 3-4× more. This is because generating text (output) requires significantly more compute per token than reading text (input), since the model must run a full forward pass for each new token it generates. This is why a summarization task (long input, short output) costs very differently from a content generation task (short input, long output) — even at the same total token count.

📄 Summarization (input-heavy)

3000 input / 200 output tokens — costs are dominated by the (cheaper) input rate, so total cost stays lower even with a large document.

✍️ Content Generation (output-heavy)

200 input / 800 output tokens — costs are dominated by the (pricier) output rate, so this workload costs more per request despite fewer total tokens.

Ways to Reduce AI API Costs

🎯

Use a smaller model where possible

Not every task needs GPT-4-class reasoning. Classification, extraction, and simple formatting tasks often work fine on GPT-4o mini, Claude Haiku, or Gemini Flash at a fraction of the cost.

📦

Enable prompt caching

Most providers now offer caching for repeated system prompts or context — this can cut input costs by 50-90% for workloads with a large, stable prompt prefix. Not modeled in this calculator; check your provider's caching pricing.

📊

Use batch APIs for non-real-time work

OpenAI and Anthropic both offer batch processing (results within 24 hours) at roughly half the price of synchronous calls — ideal for bulk summarization, data labeling, or offline generation.

✂️

Trim unnecessary context

Long system prompts, unused few-shot examples, and verbose instructions all cost money on every single request. Audit and trim your prompts regularly.

🔄

Cap output length

Setting a reasonable max_tokens limit prevents runaway generation costs from occasional verbose responses.

Frequently Asked Questions

How accurate is this calculator?

The math is exact given your inputs — token counts multiplied by published per-model pricing. The pricing data itself is a snapshot and can go stale as providers update rates, so always verify current pricing before finalizing a budget.

Why isn't every model included?

Only models with confidently verified, published pricing are listed. Brand-new models are added once official pricing is confirmed rather than estimated — an inaccurate cost projection is worse than none at all.

Does this include prompt caching or batch discounts?

No. This calculator uses standard synchronous input/output pricing. Prompt caching and batch processing can reduce real-world costs significantly — check your provider's pricing page for those specific discounts.

How do I estimate tokens if I don't know my exact usage?

Use one of the built-in presets (chatbot, content generation, code assistant, summarization) as a starting point, or paste your actual prompts into the Token Counter tool for an exact count.

Why does output cost more than input for every model?

Generating text requires more compute per token than reading it, since the model must run inference for each new token produced. This is standard across all major providers — output pricing is typically 3-4× the input rate.

Get Exact Token Counts First

Use the Token Counter to get an exact BPE token count for your actual prompts, then bring those numbers back here to model real costs across every provider.

Token Counter

Explore All Tools

134 free tools — no signup required

All 134 tools are free · No signup · No ads