AI API Cost Calculator
Compare API pricing across GPT, Claude, Gemini, Mistral, and DeepSeek — enter your usage and see monthly cost projections instantly. Free, no sign-up.
Not sure how many tokens your prompts use?
Get an exact token count with the Token Counter tool
Your Usage
Quick Presets
Models to Compare
OpenAI
Anthropic
DeepSeek
Currently Using (optional)
See potential savings if you switch to the cheapest option.
Switch to Gemini 1.5 Flash and save
$123.67/month
$1,484/year compared to GPT-4o
Monthly Cost Comparison
| Model | Per Request | Daily | Monthly | Annual |
|---|---|---|---|---|
Gemini 1.5 FlashCHEAPEST | $0.00013 | $0.1275 | $3.82 | $46.54 |
GPT-4o mini | $0.00025 | $0.2550 | $7.65 | $93.07 |
DeepSeek V3 | $0.00047 | $0.4650 | $13.95 | $169.73 |
GPT-4oCURRENT | $0.00425 | $4.25 | $127.50 | $1,551 |
Claude 3.5 Sonnet | $0.00600 | $6.00 | $180.00 | $2,190 |
Pricing shown is a snapshot of published rates and does not include prompt caching, batch discounts, or fine-tuning costs. Always verify current pricing on the provider's official page before budgeting.
Support Our Free Tools
If you find this calculator helpful, please consider supporting our work. Your contribution helps us build and maintain these free tools for everyone.
Buy me a coffeeFree AI API Cost Calculator — Compare 16 Models Side by Side
Choosing between GPT-4o, Claude, Gemini, Mistral, and DeepSeek isn't just about capability — cost per token varies by up to 800× between the cheapest and most expensive models. This calculator lets you enter your actual usage pattern (tokens per request, requests per day) and instantly compare projected monthly costs across every major provider, ranked from cheapest to most expensive.
Not sure how many tokens your prompts actually use? Get an exact count with the Token Counter tool, then bring those numbers back here to model your real costs.
Why Input and Output Pricing Differ
Every major provider charges more for output tokens than input tokens — typically 3-4× more. This is because generating text (output) requires significantly more compute per token than reading text (input), since the model must run a full forward pass for each new token it generates. This is why a summarization task (long input, short output) costs very differently from a content generation task (short input, long output) — even at the same total token count.
📄 Summarization (input-heavy)
3000 input / 200 output tokens — costs are dominated by the (cheaper) input rate, so total cost stays lower even with a large document.
✍️ Content Generation (output-heavy)
200 input / 800 output tokens — costs are dominated by the (pricier) output rate, so this workload costs more per request despite fewer total tokens.
Ways to Reduce AI API Costs
Use a smaller model where possible
Not every task needs GPT-4-class reasoning. Classification, extraction, and simple formatting tasks often work fine on GPT-4o mini, Claude Haiku, or Gemini Flash at a fraction of the cost.
Enable prompt caching
Most providers now offer caching for repeated system prompts or context — this can cut input costs by 50-90% for workloads with a large, stable prompt prefix. Not modeled in this calculator; check your provider's caching pricing.
Use batch APIs for non-real-time work
OpenAI and Anthropic both offer batch processing (results within 24 hours) at roughly half the price of synchronous calls — ideal for bulk summarization, data labeling, or offline generation.
Trim unnecessary context
Long system prompts, unused few-shot examples, and verbose instructions all cost money on every single request. Audit and trim your prompts regularly.
Cap output length
Setting a reasonable max_tokens limit prevents runaway generation costs from occasional verbose responses.
Frequently Asked Questions
How accurate is this calculator?
The math is exact given your inputs — token counts multiplied by published per-model pricing. The pricing data itself is a snapshot and can go stale as providers update rates, so always verify current pricing before finalizing a budget.
Why isn't every model included?
Only models with confidently verified, published pricing are listed. Brand-new models are added once official pricing is confirmed rather than estimated — an inaccurate cost projection is worse than none at all.
Does this include prompt caching or batch discounts?
No. This calculator uses standard synchronous input/output pricing. Prompt caching and batch processing can reduce real-world costs significantly — check your provider's pricing page for those specific discounts.
How do I estimate tokens if I don't know my exact usage?
Use one of the built-in presets (chatbot, content generation, code assistant, summarization) as a starting point, or paste your actual prompts into the Token Counter tool for an exact count.
Why does output cost more than input for every model?
Generating text requires more compute per token than reading it, since the model must run inference for each new token produced. This is standard across all major providers — output pricing is typically 3-4× the input rate.
Get Exact Token Counts First
Use the Token Counter to get an exact BPE token count for your actual prompts, then bring those numbers back here to model real costs across every provider.
Explore All Tools
134 free tools — no signup required
All 134 tools are free · No signup · No ads
