DevKits

AI API Cost Calculator (GPT-4o, Claude, Gemini, DeepSeek)

Estimate monthly LLM API cost across GPT-4o, Claude 3.5, Gemini, DeepSeek and more. Compare providers side-by-side with prompt caching and batch discounts.

Last updated:

273 chars · 9 lines

Token counts are estimated locally (±3% vs official tokenizers). Text is never sent to any server.

Cheapest option at this workload

Gemini 1.5 Flash 8B$0.9855/mo· $0.000033 per request · ~76 in / 200 out tokens
ModelPer request
googleGemini 1.5 Flash 8B
$0.037/M$0.150/M$0.000033$0.98551049K
googleGemini 1.5 Flash
$0.075/M$0.300/M$0.000066$1.971049K
googleGemini 2.0 Flash
$0.100/M$0.400/M$0.000088$2.631049K
openaiGPT-4o mini
$0.150/M$0.600/M$0.000131$3.93128K
mistralCodestral
$0.300/M$0.900/M$0.000202$6.07256K
deepseekDeepSeek V3 (chat)
$0.270/M$1.10/M$0.000240$7.2066K
metaLlama 3.3 70B (Together)
$0.880/M$0.880/M$0.000241$7.23131K
anthropicClaude 3 Haiku
$0.250/M$1.25/M$0.000270$8.09200K
openaiGPT-3.5 Turbo
$0.500/M$1.50/M$0.000337$10.1116K
deepseekDeepSeek R1 (reasoner)
$0.550/M$2.19/M$0.000479$14.3666K
anthropicClaude 3.5 Haiku
$0.800/M$4.00/M$0.000863$25.90200K
metaLlama 3.1 405B (Together)
$3.50/M$3.50/M$0.000959$28.77131K
openaio3-mini
$1.10/M$4.40/M$0.000961$28.84200K
googleGemini 1.5 Pro
$1.25/M$5.00/M$0.001095$32.852097K
mistralMistral Large 2
$2.00/M$6.00/M$0.001348$40.44128K
xaiGrok 2
$2.00/M$10.00/M$0.002148$64.44131K
xaiGrok 2 Vision
$2.00/M$10.00/M$0.002148$64.4433K
openaiGPT-4o
$2.50/M$10.00/M$0.002185$65.55128K
openaio1-mini
$3.00/M$12.00/M$0.002622$78.66128K
anthropicClaude 3.5 Sonnet
$3.00/M$15.00/M$0.003237$97.11200K
openaiGPT-4 Turbo
$10.00/M$30.00/M$0.006740$202128K
openaio1
$15.00/M$60.00/M$0.0131$393200K
anthropicClaude 3 Opus
$15.00/M$75.00/M$0.0162$486200K
Pricing source: fallback· last fetched 17mo ago

Estimate the cost of an OpenAI / Anthropic / Google API call by choosing a model and entering token counts (input, output, cache read, batch). Compare multiple models side-by-side. Pricing is fetched daily from the public LiteLLM catalog.

Tips & Best Practices

  • Input tokens are typically 3-5× cheaper than output tokens; long system prompts hurt total cost most when the response is short.
  • Cache reads (Anthropic / OpenAI prompt caching) can cut input cost by 90%+ for repeated system prompts.
  • Batch API (OpenAI) is 50% cheaper but responses are delivered asynchronously within 24 hours.

Frequently Asked Questions

Where do the prices come from?

Prices are refreshed daily from the BerriAI/litellm public pricing catalog, which itself mirrors each provider's official pricing page. When our automated fetch fails we fall back to a manually maintained snapshot so the tool never shows blank values.

How accurate is the token count?

Token counts use a fast statistical estimator that matches official tokenizers within ±3% for typical English/code prompts. For billing-critical estimates, verify with the provider's own tokenizer.

Does the calculator support prompt caching and batch discounts?

Yes. Toggle 'Cached input' to price the input tokens at the cache-read rate (available on GPT-4o, Claude 3.5, DeepSeek and others). Toggle 'Batch API' to apply the 50%-off batch rate where the provider supports it.

Is my prompt sent to any server?

No. Token estimation runs entirely in your browser. The tool only fetches the public pricing catalog from our server; your prompt text never leaves the tab.

Related Tools