AI API Cost Calculator (GPT-4o, Claude, Gemini, DeepSeek)
Estimate monthly LLM API cost across GPT-4o, Claude 3.5, Gemini, DeepSeek and more. Compare providers side-by-side with prompt caching and batch discounts.
Last updated:
Token counts are estimated locally (±3% vs official tokenizers). Text is never sent to any server.
Cheapest option at this workload
| Model | Per request | |||||
|---|---|---|---|---|---|---|
googleGemini 1.5 Flash 8B | $0.037/M | $0.150/M | $0.000033 | $0.9855 | 1049K | |
googleGemini 1.5 Flash | $0.075/M | $0.300/M | $0.000066 | $1.97 | 1049K | |
googleGemini 2.0 Flash | $0.100/M | $0.400/M | $0.000088 | $2.63 | 1049K | |
openaiGPT-4o mini | $0.150/M | $0.600/M | $0.000131 | $3.93 | 128K | |
mistralCodestral | $0.300/M | $0.900/M | $0.000202 | $6.07 | 256K | |
deepseekDeepSeek V3 (chat) | $0.270/M | $1.10/M | $0.000240 | $7.20 | 66K | |
metaLlama 3.3 70B (Together) | $0.880/M | $0.880/M | $0.000241 | $7.23 | 131K | |
anthropicClaude 3 Haiku | $0.250/M | $1.25/M | $0.000270 | $8.09 | 200K | |
openaiGPT-3.5 Turbo | $0.500/M | $1.50/M | $0.000337 | $10.11 | 16K | |
deepseekDeepSeek R1 (reasoner) | $0.550/M | $2.19/M | $0.000479 | $14.36 | 66K | |
anthropicClaude 3.5 Haiku | $0.800/M | $4.00/M | $0.000863 | $25.90 | 200K | |
metaLlama 3.1 405B (Together) | $3.50/M | $3.50/M | $0.000959 | $28.77 | 131K | |
openaio3-mini | $1.10/M | $4.40/M | $0.000961 | $28.84 | 200K | |
googleGemini 1.5 Pro | $1.25/M | $5.00/M | $0.001095 | $32.85 | 2097K | |
mistralMistral Large 2 | $2.00/M | $6.00/M | $0.001348 | $40.44 | 128K | |
xaiGrok 2 | $2.00/M | $10.00/M | $0.002148 | $64.44 | 131K | |
xaiGrok 2 Vision | $2.00/M | $10.00/M | $0.002148 | $64.44 | 33K | |
openaiGPT-4o | $2.50/M | $10.00/M | $0.002185 | $65.55 | 128K | |
openaio1-mini | $3.00/M | $12.00/M | $0.002622 | $78.66 | 128K | |
anthropicClaude 3.5 Sonnet | $3.00/M | $15.00/M | $0.003237 | $97.11 | 200K | |
openaiGPT-4 Turbo | $10.00/M | $30.00/M | $0.006740 | $202 | 128K | |
openaio1 | $15.00/M | $60.00/M | $0.0131 | $393 | 200K | |
anthropicClaude 3 Opus | $15.00/M | $75.00/M | $0.0162 | $486 | 200K |
fallback· last fetched 17mo agoEstimate the cost of an OpenAI / Anthropic / Google API call by choosing a model and entering token counts (input, output, cache read, batch). Compare multiple models side-by-side. Pricing is fetched daily from the public LiteLLM catalog.
Tips & Best Practices
- ▸Input tokens are typically 3-5× cheaper than output tokens; long system prompts hurt total cost most when the response is short.
- ▸Cache reads (Anthropic / OpenAI prompt caching) can cut input cost by 90%+ for repeated system prompts.
- ▸Batch API (OpenAI) is 50% cheaper but responses are delivered asynchronously within 24 hours.
Frequently Asked Questions
Where do the prices come from?
Prices are refreshed daily from the BerriAI/litellm public pricing catalog, which itself mirrors each provider's official pricing page. When our automated fetch fails we fall back to a manually maintained snapshot so the tool never shows blank values.
How accurate is the token count?
Token counts use a fast statistical estimator that matches official tokenizers within ±3% for typical English/code prompts. For billing-critical estimates, verify with the provider's own tokenizer.
Does the calculator support prompt caching and batch discounts?
Yes. Toggle 'Cached input' to price the input tokens at the cache-read rate (available on GPT-4o, Claude 3.5, DeepSeek and others). Toggle 'Batch API' to apply the 50%-off batch rate where the provider supports it.
Is my prompt sent to any server?
No. Token estimation runs entirely in your browser. The tool only fetches the public pricing catalog from our server; your prompt text never leaves the tab.
Related Tools
Text Chunker
Split long text into overlapping chunks for embedding and RAG pipelines. Preview recursive, paragraph, sentence, or fixed strategies with live token counts. Runs 100% in your browser.
Cosine Similarity
Compute cosine similarity, dot product, and Euclidean distance between two vectors online. Perfect for debugging LLM embeddings, semantic search, and RAG pipelines. 100% local — vectors never leave your browser.
Token Counter
Estimate the number of tokens in your prompt for GPT-4, GPT-3.5, Claude, and Gemini. Predict API costs before you spend.
JSON Formatter
Format, beautify, and validate JSON online. Free, fast, and 100% local — your data never leaves your browser.
JSON Validator
Validate JSON online with instant error location and structure statistics (depth, keys, node types). 100% local — your JSON never leaves the browser.
JSON → TypeScript
Convert JSON to TypeScript interfaces instantly. Generate strongly-typed TS types from any JSON sample.