LLM API Pricing Comparison (2026)
Input, output, and cached-input prices per 1M tokens, plus context window and max output, for 100+ models across every major provider. Search, filter by provider, and sort by any column to find the cheapest model that fits your context needs.
16 models
| Provider | Cached in $/1M | Max output | ||||
|---|---|---|---|---|---|---|
| Gemini 1.5 Flash | $0.07 | $0.00 | — | 8K | 8K | |
| Gemini 2.0 Flash | $0.10 | $0.40 | $0.03 | 1.0M | 8K | |
| GPT-4o mini | OpenAI | $0.15 | $0.60 | $0.07 | 128K | 16K |
| Claude 3 Haiku | Anthropic | $0.25 | $1.25 | $0.03 | 200K | 4K |
| DeepSeek V3 (chat) | DeepSeek | $0.28 | $0.42 | $0.03 | 131K | 8K |
| DeepSeek R1 (reasoner) | DeepSeek | $0.28 | $0.42 | $0.03 | 131K | 66K |
| Codestral | Mistral | $0.30 | $0.90 | $0.03 | 128K | 128K |
| Mistral Large 2 | Mistral | $0.50 | $1.50 | $0.05 | 262K | 262K |
| GPT-3.5 Turbo | OpenAI | $0.50 | $1.50 | — | 16K | 4K |
| Llama 3.3 70B (Together) | Meta | $1.04 | $1.04 | — | 131K | 4K |
| o3-mini | OpenAI | $1.10 | $4.40 | $0.55 | 200K | 100K |
| GPT-4o | OpenAI | $2.50 | $10.00 | $1.25 | 128K | 16K |
| Llama 3.1 405B (Together) | Meta | $3.50 | $3.50 | — | 131K | 4K |
| GPT-4 Turbo | OpenAI | $10.00 | $30.00 | — | 128K | 4K |
| Claude 3 Opus | Anthropic | $15.00 | $75.00 | $1.50 | 200K | 4K |
| o1 | OpenAI | $15.00 | $60.00 | $7.50 | 200K | 100K |
Prices in USD per 1M tokens. Source: litellm · last updated Sep 2, 2026, 3:40 AM. Need to estimate a real workload? Use the AI Cost Calculator.