AI Cost Estimator
Estimate monthly or annual LLM API costs per model from workload presets, request volume, token counts and prompt-cache hit rate.
AI Cost Estimator
Use case preset
Workload configuration
Cached portion is billed at each model's published cached input rate
Lowest monthly cost: GPT-OSS 20B (Groq) at $5.63
Monthly cost comparison
Detailed cost breakdown
| GPT-OSS 20B (Groq)Cheapest | Groq | $0.037 | $0.150 | $0.188 | $5.63 |
| GPT-5-nano | OpenAI | $0.025 | $0.200 | $0.225 | $6.75 |
| GPT-4.1-nano | OpenAI | $0.050 | $0.200 | $0.250 | $7.50 |
| Gemini 2.5 Flash-Lite | $0.050 | $0.200 | $0.250 | $7.50 | |
| GPT-6 Luna | OpenAI | $0.050 | $0.250 | $0.300 | $9.00 |
| Qwen3.8-Flash | Qwen | $0.075 | $0.235 | $0.310 | $9.30 |
| Mistral Small 4 | Mistral | $0.075 | $0.300 | $0.375 | $11.25 |
| GPT-OSS 120B (Groq) | Groq | $0.075 | $0.300 | $0.375 | $11.25 |
| Codestral 25.08 | Mistral | $0.150 | $0.450 | $0.600 | $18.00 |
| GPT-5.6 Luna | OpenAI | $0.100 | $0.600 | $0.700 | $21.00 |
| GPT-5.4 Nano | OpenAI | $0.100 | $0.625 | $0.725 | $21.75 |
| DeepSeek V4.1 Flash | DeepSeek | $0.150 | $0.600 | $0.750 | $22.50 |
| GPT-4.1-mini | OpenAI | $0.200 | $0.800 | $1.00 | $30.00 |
| Mistral Large 3 | Mistral | $0.250 | $0.750 | $1.00 | $30.00 |
| GPT-5-mini | OpenAI | $0.125 | $1.00 | $1.13 | $33.75 |
| Gemini 3.5 Flash-Lite | $0.150 | $1.25 | $1.40 | $42.00 | |
| Gemini 2.5 Flash | $0.150 | $1.25 | $1.40 | $42.00 | |
| Grok Build 0.1 | xAI | $0.500 | $1.00 | $1.50 | $45.00 |
| Grok 4.3 | xAI | $0.625 | $1.25 | $1.88 | $56.25 |
| Grok 4.20 | xAI | $0.625 | $1.25 | $1.88 | $56.25 |
| Gemini 3.8 Flash | $0.375 | $1.88 | $2.25 | $67.50 | |
| Gemini 3.7 Flash | $0.375 | $1.88 | $2.25 | $67.50 | |
| Gemini 3.6 Flash | $0.375 | $1.88 | $2.25 | $67.50 | |
| GPT-5.4 Mini | OpenAI | $0.375 | $2.25 | $2.63 | $78.75 |
| DeepSeek V4 Pro | DeepSeek | $0.660 | $1.98 | $2.64 | $79.20 |
| o4-mini | OpenAI | $0.550 | $2.20 | $2.75 | $82.50 |
| o3-mini | OpenAI | $0.550 | $2.20 | $2.75 | $82.50 |
| Muse Spark 1.3 | Meta | $0.625 | $2.13 | $2.75 | $82.50 |
| Claude Haiku 4.5 | Anthropic | $0.500 | $2.50 | $3.00 | $90.00 |
| Grok 4.7 | xAI | $1.00 | $3.00 | $4.00 | $120.00 |
| Grok 4.6 | xAI | $1.00 | $3.00 | $4.00 | $120.00 |
| Grok 4.5 | xAI | $1.00 | $3.00 | $4.00 | $120.00 |
| Qwen3.8-Max | Qwen | $1.00 | $3.00 | $4.00 | $120.00 |
| Mistral Medium 3.5 | Mistral | $0.750 | $3.75 | $4.50 | $135.00 |
| o3 | OpenAI | $1.00 | $4.00 | $5.00 | $150.00 |
| GPT-4.1 | OpenAI | $1.00 | $4.00 | $5.00 | $150.00 |
| Gemini 3.5 Flash | $0.750 | $4.50 | $5.25 | $157.50 | |
| GPT-5.1 | OpenAI | $0.625 | $5.00 | $5.63 | $168.75 |
| GPT-5 | OpenAI | $0.625 | $5.00 | $5.63 | $168.75 |
| Gemini 2.5 Pro | $0.625 | $5.00 | $5.63 | $168.75 | |
| Claude Sonnet 5.5 | Anthropic | $1.00 | $5.00 | $6.00 | $180.00 |
| Claude Sonnet 5 | Anthropic | $1.00 | $5.00 | $6.00 | $180.00 |
| GPT-6.1 Sol | OpenAI | $1.00 | $5.00 | $6.00 | $180.00 |
| GPT-6 Sol | OpenAI | $1.00 | $5.00 | $6.00 | $180.00 |
| GPT-5.6 Terra | OpenAI | $1.00 | $6.00 | $7.00 | $210.00 |
| Gemini 3.1 Pro Preview | $1.00 | $6.00 | $7.00 | $210.00 | |
| GPT-5.3-Codex | OpenAI | $0.875 | $7.00 | $7.88 | $236.25 |
| GPT-5.2 | OpenAI | $0.875 | $7.00 | $7.88 | $236.25 |
| GPT-5.4 | OpenAI | $1.25 | $7.50 | $8.75 | $262.50 |
| Claude Sonnet 4.6 | Anthropic | $1.50 | $7.50 | $9.00 | $270.00 |
| Claude Sonnet 4.5 | Anthropic | $1.50 | $7.50 | $9.00 | $270.00 |
| Claude Opus 5.5 | Anthropic | $2.00 | $10.00 | $12.00 | $360.00 |
| GPT-5.6 Sol | OpenAI | $2.00 | $10.00 | $12.00 | $360.00 |
| Claude Opus 5 | Anthropic | $2.50 | $12.50 | $15.00 | $450.00 |
| Claude Opus 4.8 | Anthropic | $2.50 | $12.50 | $15.00 | $450.00 |
| Claude Opus 4.7 | Anthropic | $2.50 | $12.50 | $15.00 | $450.00 |
| Claude Opus 4.6 | Anthropic | $2.50 | $12.50 | $15.00 | $450.00 |
| Claude Opus 4.5 | Anthropic | $2.50 | $12.50 | $15.00 | $450.00 |
| GPT-5.5 | OpenAI | $2.50 | $15.00 | $17.50 | $525.00 |
| Claude Fable 5.1 | Anthropic | $5.00 | $25.00 | $30.00 | $900.00 |
| Claude Fable 5 | Anthropic | $5.00 | $25.00 | $30.00 | $900.00 |
| GPT-6 Astra | OpenAI | $5.00 | $25.00 | $30.00 | $900.00 |
| o3-pro | OpenAI | $10.00 | $40.00 | $50.00 | $1,500.00 |
| GPT-5.2 Pro | OpenAI | $10.50 | $84.00 | $94.50 | $2,835.00 |
| GPT-5.5 Pro | OpenAI | $15.00 | $90.00 | $105.00 | $3,150.00 |
| GPT-5.4 ProPriciest | OpenAI | $15.00 | $90.00 | $105.00 | $3,150.00 |
Cost vs. capability insights
Compared to the cheapest option (GPT-OSS 20B (Groq)):
Workload summary
Pricing data last updated: October 3, 2026. Prices per 1M tokens. Estimates are approximate and may vary with actual usage patterns.
What this tool does
LLM API pricing in 2026 spans from $0.10 input / $0.50 output per million tokens (GPT-6 Luna) up to $10 input / $50 output (Claude Fable 5.1), across 23 current models from 9 providers. Flagship rates: GPT-6 Astra $10 input / $50 output, Claude Fable 5.1 $10 input / $50 output, Gemini 3.1 Pro Preview $2 input / $12 output. Every price in the table below is verified against the vendor's official pricing page. This estimator turns those per-token rates into monthly workload costs.
Updated . Provided as is. Check the output before you rely on it in production.
How to use AI Cost Estimator
- 1
Select your AI models
Choose the models you use, such as GPT, Claude, Gemini, Grok or DeepSeek. The tool displays current pricing per 1M input and output tokens. Add multiple models if you're comparing or using a mix.
- 2
Estimate token counts
Input your expected usage: requests per day, average input tokens (rule of thumb: 1 token is about 4 characters of English), average output tokens, and billing days per month.
- 3
Set a cache hit rate
Enter the share of input tokens you expect to land on a warm prompt cache. That share is repriced at each model's published cached-input rate; models with no published cached rate stay at the full input price.
- 4
Calculate total and per-request costs
The table gives daily, monthly and annual cost per model, split into input and output, and marks the cheapest and priciest rows. Switch between the monthly and annual view to compare against your budget.
Questions and answers
What is AI Cost Estimator?
Does AI Cost Estimator store or send my data?
How accurate are the cost estimates?
For AI agents: how to call this tool
Machine-readable contract, endpoints and examples. Humans can ignore this section.
Best Path For Builders
Browser workflow
Runs instantly in the browser with private local processing and copy/export-ready output.
Browser Workflow
This tool is optimized for instant in-browser execution with local data handling. Run it here and copy/export the output directly.
/ai-cost-estimator/
For automation planning, fetch the canonical contract at /api/tool/ai-cost-estimator.json.
How much do LLM APIs cost per million tokens across providers in 2026?
Current list prices per 1 million tokens for 23 models across Anthropic, OpenAI, Google, DeepSeek, xAI, Mistral, Meta, Qwen, Groq. Input and output tokens are billed separately, and output is typically several times more expensive — for chat-style workloads with long answers, output cost usually dominates the bill.
| Model | Provider | Input / 1M | Output / 1M | Tier |
|---|---|---|---|---|
| Claude Fable 5.1 | Anthropic | $10 | $50 | premium |
| Claude Opus 5.5 | Anthropic | $4 | $20 | premium |
| Claude Sonnet 5.5 | Anthropic | $2 | $10 | mid |
| Claude Haiku 4.5 | Anthropic | $1 | $5 | budget |
| GPT-6 Astra | OpenAI | $10 | $50 | premium |
| GPT-6.1 Sol | OpenAI | $2 | $10 | mid |
| GPT-6 Luna | OpenAI | $0.10 | $0.50 | budget |
| GPT-5.5 | OpenAI | $5 | $30 | premium |
| GPT-5.4 Mini | OpenAI | $0.75 | $4.50 | mid |
| Gemini 3.8 Flash | $0.75 | $3.75 | mid | |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | budget | |
| Gemini 3.1 Pro Preview | $2 | $12 | premium | |
| DeepSeek V4.1 Flash | DeepSeek | $0.30 | $1.20 | budget |
| DeepSeek V4 Pro | DeepSeek | $1.32 | $3.96 | mid |
| Grok 4.7 | xAI | $2 | $6 | premium |
| Grok 4.3 | xAI | $1.25 | $2.50 | mid |
| Mistral Medium 3.5 | Mistral | $1.50 | $7.50 | premium |
| Mistral Small 4 | Mistral | $0.15 | $0.60 | budget |
| Mistral Large 3 | Mistral | $0.50 | $1.50 | mid |
| Muse Spark 1.3 | Meta | $1.25 | $4.25 | mid |
| Qwen3.8-Max | Qwen | $2 | $6 | premium |
| Qwen3.8-Flash | Qwen | $0.15 | $0.47 | budget |
| GPT-OSS 120B (Groq) | Groq | $0.15 | $0.60 | budget |
Sources: official pricing pages of OpenAI, Anthropic, Google AI, xAI, DeepSeek, Mistral, and Groq. Lineup verified 2026-10-03; data refreshed October 3, 2026. Prices in USD; caching, batch, and long-context tiers can change the effective rate.
Which provider is cheapest per million tokens?
Among current frontier-lab models, GPT-6 Luna is the cheapest in this dataset at $0.10 input / $0.50 output per million tokens. Open-weight models served via Groq undercut most closed models for high-volume pipelines. At the premium end, Claude Fable 5.1 tops the table at $10 input / $50 output. The gap between the cheapest and the most expensive model is well over two orders of magnitude, which is why model routing — sending easy requests to budget tiers — is usually the single largest cost lever.
How do I estimate a real monthly bill from these rates?
- Estimate average input and output tokens per request. As a rule of thumb, 750 English words is roughly 1,000 tokens on GPT-family tokenizers, so a 1,000-word prompt is closer to 1,300 tokens.
- Multiply by requests per day and by 30, then apply the per-million rates from the table separately for input and output.
- Account for prompt caching where supported — repeated system prompts and context can bill at a fraction of the input rate.
- Use the estimator above to run these numbers per model and compare providers side by side; the calculation runs client-side.