Skip to content

AI Cost Estimator

Estimate monthly or annual LLM API costs per model from workload presets, request volume, token counts and prompt-cache hit rate.

AI Cost Estimator

Use case preset

Workload configuration

Cached portion is billed at each model's published cached input rate

Lowest monthly cost: GPT-OSS 20B (Groq) at $5.63

Total tokens / month
30.00M
Cheapest (Monthly)
$5.63
GPT-OSS 20B (Groq)
Most expensive (Monthly)
$3,150.00
GPT-5.4 Pro
Average (Monthly)
$333.46
Across 66 models

Monthly cost comparison

GPT-OSS 20B (Groq)Groq
$5.63
GPT-5-nanoOpenAI
$6.75
GPT-4.1-nanoOpenAI
$7.50
Gemini 2.5 Flash-LiteGoogle
$7.50
GPT-6 LunaOpenAI
$9.00
Qwen3.8-FlashQwen
$9.30
Mistral Small 4Mistral
$11.25
GPT-OSS 120B (Groq)Groq
$11.25
Codestral 25.08Mistral
$18.00
GPT-5.6 LunaOpenAI
$21.00
GPT-5.4 NanoOpenAI
$21.75
DeepSeek V4.1 FlashDeepSeek
$22.50
GPT-4.1-miniOpenAI
$30.00
Mistral Large 3Mistral
$30.00
GPT-5-miniOpenAI
$33.75
Gemini 3.5 Flash-LiteGoogle
$42.00
Gemini 2.5 FlashGoogle
$42.00
Grok Build 0.1xAI
$45.00
Grok 4.3xAI
$56.25
Grok 4.20xAI
$56.25
Gemini 3.8 FlashGoogle
$67.50
Gemini 3.7 FlashGoogle
$67.50
Gemini 3.6 FlashGoogle
$67.50
GPT-5.4 MiniOpenAI
$78.75
DeepSeek V4 ProDeepSeek
$79.20
o4-miniOpenAI
$82.50
o3-miniOpenAI
$82.50
Muse Spark 1.3Meta
$82.50
Claude Haiku 4.5Anthropic
$90.00
Grok 4.7xAI
$120.00
Grok 4.6xAI
$120.00
Grok 4.5xAI
$120.00
Qwen3.8-MaxQwen
$120.00
Mistral Medium 3.5Mistral
$135.00
o3OpenAI
$150.00
GPT-4.1OpenAI
$150.00
Gemini 3.5 FlashGoogle
$157.50
GPT-5.1OpenAI
$168.75
GPT-5OpenAI
$168.75
Gemini 2.5 ProGoogle
$168.75
Claude Sonnet 5.5Anthropic
$180.00
Claude Sonnet 5Anthropic
$180.00
GPT-6.1 SolOpenAI
$180.00
GPT-6 SolOpenAI
$180.00
GPT-5.6 TerraOpenAI
$210.00
Gemini 3.1 Pro PreviewGoogle
$210.00
GPT-5.3-CodexOpenAI
$236.25
GPT-5.2OpenAI
$236.25
GPT-5.4OpenAI
$262.50
Claude Sonnet 4.6Anthropic
$270.00
Claude Sonnet 4.5Anthropic
$270.00
Claude Opus 5.5Anthropic
$360.00
GPT-5.6 SolOpenAI
$360.00
Claude Opus 5Anthropic
$450.00
Claude Opus 4.8Anthropic
$450.00
Claude Opus 4.7Anthropic
$450.00
Claude Opus 4.6Anthropic
$450.00
Claude Opus 4.5Anthropic
$450.00
GPT-5.5OpenAI
$525.00
Claude Fable 5.1Anthropic
$900.00
Claude Fable 5Anthropic
$900.00
GPT-6 AstraOpenAI
$900.00
o3-proOpenAI
$1,500.00
GPT-5.2 ProOpenAI
$2,835.00
GPT-5.5 ProOpenAI
$3,150.00
GPT-5.4 ProOpenAI
$3,150.00

Detailed cost breakdown

GPT-OSS 20B (Groq)CheapestGroq$0.037$0.150$0.188$5.63
GPT-5-nanoOpenAI$0.025$0.200$0.225$6.75
GPT-4.1-nanoOpenAI$0.050$0.200$0.250$7.50
Gemini 2.5 Flash-LiteGoogle$0.050$0.200$0.250$7.50
GPT-6 LunaOpenAI$0.050$0.250$0.300$9.00
Qwen3.8-FlashQwen$0.075$0.235$0.310$9.30
Mistral Small 4Mistral$0.075$0.300$0.375$11.25
GPT-OSS 120B (Groq)Groq$0.075$0.300$0.375$11.25
Codestral 25.08Mistral$0.150$0.450$0.600$18.00
GPT-5.6 LunaOpenAI$0.100$0.600$0.700$21.00
GPT-5.4 NanoOpenAI$0.100$0.625$0.725$21.75
DeepSeek V4.1 FlashDeepSeek$0.150$0.600$0.750$22.50
GPT-4.1-miniOpenAI$0.200$0.800$1.00$30.00
Mistral Large 3Mistral$0.250$0.750$1.00$30.00
GPT-5-miniOpenAI$0.125$1.00$1.13$33.75
Gemini 3.5 Flash-LiteGoogle$0.150$1.25$1.40$42.00
Gemini 2.5 FlashGoogle$0.150$1.25$1.40$42.00
Grok Build 0.1xAI$0.500$1.00$1.50$45.00
Grok 4.3xAI$0.625$1.25$1.88$56.25
Grok 4.20xAI$0.625$1.25$1.88$56.25
Gemini 3.8 FlashGoogle$0.375$1.88$2.25$67.50
Gemini 3.7 FlashGoogle$0.375$1.88$2.25$67.50
Gemini 3.6 FlashGoogle$0.375$1.88$2.25$67.50
GPT-5.4 MiniOpenAI$0.375$2.25$2.63$78.75
DeepSeek V4 ProDeepSeek$0.660$1.98$2.64$79.20
o4-miniOpenAI$0.550$2.20$2.75$82.50
o3-miniOpenAI$0.550$2.20$2.75$82.50
Muse Spark 1.3Meta$0.625$2.13$2.75$82.50
Claude Haiku 4.5Anthropic$0.500$2.50$3.00$90.00
Grok 4.7xAI$1.00$3.00$4.00$120.00
Grok 4.6xAI$1.00$3.00$4.00$120.00
Grok 4.5xAI$1.00$3.00$4.00$120.00
Qwen3.8-MaxQwen$1.00$3.00$4.00$120.00
Mistral Medium 3.5Mistral$0.750$3.75$4.50$135.00
o3OpenAI$1.00$4.00$5.00$150.00
GPT-4.1OpenAI$1.00$4.00$5.00$150.00
Gemini 3.5 FlashGoogle$0.750$4.50$5.25$157.50
GPT-5.1OpenAI$0.625$5.00$5.63$168.75
GPT-5OpenAI$0.625$5.00$5.63$168.75
Gemini 2.5 ProGoogle$0.625$5.00$5.63$168.75
Claude Sonnet 5.5Anthropic$1.00$5.00$6.00$180.00
Claude Sonnet 5Anthropic$1.00$5.00$6.00$180.00
GPT-6.1 SolOpenAI$1.00$5.00$6.00$180.00
GPT-6 SolOpenAI$1.00$5.00$6.00$180.00
GPT-5.6 TerraOpenAI$1.00$6.00$7.00$210.00
Gemini 3.1 Pro PreviewGoogle$1.00$6.00$7.00$210.00
GPT-5.3-CodexOpenAI$0.875$7.00$7.88$236.25
GPT-5.2OpenAI$0.875$7.00$7.88$236.25
GPT-5.4OpenAI$1.25$7.50$8.75$262.50
Claude Sonnet 4.6Anthropic$1.50$7.50$9.00$270.00
Claude Sonnet 4.5Anthropic$1.50$7.50$9.00$270.00
Claude Opus 5.5Anthropic$2.00$10.00$12.00$360.00
GPT-5.6 SolOpenAI$2.00$10.00$12.00$360.00
Claude Opus 5Anthropic$2.50$12.50$15.00$450.00
Claude Opus 4.8Anthropic$2.50$12.50$15.00$450.00
Claude Opus 4.7Anthropic$2.50$12.50$15.00$450.00
Claude Opus 4.6Anthropic$2.50$12.50$15.00$450.00
Claude Opus 4.5Anthropic$2.50$12.50$15.00$450.00
GPT-5.5OpenAI$2.50$15.00$17.50$525.00
Claude Fable 5.1Anthropic$5.00$25.00$30.00$900.00
Claude Fable 5Anthropic$5.00$25.00$30.00$900.00
GPT-6 AstraOpenAI$5.00$25.00$30.00$900.00
o3-proOpenAI$10.00$40.00$50.00$1,500.00
GPT-5.2 ProOpenAI$10.50$84.00$94.50$2,835.00
GPT-5.5 ProOpenAI$15.00$90.00$105.00$3,150.00
GPT-5.4 ProPriciestOpenAI$15.00$90.00$105.00$3,150.00

Cost vs. capability insights

Compared to the cheapest option (GPT-OSS 20B (Groq)):

GPT-5-nano+20% cost|cost-efficient|+$1.13/mo
GPT-4.1-nano+33% cost|cost-efficient|+$1.88/mo
Gemini 2.5 Flash-Lite+33% cost|cost-efficient|+$1.88/mo
GPT-6 Luna+60% cost|cost-efficient|+$3.38/mo
Qwen3.8-Flash+65% cost|cost-efficient|+$3.68/mo

Workload summary

Requests/day1.0K
Tokens/request500 in + 500 out
Tokens/day1.00M
Tokens/month30.00M

Pricing data last updated: October 3, 2026. Prices per 1M tokens. Estimates are approximate and may vary with actual usage patterns.

What this tool does

LLM API pricing in 2026 spans from $0.10 input / $0.50 output per million tokens (GPT-6 Luna) up to $10 input / $50 output (Claude Fable 5.1), across 23 current models from 9 providers. Flagship rates: GPT-6 Astra $10 input / $50 output, Claude Fable 5.1 $10 input / $50 output, Gemini 3.1 Pro Preview $2 input / $12 output. Every price in the table below is verified against the vendor's official pricing page. This estimator turns those per-token rates into monthly workload costs.

Updated . Provided as is. Check the output before you rely on it in production.

How to use AI Cost Estimator

  1. 1

    Select your AI models

    Choose the models you use, such as GPT, Claude, Gemini, Grok or DeepSeek. The tool displays current pricing per 1M input and output tokens. Add multiple models if you're comparing or using a mix.

  2. 2

    Estimate token counts

    Input your expected usage: requests per day, average input tokens (rule of thumb: 1 token is about 4 characters of English), average output tokens, and billing days per month.

  3. 3

    Set a cache hit rate

    Enter the share of input tokens you expect to land on a warm prompt cache. That share is repriced at each model's published cached-input rate; models with no published cached rate stay at the full input price.

  4. 4

    Calculate total and per-request costs

    The table gives daily, monthly and annual cost per model, split into input and output, and marks the cheapest and priciest rows. Switch between the monthly and annual view to compare against your budget.

Questions and answers

What is AI Cost Estimator?
LLM API cost is set by tokens: each request bills its input and output at the model's per-million rate. This estimator starts from a workload preset such as a support chatbot or RAG pipeline, applies volume and a cache-hit share at each model's cached rate, and ranks models by monthly or annual cost.
Does AI Cost Estimator store or send my data?
No. All processing happens entirely in your browser. Your workload data never leaves your device — nothing is sent to any server.
How accurate are the cost estimates?
Cost estimates use current published API pricing from each provider. Actual costs may vary based on factors like caching, batching discounts, and token count variations, but the estimates give you a reliable baseline for budgeting.
For AI agents: how to call this tool

Machine-readable contract, endpoints and examples. Humans can ignore this section.

Best Path For Builders

Browser workflow

Runs instantly in the browser with private local processing and copy/export-ready output.

Browser Workflow

This tool is optimized for instant in-browser execution with local data handling. Run it here and copy/export the output directly.

/ai-cost-estimator/

For automation planning, fetch the canonical contract at /api/tool/ai-cost-estimator.json.

How much do LLM APIs cost per million tokens across providers in 2026?

Current list prices per 1 million tokens for 23 models across Anthropic, OpenAI, Google, DeepSeek, xAI, Mistral, Meta, Qwen, Groq. Input and output tokens are billed separately, and output is typically several times more expensive — for chat-style workloads with long answers, output cost usually dominates the bill.

Model Provider Input / 1M Output / 1M Tier
Claude Fable 5.1 Anthropic $10 $50 premium
Claude Opus 5.5 Anthropic $4 $20 premium
Claude Sonnet 5.5 Anthropic $2 $10 mid
Claude Haiku 4.5 Anthropic $1 $5 budget
GPT-6 Astra OpenAI $10 $50 premium
GPT-6.1 Sol OpenAI $2 $10 mid
GPT-6 Luna OpenAI $0.10 $0.50 budget
GPT-5.5 OpenAI $5 $30 premium
GPT-5.4 Mini OpenAI $0.75 $4.50 mid
Gemini 3.8 Flash Google $0.75 $3.75 mid
Gemini 3.5 Flash-Lite Google $0.30 $2.50 budget
Gemini 3.1 Pro Preview Google $2 $12 premium
DeepSeek V4.1 Flash DeepSeek $0.30 $1.20 budget
DeepSeek V4 Pro DeepSeek $1.32 $3.96 mid
Grok 4.7 xAI $2 $6 premium
Grok 4.3 xAI $1.25 $2.50 mid
Mistral Medium 3.5 Mistral $1.50 $7.50 premium
Mistral Small 4 Mistral $0.15 $0.60 budget
Mistral Large 3 Mistral $0.50 $1.50 mid
Muse Spark 1.3 Meta $1.25 $4.25 mid
Qwen3.8-Max Qwen $2 $6 premium
Qwen3.8-Flash Qwen $0.15 $0.47 budget
GPT-OSS 120B (Groq) Groq $0.15 $0.60 budget

Sources: official pricing pages of OpenAI, Anthropic, Google AI, xAI, DeepSeek, Mistral, and Groq. Lineup verified 2026-10-03; data refreshed October 3, 2026. Prices in USD; caching, batch, and long-context tiers can change the effective rate.

Which provider is cheapest per million tokens?

Among current frontier-lab models, GPT-6 Luna is the cheapest in this dataset at $0.10 input / $0.50 output per million tokens. Open-weight models served via Groq undercut most closed models for high-volume pipelines. At the premium end, Claude Fable 5.1 tops the table at $10 input / $50 output. The gap between the cheapest and the most expensive model is well over two orders of magnitude, which is why model routing — sending easy requests to budget tiers — is usually the single largest cost lever.

How do I estimate a real monthly bill from these rates?

  1. Estimate average input and output tokens per request. As a rule of thumb, 750 English words is roughly 1,000 tokens on GPT-family tokenizers, so a 1,000-word prompt is closer to 1,300 tokens.
  2. Multiply by requests per day and by 30, then apply the per-million rates from the table separately for input and output.
  3. Account for prompt caching where supported — repeated system prompts and context can bill at a fraction of the input rate.
  4. Use the estimator above to run these numbers per model and compare providers side by side; the calculation runs client-side.