Prompt Cache ROI Calculator
Compute prompt-caching breakeven and monthly savings from your static prefix, query tokens and call volume for every current model with cache pricing.
Prompt Cache ROI Calculator
Prompt structure
Fraction of calls landing on a warm cache. 100% = always warm; 0% = every call writes fresh.
Selected model
Cache read at 3% of input; first write 1.25x input.
Monthly savings $3798, 50.6 percent
| Model | No-cache / mo | Cached / mo | Savings | Savings % | Breakeven |
|---|---|---|---|---|---|
Claude Fable 5.1 Anthropic | $7500 | $3702 | $3798 | 50.6% | 1 calls |
Claude Opus 5.5 Anthropic | $3000 | $1522 | $1478 | 49.3% | 1 calls |
Claude Sonnet 5.5 Anthropic | $1500 | $802 | $698 | 46.6% | 1 calls |
Claude Haiku 4.5 Anthropic | $750 | $401 | $349 | 46.6% | 1 calls |
GPT-6 Astra OpenAI | $7500 | $4008 | $3492 | 46.6% | 1 calls |
GPT-6.1 Sol OpenAI | $1500 | $761 | $739 | 49.3% | 1 calls |
GPT-6 Sol OpenAI | $1500 | $802 | $698 | 46.6% | 1 calls |
GPT-6 Luna OpenAI | $75.00 | $40.08 | $34.92 | 46.6% | 1 calls |
GPT-5.6 Sol OpenAI | $3000 | $1603 | $1397 | 46.6% | 1 calls |
GPT-5.6 Terra OpenAI | $1596 | $898 | $698 | 43.8% | 1 calls |
GPT-5.6 Luna OpenAI | $160 | $89.76 | $69.84 | 43.8% | 1 calls |
GPT-5.5 OpenAI | $3990 | $2154 | $1836 | 46.0% | — |
GPT-5.4 OpenAI | $1995 | $1077 | $918 | 46.0% | — |
GPT-5.4 Mini OpenAI | $599 | $323 | $275 | 46.0% | 1 calls |
Gemini 3.8 Flash Google | $563 | $287 | $275 | 49.0% | 1 calls |
Gemini 3.7 Flash Google | $563 | $287 | $275 | 49.0% | 1 calls |
Gemini 3.6 Flash Google | $563 | $287 | $275 | 49.0% | 1 calls |
Gemini 3.5 Flash Google | $1197 | $646 | $551 | 46.0% | 1 calls |
Gemini 3.5 Flash-Lite Google | $273 | $163 | $110 | 40.4% | — |
Gemini 3.1 Pro Preview Google | $1596 | $862 | $734 | 46.0% | — |
DeepSeek V4.1 Flash DeepSeek | $211 | $90.65 | $120 | 57.0% | — |
DeepSeek V4 Pro DeepSeek | $863 | $343 | $521 | 60.3% | — |
Grok 4.7 xAI | $1308 | $696 | $612 | 46.8% | — |
Grok 4.6 xAI | $1308 | $696 | $612 | 46.8% | — |
Grok 4.5 xAI | $1308 | $614 | $694 | 53.0% | — |
Grok 4.3 xAI | $758 | $329 | $428 | 56.6% | — |
Grok 4.20 xAI | $758 | $329 | $428 | 56.6% | — |
Grok Build 0.1 xAI | $606 | $280 | $326 | 53.9% | — |
Muse Spark 1.3 Meta | $842 | $393 | $449 | 53.3% | — |
Qwen3.8-Max Qwen | $1308 | $630 | $678 | 51.8% | 1 calls |
Qwen3.8-Flash Qwen | $99.06 | $47.99 | $51.07 | 51.6% | 1 calls |
Without caching, every call pays the full input rate on (static + query) tokens. With caching, the first call writes the static prefix at that model's cache-write rate, and every subsequent warm read swaps the input rate for that model's cached-read rate — both taken from the table above, per model, not from a flat per-vendor discount. Hit rate scales how many of your monthly calls land on a warm prefix.
The modelled write premium is the 5-minute cache (1.25x input on vendors that charge one). Anthropic's 1-hour cache costs 2x input to write and is not modelled here. A prefix shorter than a published vendor minimum (Anthropic 1,024, OpenAI 1,024 tokens) is never cached and shows zero savings; vendors that publish no minimum are not gated here, so check their docs before assuming a short prefix caches.
What this tool does
The Prompt Cache ROI Calculator shows whether prompt caching pays off for your workload. Enter your static prefix size, per-query tokens, calls per day, and model, and it computes the caching breakeven and projected monthly savings versus no cache for every current Anthropic, OpenAI, Google and DeepSeek model with a published cache-read price, using list prices read from the vendor pricing pages on the date shown in the tool.
Updated . Provided as is. Check the output before you rely on it in production.
How to use Prompt Cache ROI Calculator
- 1
Describe your prompt structure
Enter the static prefix size (system, few-shot, retrieved context that stays constant), per-call query tokens, average output tokens, and how many calls per day you expect. Each field accepts plain integers.
- 2
Set the cache hit rate
The slider controls what fraction of monthly calls land on a warm prefix. Real-world hit rates depend on TTL and traffic shape — try 80-95% for steady traffic and lower values for spiky workloads.
- 3
Pick a model
Each model card shows input, cached read, cache write where the vendor charges one, and output prices per million tokens. The list is every current model with a published cache-read price, read from the vendor pricing pages on the date shown.
- 4
Read savings and breakeven
Headline tiles show no-cache vs cached monthly cost, dollar savings, and how many warm reads pay back the first cache write. The full comparison table runs the same calculation across every embedded model.
Questions and answers
Where does the pricing come from?
How is breakeven calculated?
What does cache hit rate mean here?
Does it send my data to a server?
Why do savings differ between vendors?
For AI agents: how to call this tool
Machine-readable contract, endpoints and examples. Humans can ignore this section.
Best Path For Builders
Browser workflow
Runs instantly in the browser with private local processing and copy/export-ready output.
Browser Workflow
This tool is optimized for instant in-browser execution with local data handling. Run it here and copy/export the output directly.
/prompt-cache-roi-calculator/
For automation planning, fetch the canonical contract at /api/tool/prompt-cache-roi-calculator.json.