Skip to content

Prompt Cache ROI Calculator

Compute prompt-caching breakeven and monthly savings from your static prefix, query tokens and call volume for every current model with cache pricing.

Prompt Cache ROI Calculator

Prompt structure

Fraction of calls landing on a warm cache. 100% = always warm; 0% = every call writes fresh.

Selected model

Input$10.00 / 1M
Cache write$12.50 / 1M
Cached read$0.2500 / 1M
Output$50.00 / 1M

Cache read at 3% of input; first write 1.25x input.

Monthly savings $3798, 50.6 percent

No-cache / month
$7500
Cached / month
$3702
Monthly savings
$3798
50.6%
Breakeven calls
1
~0.0h spread
Calls / month
60,000
ModelNo-cache / moCached / moSavingsSavings %Breakeven
Claude Fable 5.1
Anthropic
$7500$3702$379850.6%1 calls
Claude Opus 5.5
Anthropic
$3000$1522$147849.3%1 calls
Claude Sonnet 5.5
Anthropic
$1500$802$69846.6%1 calls
Claude Haiku 4.5
Anthropic
$750$401$34946.6%1 calls
GPT-6 Astra
OpenAI
$7500$4008$349246.6%1 calls
GPT-6.1 Sol
OpenAI
$1500$761$73949.3%1 calls
GPT-6 Sol
OpenAI
$1500$802$69846.6%1 calls
GPT-6 Luna
OpenAI
$75.00$40.08$34.9246.6%1 calls
GPT-5.6 Sol
OpenAI
$3000$1603$139746.6%1 calls
GPT-5.6 Terra
OpenAI
$1596$898$69843.8%1 calls
GPT-5.6 Luna
OpenAI
$160$89.76$69.8443.8%1 calls
GPT-5.5
OpenAI
$3990$2154$183646.0%—
GPT-5.4
OpenAI
$1995$1077$91846.0%—
GPT-5.4 Mini
OpenAI
$599$323$27546.0%1 calls
Gemini 3.8 Flash
Google
$563$287$27549.0%1 calls
Gemini 3.7 Flash
Google
$563$287$27549.0%1 calls
Gemini 3.6 Flash
Google
$563$287$27549.0%1 calls
Gemini 3.5 Flash
Google
$1197$646$55146.0%1 calls
Gemini 3.5 Flash-Lite
Google
$273$163$11040.4%—
Gemini 3.1 Pro Preview
Google
$1596$862$73446.0%—
DeepSeek V4.1 Flash
DeepSeek
$211$90.65$12057.0%—
DeepSeek V4 Pro
DeepSeek
$863$343$52160.3%—
Grok 4.7
xAI
$1308$696$61246.8%—
Grok 4.6
xAI
$1308$696$61246.8%—
Grok 4.5
xAI
$1308$614$69453.0%—
Grok 4.3
xAI
$758$329$42856.6%—
Grok 4.20
xAI
$758$329$42856.6%—
Grok Build 0.1
xAI
$606$280$32653.9%—
Muse Spark 1.3
Meta
$842$393$44953.3%—
Qwen3.8-Max
Qwen
$1308$630$67851.8%1 calls
Qwen3.8-Flash
Qwen
$99.06$47.99$51.0751.6%1 calls
How the math works

Without caching, every call pays the full input rate on (static + query) tokens. With caching, the first call writes the static prefix at that model's cache-write rate, and every subsequent warm read swaps the input rate for that model's cached-read rate — both taken from the table above, per model, not from a flat per-vendor discount. Hit rate scales how many of your monthly calls land on a warm prefix.

The modelled write premium is the 5-minute cache (1.25x input on vendors that charge one). Anthropic's 1-hour cache costs 2x input to write and is not modelled here. A prefix shorter than a published vendor minimum (Anthropic 1,024, OpenAI 1,024 tokens) is never cached and shows zero savings; vendors that publish no minimum are not gated here, so check their docs before assuming a short prefix caches.

What this tool does

The Prompt Cache ROI Calculator shows whether prompt caching pays off for your workload. Enter your static prefix size, per-query tokens, calls per day, and model, and it computes the caching breakeven and projected monthly savings versus no cache for every current Anthropic, OpenAI, Google and DeepSeek model with a published cache-read price, using list prices read from the vendor pricing pages on the date shown in the tool.

Updated . Provided as is. Check the output before you rely on it in production.

How to use Prompt Cache ROI Calculator

  1. 1

    Describe your prompt structure

    Enter the static prefix size (system, few-shot, retrieved context that stays constant), per-call query tokens, average output tokens, and how many calls per day you expect. Each field accepts plain integers.

  2. 2

    Set the cache hit rate

    The slider controls what fraction of monthly calls land on a warm prefix. Real-world hit rates depend on TTL and traffic shape — try 80-95% for steady traffic and lower values for spiky workloads.

  3. 3

    Pick a model

    Each model card shows input, cached read, cache write where the vendor charges one, and output prices per million tokens. The list is every current model with a published cache-read price, read from the vendor pricing pages on the date shown.

  4. 4

    Read savings and breakeven

    Headline tiles show no-cache vs cached monthly cost, dollar savings, and how many warm reads pay back the first cache write. The full comparison table runs the same calculation across every embedded model.

Questions and answers

Where does the pricing come from?
The public list prices on each vendor's pricing page, per million tokens, for every current model that publishes a cache-read price, from Anthropic, OpenAI, Google, DeepSeek, xAI, Qwen and Meta. The verification date is shown on the page; rows cover input, cached read, cache write where charged, and output.
How is breakeven calculated?
Breakeven is the number of warm calls needed to repay one cache write: the write premium on the static prefix divided by the saving each cached read makes over uncached input. Models with a write charge, such as Anthropic's and the newest OpenAI models at 1.25x input, need a few reads; models without one break even at once.
What does cache hit rate mean here?
The fraction of your monthly calls that land on a warm prefix. Cold calls pay the full write rate; warm calls pay the cached read rate. Adjust the slider to see how your TTL and traffic shape change projected savings.
Does it send my data to a server?
No. All math happens in your browser using the embedded pricing table. Token counts, cache assumptions, and cost projections never leave the page.
Why do savings differ between vendors?
Cache reads cost about 10% of input on Anthropic and OpenAI (2.5% on Claude Fable 5.1, 5% on GPT-6.1 Sol) and 2–3% on DeepSeek. Anthropic and current OpenAI models such as GPT-6 Astra charge a 1.25x write premium on the first call. Heavy reuse pays it back quickly; a single warm read may not.
For AI agents: how to call this tool

Machine-readable contract, endpoints and examples. Humans can ignore this section.

Best Path For Builders

Browser workflow

Runs instantly in the browser with private local processing and copy/export-ready output.

Browser Workflow

This tool is optimized for instant in-browser execution with local data handling. Run it here and copy/export the output directly.

/prompt-cache-roi-calculator/

For automation planning, fetch the canonical contract at /api/tool/prompt-cache-roi-calculator.json.