LLM Latency Estimator
Estimate time to first token, generation time and total latency for current LLMs from your token counts, with a UX pattern for each wait.
LLM Latency Estimator
Model comparison
| UX hint | ||||||
|---|---|---|---|---|---|---|
| claude-haiku-4-5 | Anthropic | 260ms | 556ms | 816ms | 180 | Spinner |
| gpt-6-luna | OpenAI | 260ms | 556ms | 816ms | 180 | Spinner |
| gemini-3.5-flash-lite | 260ms | 556ms | 816ms | 180 | Spinner | |
| deepseek-flash | DeepSeek | 260ms | 556ms | 816ms | 180 | Spinner |
| mistral-small-2603 | Mistral | 260ms | 556ms | 816ms | 180 | Spinner |
| qwen3.8-flash | Qwen | 260ms | 556ms | 816ms | 180 | Spinner |
| openai/gpt-oss-120b | Groq | 260ms | 556ms | 816ms | 180 | Spinner |
| claude-sonnet-5-5 | Anthropic | 510ms | 833ms | 1.3s | 120 | Spinner |
| gpt-6.1-sol | OpenAI | 510ms | 833ms | 1.3s | 120 | Spinner |
| gpt-5.4-mini | OpenAI | 510ms | 833ms | 1.3s | 120 | Spinner |
| gemini-3.8-flash | 510ms | 833ms | 1.3s | 120 | Spinner | |
| deepseek-v4-pro | DeepSeek | 510ms | 833ms | 1.3s | 120 | Spinner |
| grok-4.3 | xAI | 510ms | 833ms | 1.3s | 120 | Spinner |
| mistral-large-2512 | Mistral | 510ms | 833ms | 1.3s | 120 | Spinner |
| muse-spark-1.3 | Meta | 510ms | 833ms | 1.3s | 120 | Spinner |
| claude-fable-5-1 | Anthropic | 910ms | 1.4s | 2.3s | 70 | Stream |
| claude-opus-5-5 | Anthropic | 910ms | 1.4s | 2.3s | 70 | Stream |
| gpt-6-astra | OpenAI | 910ms | 1.4s | 2.3s | 70 | Stream |
| gpt-5.5 | OpenAI | 910ms | 1.4s | 2.3s | 70 | Stream |
| gemini-3.1-pro-preview | 910ms | 1.4s | 2.3s | 70 | Stream | |
| grok-4.7 | xAI | 910ms | 1.4s | 2.3s | 70 | Stream |
| mistral-medium-3504 | Mistral | 910ms | 1.4s | 2.3s | 70 | Stream |
| qwen3.8-max | Qwen | 910ms | 1.4s | 2.3s | 70 | Stream |
Estimates based on typical API latencies. Actual performance varies by load, region, prompt complexity, and provider infrastructure. TTFB includes additional 10ms for 500 input token processing overhead.
Updated . Provided as is. Check the output before you rely on it in production.
How to use LLM Latency Estimator
- 1
Select a model
Choose one of the current models offered across OpenAI, Anthropic, Google, xAI, DeepSeek, Mistral and Groq. Speed assumptions follow each model's price tier.
- 2
Enter token counts
Set the expected input token count (your prompt) and output token count (the response), or use a quick preset for common scenarios.
- 3
Read the latency estimate
See the estimated time to first token, generation time and total latency. The UX badge suggests whether to stream, show a spinner or run a background job.
- 4
Compare across models
The comparison table runs the same input and output sizes across every model, sorted from fastest to slowest.
Questions and answers
What is LLM Latency Estimator?
How accurate are the latency estimates?
Does it send data to a server?
What UX recommendations does it provide?
For AI agents: how to call this tool
Machine-readable contract, endpoints and examples. Humans can ignore this section.
Best Path For Builders
Browser workflow
Runs instantly in the browser with private local processing and copy/export-ready output.
Browser Workflow
This tool is optimized for instant in-browser execution with local data handling. Run it here and copy/export the output directly.
/llm-latency-estimator/
For automation planning, fetch the canonical contract at /api/tool/llm-latency-estimator.json.