Skip to content

LLM Token Counter

Count tokens for GPT, Claude, Gemini, Grok, DeepSeek, Mistral, Qwen and Muse models and estimate API cost. Also available as a REST API.

LLM Token Counter

265 characters, 48 words, 3 lines

Characters
265
Words
48
Lines
3

Token counts

GPT
~67
estimate
Claude
~76
estimate
Gemini
~67
estimate
DeepSeek
~67
estimate
Mistral
~67
estimate
Grok
~67
estimate
Muse
~67
estimate
Qwen
~67
estimate

OpenAI GPT and o-series counts come from the real o200k_base tokenizer. Vendors that publish no browser tokenizer are estimated from average characters per token, so those numbers vary with content. Loading the tokenizer…

Cost (for this text as input)

TokensMethod
GPT-5-nano~67estimate$0.0000$0.0000
GPT-OSS 20B (Groq)~67estimate$0.0000$0.0000
GPT-6 Luna~67estimate$0.0000$0.0000
GPT-4.1-nano~67estimate$0.0000$0.0000
Gemini 2.5 Flash-Lite~67estimate$0.0000$0.0000
Mistral Small 4~67estimate$0.0000$0.0000
Qwen3.8-Flash~67estimate$0.0000$0.0000
GPT-OSS 120B (Groq)~67estimate$0.0000$0.0000
GPT-5.6 Luna~67estimate$0.0000$0.0000
GPT-5.4 Nano~67estimate$0.0000$0.0000
GPT-5-mini~67estimate$0.0000$0.0001
Gemini 3.5 Flash-Lite~67estimate$0.0000$0.0002
Gemini 2.5 Flash~67estimate$0.0000$0.0002
DeepSeek V4.1 Flash~67estimate$0.0000$0.0000
Codestral 25.08~67estimate$0.0000$0.0000
GPT-4.1-mini~67estimate$0.0000$0.0001
Mistral Large 3~67estimate$0.0000$0.0001
GPT-5.4 Mini~67estimate$0.0000$0.0003
Gemini 3.8 Flash~67estimate$0.0000$0.0003
Gemini 3.7 Flash~67estimate$0.0000$0.0003
Gemini 3.6 Flash~67estimate$0.0000$0.0003
Grok Build 0.1~67estimate$0.0000$0.0001
o4-mini~67estimate$0.0000$0.0003
o3-mini~67estimate$0.0000$0.0003
Claude Haiku 4.5~76estimate$0.0000$0.0004
GPT-5.1~67estimate$0.0000$0.0007
GPT-5~67estimate$0.0000$0.0007
Gemini 2.5 Pro~67estimate$0.0000$0.0007
Grok 4.3~67estimate$0.0000$0.0002
Grok 4.20~67estimate$0.0000$0.0002
Muse Spark 1.3~67estimate$0.0000$0.0003
DeepSeek V4 Pro~67estimate$0.0000$0.0003
Gemini 3.5 Flash~67estimate$0.0001$0.0006
Mistral Medium 3.5~67estimate$0.0001$0.0005
GPT-5.3-Codex~67estimate$0.0001$0.0009
GPT-5.2~67estimate$0.0001$0.0009
GPT-6.1 Sol~67estimate$0.0001$0.0007
GPT-6 Sol~67estimate$0.0001$0.0007
GPT-5.6 Terra~67estimate$0.0001$0.0008
o3~67estimate$0.0001$0.0005
GPT-4.1~67estimate$0.0001$0.0005
Gemini 3.1 Pro Preview~67estimate$0.0001$0.0008
Grok 4.7~67estimate$0.0001$0.0004
Grok 4.6~67estimate$0.0001$0.0004
Grok 4.5~67estimate$0.0001$0.0004
Qwen3.8-Max~67estimate$0.0001$0.0004
Claude Sonnet 5.5~76estimate$0.0002$0.0008
Claude Sonnet 5~76estimate$0.0002$0.0008
GPT-5.4~67estimate$0.0002$0.0010
Claude Sonnet 4.6~76estimate$0.0002$0.0011
Claude Sonnet 4.5~76estimate$0.0002$0.0011
GPT-5.6 Sol~67estimate$0.0003$0.0013
Claude Opus 5.5~76estimate$0.0003$0.0015
GPT-5.5~67estimate$0.0003$0.0020
Claude Opus 5~76estimate$0.0004$0.0019
Claude Opus 4.8~76estimate$0.0004$0.0019
Claude Opus 4.7~76estimate$0.0004$0.0019
Claude Opus 4.6~76estimate$0.0004$0.0019
Claude Opus 4.5~76estimate$0.0004$0.0019
GPT-6 Astra~67estimate$0.0007$0.0034
Claude Fable 5.1~76estimate$0.0008$0.0038
Claude Fable 5~76estimate$0.0008$0.0038
o3-pro~67estimate$0.0013$0.0054
GPT-5.2 Pro~67estimate$0.0014$0.0113
GPT-5.5 Pro~67estimate$0.0020$0.0121
GPT-5.4 Pro~67estimate$0.0020$0.0121

Prices per 1M tokens. Input cost = cost if this text is sent as input. Output cost = cost if this text were generated as output. Data as of October 3, 2026.

Quick cost calculator

Enter a token count to compare costs across models

What this tool does

You can count tokens for GPT-5 and Claude prompts offline with this token counter. Every OpenAI model listed here is counted exactly, with the o200k_base tokenizer running in your browser — the same tokenizer OpenAI bills on. Claude, Gemini, Grok, DeepSeek and Mistral publish no tokenizer, so those rows are an estimate from calibrated characters-per-token ratios — 3.5 for Claude, 4.0 elsewhere — and every row says which of the two it is. No prompt text is ever sent to a server. For a billing-exact number on an estimated family, check the provider's own tokenizer.

Updated . Provided as is. Check the output before you rely on it in production.

How to use LLM Token Counter

  1. 1

    Paste text to count

    Paste any content: prompt text, article, code snippet, or conversation. OpenAI GPT and o-series rows are counted with the real o200k_base tokenizer, so those numbers are exact. Claude, Gemini, Grok, DeepSeek and Mistral publish no browser tokenizer, so their rows are character-ratio estimates.

  2. 2

    Compare the per-family counts

    The same text is counted for every tokenizer family at once, so you can see how far the estimates diverge. There is no model to pick: all families are shown side by side.

  3. 3

    See token estimate and cost breakdown

    Get total token count, approximate cost at current pricing, and cost per 1M tokens. Useful for budgeting API calls.

  4. 4

    Load a file instead of pasting

    Upload a text file to count it without pasting. To compare several prompts, count them one at a time and note each result; there is no multi-prompt batch mode.

  5. 5

    Test text reduction strategies

    Use the counter to compare token usage before and after shortening. Remove redundant words or summarize to optimize API costs.

Questions and answers

What is LLM Token Counter?
A token is the unit of text a language model reads and bills by, often a word or part of one. This counter gives exact o200k_base counts for OpenAI models and character-ratio estimates for Claude, Gemini, Grok, DeepSeek, Mistral, Qwen and Muse, and prices the count across models.
Why do token counts differ between models?
Each model family uses its own tokenizer with a different vocabulary and encoding rules, so the same text produces different token counts, which directly affects API cost.

REST API

Base URL

https://aidevhub.io/api/token-counter/

No authentication, fair use. Abusive traffic is throttled at the edge; responses carry no quota headers. CORS enabled. OpenAPI spec

Endpoints

GET /api/token-counter/ Count tokens and estimate costs
POST /api/token-counter/ Count tokens and estimate costs

Example

curl "https://aidevhub.io/api/token-counter/?text=Hello+world"

Example Response

{
  "text_length": 13,
  "tokens": {
    "GPT-6 Astra": {
      "model": "GPT-6 Astra",
      "tokens": 4,
      "method": "exact",
      "cost": {
        "input": 0.000039999999999999996,
        "output": 0.00019999999999999998
      }
    }
  },
  "model_count": 39
}
For AI agents: how to call this tool

Machine-readable contract, endpoints and examples. Humans can ignore this section.

Best Path For Builders

Dedicated API endpoint

Deterministic outputs, machine-safe contracts, and production-ready examples.

Dedicated API

https://aidevhub.io/api/token-counter/

OpenAPI: https://aidevhub.io/api/openapi.yaml

GET /api/token-counter/ Count tokens and estimate costs
POST /api/token-counter/ Count tokens and estimate costs

How do I count tokens for GPT and Claude prompts with an offline tokenizer?

Paste your prompt into the counter above. It computes character and word counts, then estimates tokens per model family using the characters-per-token ratios below — all in the browser, with no API call and no network dependency after the page loads. This matters when prompts contain confidential material, or when you need counts inside an offline or air-gapped workflow.

Characters-per-token ratios by model family

Model family Characters per token Est. tokens per 1,000 characters
GPT 4.0 ~250
Claude 3.5 ~286
Gemini 4.0 ~250
DeepSeek 4.0 ~250
Mistral 4.0 ~250
Grok 4.0 ~250
Muse 4.0 ~250
Qwen 4.0 ~250

Ratios apply to typical English prose; data as of October 3, 2026. A practical consequence: the same prompt usually produces more tokens on Claude (3.5 chars/token) than on GPT-family models (4.0 chars/token).

How accurate is a ratio-based token estimate?

A characters-per-token ratio is an estimate, not a byte-pair-encoding run. It is closest for plain English text and drifts for source code, dense punctuation, non-Latin scripts, and unusual whitespace, where real tokenizers split text differently. Use ratio estimates for sizing prompts, comparing models, and budgeting; use the provider's own tokenizer when you need billing-exact counts.

When do I need an exact count instead?

  1. Enforcing hard context-window limits, where an overshoot causes a request to be rejected or silently truncated.
  2. Metering or invoicing customers on token usage.
  3. For OpenAI models, run the open-source tiktoken library locally; for Anthropic models, call the count-tokens endpoint of the API.
  4. For everything before that point — drafting, comparing, budgeting — a client-side estimate is faster and keeps the prompt on your machine.