Tokenizer comparison
A token is not a universal unit. The counts below are computed in your browser with real tokenizers — nothing is estimated by ratio.
- Characters
- 216
- UTF-8 bytes
- 255
- Non-ASCII characters
- 25
o200k_base
Used by current OpenAI GPT-4o/o-series family. Exact count, computed locally.
Comparing models on price per token assumes a token is a token.⏎It is not: "internationalisation" splits differently under each BPE.⏎код на русском · 日本語のテキスト · 🙂🙂🙂⏎function f(x){return x.map(v=>v*2).filter(Boolean)}
cl100k_base
Earlier OpenAI encoding (GPT-3.5/4). Exact count, computed locally.
Comparing models on price per token assumes a token is a token.⏎It is not: "internationalisation" splits differently under each BPE.⏎код на русском · 日本語のテキスト · 🙂🙂🙂⏎function f(x){return x.map(v=>v*2).filter(Boolean)}
Tokenizers we do not have
Anthropic, Gemini, Llama, Qwen, Mistral and GLM use different tokenizers. Exact counts for those families require either the vendor token-counting endpoints or the model's own tokenizer files, neither of which is wired up yet. Rather than apply a fudge factor, we show nothing for them. Until that lands, treat cross-family comparisons on token counts alone as approximate, and prefer cost per 1k characters: $ per 1k chars = price per 1M tokens ÷ 1000 × (1k chars ÷ chars-per-token).