Tokenizer comparison

A token is not a universal unit. The counts below are computed in your browser with real tokenizers — nothing is estimated by ratio.

Characters
216
UTF-8 bytes
255
Non-ASCII characters
25

o200k_base

60 tokens · 3.60 chars/token

Used by current OpenAI GPT-4o/o-series family. Exact count, computed locally.

Comparing models on price per token assumes a token is a token.⏎It is not: "internationalisation" splits differently under each BPE.⏎код на русском · 日本語のテキスト · 🙂🙂🙂⏎function f(x){return x.map(v=>v*2).filter(Boolean)}

cl100k_base

68 tokens · 3.18 chars/token

Earlier OpenAI encoding (GPT-3.5/4). Exact count, computed locally.

Comparing models on price per token assumes a token is a token.⏎It is not: "internationalisation" splits differently under each BPE.⏎код на русском · 日本語のテキスト · 🙂🙂🙂⏎function f(x){return x.map(v=>v*2).filter(Boolean)}

Tokenizers we do not have

Anthropic, Gemini, Llama, Qwen, Mistral and GLM use different tokenizers. Exact counts for those families require either the vendor token-counting endpoints or the model's own tokenizer files, neither of which is wired up yet. Rather than apply a fudge factor, we show nothing for them. Until that lands, treat cross-family comparisons on token counts alone as approximate, and prefer cost per 1k characters: $ per 1k chars = price per 1M tokens ÷ 1000 × (1k chars ÷ chars-per-token).