Token cost by language
The same text takes a different number of tokens in each language. We measured 30 languages with two OpenAI tokenizers, so you can see how much more (or less) a message costs outside English.
What the data shows
On the newer o200k_base encoding, the same text costs between 1.02× (Chinese (Simplified)) and 1.93× (Greek) of the English token count. On the older cl100k_base encoding the range is 1.20× (Portuguese) to 4.99× (Bengali).
The newer encoding helps most where the older one was weakest: Bengali needs 69% fewer tokens on o200k_base, while Portuguese, which was already efficient, saves 8%. Latin-script languages cluster close to English on both encodings. Scripts such as Bengali, Greek, Thai and Hindi were the most expensive on cl100k_base and improved the most.
Two things pull in opposite directions for East Asian languages. They have the fewest characters per token (Chinese (Traditional) has 1.20 on o200k_base), but they also need far fewer characters to say the same thing, so their total cost lands closer to English than the characters-per-token figure suggests. That is why we compare total tokens against the English version, not characters per token alone.
Cost relative to English
1.00× means the same number of tokens as the English version. The table below holds the same numbers.
All 30 languages
| Language | Script | Characters | Tokens o200k_base | Tokens cl100k_base | Chars / token o200k_base | × English o200k_base | × English cl100k_base | Saved by o200k_base |
|---|---|---|---|---|---|---|---|---|
| EnglishEnglish | Latin | 902 | 200 | 200 | 4.51 | baseline | baseline | – |
| GreekΕλληνικά | Greek | 1,089 | 386 | 924 | 2.82 | 1.93× | 4.62× | 58% |
| HungarianMagyar | Latin | 1,013 | 353 | 416 | 2.87 | 1.76× | 2.08× | 15% |
| Bengaliবাংলা | Bengali | 849 | 309 | 998 | 2.75 | 1.54× | 4.99× | 69% |
| UkrainianУкраїнська | Cyrillic | 967 | 304 | 514 | 3.18 | 1.52× | 2.57× | 41% |
| Japanese日本語 | Han + Kana | 455 | 301 | 402 | 1.51 | 1.50× | 2.01× | 25% |
| Thaiไทย | Thai | 773 | 301 | 710 | 2.57 | 1.50× | 3.55× | 58% |
| Hindiहिन्दी | Devanagari | 910 | 298 | 878 | 3.05 | 1.49× | 4.39× | 66% |
| FinnishSuomi | Latin | 954 | 295 | 361 | 3.23 | 1.48× | 1.80× | 18% |
| CzechČeština | Latin | 874 | 291 | 380 | 3.00 | 1.46× | 1.90× | 23% |
| RomanianRomână | Latin | 963 | 273 | 316 | 3.53 | 1.36× | 1.58× | 14% |
| PolishPolski | Latin | 915 | 267 | 313 | 3.43 | 1.33× | 1.56× | 15% |
| Chinese (Traditional)繁體中文 | Han | 317 | 265 | 397 | 1.20 | 1.32× | 1.99× | 33% |
| NorwegianNorsk | Latin | 943 | 263 | 301 | 3.59 | 1.31× | 1.50× | 13% |
| Korean한국어 | Hangul | 474 | 263 | 382 | 1.80 | 1.31× | 1.91× | 31% |
| TurkishTürkçe | Latin | 930 | 262 | 342 | 3.55 | 1.31× | 1.71× | 23% |
| FilipinoFilipino | Latin | 1,013 | 262 | 315 | 3.87 | 1.31× | 1.57× | 17% |
| SwedishSvenska | Latin | 931 | 258 | 295 | 3.61 | 1.29× | 1.48× | 13% |
| DanishDansk | Latin | 935 | 257 | 293 | 3.64 | 1.28× | 1.47× | 12% |
| GermanDeutsch | Latin | 1,115 | 252 | 286 | 4.42 | 1.26× | 1.43× | 12% |
| FrenchFrançais | Latin | 1,082 | 244 | 278 | 4.43 | 1.22× | 1.39× | 12% |
| RussianРусский | Cyrillic | 951 | 243 | 387 | 3.91 | 1.22× | 1.94× | 37% |
| VietnameseTiếng Việt | Latin | 867 | 239 | 391 | 3.63 | 1.20× | 1.96× | 39% |
| MalayBahasa Melayu | Latin | 991 | 235 | 288 | 4.22 | 1.18× | 1.44× | 18% |
| ItalianItaliano | Latin | 918 | 231 | 264 | 3.97 | 1.16× | 1.32× | 13% |
| SpanishEspañol | Latin | 920 | 224 | 248 | 4.11 | 1.12× | 1.24× | 10% |
| DutchNederlands | Latin | 928 | 220 | 268 | 4.22 | 1.10× | 1.34× | 18% |
| PortuguesePortuguês | Latin | 942 | 219 | 239 | 4.30 | 1.09× | 1.20× | 8% |
| IndonesianBahasa Indonesia | Latin | 932 | 218 | 274 | 4.28 | 1.09× | 1.37× | 20% |
| Chinese (Simplified)简体中文 | Han | 310 | 204 | 295 | 1.52 | 1.02× | 1.48× | 31% |
How we measured it
For each language we took this site's own homepage copy: the tagline, the two "what is an online clipboard" paragraphs and the four how-to steps, joined with blank lines. We counted tokens for that text with the o200k_base and cl100k_base encodings using the open-source gpt-tokenizer library (version 4.0.0), a port of OpenAI's tiktoken. The ratio columns divide each language's count by the English count for the same encoding. The numbers are recomputed every time the site is built, on 2026-10-09.
Limits to keep in mind. The translations were written by this site and have not been reviewed by professional translators, and the wording chosen affects length and therefore token count. The sample is one kind of text, short marketing and instructional copy, and about 902 characters in English. Code, legal text, names and numbers behave differently. Treat the ratios as a guide to the size of the effect, and measure your own text for any decision that costs money.
Show the English sample that was counted
Copy text on one device and paste it on another — no app, no account. A clipboard is the temporary memory your computer or phone uses when you copy and paste. It only works on a single device. An online clipboard does the same job over the internet, so you can copy text on your laptop and paste it on your phone, a work computer, a friend’s tablet or a smart TV. WebClipboard is perfect for moving links, notes, addresses, code snippets, Wi-Fi details or any other text without emailing yourself or installing an app. It is fast, free and works in every language. Type or paste your text into the box and choose how long it should be kept. Click “Save to clipboard” to get a 6-digit code, a short link and a QR code. On your other device, open this website and enter the code — or simply scan the QR code. Copy the text or download it as a .txt file. It disappears automatically when it expires.
Paste it into the token calculator to reproduce the English counts. The other languages' samples are the same copy in those languages, visible on each language's homepage.
Cite this data
You are welcome to quote or reuse these figures. Please link back to this page.
WebClipboard. "Token cost by language: 30 languages, two OpenAI tokenizers." 2026-10-09. https://webclipboard.online/token-cost-by-language/
Frequently asked questions
Why do some languages need more tokens than English?
A tokenizer splits text using a vocabulary of common pieces learned from large amounts of text. Languages and scripts that are better represented in that text get longer, more efficient pieces, while others are cut into smaller ones. How long the translation is also matters, so the ratio reflects both effects.
Which tokenizer is better for non-English text?
In this sample, o200k_base needed 34% fewer tokens than cl100k_base across the 29 non-English languages in total. The saving ranged from 8% for Portuguese to 69% for Bengali.
How was this measured?
We took the site's own copy for each language (the tagline, two about paragraphs and four how-to steps), counted its tokens with the o200k_base and cl100k_base encodings using gpt-tokenizer 4.0.0, and divided each language's count by the English count. The English sample is 902 characters and 200 tokens with o200k_base.
Why does Chinese look cheap here when it has so few characters per token?
Chinese packs a lot of meaning into each character, so the same copy is far shorter. Chinese (Simplified) has 1.52 characters per token on o200k_base, close to the lowest in the table, but 310 characters in total against 902 for English, so its overall cost lands at 1.02× English. Compare the total tokens, not the characters per token.
Does this apply to Claude, Gemini, Llama and other models?
Not directly. Other providers train their own tokenizers, so counts and ratios differ. These tables describe two OpenAI encodings only. Use your provider's own token counter before relying on a number for billing.
Can I use or cite this data?
Yes. You are welcome to cite it or reuse the figures; please link back to this page. The full data is available as CSV and JSON below.