Data · measured 2026-10-09

Token cost by language

The same text takes a different number of tokens in each language. We measured 30 languages with two OpenAI tokenizers, so you can see how much more (or less) a message costs outside English.

Median cost vs English1.31×o200k_base. With cl100k_base it is 1.58×.
Highest, o200k_base1.93×Greek
Highest, cl100k_base4.99×Bengali
Within 25% of English10 of 29on o200k_base, versus 2 on cl100k_base
Saved by o200k_base34%fewer tokens than cl100k_base across all 29 non-English languages

What the data shows

On the newer o200k_base encoding, the same text costs between 1.02× (Chinese (Simplified)) and 1.93× (Greek) of the English token count. On the older cl100k_base encoding the range is 1.20× (Portuguese) to 4.99× (Bengali).

The newer encoding helps most where the older one was weakest: Bengali needs 69% fewer tokens on o200k_base, while Portuguese, which was already efficient, saves 8%. Latin-script languages cluster close to English on both encodings. Scripts such as Bengali, Greek, Thai and Hindi were the most expensive on cl100k_base and improved the most.

Two things pull in opposite directions for East Asian languages. They have the fewest characters per token (Chinese (Traditional) has 1.20 on o200k_base), but they also need far fewer characters to say the same thing, so their total cost lands closer to English than the characters-per-token figure suggests. That is why we compare total tokens against the English version, not characters per token alone.

Chart

Cost relative to English

1.00× means the same number of tokens as the English version. The table below holds the same numbers.

o200k_base cl100k_base
Table

All 30 languages

Language Script Characters Tokens
o200k_base
Tokens
cl100k_base
Chars / token
o200k_base
× English
o200k_base
× English
cl100k_base
Saved by
o200k_base
EnglishEnglish Latin 902 200 200 4.51 baseline baseline –
GreekΕλληνικά Greek 1,089 386 924 2.82 1.93× 4.62× 58%
HungarianMagyar Latin 1,013 353 416 2.87 1.76× 2.08× 15%
Bengaliবাংলা Bengali 849 309 998 2.75 1.54× 4.99× 69%
UkrainianУкраїнська Cyrillic 967 304 514 3.18 1.52× 2.57× 41%
Japanese日本語 Han + Kana 455 301 402 1.51 1.50× 2.01× 25%
Thaiไทย Thai 773 301 710 2.57 1.50× 3.55× 58%
Hindiहिन्दी Devanagari 910 298 878 3.05 1.49× 4.39× 66%
FinnishSuomi Latin 954 295 361 3.23 1.48× 1.80× 18%
CzechČeština Latin 874 291 380 3.00 1.46× 1.90× 23%
RomanianRomână Latin 963 273 316 3.53 1.36× 1.58× 14%
PolishPolski Latin 915 267 313 3.43 1.33× 1.56× 15%
Chinese (Traditional)繁體中文 Han 317 265 397 1.20 1.32× 1.99× 33%
NorwegianNorsk Latin 943 263 301 3.59 1.31× 1.50× 13%
Korean한국어 Hangul 474 263 382 1.80 1.31× 1.91× 31%
TurkishTürkçe Latin 930 262 342 3.55 1.31× 1.71× 23%
FilipinoFilipino Latin 1,013 262 315 3.87 1.31× 1.57× 17%
SwedishSvenska Latin 931 258 295 3.61 1.29× 1.48× 13%
DanishDansk Latin 935 257 293 3.64 1.28× 1.47× 12%
GermanDeutsch Latin 1,115 252 286 4.42 1.26× 1.43× 12%
FrenchFrançais Latin 1,082 244 278 4.43 1.22× 1.39× 12%
RussianРусский Cyrillic 951 243 387 3.91 1.22× 1.94× 37%
VietnameseTiếng Việt Latin 867 239 391 3.63 1.20× 1.96× 39%
MalayBahasa Melayu Latin 991 235 288 4.22 1.18× 1.44× 18%
ItalianItaliano Latin 918 231 264 3.97 1.16× 1.32× 13%
SpanishEspañol Latin 920 224 248 4.11 1.12× 1.24× 10%
DutchNederlands Latin 928 220 268 4.22 1.10× 1.34× 18%
PortuguesePortuguês Latin 942 219 239 4.30 1.09× 1.20× 8%
IndonesianBahasa Indonesia Latin 932 218 274 4.28 1.09× 1.37× 20%
Chinese (Simplified)简体中文 Han 310 204 295 1.52 1.02× 1.48× 31%
Method

How we measured it

For each language we took this site's own homepage copy: the tagline, the two "what is an online clipboard" paragraphs and the four how-to steps, joined with blank lines. We counted tokens for that text with the o200k_base and cl100k_base encodings using the open-source gpt-tokenizer library (version 4.0.0), a port of OpenAI's tiktoken. The ratio columns divide each language's count by the English count for the same encoding. The numbers are recomputed every time the site is built, on 2026-10-09.

Limits to keep in mind. The translations were written by this site and have not been reviewed by professional translators, and the wording chosen affects length and therefore token count. The sample is one kind of text, short marketing and instructional copy, and about 902 characters in English. Code, legal text, names and numbers behave differently. Treat the ratios as a guide to the size of the effect, and measure your own text for any decision that costs money.

Show the English sample that was counted
Copy text on one device and paste it on another — no app, no account.

A clipboard is the temporary memory your computer or phone uses when you copy and paste. It only works on a single device. An online clipboard does the same job over the internet, so you can copy text on your laptop and paste it on your phone, a work computer, a friend’s tablet or a smart TV.

WebClipboard is perfect for moving links, notes, addresses, code snippets, Wi-Fi details or any other text without emailing yourself or installing an app. It is fast, free and works in every language.

Type or paste your text into the box and choose how long it should be kept.

Click “Save to clipboard” to get a 6-digit code, a short link and a QR code.

On your other device, open this website and enter the code — or simply scan the QR code.

Copy the text or download it as a .txt file. It disappears automatically when it expires.

Paste it into the token calculator to reproduce the English counts. The other languages' samples are the same copy in those languages, visible on each language's homepage.

Cite

Cite this data

You are welcome to quote or reuse these figures. Please link back to this page.

WebClipboard. "Token cost by language: 30 languages, two OpenAI tokenizers." 2026-10-09. https://webclipboard.online/token-cost-by-language/
FAQ

Frequently asked questions

Why do some languages need more tokens than English?

A tokenizer splits text using a vocabulary of common pieces learned from large amounts of text. Languages and scripts that are better represented in that text get longer, more efficient pieces, while others are cut into smaller ones. How long the translation is also matters, so the ratio reflects both effects.

Which tokenizer is better for non-English text?

In this sample, o200k_base needed 34% fewer tokens than cl100k_base across the 29 non-English languages in total. The saving ranged from 8% for Portuguese to 69% for Bengali.

How was this measured?

We took the site's own copy for each language (the tagline, two about paragraphs and four how-to steps), counted its tokens with the o200k_base and cl100k_base encodings using gpt-tokenizer 4.0.0, and divided each language's count by the English count. The English sample is 902 characters and 200 tokens with o200k_base.

Why does Chinese look cheap here when it has so few characters per token?

Chinese packs a lot of meaning into each character, so the same copy is far shorter. Chinese (Simplified) has 1.52 characters per token on o200k_base, close to the lowest in the table, but 310 characters in total against 902 for English, so its overall cost lands at 1.02× English. Compare the total tokens, not the characters per token.

Does this apply to Claude, Gemini, Llama and other models?

Not directly. Other providers train their own tokenizers, so counts and ratios differ. These tables describe two OpenAI encodings only. Use your provider's own token counter before relying on a number for billing.

Can I use or cite this data?

Yes. You are welcome to cite it or reuse the figures; please link back to this page. The full data is available as CSV and JSON below.

Related

Keep reading

Share this tool