~/tokens · o200k_base · runs in your browser

o200k_base token counter

Count tokens with the o200k_base encoding, the one used by GPT-4o, GPT-4.1, GPT-5 and the o-series. Paste your text and see the exact count, the token split and the cost at your prices.

Tokens0
Characters0
Words0
Chars / token–
Tokens / word–

Estimate the cost

Enter the prices from your provider's pricing page, in price per 1 million tokens. Nothing is pre-filled because prices change.

Input cost–for the pasted text
Output cost–500 output tokens
Per request–input + output
Total–

Add at least one price above to see a cost.

Does it fit?

Your text as a share of common context window sizes. The window also has to hold the reply.

How the text is split

Each shaded block is one token (or a few tokens that make up one character). Hover a block to see its id.

Your text is processed in your browser and is never uploaded. OpenAI tokenizer counts are exact for the text you enter. Chat messages, tool definitions and images add a few extra tokens per request that this tool does not include, and other providers use different tokenizers, so treat totals as estimates.
01

How to count o200k_base tokens

  1. Paste your text, prompt, code or document into the box above. The count updates as you type.
  2. Read the exact o200k_base token count, the characters-per-token ratio and how much of common context windows it uses.
  3. Hover a shaded block to see its token id, and watch how spaces, case, numbers and non-English text are split.
  4. Optionally enter your provider's input and output prices per 1 million tokens to estimate the cost of a request.
02

What is o200k_base?

o200k_base is one of the tokenizer vocabularies published by OpenAI. It was built with byte-pair encoding, which repeatedly merges the most frequent neighbouring pieces of text until the vocabulary reaches its target size. o200k_base has about 200,000 entries (the library reports a vocabulary size of 200,006, which includes special tokens). Every piece of text is mapped to a list of integer ids from this table, and the model reads and writes those ids, not letters or words.

The gpt-tokenizer library documents o200k_base as the encoding for GPT-5, GPT-4.1 and GPT-4o models and for the o-series reasoning models. Models the library does not list explicitly also default to it.

The counter on this page is GPT-4o and newer-specific: it applies the o200k_base merge rules only. Because this vocabulary is larger, it usually represents code, emoji and many non-English scripts in fewer tokens than the older cl100k_base. For a broader explanation, read how tokenizers work.

Models mapped to cl100k_base in the library (23)

Generated from the gpt-tokenizer 4.0.0 model table. Models the library does not list explicitly default to o200k_base. Model lists change, so confirm in your provider's documentation.

gpt-3.5, gpt-3.5-0301, gpt-3.5-turbo, gpt-3.5-turbo-0125, gpt-3.5-turbo-0613, gpt-3.5-turbo-1106, gpt-3.5-turbo-16k-0613, gpt-3.5-turbo-instruct, gpt-4, gpt-4-0125-preview, gpt-4-0314, gpt-4-0613, gpt-4-1106-preview, gpt-4-1106-vision-preview, gpt-4-32k, gpt-4-turbo, gpt-4-turbo-2024-04-09, gpt-4-turbo-preview, text-embedding-3-large, text-embedding-3-small, text-embedding-ada-002, babbage-002, davinci-002

03

Token counts for common text

Measured with the real encodings when this site was built. Token ids are shown for o200k_base; the last column shows the same text in cl100k_base.

Texto200k_base tokenscl100k_base tokenso200k_base token ids
hello world 2 2 24912, 2375
Hello 1 1 13225
HELLO 2 2 111642, 2699
tokenization 2 2 10346, 2860
1234567 3 3 7633, 19354, 22
https://webclipboard.online/ 6 6 4172, 1684, 4116, 163134, 107641, 14
jane.doe@example.com 6 6 73, 1986, 1380, 5578, 81309, 1136
🌍 2 3 64364, 235
こんにちは世界 2 4 95839, 28428
Привет, мир 4 6 23881, 131903, 11, 37934
नमस्ते दुनिया 5 13 998, 1637, 14681, 628, 64593
8 spaces in a row 1 1 269
04

o200k_base across languages

We counted the same website copy in 30 languages. With o200k_base, the median non-English language costs 1.31× the English token count, and 10 of 29 languages stay within 25% of English. The most expensive was Greek at 1.93×. Compared with cl100k_base, o200k_base needed 34% fewer tokens in total across the non-English languages.

See every language, the chart and the downloadable data on token cost by language.

05

Count o200k_base tokens in code

To count tokens inside your own application, use the same encoding. In JavaScript with gpt-tokenizer, which this page uses:

import { encode, decode } from 'gpt-tokenizer/encoding/o200k_base';

const ids = encode('hello world');   // [24912, 2375]
console.log(ids.length);              // 2
console.log(decode(ids));             // 'hello world'

In Python with OpenAI's tiktoken:

import tiktoken

enc = tiktoken.get_encoding("o200k_base")
ids = enc.encode("hello world")
print(len(ids))   # 2
06

What this counter does not include

The number shown is the exact token count of the text you pasted, for o200k_base. A real API request usually contains more than that:

  • Message formatting. Roles and separators add a few tokens per message.
  • System instructions and tool definitions. These are billed as input on every request.
  • Images, audio and files. They are converted to tokens by the provider.
  • Hidden reasoning. Some models bill internal thinking as output tokens.
  • Other providers. Models from other companies use different tokenizers, so this count is only a guide for them.

Learn how those pieces turn into a bill in how LLM API pricing works.

07

Frequently asked questions

What is o200k_base?

o200k_base is a byte-pair-encoding vocabulary with about 200,000 entries. A tokenizer built on it turns text into integer token ids, and language models that use it read and write those ids.

Which models use o200k_base?

The gpt-tokenizer library documents o200k_base as the encoding for GPT-5, GPT-4.1 and GPT-4o models and for the o-series reasoning models. Models the library does not list explicitly also default to it. Check your provider's documentation for the model you are using, since model lists change.

Is this o200k_base counter exact?

The counts come from gpt-tokenizer 4.0.0, a port of OpenAI's tiktoken that is tested for compatibility with it, so the token count of the text you paste is exact for o200k_base. Message wrappers, tool definitions and images add extra tokens in a real API request that are not included here.

Is my text sent to a server?

No. The tokenizer runs in your browser. Only the tokenizer vocabulary is downloaded, once, and your text never leaves your device.

How is o200k_base different from cl100k_base?

cl100k_base has about 100,000 entries. Across the 29 non-English languages we measured, o200k_base costs a median of 1.31× the English token count, against 1.58× for cl100k_base. See the full table on the token cost by language page.

Can I count o200k_base tokens in code?

Yes. The examples on this page show how to count them in JavaScript with gpt-tokenizer and in Python with tiktoken.

08

More token tools and guides

Share this tool