cl100k_base tokenizer online
Count tokens with the cl100k_base encoding used by GPT-4, GPT-3.5 and the text-embedding-3 models. Paste your text and see the exact count, the token split and the cost at your prices.
Only the first 1,000,000 characters are counted.
Estimate the cost
Enter the prices from your provider's pricing page, in price per 1 million tokens. Nothing is pre-filled because prices change.
Add at least one price above to see a cost.
Does it fit?
Your text as a share of common context window sizes. The window also has to hold the reply.
How the text is split
Each shaded block is one token (or a few tokens that make up one character). Hover a block to see its id.
How to count cl100k_base tokens
- Paste your text, prompt, code or document into the box above. The count updates as you type.
- Read the exact cl100k_base token count, the characters-per-token ratio and how much of common context windows it uses.
- Hover a shaded block to see its token id, and watch how spaces, case, numbers and non-English text are split.
- Optionally enter your provider's input and output prices per 1 million tokens to estimate the cost of a request.
What is cl100k_base?
cl100k_base is one of the tokenizer vocabularies published by OpenAI. It was built with byte-pair encoding, which repeatedly merges the most frequent neighbouring pieces of text until the vocabulary reaches its target size. cl100k_base has about 100,000 entries (the library reports a vocabulary size of 100,264, which includes special tokens). Every piece of text is mapped to a list of integer ids from this table, and the model reads and writes those ids, not letters or words.
The gpt-tokenizer library documents gpt-4-* and gpt-3.5-* models as using cl100k_base, and its model table also maps the text-embedding-3 and text-embedding-ada-002 embedding models to it.
The counter on this page is GPT-4 and GPT-3.5-specific: it applies the cl100k_base merge rules only. It is the older of the two common vocabularies and usually needs more tokens than o200k_base for code, emoji and many non-English scripts. For a broader explanation, read how tokenizers work.
Models mapped to cl100k_base in the library (23)
Generated from the gpt-tokenizer 4.0.0 model table. Models the library does not list explicitly default to o200k_base. Model lists change, so confirm in your provider's documentation.
gpt-3.5, gpt-3.5-0301, gpt-3.5-turbo, gpt-3.5-turbo-0125, gpt-3.5-turbo-0613, gpt-3.5-turbo-1106, gpt-3.5-turbo-16k-0613, gpt-3.5-turbo-instruct, gpt-4, gpt-4-0125-preview, gpt-4-0314, gpt-4-0613, gpt-4-1106-preview, gpt-4-1106-vision-preview, gpt-4-32k, gpt-4-turbo, gpt-4-turbo-2024-04-09, gpt-4-turbo-preview, text-embedding-3-large, text-embedding-3-small, text-embedding-ada-002, babbage-002, davinci-002
Token counts for common text
Measured with the real encodings when this site was built. Token ids are shown for cl100k_base; the last column shows the same text in o200k_base.
| Text | cl100k_base tokens | o200k_base tokens | cl100k_base token ids |
|---|---|---|---|
| hello world | 2 | 2 | 15339, 1917 |
| Hello | 1 | 1 | 9906 |
| HELLO | 2 | 2 | 51812, 1623 |
| tokenization | 2 | 2 | 5963, 2065 |
| 1234567 | 3 | 3 | 4513, 10961, 22 |
| https://webclipboard.online/ | 6 | 6 | 2485, 1129, 2984, 71948, 68719, 14 |
| jane.doe@example.com | 6 | 6 | 73, 2194, 962, 4748, 36587, 916 |
| 🌍 | 3 | 2 | 9468, 234, 235 |
| こんにちは世界 | 4 | 2 | 90115, 3574, 244, 98220 |
| Привет, мир | 6 | 4 | 54745, 28089, 8341, 11, 11562, 78746 |
| नमस्ते दुनिया | 13 | 5 | 61196, 88344, 79468, 31584, 97, 35470, 15272, 99, … |
| 8 spaces in a row | 1 | 1 | 260 |
cl100k_base across languages
We counted the same website copy in 30 languages. With cl100k_base, the median non-English language costs 1.58× the English token count, and 2 of 29 languages stay within 25% of English. The most expensive was Bengali at 4.99×. Compared with o200k_base, cl100k_base needed more tokens for almost every non-English language, and Bengali reached 4.99× the English count.
See every language, the chart and the downloadable data on token cost by language.
Count cl100k_base tokens in code
To count tokens inside your own application, use the same encoding. In JavaScript with gpt-tokenizer, which this page uses:
import { encode, decode } from 'gpt-tokenizer/encoding/cl100k_base';
const ids = encode('hello world'); // [15339, 1917]
console.log(ids.length); // 2
console.log(decode(ids)); // 'hello world' In Python with OpenAI's tiktoken:
import tiktoken
enc = tiktoken.get_encoding("cl100k_base")
ids = enc.encode("hello world")
print(len(ids)) # 2 What this counter does not include
The number shown is the exact token count of the text you pasted, for cl100k_base. A real API request usually contains more than that:
- Message formatting. Roles and separators add a few tokens per message.
- System instructions and tool definitions. These are billed as input on every request.
- Images, audio and files. They are converted to tokens by the provider.
- Hidden reasoning. Some models bill internal thinking as output tokens.
- Other providers. Models from other companies use different tokenizers, so this count is only a guide for them.
Learn how those pieces turn into a bill in how LLM API pricing works.
Frequently asked questions
What is cl100k_base?
cl100k_base is a byte-pair-encoding vocabulary with about 100,000 entries. A tokenizer built on it turns text into integer token ids, and language models that use it read and write those ids.
Which models use cl100k_base?
The gpt-tokenizer library documents gpt-4-* and gpt-3.5-* models as using cl100k_base, and its model table also maps the text-embedding-3 and text-embedding-ada-002 embedding models to it. Check your provider's documentation for the model you are using, since model lists change.
Is this cl100k_base counter exact?
The counts come from gpt-tokenizer 4.0.0, a port of OpenAI's tiktoken that is tested for compatibility with it, so the token count of the text you paste is exact for cl100k_base. Message wrappers, tool definitions and images add extra tokens in a real API request that are not included here.
Is my text sent to a server?
No. The tokenizer runs in your browser. Only the tokenizer vocabulary is downloaded, once, and your text never leaves your device.
How is cl100k_base different from o200k_base?
o200k_base has about 200,000 entries. Across the 29 non-English languages we measured, cl100k_base costs a median of 1.58× the English token count, against 1.31× for o200k_base. See the full table on the token cost by language page.
Can I count cl100k_base tokens in code?
Yes. The examples on this page show how to count them in JavaScript with gpt-tokenizer and in Python with tiktoken.