~/tokens · cl100k_base · runs in your browser

cl100k_base tokenizer online

Count tokens with the cl100k_base encoding used by GPT-4, GPT-3.5 and the text-embedding-3 models. Paste your text and see the exact count, the token split and the cost at your prices.

Tokens0
Characters0
Words0
Chars / token–
Tokens / word–

Estimate the cost

Enter the prices from your provider's pricing page, in price per 1 million tokens. Nothing is pre-filled because prices change.

Input cost–for the pasted text
Output cost–500 output tokens
Per request–input + output
Total–

Add at least one price above to see a cost.

Does it fit?

Your text as a share of common context window sizes. The window also has to hold the reply.

How the text is split

Each shaded block is one token (or a few tokens that make up one character). Hover a block to see its id.

Your text is processed in your browser and is never uploaded. OpenAI tokenizer counts are exact for the text you enter. Chat messages, tool definitions and images add a few extra tokens per request that this tool does not include, and other providers use different tokenizers, so treat totals as estimates.
01

How to count cl100k_base tokens

  1. Paste your text, prompt, code or document into the box above. The count updates as you type.
  2. Read the exact cl100k_base token count, the characters-per-token ratio and how much of common context windows it uses.
  3. Hover a shaded block to see its token id, and watch how spaces, case, numbers and non-English text are split.
  4. Optionally enter your provider's input and output prices per 1 million tokens to estimate the cost of a request.
02

What is cl100k_base?

cl100k_base is one of the tokenizer vocabularies published by OpenAI. It was built with byte-pair encoding, which repeatedly merges the most frequent neighbouring pieces of text until the vocabulary reaches its target size. cl100k_base has about 100,000 entries (the library reports a vocabulary size of 100,264, which includes special tokens). Every piece of text is mapped to a list of integer ids from this table, and the model reads and writes those ids, not letters or words.

The gpt-tokenizer library documents gpt-4-* and gpt-3.5-* models as using cl100k_base, and its model table also maps the text-embedding-3 and text-embedding-ada-002 embedding models to it.

The counter on this page is GPT-4 and GPT-3.5-specific: it applies the cl100k_base merge rules only. It is the older of the two common vocabularies and usually needs more tokens than o200k_base for code, emoji and many non-English scripts. For a broader explanation, read how tokenizers work.

Models mapped to cl100k_base in the library (23)

Generated from the gpt-tokenizer 4.0.0 model table. Models the library does not list explicitly default to o200k_base. Model lists change, so confirm in your provider's documentation.

gpt-3.5, gpt-3.5-0301, gpt-3.5-turbo, gpt-3.5-turbo-0125, gpt-3.5-turbo-0613, gpt-3.5-turbo-1106, gpt-3.5-turbo-16k-0613, gpt-3.5-turbo-instruct, gpt-4, gpt-4-0125-preview, gpt-4-0314, gpt-4-0613, gpt-4-1106-preview, gpt-4-1106-vision-preview, gpt-4-32k, gpt-4-turbo, gpt-4-turbo-2024-04-09, gpt-4-turbo-preview, text-embedding-3-large, text-embedding-3-small, text-embedding-ada-002, babbage-002, davinci-002

03

Token counts for common text

Measured with the real encodings when this site was built. Token ids are shown for cl100k_base; the last column shows the same text in o200k_base.

Textcl100k_base tokenso200k_base tokenscl100k_base token ids
hello world 2 2 15339, 1917
Hello 1 1 9906
HELLO 2 2 51812, 1623
tokenization 2 2 5963, 2065
1234567 3 3 4513, 10961, 22
https://webclipboard.online/ 6 6 2485, 1129, 2984, 71948, 68719, 14
jane.doe@example.com 6 6 73, 2194, 962, 4748, 36587, 916
🌍 3 2 9468, 234, 235
こんにちは世界 4 2 90115, 3574, 244, 98220
Привет, мир 6 4 54745, 28089, 8341, 11, 11562, 78746
नमस्ते दुनिया 13 5 61196, 88344, 79468, 31584, 97, 35470, 15272, 99, …
8 spaces in a row 1 1 260
04

cl100k_base across languages

We counted the same website copy in 30 languages. With cl100k_base, the median non-English language costs 1.58× the English token count, and 2 of 29 languages stay within 25% of English. The most expensive was Bengali at 4.99×. Compared with o200k_base, cl100k_base needed more tokens for almost every non-English language, and Bengali reached 4.99× the English count.

See every language, the chart and the downloadable data on token cost by language.

05

Count cl100k_base tokens in code

To count tokens inside your own application, use the same encoding. In JavaScript with gpt-tokenizer, which this page uses:

import { encode, decode } from 'gpt-tokenizer/encoding/cl100k_base';

const ids = encode('hello world');   // [15339, 1917]
console.log(ids.length);              // 2
console.log(decode(ids));             // 'hello world'

In Python with OpenAI's tiktoken:

import tiktoken

enc = tiktoken.get_encoding("cl100k_base")
ids = enc.encode("hello world")
print(len(ids))   # 2
06

What this counter does not include

The number shown is the exact token count of the text you pasted, for cl100k_base. A real API request usually contains more than that:

  • Message formatting. Roles and separators add a few tokens per message.
  • System instructions and tool definitions. These are billed as input on every request.
  • Images, audio and files. They are converted to tokens by the provider.
  • Hidden reasoning. Some models bill internal thinking as output tokens.
  • Other providers. Models from other companies use different tokenizers, so this count is only a guide for them.

Learn how those pieces turn into a bill in how LLM API pricing works.

07

Frequently asked questions

What is cl100k_base?

cl100k_base is a byte-pair-encoding vocabulary with about 100,000 entries. A tokenizer built on it turns text into integer token ids, and language models that use it read and write those ids.

Which models use cl100k_base?

The gpt-tokenizer library documents gpt-4-* and gpt-3.5-* models as using cl100k_base, and its model table also maps the text-embedding-3 and text-embedding-ada-002 embedding models to it. Check your provider's documentation for the model you are using, since model lists change.

Is this cl100k_base counter exact?

The counts come from gpt-tokenizer 4.0.0, a port of OpenAI's tiktoken that is tested for compatibility with it, so the token count of the text you paste is exact for cl100k_base. Message wrappers, tool definitions and images add extra tokens in a real API request that are not included here.

Is my text sent to a server?

No. The tokenizer runs in your browser. Only the tokenizer vocabulary is downloaded, once, and your text never leaves your device.

How is cl100k_base different from o200k_base?

o200k_base has about 200,000 entries. Across the 29 non-English languages we measured, cl100k_base costs a median of 1.58× the English token count, against 1.31× for o200k_base. See the full table on the token cost by language page.

Can I count cl100k_base tokens in code?

Yes. The examples on this page show how to count them in JavaScript with gpt-tokenizer and in Python with tiktoken.

08

More token tools and guides

Share this tool