Tokenization Visualizer
Run the Tokenization Visualizer MicroSim Fullscreen
Edit in the p5.js Editor
About This MicroSim
A language model never reads words. It reads tokens, and a token is often smaller than a word. This MicroSim shows one sentence three ways:
- The text you typed, as plain text.
- How you see it: one box per word.
- How the model sees it: one colored chip per token.
Under the chips, a word count and a token count sit side by side. The token count is almost always the larger number, and it is the number that a model's cost and context-window capacity are measured in.
The split is an illustration, not a real tokenizer
The splitter in this MicroSim is a small set of hand-written rules. It is not the tokenizer of any real model, and its counts will not match a real model's counts. Real tokenizers, such as byte-pair encoding, learn their vocabulary of pieces from training text. The rules here only reproduce the general pattern, so use the counts to compare sentences with each other, not to estimate a bill.
The rules the illustrative splitter follows:
| Rule | Example |
|---|---|
| Short or very common words stay whole | into is 1 token |
| A familiar word ending splits off | Tokenization becomes Token + iz + ation |
| A familiar word beginning splits off | unbelievable becomes un + believ + able |
| A long, unfamiliar stretch of letters breaks into short fragments | establish becomes esta + blish |
| Punctuation marks are tokens of their own | . is 1 token |
| Long numbers break into groups of up to three digits | 12345 becomes 123 + 45 |
How to Use
- Read the example sentence, then compare the five word boxes with the eight token chips.
- Click any token chip. An infobox shows the token's text, its position index, and one sentence about why it split where it did.
- Type or paste your own sentence (up to 120 characters). The word count updates as you type and the token count changes to a question mark.
- Predict the token count, then select Tokenize (or press Enter) to check. The word boxes break apart into token chips in one quick transition.
- Select Reset to bring back the example sentence.
Try a sentence of short everyday words, then a sentence with long technical words. The tokens-per-word number shows the difference.
Iframe Embed Code
You can add this MicroSim to any web page by adding this to your HTML:
1 2 3 4 | |
Lesson Plan
Audience
Adult learners, educators, and instructional designers who are new to large language models. No programming background is needed.
Learning Objective
Explain why a language model's cost and capacity are measured in tokens rather than words. (Bloom's Taxonomy: Understand)
Duration
10-15 minutes
Prerequisites
- Knows that a large language model predicts the next unit of text
- Has read the definitions of token and tokenization in Chapter 1
Activities
- Exploration (5 min): Compare the word row and the token row for the
example sentence. Click each chip in
Token+iz+ationand read why it split. - Guided Practice (5 min): Type three sentences: one made of short common words, one with long technical words, and one with a long number. For each, write down a predicted token count before selecting Tokenize.
- Assessment (5 min): In two or three sentences, explain to a colleague why a 1,000-word document does not cost "1,000 units" to send to a model.
Assessment
- The learner states that a model processes tokens, not words.
- The learner predicts that long or unusual words produce more tokens than short common words, and confirms it with their own sentence.
- The learner explains that limits and prices follow the token count, so the word count underestimates both.
Discussion Questions
- Which of your sentences had the highest tokens-per-word number? Why?
- Why might a tokenizer keep a common word whole but break a rare word apart?
- This splitter is hand-written. What would a real tokenizer need in order to decide where to split?
References
- Byte pair encoding - Wikipedia - The subword method behind many real tokenizers, which this MicroSim only approximates.
- Large language model - Wikipedia - Background on tokenization and context windows in language models.
- OpenAI Tokenizer - OpenAI - A real tokenizer you can use to compare actual token splits with the illustrative ones shown here.
- p5.js Reference - p5.js - Documentation for the library used to build this MicroSim.