Tokens -- The Building Blocks of Language
Knowledge
When you talk to an LLM, it does not read your message word by word but breaks it down into smaller units: tokens. A token is a subword unit -- roughly three quarters of a word. Some short words like "AI" or "is" are a single token. Longer words like "Artificial" are broken into 2-3 tokens.
Why not just use whole words? Because every language has millions of possible words (think of compound words like "Donaudampfschifffahrtsgesellschaft" in German). By breaking text into tokens, the model can work with a manageable vocabulary of 30,000-100,000 units.
Understanding
To give you a sense of token quantities, here are some everyday comparisons:
- A text message: about 25-30 tokens
- A short email: about 250-300 tokens
- A newspaper article: about 1,000-2,000 tokens
- A novel: about 100,000 tokens
Tokens matter for a practical reason: costs are calculated per token. Every token you send to an LLM (input) and every token it sends back (output) costs money. Additionally, there is a token limit -- the maximum number of tokens for input and output combined.
*What Does That Cost in Practice?
With common LLMs, 1,000 tokens cost about 0.001 to 0.01 euros. A typical conversation with 10 messages costs only a few cents. With free offerings like ChatGPT Free or Claude Free, you pay nothing -- the provider covers the costs.
Original text
Artificial Intelligence is changing our world
Token breakdown (simulated)
Tokens
14
Characters
45
Approximate cost (GPT-4)
$0.00028
Note: This is a simplified simulation. Real tokenizers (like BPE) work differently.
Apply
If you use an LLM and get a very long response that breaks off mid-sentence, you now know why: the token limit was reached. You can simply write "Please continue" and the LLM will pick up where it left off.
Estimate the approximate token count:
Reflect
Why are costs for LLMs calculated per token and not per word?