Zum Inhalt springen

Tokens -- The Building Blocks of Language

Knowledge

When you talk to an LLM, it does not read your message word by word but breaks it down into smaller units: tokens. A token is a subword unit -- roughly three quarters of a word. Some short words like "AI" or "is" are a single token. Longer words like "Artificial" are broken into 2-3 tokens.

Why not just use whole words? Because every language has millions of possible words (think of compound words like "Donaudampfschifffahrtsgesellschaft" in German). By breaking text into tokens, the model can work with a manageable vocabulary of 30,000-100,000 units.

Understanding

To give you a sense of token quantities, here are some everyday comparisons:

  • A text message: about 25-30 tokens
  • A short email: about 250-300 tokens
  • A newspaper article: about 1,000-2,000 tokens
  • A novel: about 100,000 tokens

Tokens matter for a practical reason: costs are calculated per token. Every token you send to an LLM (input) and every token it sends back (output) costs money. Additionally, there is a token limit -- the maximum number of tokens for input and output combined.

*What Does That Cost in Practice?

With common LLMs, 1,000 tokens cost about 0.001 to 0.01 euros. A typical conversation with 10 messages costs only a few cents. With free offerings like ChatGPT Free or Claude Free, you pay nothing -- the provider covers the costs.

Original text

Artificial Intelligence is changing our world

Token breakdown (simulated)

Artificial·Intelligence·is·changing·our·world

Tokens

14

Characters

45

Approximate cost (GPT-4)

$0.00028

Note: This is a simplified simulation. Real tokenizers (like BPE) work differently.

Apply

If you use an LLM and get a very long response that breaks off mid-sentence, you now know why: the token limit was reached. You can simply write "Please continue" and the LLM will pick up where it left off.

Estimate the approximate token count:

A text message has approximately tokens.

Reflect

Why are costs for LLMs calculated per token and not per word?