Zum Inhalt springen

Training — How an LLM Learns

Knowledge

An LLM like ChatGPT or Claude is not programmed to answer questions. It is trained — in multiple phases that build on each other. This process takes weeks to months and requires thousands of specialized computer chips (GPUs). But the result is impressive: a system that can understand and generate human language.

The training consists of two central phases:

Phase 1: Pre-Training — Learning Language

In the first phase, the model reads enormous amounts of text from the internet: books, Wikipedia articles, websites, programming code, scientific papers, and much more. We are talking about trillions of words — more text than a human could read in a thousand lifetimes.

In doing so, the model does not memorize facts. Instead, it learns statistical patterns: Which words are likely to follow each other? Which sentence structures are common? How are concepts related? For example, the model learns that after "The capital of Germany is" the word "Berlin" is very likely to follow — not because it stored this as a fact, but because it has seen this pattern thousands of times.

A helpful analogy: Imagine a child learning language by listening to its surroundings. It hears thousands of sentences and gradually recognizes patterns: "I am", "you are", "he is". Nobody explains the grammar rules to the child — it learns them from the patterns. An LLM learns in pre-training in exactly the same way.

After pre-training, the model can complete texts impressively well. But it is not yet a helpful assistant. It is more like a text machine that simply keeps writing — without regard for whether the answer is helpful, correct, or safe.

Phase 2: RLHF — Learning Manners

RLHF stands for Reinforcement Learning from Human Feedback. In this phase, the model learns not just to complete text, but to give helpful, harmless, and honest answers.

Here is how it works: humans ask the model questions and receive several different answers. They then rate these answers: Which is the most helpful? Which is the safest? Which is the most accurate? From these ratings, the model learns what kind of answers are preferred.

Extending the analogy: If pre-training is like "learning to speak," then RLHF is like "learning manners." The child first learned to talk. Now it learns how to talk: being polite, giving helpful answers, not saying dangerous things, and admitting when it does not know something.

iWhy RLHF Is So Important

Without RLHF, an LLM would just be a text completion machine. In response to the question "How can I help someone who is sad?", it might simply continue writing a Wikipedia article about depression, instead of giving an empathetic, practical answer. RLHF is the crucial step that turns a text machine into a helpful assistant.

The Training Process of an LLM

Understanding

Why can an LLM not yet be a good assistant after pre-training alone?

The difference between the two phases shows in everyday use. When an LLM answers a question particularly clearly and in a structured way, recognizes dangers, or honestly says "I don't know" — that is a result of RLHF. The language ability itself comes from pre-training.

A concrete example: if you ask an LLM "How do I bake a cake?", the model after pre-training alone might output text that starts with a cake recipe, then transitions into a blog post about bakeries, and finally ends up at a newspaper article about wheat prices. After RLHF, it gives you a clear, helpful step-by-step guide.

Apply

The next time you use an LLM, watch for these signs that RLHF has had an effect:

  • The model structures its answer clearly (headings, bullet points, steps)
  • It warns you about dangerous or sensitive topics
  • It says "I'm not sure" instead of inventing false information
  • It asks for clarification when your question is unclear

Outlook: RLHF was a breakthrough, but research continues. Newer methods like Constitutional AI (developed by Anthropic, the company behind Claude) let the model partially evaluate itself based on a set of principles. Another method called DPO (Direct Preference Optimization) simplifies the training process. These methods are explained in more detail in the Advanced path.

Reflect

A friend says: 'ChatGPT knows everything because it has read the entire internet.' What do you reply?