Zum Inhalt springen

Reasoning Models

iAs of: September 2026

Specific model names (o1, o3, Claude Opus 5, GPT-5.6, etc.) reflect the state as of September 2026. The concepts (Thinking Tokens, Chain-of-Thought, Extended Thinking) are long-lived — the example models become outdated quickly.

Knowledge

Until 2024, LLMs could only "think" in a structured way if you explicitly prompted them to: "Think step by step" (Chain-of-Thought prompting). That changed fundamentally with reasoning models -- models that natively work through an internal thinking process before responding.

The key reasoning models:

  • OpenAI o1 / o3: The first models with native chain-of-thought. They use "thinking tokens" that are processed internally but not directly shown to the user.
  • Claude Opus 5 with Extended Thinking: Anthropic's approach -- the thinking process is returned as summarized thinking (Adaptive Thinking).
  • GPT-5.6: Offers a configurable thinking mode with different intensity levels.

iThinking Tokens

Thinking tokens are tokens the model generates during its thinking process. They count toward token usage but are not included in the output. With strong reasoning models, this can amount to thousands of tokens before the model writes its actual response.

Understand

Prompt-based CoT vs. Native Reasoning

Prompt-based Chain-of-Thought (2022-2024):

Prompt: "Think step by step: If a train departs at 2:00 PM
and takes 3.5 hours, when does it arrive?"

The model follows an instruction. The quality of reasoning depends heavily on the prompt. If you forget "think step by step", the model might jump straight to the answer and make mistakes.

Native Reasoning (from 2024 onwards):

The model thinks automatically, regardless of the prompt. It internally breaks complex problems into sub-steps, checks intermediate results, and self-corrects. This happens in the thinking tokens.

The difference is fundamental: it's like the difference between someone who only thinks when told to and someone who naturally reflects.

When to use which model?

TaskStandard LLMReasoning Model
Writing an emailSufficientOverkill
Solving a math problemError-proneSignificantly better
Code debuggingBasicExcellent
Creative writingGoodSlower, not necessary
Logic puzzlesOften wrongReliable
Simple questionsFast & cheapToo expensive & slow

*Cost-Benefit

Reasoning models consume significantly more tokens (and therefore money) than standard models. For a simple summary, a reasoning model can use 10x the tokens. Use reasoning models specifically for tasks that require genuine thinking.

Apply

When working with Claude Opus 5 or GPT-5.6 Sol, observe the thinking process:

  1. Ask a simple question: "What is the capital of Germany?" -- The model responds immediately, barely any thinking tokens.
  2. Ask a complex question: "Explain why the number 0.999... equals 1, using three different proof methods." -- You can see how the model explores approaches, discards some, and refines others.

Observing this trains your intuition for when native reasoning provides real value.

Reflect

A company uses a reasoning model for its chatbot that answers simple FAQ questions. What is the main problem?