Plan-and-Execute Pattern
Knowledge
ReAct is powerful but expensive: every thinking step consumes tokens on a high-performance (and costly) model. The plan-and-execute pattern solves this problem elegantly: an expensive model plans, a cheap model executes.
The Architecture
- Planner (e.g., Claude Opus, GPT-5.6) -- Creates a detailed execution plan once
- Executor (e.g., Claude Haiku 4.5, GPT-5 mini) -- Carries out the individual steps of the plan
- Evaluator (optional, the expensive model again) -- Reviews the overall result at the end
Why 90% Cost Savings Are Possible
Let's work through a concrete example:
| Aspect | Pure ReAct (Opus) | Plan-and-Execute |
|---|---|---|
| Planning phase | 5 thinking steps x Opus | 1 planning step x Opus |
| Execution | 10 tool calls x Opus | 10 tool calls x Haiku |
| Evaluation | -- | 1 review step x Opus |
| Opus tokens | ~15,000 | ~3,000 |
| Haiku tokens | 0 | ~8,000 |
| Estimated cost | ~$0.075 | ~$0.023 |
The trick: executing individual steps doesn't require brilliant reasoning. "Call this API with these parameters" is something a small, cheap model can handle just as well. The expensive model is reserved for strategic planning and final quality control.
iPrice Comparison (as of 2026)
Claude Opus: ~$5/1M input tokens. Claude Haiku: ~$1/1M input tokens. That's a factor of 5. Even if the planner needs somewhat more tokens, the savings are enormous -- especially on output token costs ($25 vs. $5/1M).
Understand
How the Plan Works
Phase 1: Planning (expensive model)
The planner receives the task and the list of available tools. It creates a structured plan:
{
"plan": [
{
"step": 1,
"action": "search_knowledge_base",
"params": { "query": "return policy premium customers" },
"reason": "Load the current policy first"
},
{
"step": 2,
"action": "get_customer_info",
"params": { "email": "customer@example.com" },
"reason": "Check customer status (premium or standard)"
},
{
"step": 3,
"action": "check_order_status",
"params": { "order_id": "ORD-2025-4832" },
"depends_on": [1, 2],
"reason": "Review order with context from steps 1 and 2"
},
{
"step": 4,
"action": "compose_response",
"params": { "template": "return_approval" },
"depends_on": [3],
"reason": "Compose response based on all gathered information"
}
]
}
Phase 2: Execution (cheap model)
The executor works through the plan step by step. It doesn't need its own reasoning -- it simply follows the instructions and makes the tool calls.
Phase 3: Evaluation (expensive model, optional)
At the end, the evaluator checks: Was the task solved correctly? Are the results consistent? If not, it can adjust the plan and start another round.
When to Use Plan-and-Execute vs. ReAct?
| Criterion | ReAct | Plan-and-Execute |
|---|---|---|
| Task complexity | Medium | High |
| Predictability of steps | Low (steps depend on results) | High (steps can be planned in advance) |
| Cost | Higher (everything on the expensive model) | Lower (execution on the cheap model) |
| Flexibility | High (can change course at any time) | Medium (plan needs adjusting when surprises arise) |
| Latency | Medium | Lower (steps can be parallelized) |
Use ReAct when the next steps heavily depend on the results and you need flexibility.
Use Plan-and-Execute when the task structure is predictable and you want to optimize costs.
*Hybrid Approach
In practice, many systems combine both patterns: plan-and-execute for the overall structure, but within individual steps the executor can use a mini ReAct loop when unexpected results come up.
Apply
An agent is supposed to summarize and categorize 200 news articles every morning. The format is always the same. Which pattern makes the most sense here?
Reflect
Plan-and-Execute separates thinking from doing -- a planner model creates the plan, a cheap executor model carries it out. For predictable, recurring tasks, this pattern saves up to 90% in costs compared to ReAct. The art lies in recognizing when steps are predictable enough.