Calculating Costs
iAs of: September 2026
Token prices and model names reflect the state as of September 2026. Note: since Opus 4.7, the Opus line uses a new tokenizer that produces around 30 % more tokens per request than older models — factor in a surcharge when migrating.
Knowledge
Your agent works -- it thinks, acts, and delivers results. But what does it all cost? Due to their loop, agents consume significantly more tokens than a single chat call. In this section, you'll learn to estimate costs realistically and optimize them.
Why Agents Are More Expensive Than Chatbots
A simple chatbot call: 1 request, ~500 input tokens, ~500 output tokens. An agent for the same task: 3-8 requests (Think-Act-Observe per iteration), each with a growing context (the entire conversation history is sent along every time).
Example calculation for a code review agent:
| Iteration | Input Tokens | Output Tokens | Cost (Claude Sonnet) |
|---|---|---|---|
| 1: Think + Act | 800 | 200 | $0.005 |
| 2: Observe + Think + Act | 1,500 | 300 | $0.009 |
| 3: Observe + Think + Act | 2,400 | 250 | $0.011 |
| 4: Observe + Final Answer | 3,200 | 500 | $0.017 |
| Total | 7,900 | 1,250 | $0.042 |
A single agent run costs about 4 cents. Sounds like nothing -- but scale it up: 100 code reviews per day is $4.20/day or ~$126/month.
Analyzing Token Consumption
The biggest cost drivers for agents:
- Context growth: Each iteration sends the entire conversation history. Iteration 5 has five times the input of iteration 1.
- Unnecessary tool calls: A poorly configured agent calls tools it doesn't need.
- Large tool results: A
read_fileon a 1,000-line file sends the entire contents as tokens. - Missing termination condition: The agent keeps running even though the task was completed long ago.
Understand
Plan-and-Execute for Cost Optimization
You know plan-and-execute from Module 04. It's also the best strategy for cost optimization:
Without plan-and-execute (pure ReAct):
- Agent tries, observes, thinks, tries again -- many iterations
- Context grows with every iteration
- Typical: 5-10 iterations per task
With plan-and-execute:
- A large model creates the plan (1 call)
- A small, cheap model executes each step (several cheap calls)
- Typical: 1 expensive + 3-5 cheap calls
Cost savings: Often 50-70% compared to pure ReAct with a large model.
More Optimization Strategies
- Trim context: Summarize old tool results instead of sending them in full
- Limit tool results: Read only the first 100 lines of a file, not all 1,000
- Caching: Cache repeated requests to the same tool
- Model routing: Route simple steps to a cheap model, complex ones to an expensive model
Try It Out: The Cost Calculator
With the following dashboard, you can calculate the costs for your agent. Choose a model, enter the expected usage, and see what ends up on the bill at the end of the month.
Cost Dashboard
Token Distribution
Cost Breakdown / Day
Projection
Per Day
$1.44
Per Week
$7.21
Per Month
$31.72
Apply
Your agent consumes 50,000 tokens per task and is called 200 times per day. What is the most effective measure to reduce costs?
Reflect
Cost management isn't an afterthought -- it's a central part of agent architecture. The choice between ReAct and plan-and-execute, model selection, and context management determine whether your agent costs $50 or $500 per month. A good agent isn't just one that delivers the best results -- it's one that delivers the best results within budget.