Zum Inhalt springen

Calculating Costs

iAs of: September 2026

Token prices and model names reflect the state as of September 2026. Note: since Opus 4.7, the Opus line uses a new tokenizer that produces around 30 % more tokens per request than older models — factor in a surcharge when migrating.

Knowledge

Your agent works -- it thinks, acts, and delivers results. But what does it all cost? Due to their loop, agents consume significantly more tokens than a single chat call. In this section, you'll learn to estimate costs realistically and optimize them.

Why Agents Are More Expensive Than Chatbots

A simple chatbot call: 1 request, ~500 input tokens, ~500 output tokens. An agent for the same task: 3-8 requests (Think-Act-Observe per iteration), each with a growing context (the entire conversation history is sent along every time).

Example calculation for a code review agent:

IterationInput TokensOutput TokensCost (Claude Sonnet)
1: Think + Act800200$0.005
2: Observe + Think + Act1,500300$0.009
3: Observe + Think + Act2,400250$0.011
4: Observe + Final Answer3,200500$0.017
Total7,9001,250$0.042

A single agent run costs about 4 cents. Sounds like nothing -- but scale it up: 100 code reviews per day is $4.20/day or ~$126/month.

Analyzing Token Consumption

The biggest cost drivers for agents:

  1. Context growth: Each iteration sends the entire conversation history. Iteration 5 has five times the input of iteration 1.
  2. Unnecessary tool calls: A poorly configured agent calls tools it doesn't need.
  3. Large tool results: A read_file on a 1,000-line file sends the entire contents as tokens.
  4. Missing termination condition: The agent keeps running even though the task was completed long ago.

Understand

Plan-and-Execute for Cost Optimization

You know plan-and-execute from Module 04. It's also the best strategy for cost optimization:

Without plan-and-execute (pure ReAct):

  • Agent tries, observes, thinks, tries again -- many iterations
  • Context grows with every iteration
  • Typical: 5-10 iterations per task

With plan-and-execute:

  • A large model creates the plan (1 call)
  • A small, cheap model executes each step (several cheap calls)
  • Typical: 1 expensive + 3-5 cheap calls

Cost savings: Often 50-70% compared to pure ReAct with a large model.

More Optimization Strategies

  • Trim context: Summarize old tool results instead of sending them in full
  • Limit tool results: Read only the first 100 lines of a file, not all 1,000
  • Caching: Cache repeated requests to the same tool
  • Model routing: Route simple steps to a cheap model, complex ones to an expensive model

Try It Out: The Cost Calculator

With the following dashboard, you can calculate the costs for your agent. Choose a model, enter the expected usage, and see what ends up on the bill at the end of the month.

Cost Dashboard

103 Requests/Day
Moderate

Token Distribution

Input (67%)Output (33%)

Cost Breakdown / Day

InputOutput
$0.41$1.03

Projection

Per Day

$1.44

Per Week

$7.21

Per Month

$31.72

Apply

Your agent consumes 50,000 tokens per task and is called 200 times per day. What is the most effective measure to reduce costs?

Reflect

Cost management isn't an afterthought -- it's a central part of agent architecture. The choice between ReAct and plan-and-execute, model selection, and context management determine whether your agent costs $50 or $500 per month. A good agent isn't just one that delivers the best results -- it's one that delivers the best results within budget.