Production Costs
iAs of: September 2026
Token prices and model names reflect the state as of September 2026 (Anthropic tariffs). Verify the official price list before doing cost calculations — providers adjust tariffs regularly, and since Opus 4.7 the Opus line uses a new tokenizer that produces around 30 % more tokens per request than older models.
Knowledge
"What does a workday with AI agents cost?" -- every manager asks this question, and the answer is rarely as simple as "the API price times the number of calls." Production costs include direct token costs, infrastructure, monitoring, and the often underestimated factor: human oversight.
The Four Cost Layers
| Layer | Examples | Share of Total Cost |
|---|---|---|
| Direct Token Costs | API calls to LLM providers | 40-60% |
| Infrastructure | Servers, Docker, databases, storage | 15-25% |
| Monitoring & Ops | Logging tools, alerting, dashboards | 5-10% |
| Human Oversight | Reviewer time, escalation handling | 15-30% |
Realistic Cost Calculation: One Workday
Let's take a concrete scenario: A multi-agent system for automated code review that processes 50 pull requests on a typical workday.
Agent Configuration:
- Planner: Claude Opus (1 call/PR, ~2,000 input tokens, ~500 output tokens)
- Executor: Claude Haiku (5 calls/PR, ~1,500 input tokens, ~300 output tokens)
- Reviewer: Claude Sonnet (1 call/PR, ~3,000 input tokens, ~800 output tokens)
Daily calculation for 50 PRs:
| Agent | Calls/Day | Input Tokens | Output Tokens | Cost/Day |
|---|---|---|---|---|
| Planner (Opus) | 50 | 100,000 | 25,000 | $1.13 |
| Executor (Haiku) | 250 | 375,000 | 75,000 | $0.75 |
| Reviewer (Sonnet) | 50 | 150,000 | 40,000 | $1.05 |
| Total | 350 | 625,000 | 140,000 | $2.93/day |
That's roughly $64/month in token costs alone. With infrastructure and monitoring, you're looking at approximately $120-170/month.
Understanding
Cost Models Compared
There are three common strategies for managing agent costs:
1. Pay-per-Use (Standard)
- You pay per token as the API charges
- Predictable with stable volume
- Risk: Unexpected spikes (agent loop explosions)
2. Batched Processing
- Anthropic and OpenAI offer batch APIs with 50% discount
- Tasks are collected and processed with a delay
- Ideal for non-time-critical tasks (e.g., nightly code reviews)
3. Tiered Model Routing
- Simple tasks to Haiku, medium to Sonnet, complex to Opus
- A router agent or rule-based system decides
- Saves 30-50% compared to a uniform model
def route_to_model(task_complexity: str) -> str:
"""Routes tasks to the appropriate model based on complexity."""
routing = {
"simple": "claude-haiku-4-5", # Formatting, simple checks
"medium": "claude-sonnet-5", # Code review, analysis
"complex": "claude-opus-5", # Architecture decisions
}
return routing.get(task_complexity, "claude-sonnet-5")
def estimate_daily_cost(
tasks_per_day: int,
avg_tokens_per_task: int,
model: str
) -> float:
"""Estimates daily costs for an agent."""
pricing = {
"claude-opus-5": 5.0,
"claude-sonnet-5": 3.0,
"claude-haiku-4-5": 1.0,
}
price_per_million = pricing.get(model, 3.0)
total_tokens = tasks_per_day * avg_tokens_per_task
return (total_tokens / 1_000_000) * price_per_million
Budgeting: The 80/20 Rule
In practice: 80% of costs come from 20% of tasks. These are typically complex tasks where the agent needs many iterations or where the Planner uses an expensive model.
Optimization levers (sorted by effectiveness):
- Context Trimming: Summarize old results instead of sending them in full (saves 30-40%)
- Model Routing: Route simple steps to a cheap model (saves 20-30%)
- Batch Processing: Collect non-time-critical tasks (saves 50% on those tasks)
- Caching: Cache repeated requests (saves 10-20%)
- Plan Optimization: Instruct the Planner to create compact plans (saves 15-25%)
Apply
Your agent system costs $300/month. The Planner (Opus) causes $180 of that, despite making only 15% of the calls. Which optimization has the biggest impact?
Reflect
Cost management for AI agents isn't a one-time setup but a continuous process. Prices change, new models appear, and your usage patterns evolve. The key is good monitoring (previous section) that enables you to optimize based on data rather than guessing. An agent that costs $100/month and saves a developer 2 hours/day is an excellent investment. An agent that costs $500/month and saves 30 minutes is not.