Zum Inhalt springen

Production Costs

iAs of: September 2026

Token prices and model names reflect the state as of September 2026 (Anthropic tariffs). Verify the official price list before doing cost calculations — providers adjust tariffs regularly, and since Opus 4.7 the Opus line uses a new tokenizer that produces around 30 % more tokens per request than older models.

Knowledge

"What does a workday with AI agents cost?" -- every manager asks this question, and the answer is rarely as simple as "the API price times the number of calls." Production costs include direct token costs, infrastructure, monitoring, and the often underestimated factor: human oversight.

The Four Cost Layers

LayerExamplesShare of Total Cost
Direct Token CostsAPI calls to LLM providers40-60%
InfrastructureServers, Docker, databases, storage15-25%
Monitoring & OpsLogging tools, alerting, dashboards5-10%
Human OversightReviewer time, escalation handling15-30%

Realistic Cost Calculation: One Workday

Let's take a concrete scenario: A multi-agent system for automated code review that processes 50 pull requests on a typical workday.

Agent Configuration:

  • Planner: Claude Opus (1 call/PR, ~2,000 input tokens, ~500 output tokens)
  • Executor: Claude Haiku (5 calls/PR, ~1,500 input tokens, ~300 output tokens)
  • Reviewer: Claude Sonnet (1 call/PR, ~3,000 input tokens, ~800 output tokens)

Daily calculation for 50 PRs:

AgentCalls/DayInput TokensOutput TokensCost/Day
Planner (Opus)50100,00025,000$1.13
Executor (Haiku)250375,00075,000$0.75
Reviewer (Sonnet)50150,00040,000$1.05
Total350625,000140,000$2.93/day

That's roughly $64/month in token costs alone. With infrastructure and monitoring, you're looking at approximately $120-170/month.

Understanding

Cost Models Compared

There are three common strategies for managing agent costs:

1. Pay-per-Use (Standard)

  • You pay per token as the API charges
  • Predictable with stable volume
  • Risk: Unexpected spikes (agent loop explosions)

2. Batched Processing

  • Anthropic and OpenAI offer batch APIs with 50% discount
  • Tasks are collected and processed with a delay
  • Ideal for non-time-critical tasks (e.g., nightly code reviews)

3. Tiered Model Routing

  • Simple tasks to Haiku, medium to Sonnet, complex to Opus
  • A router agent or rule-based system decides
  • Saves 30-50% compared to a uniform model
def route_to_model(task_complexity: str) -> str:
    """Routes tasks to the appropriate model based on complexity."""
    routing = {
        "simple": "claude-haiku-4-5",     # Formatting, simple checks
        "medium": "claude-sonnet-5",     # Code review, analysis
        "complex": "claude-opus-5",      # Architecture decisions
    }
    return routing.get(task_complexity, "claude-sonnet-5")

def estimate_daily_cost(
    tasks_per_day: int,
    avg_tokens_per_task: int,
    model: str
) -> float:
    """Estimates daily costs for an agent."""
    pricing = {
        "claude-opus-5": 5.0,
        "claude-sonnet-5": 3.0,
        "claude-haiku-4-5": 1.0,
    }
    price_per_million = pricing.get(model, 3.0)
    total_tokens = tasks_per_day * avg_tokens_per_task
    return (total_tokens / 1_000_000) * price_per_million

Budgeting: The 80/20 Rule

In practice: 80% of costs come from 20% of tasks. These are typically complex tasks where the agent needs many iterations or where the Planner uses an expensive model.

Optimization levers (sorted by effectiveness):

  1. Context Trimming: Summarize old results instead of sending them in full (saves 30-40%)
  2. Model Routing: Route simple steps to a cheap model (saves 20-30%)
  3. Batch Processing: Collect non-time-critical tasks (saves 50% on those tasks)
  4. Caching: Cache repeated requests (saves 10-20%)
  5. Plan Optimization: Instruct the Planner to create compact plans (saves 15-25%)

Apply

Your agent system costs $300/month. The Planner (Opus) causes $180 of that, despite making only 15% of the calls. Which optimization has the biggest impact?

Reflect

Cost management for AI agents isn't a one-time setup but a continuous process. Prices change, new models appear, and your usage patterns evolve. The key is good monitoring (previous section) that enables you to optimize based on data rather than guessing. An agent that costs $100/month and saves a developer 2 hours/day is an excellent investment. An agent that costs $500/month and saves 30 minutes is not.