Zum Inhalt springen

Agent Monitoring

iAs of: September 2026

Model IDs and token prices in the code examples reflect the state as of September 2026 (Anthropic tariffs). The monitoring concepts are long-lived — only update the model IDs and prices before running in production.

Knowledge

A production agent without monitoring is like a car without a dashboard: You don't know how fast you're going, how much fuel you're burning, or whether the engine is overheating. In this section, you'll build a monitoring system that shows you at all times what your multi-agent system is doing.

The Four Pillars of Agent Monitoring

PillarWhat Is Measured?Why Is It Important?
Token TrackingInput/output tokens per agent, per taskDetect cost explosions early
Error HandlingError types, frequency, affected agentsEnsure system stability
LoggingEvery agent action with timestamp and contextDebugging and audit trail
AlertingThresholds for cost, error rate, latencyAutomatic notification on problems

Implementing Token Tracking

Every API call returns the tokens consumed. You need to capture these systematically:

from dataclasses import dataclass, field
from datetime import datetime

@dataclass
class TokenUsage:
    agent_name: str
    model: str
    input_tokens: int
    output_tokens: int
    timestamp: str = field(default_factory=lambda: datetime.now().isoformat())

    @property
    def total_tokens(self) -> int:
        return self.input_tokens + self.output_tokens

    @property
    def estimated_cost(self) -> float:
        """Estimates cost based on the model."""
        pricing = {
            "claude-opus-5": (5.0, 25.0),      # $/1M tokens (input, output)
            "claude-sonnet-5": (3.0, 15.0),
            "claude-haiku-4-5": (1.0, 5.0),
        }
        input_price, output_price = pricing.get(self.model, (3.0, 15.0))
        return (
            (self.input_tokens / 1_000_000) * input_price +
            (self.output_tokens / 1_000_000) * output_price
        )

class TokenTracker:
    def __init__(self):
        self.records: list[TokenUsage] = []

    def record(self, agent_name: str, model: str, response) -> TokenUsage:
        usage = TokenUsage(
            agent_name=agent_name,
            model=model,
            input_tokens=response.usage.input_tokens,
            output_tokens=response.usage.output_tokens
        )
        self.records.append(usage)
        return usage

    def total_cost(self) -> float:
        return sum(r.estimated_cost for r in self.records)

    def cost_by_agent(self) -> dict[str, float]:
        costs: dict[str, float] = {}
        for r in self.records:
            costs[r.agent_name] = costs.get(r.agent_name, 0) + r.estimated_cost
        return costs

Understanding

Error Handling with Retry Logic

In production, API calls fail -- rate limits, timeouts, server errors. Your system must handle this:

import time
import logging
from functools import wraps

logger = logging.getLogger("agent-system")

def retry_with_backoff(max_retries: int = 3, base_delay: float = 1.0):
    """Decorator for automatic retries with exponential backoff."""
    def decorator(func):
        @wraps(func)
        def wrapper(*args, **kwargs):
            for attempt in range(max_retries):
                try:
                    return func(*args, **kwargs)
                except Exception as e:
                    if attempt == max_retries - 1:
                        logger.error(f"All {max_retries} attempts failed: {e}")
                        raise
                    delay = base_delay * (2 ** attempt)
                    logger.warning(
                        f"Attempt {attempt + 1} failed: {e}. "
                        f"Retry in {delay}s..."
                    )
                    time.sleep(delay)
        return wrapper
    return decorator

# Usage in the ProductionAgent
class ProductionAgent:
    @retry_with_backoff(max_retries=3, base_delay=1.0)
    def process(self, task: str) -> str:
        response = client.messages.create(
            model=self.model,
            max_tokens=2048,
            system=self.system_prompt,
            messages=[{"role": "user", "content": task}]
        )
        # Token tracking
        self.tracker.record(self.name, self.model, response)
        return response.content[0].text

Alerting System

Define thresholds that trigger notifications when exceeded:

@dataclass
class AlertRule:
    name: str
    metric: str              # "cost_per_hour", "error_rate", "latency_p95"
    threshold: float
    window_minutes: int = 60
    action: str = "log"      # "log", "email", "slack", "pagerduty"

PRODUCTION_ALERTS = [
    AlertRule(name="Cost Alert", metric="cost_per_hour", threshold=5.0),
    AlertRule(name="Error Alert", metric="error_rate", threshold=0.1),
    AlertRule(name="Latency Alert", metric="latency_p95", threshold=30.0),
]

The Production Dashboard

The following dashboard shows you token consumption, costs, and error rates of your multi-agent system at a glance:

Production Agent Dashboard

Token Usage/Day

59.6k

↓ -3.1%

Cost/Day

$30.00

↑ +5.2%

Error Rate

1.7%

↑ +0.3%

Avg Latency

330ms

↓ -8ms

Token Usage (7 Days)

InputOutput
Token Usage over 7 Days0k13k26k39k52k66kMonTueWedThuFriSatSun

Cost per Agent ($/Day)

Cost per Agent$0$4$8$12$16Planner$14.0Executor$9.0Reviewer$5.0Monitor$2.0
Total: $30.00/Tag

Error Log (last 5)

05:08:52 AMExecutorRequest timed out after 30s
05:02:52 AMPlannerRate limit exceeded, retry in 60s
04:52:52 AMExecutorOpenAI API returned 503
04:50:52 AMReviewerFailed to parse JSON response
04:44:52 AMReviewerOpenAI API returned 503

Apply

Your production dashboard shows that the Executor agent causes 80% of total costs despite using the cheapest model. What is the most likely cause?

Reflect

Monitoring isn't an optional add-on -- it's the foundation for every production decision. Without token tracking, you don't know if your budget holds. Without error logging, you can't find bugs. Without alerting, you notice problems only when it's too late. Invest the time in good monitoring -- it's the cheapest insurance you can buy for your agent system.