Agent Monitoring
iAs of: September 2026
Model IDs and token prices in the code examples reflect the state as of September 2026 (Anthropic tariffs). The monitoring concepts are long-lived — only update the model IDs and prices before running in production.
Knowledge
A production agent without monitoring is like a car without a dashboard: You don't know how fast you're going, how much fuel you're burning, or whether the engine is overheating. In this section, you'll build a monitoring system that shows you at all times what your multi-agent system is doing.
The Four Pillars of Agent Monitoring
| Pillar | What Is Measured? | Why Is It Important? |
|---|---|---|
| Token Tracking | Input/output tokens per agent, per task | Detect cost explosions early |
| Error Handling | Error types, frequency, affected agents | Ensure system stability |
| Logging | Every agent action with timestamp and context | Debugging and audit trail |
| Alerting | Thresholds for cost, error rate, latency | Automatic notification on problems |
Implementing Token Tracking
Every API call returns the tokens consumed. You need to capture these systematically:
from dataclasses import dataclass, field
from datetime import datetime
@dataclass
class TokenUsage:
agent_name: str
model: str
input_tokens: int
output_tokens: int
timestamp: str = field(default_factory=lambda: datetime.now().isoformat())
@property
def total_tokens(self) -> int:
return self.input_tokens + self.output_tokens
@property
def estimated_cost(self) -> float:
"""Estimates cost based on the model."""
pricing = {
"claude-opus-5": (5.0, 25.0), # $/1M tokens (input, output)
"claude-sonnet-5": (3.0, 15.0),
"claude-haiku-4-5": (1.0, 5.0),
}
input_price, output_price = pricing.get(self.model, (3.0, 15.0))
return (
(self.input_tokens / 1_000_000) * input_price +
(self.output_tokens / 1_000_000) * output_price
)
class TokenTracker:
def __init__(self):
self.records: list[TokenUsage] = []
def record(self, agent_name: str, model: str, response) -> TokenUsage:
usage = TokenUsage(
agent_name=agent_name,
model=model,
input_tokens=response.usage.input_tokens,
output_tokens=response.usage.output_tokens
)
self.records.append(usage)
return usage
def total_cost(self) -> float:
return sum(r.estimated_cost for r in self.records)
def cost_by_agent(self) -> dict[str, float]:
costs: dict[str, float] = {}
for r in self.records:
costs[r.agent_name] = costs.get(r.agent_name, 0) + r.estimated_cost
return costs
Understanding
Error Handling with Retry Logic
In production, API calls fail -- rate limits, timeouts, server errors. Your system must handle this:
import time
import logging
from functools import wraps
logger = logging.getLogger("agent-system")
def retry_with_backoff(max_retries: int = 3, base_delay: float = 1.0):
"""Decorator for automatic retries with exponential backoff."""
def decorator(func):
@wraps(func)
def wrapper(*args, **kwargs):
for attempt in range(max_retries):
try:
return func(*args, **kwargs)
except Exception as e:
if attempt == max_retries - 1:
logger.error(f"All {max_retries} attempts failed: {e}")
raise
delay = base_delay * (2 ** attempt)
logger.warning(
f"Attempt {attempt + 1} failed: {e}. "
f"Retry in {delay}s..."
)
time.sleep(delay)
return wrapper
return decorator
# Usage in the ProductionAgent
class ProductionAgent:
@retry_with_backoff(max_retries=3, base_delay=1.0)
def process(self, task: str) -> str:
response = client.messages.create(
model=self.model,
max_tokens=2048,
system=self.system_prompt,
messages=[{"role": "user", "content": task}]
)
# Token tracking
self.tracker.record(self.name, self.model, response)
return response.content[0].text
Alerting System
Define thresholds that trigger notifications when exceeded:
@dataclass
class AlertRule:
name: str
metric: str # "cost_per_hour", "error_rate", "latency_p95"
threshold: float
window_minutes: int = 60
action: str = "log" # "log", "email", "slack", "pagerduty"
PRODUCTION_ALERTS = [
AlertRule(name="Cost Alert", metric="cost_per_hour", threshold=5.0),
AlertRule(name="Error Alert", metric="error_rate", threshold=0.1),
AlertRule(name="Latency Alert", metric="latency_p95", threshold=30.0),
]
The Production Dashboard
The following dashboard shows you token consumption, costs, and error rates of your multi-agent system at a glance:
Production Agent Dashboard
Token Usage/Day
59.6k
↓ -3.1%Cost/Day
$30.00
↑ +5.2%Error Rate
1.7%
↑ +0.3%Avg Latency
330ms
↓ -8msToken Usage (7 Days)
Cost per Agent ($/Day)
Error Log (last 5)
Apply
Your production dashboard shows that the Executor agent causes 80% of total costs despite using the cheapest model. What is the most likely cause?
Reflect
Monitoring isn't an optional add-on -- it's the foundation for every production decision. Without token tracking, you don't know if your budget holds. Without error logging, you can't find bugs. Without alerting, you notice problems only when it's too late. Invest the time in good monitoring -- it's the cheapest insurance you can buy for your agent system.