Zum Inhalt springen

Hands-On: Production Agent

iAs of: May 2026

Model names and token prices in the code examples reflect the state as of May 2026 (Anthropic tariffs). The concepts (routing, monitoring, multi-agent) are stable — only update the model IDs and prices before running in production.

From Prototype to Production Agent

In the previous modules, you've designed multi-agent systems, mastered RAG architectures, understood fine-tuning, and internalized AI safety. Now you're bringing it all together -- building an agent that doesn't just work, but runs reliably, securely, and cost-efficiently in production.

The difference between a prototype and a production agent is like the difference between a sketch and a load-bearing building: The core idea is the same, but the requirements for stability, maintainability, and security are fundamentally different.

*Learning Objective

After this module, you'll be able to set up a multi-agent system for production use, implement monitoring and alerting, calculate and optimize costs, configure secure execution environments with Docker, and apply human-in-the-loop patterns. You'll be operating at Bloom's levels of "Evaluate" and "Create."

What Separates Production from Development?

In development, an agent is successful when it solves the task. In production, it must additionally:

Be reliable: It can't crash after 2 hours because an API call fails. Error handling, retries, and fallbacks are mandatory.

Be observable: You must know at all times what your agent is doing, how many tokens it's consuming, and whether errors are occurring. Without monitoring, you're flying blind.

Be secure: An agent with access to file systems, databases, or APIs can cause enormous damage if it runs unchecked. Sandboxing and permission boundaries aren't optional extras -- they're requirements.

Be cost-efficient: An agent that costs $500/month when $80 would suffice is not a good agent -- regardless of how good its results are.

Enable human control: Not every decision should be made by an agent alone. Human-in-the-loop patterns define when a human needs to intervene.

Prerequisites

Before you start, make sure you've completed the following modules:

  • Module 01 -- Multi-Agent Systems: You know orchestration patterns and agent communication
  • Module 02 -- RAG Deep Dive: You understand context-based architectures
  • Module 03 -- Fine-Tuning & Custom Models: You know when a custom model makes sense
  • Module 04 -- AI Safety & Governance: You know the risks and regulatory requirements

!Technical Requirements

For the hands-on exercises, you'll need: Python 3.10+ or Node.js 18+, Docker Desktop, an API key for Claude or GPT, and basic knowledge of container technology. Code examples are shown in Python and TypeScript.

What to Expect

This module consists of five building blocks, each one building on the last:

Step 1 -- Multi-Agent Setup: You build a planner-executor-reviewer architecture. Three agents work together, each with a clear role.

Step 2 -- Monitoring: You implement token tracking, error logging, and alerting. By the end, you'll see what your system is doing in a dashboard.

Step 3 -- Production Costs: You calculate what a workday with AI agents costs and optimize for a realistic budget.

Step 4 -- Docker Sandbox: You set up a secure execution environment where your agent can use tools without endangering the host system.

Step 5 -- Human-in-the-Loop: You define when a human needs to decide and build approval gates and escalation patterns.

Ready?

Let's get started -- with building your multi-agent system.

Reflect

In this hands-on module, you will build a complete multi-agent system for production. From architecture to monitoring, Docker sandboxing, and human-in-the-loop -- you will learn every building block. Let us get started with building your multi-agent system.