Zum Inhalt springen

Guardrails & Reality Check

Knowledge

Here's the uncomfortable truth: AI agents make mistakes. They're not deterministic programs that always do the same thing. They're built on LLMs, and LLMs are probabilistic -- they generate probable, not guaranteed correct outputs.

This means: an agent that worked perfectly yesterday can fail on the exact same task tomorrow. Not because something is broken, but because the LLM chose a different token sequence.

!Core Principle

Treat every agent output as "probably correct, but not guaranteed." Build your system as if any step could fail -- because it can.

What Can Go Wrong?

1. Hallucination in the Reasoning Chain

The agent fabricates facts and builds on them. This is especially dangerous because the rest of the chain looks logically consistent, yet rests on a false foundation.

Example: A support agent claims there's a "30-day premium return guarantee" that doesn't actually exist -- and then correctly initiates a return that should never have happened.

2. Tool Misuse

The agent calls the wrong tool or passes incorrect parameters. Particularly critical with write operations.

Example: An agent is supposed to read customer data (get_customer), but accidentally calls delete_customer because the tool names sound similar.

3. Infinite Loops

The agent gets stuck in a loop: it doesn't realize it's repeating the same action because each iteration is phrased slightly differently.

Example: The agent searches for information, doesn't find it, rephrases the search, doesn't find it again -- 50 times in a row.

4. Scope Creep

The agent interprets the task more broadly than intended and performs actions that were never planned.

Example: "Optimize the website" leads the agent to independently modify the database credentials because it sees "optimization potential."

Understand

Guardrail Strategies

Error Handling and Retry Logic

def safe_tool_call(tool, params, max_retries=3):
    for attempt in range(max_retries):
        try:
            result = tool.execute(params)
            if validate_result(result):
                return result
            # Invalid result -- retry with feedback
            params = adjust_params(params, result)
        except ToolError as e:
            if attempt == max_retries - 1:
                return fallback_response(e)
            # Wait briefly, then try again
            time.sleep(2 ** attempt)

Important: Don't just blindly retry! Give the agent feedback on why the attempt failed so it can adjust its strategy.

Output Validation

Check every agent response against defined rules before it reaches the user:

  • Schema validation -- Does the output match the expected format?
  • Fact checking -- Are the mentioned numbers, dates, and policies accurate?
  • Security check -- Does the response contain sensitive data that shouldn't be exposed?
  • Relevance check -- Does the response actually answer the question asked?

Human-in-the-Loop

For critical actions: build in human approval.

Unacceptable RiskHigh RiskLimited RiskMinimal Risk

Tap a level to see details

Risk LevelExampleStrategy
LowReading informationExecute automatically
MediumSending an emailAgent drafts, human approves
HighDeleting data, transferring moneyHuman approval + four-eyes principle

*Rule of Thumb

Anything that can't be undone needs human approval. Read operations: automatic. Write operations: with caution. Delete operations: always with approval.

Max Steps Limit

Always set an upper bound on the number of steps an agent can take. 10-20 steps are sufficient for most tasks. Without a limit, a malfunctioning agent can run indefinitely and rack up costs.

Tool Permissions

Not every agent needs access to every tool. Assign permissions following the least-privilege principle:

  • A research agent only needs read access
  • A support agent needs read and limited write access
  • Only an admin agent gets full access -- and even then, only with human-in-the-loop

Apply

The Reality Check: What You Need to Know

  1. Agents aren't reliable enough for autonomous, critical decisions -- as of 2026, they need human oversight for anything with real consequences.
  2. The error rate is dropping, but it will never hit zero -- probabilistic systems always have a residual error rate.
  3. Costs can explode -- an agent without a max-steps limit or with too many retries can rack up hundreds of dollars in API costs within minutes.
  4. Logging is mandatory -- you must log every thinking step, every tool call, and every result. Without logs, you can't debug failures.

Your agent is supposed to automatically cancel customer orders when a customer requests it. Which guardrail strategy is most important?

Reflect

Guardrails are not a luxury but a necessity. Input validation, output checks, and human-in-the-loop protect your system from the typical errors of probabilistic agents. The rule of thumb: the harder an action is to reverse, the more control you need.