Guardrails & Reality Check
Knowledge
Here's the uncomfortable truth: AI agents make mistakes. They're not deterministic programs that always do the same thing. They're built on LLMs, and LLMs are probabilistic -- they generate probable, not guaranteed correct outputs.
This means: an agent that worked perfectly yesterday can fail on the exact same task tomorrow. Not because something is broken, but because the LLM chose a different token sequence.
!Core Principle
Treat every agent output as "probably correct, but not guaranteed." Build your system as if any step could fail -- because it can.
What Can Go Wrong?
1. Hallucination in the Reasoning Chain
The agent fabricates facts and builds on them. This is especially dangerous because the rest of the chain looks logically consistent, yet rests on a false foundation.
Example: A support agent claims there's a "30-day premium return guarantee" that doesn't actually exist -- and then correctly initiates a return that should never have happened.
2. Tool Misuse
The agent calls the wrong tool or passes incorrect parameters. Particularly critical with write operations.
Example: An agent is supposed to read customer data (get_customer), but accidentally calls delete_customer because the tool names sound similar.
3. Infinite Loops
The agent gets stuck in a loop: it doesn't realize it's repeating the same action because each iteration is phrased slightly differently.
Example: The agent searches for information, doesn't find it, rephrases the search, doesn't find it again -- 50 times in a row.
4. Scope Creep
The agent interprets the task more broadly than intended and performs actions that were never planned.
Example: "Optimize the website" leads the agent to independently modify the database credentials because it sees "optimization potential."
Understand
Guardrail Strategies
Error Handling and Retry Logic
def safe_tool_call(tool, params, max_retries=3):
for attempt in range(max_retries):
try:
result = tool.execute(params)
if validate_result(result):
return result
# Invalid result -- retry with feedback
params = adjust_params(params, result)
except ToolError as e:
if attempt == max_retries - 1:
return fallback_response(e)
# Wait briefly, then try again
time.sleep(2 ** attempt)
Important: Don't just blindly retry! Give the agent feedback on why the attempt failed so it can adjust its strategy.
Output Validation
Check every agent response against defined rules before it reaches the user:
- Schema validation -- Does the output match the expected format?
- Fact checking -- Are the mentioned numbers, dates, and policies accurate?
- Security check -- Does the response contain sensitive data that shouldn't be exposed?
- Relevance check -- Does the response actually answer the question asked?
Human-in-the-Loop
For critical actions: build in human approval.
Hover or clickTap a level to see details
| Risk Level | Example | Strategy |
|---|---|---|
| Low | Reading information | Execute automatically |
| Medium | Sending an email | Agent drafts, human approves |
| High | Deleting data, transferring money | Human approval + four-eyes principle |
*Rule of Thumb
Anything that can't be undone needs human approval. Read operations: automatic. Write operations: with caution. Delete operations: always with approval.
Max Steps Limit
Always set an upper bound on the number of steps an agent can take. 10-20 steps are sufficient for most tasks. Without a limit, a malfunctioning agent can run indefinitely and rack up costs.
Tool Permissions
Not every agent needs access to every tool. Assign permissions following the least-privilege principle:
- A research agent only needs read access
- A support agent needs read and limited write access
- Only an admin agent gets full access -- and even then, only with human-in-the-loop
Apply
The Reality Check: What You Need to Know
- Agents aren't reliable enough for autonomous, critical decisions -- as of 2026, they need human oversight for anything with real consequences.
- The error rate is dropping, but it will never hit zero -- probabilistic systems always have a residual error rate.
- Costs can explode -- an agent without a max-steps limit or with too many retries can rack up hundreds of dollars in API costs within minutes.
- Logging is mandatory -- you must log every thinking step, every tool call, and every result. Without logs, you can't debug failures.
Your agent is supposed to automatically cancel customer orders when a customer requests it. Which guardrail strategy is most important?
Reflect
Guardrails are not a luxury but a necessity. Input validation, output checks, and human-in-the-loop protect your system from the typical errors of probabilistic agents. The rule of thumb: the harder an action is to reverse, the more control you need.