Zum Inhalt springen

Testing & Optimizing the Workflow

iAs of: September 2026

The cost figures are ballpark values based on September 2026 rates. Since the Sonnet price cut, real reviews land toward the lower end of the ranges shown. Still run your own numbers — cost depends heavily on PR size and iteration count.

Learn

Walking Through the Workflow

Before using the workflow in daily work, test it systematically. Here are three scenarios you should walk through:

Scenario 1 -- Local Change:

  1. Modify a TypeScript file (e.g., add a new function)
  2. Observe: The PostToolUse hooks fire automatically (Prettier formats, ESLint checks)
  3. The code is immediately clean -- without manual cleanup

Scenario 2 -- Review via Skill:

  1. Invoke /review
  2. The skill starts the code-reviewer agent in fork context
  3. The agent reads the diff, checks against the rules in CLAUDE.md, and provides structured feedback

Scenario 3 -- PR Review with GitHub:

  1. Create a PR on GitHub
  2. Invoke /review <PR-number>
  3. The MCP server fetches the PR diff, the agent reviews and provides feedback with file references

iDebugging Tip

If a step doesn't work, check in this order: 1) Is the MCP server reachable? (gh auth status) 2) Is the agent correctly defined? (Check syntax in code-reviewer.md) 3) Are the hooks correctly configured? (Validate JSON syntax in settings.json)

Calculating Costs

An important aspect of every workflow: What does it cost? Costs depend on the chosen model and the size of the reviews.

ModelPer Review (small)Per Review (large)5 Reviews/Day20 Reviews/Day
Haiku~$0.01-0.03~$0.05-0.10~$0.05-0.50~$0.20-2.00
Sonnet~$0.05-0.15~$0.15-0.40~$0.25-2.00~$1.00-8.00
Opus~$0.20-0.50~$0.50-1.50~$1.00-7.50~$4.00-30.00

*Keeping Costs Under Control

With Sonnet as the default model and 5 reviews per day, you're looking at $0.25-2.50 daily. That's significantly cheaper than manual reviews, which cost 15-30 minutes of developer time per review. Use maxTurns in the agent definition to prevent infinite loops.

A team does 10 PRs per day and uses Sonnet for reviews. What daily costs should they roughly expect?

Understand

Security Best Practices

An automated workflow that reads code and executes shell commands requires thoughtful security measures:

Token Splitting: Use different API keys for different environments. The key for local development should have different permissions than the one for CI/CD. This limits the damage if a key is compromised.

Sandboxing: Use /sandbox for OS-level isolation. The review agent can analyze code without having access to your entire system. Especially important when reviewing third-party code.

MCP Servers: Only connect trusted MCP servers. Every MCP server is a potential attack surface. Check the source before connecting a server, and give it only the minimally necessary permissions.

Hooks as Safety Net: Use exit code 2 to block dangerous actions. A pre-commit hook can, for example, prevent files with hardcoded secrets from being committed:

#!/bin/bash
# .claude/hooks/secret-check.sh
if grep -rn "PRIVATE_KEY\|SECRET_KEY\|password=" --include="*.ts" .; then
  echo "BLOCKED: Possible secrets found in code!"
  exit 2
fi
exit 0

!Principle of Least Privilege

Give the agent only the permissions it truly needs. A review agent needs read access to code and PR data, but no write access to the repository. Configure the MCP server permissions accordingly.

Iterating and Optimizing

Your first workflow won't be perfect -- and that's okay. Here are the most important levers for optimization:

Tune triggers: Are the skill triggers specific enough? If /review delivers too many false-positive findings, refine the instructions in SKILL.md. Add examples of desired and undesired feedback.

Hook order: Fast checks first. If Prettier takes 0.5 seconds and ESLint takes 5 seconds, start with Prettier. This way you get faster feedback in the success case.

Adjust agent model:

  • Haiku for simple reviews: formatting, naming conventions, TODO comments
  • Sonnet for standard reviews: bugs, performance, code quality
  • Opus for complex reviews: architecture decisions, security audits, API design

Refine the prompt: Based on review quality:

  • Too many findings? Raise the severity threshold or add "Ignore style issues."
  • Too few findings? Add more specific review criteria.
  • Unclear feedback? Explicitly request code examples for improvement suggestions.

Apply

Checklist for Your Own Workflow

Go through these points before using your workflow in production:

  • CLAUDE.md contains all relevant project and review rules
  • Agent definition is clear and specific (review criteria, output format)
  • MCP server is correctly configured and reachable
  • Hooks run without errors (test each one individually)
  • Costs are calculated and acceptable
  • Security: No secrets in the repository, sandboxing active, minimal permissions
  • Skill /review works both locally and with a PR number

Next Steps

Once your review workflow is running, you can transfer the concept to other tasks:

  • /test -- Skill that automatically generates tests for new functions
  • /docs -- Skill that creates documentation from code comments
  • /refactor -- Skill that makes improvement suggestions for existing code
  • /security -- Specialized skill just for security audits

The pattern is always the same: define the agent, create the skill, configure hooks, test, optimize.

*Workflow as Template

Your finished review workflow can serve as a template for other projects. Copy the .claude/ structure, adapt CLAUDE.md to the new project, and you immediately have a working workflow. Share good workflows with your team!

Reflect

You've now built a complete, automated code review workflow -- from project configuration through skills and hooks to security and cost optimization. The key takeaway: This workflow is not a static product but a living system that you continuously improve. Every review, every piece of feedback, and every false alarm is an opportunity to make the workflow better. That's exactly what makes Claude Code so powerful: You're not just building one-off solutions, but systems that grow with you and your team.