AI Agents

Build a Code Review Agent with Claude: Step-by-Step (2026)

Build a production code review agent using Claude - git diff ingestion, structured finding output, severity classification.

Code review is one of the most immediately useful things you can build with an AI agent, and it demonstrates a pattern that appears everywhere in production agent systems: read a structured artefact (the diff), analyse it against a rubric (review criteria), and return machine-readable output (structured findings). The version that most tutorials build stops at "ask Claude to review this code" and get back prose. The version you actually want to ship returns structured findings with severity levels, file:line references, and specific fix suggestions - output that can be posted to a PR, routed to different reviewers by severity, or fed into a CI gate. This tutorial builds the latter.

Code review agent
A code review agent is an AI agent that ingests a git diff or a set of files, analyses the changes against correctness, security, and style criteria, and returns structured findings with file path, line number, severity classification, and fix suggestion - enabling automated PR review comments and CI integration without human review of every change.

What we're building

git diff or file list
        │
        ▼
┌───────────────────────────────┐
│  Code Review Agent            │
│  (claude-sonnet-5)            │
│                               │
│  Tools:                       │
│   ├── read_diff               │
│   └── read_file               │
│                               │
│  Output via tool:             │
│   └── submit_findings         │
└───────────────┬───────────────┘
                │
                ▼
┌───────────────────────────────┐
│  Structured findings array    │
│  [{                           │
│    file, line, severity,      │
│    category, description,     │
│    suggestion                 │
│  }]                           │
└───────────────────────────────┘

Prerequisites: Node.js 18+, an Anthropic API key, and a git repository to review.

Step 1 - Define the tools

import Anthropic from '@anthropic-ai/sdk';


const client = new Anthropic({ apiKey: process.env.ANTHROPIC_API_KEY });

const tools: Anthropic.Tool[] = [
  {
    name: 'read_diff',
    description: 'Read the git diff for staged changes, the last commit, or a specific commit range. Use this first to understand what changed.',
    input_schema: {
      type: 'object' as const,
      properties: {
        target: {
          type: 'string',
          description: 'What to diff: "staged" (git diff --staged), "head" (git diff HEAD), or a commit SHA or range like "abc123..def456".',
        },
        path_filter: {
          type: 'string',
          description: 'Optional: limit diff to a specific path or glob (e.g., "src/" or "*.ts").',
        },
      },
      required: ['target'],
    },
  },
  {
    name: 'read_file',
    description: 'Read the full content of a file. Use this to see context around a changed line when the diff is not sufficient.',
    input_schema: {
      type: 'object' as const,
      properties: {
        path: { type: 'string', description: 'Relative path to the file.' },
        start_line: { type: 'number', description: 'Optional: first line to return (1-indexed).' },
        end_line: { type: 'number', description: 'Optional: last line to return (inclusive).' },
      },
      required: ['path'],
    },
  },
  {
    name: 'submit_findings',
    description: 'Submit the final structured code review findings. Call this when the review is complete.',
    input_schema: {
      type: 'object' as const,
      properties: {
        findings: {
          type: 'array',
          items: {
            type: 'object',
            properties: {
              file: { type: 'string', description: 'Relative file path.' },
              line: { type: 'number', description: 'Line number the finding refers to.' },
              severity: {
                type: 'string',
                enum: ['critical', 'high', 'medium', 'low'],
                description: 'critical: security/data loss. high: correctness. medium: performance/maintainability. low: style.',
              },
              category: {
                type: 'string',
                enum: ['security', 'correctness', 'performance', 'maintainability', 'style'],
              },
              description: { type: 'string', description: 'What the issue is and why it matters.' },
              suggestion: { type: 'string', description: 'Specific fix or improved code snippet.' },
            },
            required: ['file', 'line', 'severity', 'category', 'description', 'suggestion'],
          },
        },
        summary: {
          type: 'string',
          description: '2-3 sentence overall assessment of the changes.',
        },
        approve: {
          type: 'boolean',
          description: 'Whether the changes are approvable. False if any critical or high findings exist.',
        },
      },
      required: ['findings', 'summary', 'approve'],
    },
  },
];

Step 2 - Implement tool handlers

function readDiff(target: string, pathFilter?: string): string {
  const pathArg = pathFilter ? `-- ${pathFilter}` : ', ';

  const commands: Record = {
    staged: `git diff --staged ${pathArg}`,
    head: `git diff HEAD ${pathArg}`,
  };

  const command = commands[target] ?? `git diff ${target} ${pathArg}`;

  try {
    const output = execSync(command, { encoding: 'utf8', maxBuffer: 10 * 1024 * 1024 });
    if (!output.trim()) return 'No changes found for the specified target.';

    // Truncate very large diffs to fit in context
    const MAX_DIFF_CHARS = 40000;
    if (output.length > MAX_DIFF_CHARS) {
      return output.slice(0, MAX_DIFF_CHARS) + `

[Diff truncated - ${output.length - MAX_DIFF_CHARS} chars omitted. Use read_file for specific sections.]`;
    }

    return output;
  } catch (err) {
    return `Error reading diff: ${(err as Error).message}`;
  }
}

function readFile(path: string, startLine?: number, endLine?: number): string {
  if (!existsSync(path)) return `File not found: ${path}`;

  const content = readFileSync(path'utf8');
  const lines = content.split('
');

  if (startLine !== undefined && endLine !== undefined) {
    const slice = lines.slice(startLine - 1, endLine);
    return slice.map((line, i) => `${startLine + i}: ${line}`).join('
');
  }

  // Return with line numbers for reference
  const MAX_FILE_CHARS = 20000;
  const annotated = lines.map((line, i) => `${i + 1}: ${line}`).join('
');
  if (annotated.length > MAX_FILE_CHARS) {
    return annotated.slice(0, MAX_FILE_CHARS) + '
[File truncated]';
  }
  return annotated;
}

async function executeTool(
  name: string,
  input: Record
): Promise {
  switch (name) {
    case 'read_diff':
      return readDiff(input.target as string, input.path_filter as string | undefined);
    case 'read_file':
      return readFile(input.path as string, input.start_line as number, input.end_line as number);
    default:
      return `Unknown tool: ${name}`;
  }
}

Step 3 - The review loop

interface ReviewResult {
  findings: Finding[];
  summary: string;
  approve: boolean;
}

interface Finding {
  file: string;
  line: number;
  severity: 'critical' | 'high' | 'medium' | 'low';
  category: string;
  description: string;
  suggestion: string;
}

const REVIEW_SYSTEM_PROMPT = `You are a senior code reviewer. Your goal is to identify real issues - not nitpicks.

Review priorities (in order):
1. Security vulnerabilities: SQL injection, XSS, auth bypass, secret exposure, path traversal
2. Correctness bugs: off-by-one errors, null dereferences, race conditions, incorrect logic
3. Performance problems: N+1 queries, unnecessary allocations, blocking operations in hot paths
4. Maintainability: overly complex logic, missing error handling, inadequate tests
5. Style: only flag style issues if they introduce ambiguity or will cause bugs

Rules:
- Only report findings for lines that actually exist in the diff or that you have read with read_file
- Do not report speculative issues ("this might cause problems if...") - only actual issues in the code
- Suggest specific fixes, not general advice
- When in doubt, read the full file context with read_file before filing a finding`;

async function runCodeReview(target = 'staged'): Promise {
  const messages: Anthropic.MessageParam[] = [
    {
      role: 'user',
      content: `Review the code changes for: ${target}. Use read_diff to fetch the changes, read_file for context where needed, and submit_findings when complete.`,
    },
  ];

  const MAX_ITERATIONS = 10;

  for (let i = 0; i < MAX_ITERATIONS; i++) {
    const response = await client.messages.create({
      model: 'claude-sonnet-5',
      max_tokens: 4096,
      system: REVIEW_SYSTEM_PROMPT,
      tools,
      messages,
    });

    messages.push({ role: 'assistant', content: response.content });

    if (response.stop_reason === 'tool_use') {
      const toolCalls = response.content.filter(
        (b): b is Anthropic.ToolUseBlock => b.type === 'tool_use'
      );

      const toolResults = await Promise.all(toolCalls.map(async (call) => {
        if (call.name === 'submit_findings') {
          // Final output - return the findings structure
          return { type: 'tool_result' as const, tool_use_id: call.id, content: 'Findings submitted.' };
        }
        const result = await executeTool(call.name, call.input as Record);
        return { type: 'tool_result' as const, tool_use_id: call.id, content: result };
      }));

      messages.push({ role: 'user', content: toolResults });

      // Check if submit_findings was called
      const submission = toolCalls.find(c => c.name === 'submit_findings');
      if (submission) {
        return submission.input as ReviewResult;
      }
    }

    if (response.stop_reason === 'end_turn') {
      return { findings: [], summary: 'Review completed without structured findings.', approve: true };
    }
  }

  throw new Error('Code review agent did not complete within iteration limit');
}

Handling failure modes

  • Very large diffs - a PR with 10,000 lines of diff will exceed the context window. Partition large diffs by file and run one review agent per file, then aggregate findings. The batch processing pattern applies directly here.
  • Agent reads irrelevant files - the agent may read_file on files not in the diff. This is usually the model seeking context; it is expensive but not wrong. Cap read_file calls with a counter and return an error after 5 calls: "File read limit reached. Proceed with available context."
  • Findings with wrong line numbers - the model occasionally references a line number that is off by 1-2. Mitigate by having the read_file tool return content with explicit line number prefixes, and by including "verify line numbers using read_file before filing a finding" in the system prompt.

Integrating into CI

Run the agent as a GitHub Actions step on every PR. The target parameter accepts a commit range, so pass origin/main..HEAD to review all commits in the PR branch. Post findings as PR review comments using the GitHub API, filtering to severity critical and high for blocking reviews and posting medium/low as informational comments. For building the full agentic loop that the review agent runs on, see the agentic loop explained. For managing the cost of running code review agents across many PRs, see agent cost management strategies.