AI Agents

AI Agent Testing Strategies: Unit Tests, Evals and Integration

How to test AI agents that you cannot predict - unit testing tool handlers, integration testing the agentic loop, writing LLM evals.

Testing a deterministic function is straightforward: given input X, assert output Y. Testing an AI agent is different in a fundamental way: given input X, you cannot assert a specific output Y, because the same input may produce output Y on one run and output Z on another, both of which are correct. The standard unit-test mental model breaks down. What you can test - and what actually matters for production confidence - is a different set of properties: tool handler correctness, loop termination behaviour, output structure, and goal achievement. This post covers a four-layer testing strategy that builds meaningful confidence in agent behaviour without requiring determinism.

LLM agent evaluation (eval)
An LLM agent evaluation is a structured test that assesses an agent's output against a goal criterion rather than an exact expected value - typically using another LLM as a judge, a structured rubric, or behavioural assertions - because agent outputs are variable and correct by degree rather than by exact match.

Layer 1 - Unit test the tool handlers (fully deterministic)

Tool handlers are pure functions: they receive structured input, do something (call an API, query a database, write a file), and return a result. Test them like any other function - with mocked dependencies and exact assertions. This layer catches the bugs that would otherwise only surface during expensive agentic runs:

import { describe, it, expect, vi } from 'vitest';

// Mock the external dependencies - not the tool handler itself
vi.mock('../src/http', () => ({
  fetch: vi.fn(),
}));

describe('webSearch tool handler', () => {
  it('returns formatted results for a valid query', async () => {
    const mockFetch = vi.mocked(await import('../src/http')).fetch;
    mockFetch.mockResolvedValue({
      ok: true,
      json: async () => ({
        web: {
          results: [
            { title: 'Test Result', url: 'https://example.com', description: 'A test result' },
          ],
        },
      }),
    });

    const result = await webSearch('test query', 1);
    expect(result).toContain('[1] Test Result');
    expect(result).toContain('https://example.com');
  });

  it('throws a descriptive error when the search API returns 429', async () => {
    const mockFetch = vi.mocked(await import('../src/http')).fetch;
    mockFetch.mockResolvedValue({ ok: false, status: 429 });

    await expect(webSearch('query', 1)).rejects.toThrow('Search API error: 429');
  });

  it('caps num_results at 10 regardless of input', async () => {
    // Implementation should enforce the 1-10 range
    const mockFetch = vi.mocked(await import('../src/http')).fetch;
    mockFetch.mockResolvedValue({ ok: true, json: async () => ({ web: { results: [] } }) });

    await webSearch('query', 100); // Should cap internally
    const callUrl = mockFetch.mock.calls[0][0] as string;
    expect(callUrl).toContain('count=10');
  });
});

describe('executeTools', () => {
  it('returns is_error true when a tool handler throws', async () => {
    const blocks: Anthropic.ContentBlock[] = [{
      type: 'tool_use',
      id: 'tu_123',
      name: 'web_search',
      input: { query: 'test' },
    }];

    // Force the handler to throw
    vi.spyOn(global'fetch').mockRejectedValue(new Error('Network error'));

    const results = await executeTools(blocks);
    expect(results[0].is_error).toBe(true);
    expect(results[0].content).toContain('Network error');
  });
});

Layer 2 - Integration test the agentic loop (with mock LLM)

Integration tests verify that the loop structure itself is correct: that tool results are fed back properly, that the loop terminates on end_turn, that the max-iteration guard fires. Mock the Anthropic client so tests run instantly without API calls or cost:

import { describe, it, expect, vi } from 'vitest';


vi.mock('@anthropic-ai/sdk');

describe('agentic loop structure', () => {
  it('terminates on end_turn and returns the final text', async () => {
    const mockCreate = vi.fn().mockResolvedValue({
      stop_reason: 'end_turn',
      content: [{ type: 'text', text: 'Task complete.' }],
    });

    vi.mocked(Anthropic).mockImplementation(() => ({
      messages: { create: mockCreate },
    } as any));

    const result = await runAgentLoop('Do something', [], 10);
    expect(result).toBe('Task complete.');
    expect(mockCreate).toHaveBeenCalledTimes(1);
  });

  it('throws when max iterations are reached without end_turn', async () => {
    const mockCreate = vi.fn().mockResolvedValue({
      stop_reason: 'tool_use',
      content: [{ type: 'tool_use', id: 'tu_1', name: 'web_search', input: { query: 'x' } }],
    });

    vi.mocked(Anthropic).mockImplementation(() => ({
      messages: { create: mockCreate },
    } as any));

    await expect(runAgentLoop('Do something', [], 3)).rejects.toThrow(
      'Agent did not complete within 3 iterations'
    );
    expect(mockCreate).toHaveBeenCalledTimes(3);
  });

  it('feeds tool results back as a user message', async () => {
    const mockCreate = vi.fn()
      .mockResolvedValueOnce({
        stop_reason: 'tool_use',
        content: [{ type: 'tool_use', id: 'tu_1', name: 'web_search', input: { query: 'test' } }],
      })
      .mockResolvedValueOnce({
        stop_reason: 'end_turn',
        content: [{ type: 'text', text: 'Done.' }],
      });

    vi.mocked(Anthropic).mockImplementation(() => ({
      messages: { create: mockCreate },
    } as any));

    await runAgentLoop('Search for test', [webSearchTool], 5);

    // Second call should include the tool result as a user message
    const secondCallMessages = mockCreate.mock.calls[1][0].messages;
    const lastMessage = secondCallMessages[secondCallMessages.length - 1];
    expect(lastMessage.role).toBe('user');
    expect(Array.isArray(lastMessage.content)).toBe(true);
    expect(lastMessage.content[0].type).toBe('tool_result');
  });
});

Layer 3 - Behavioural evals (with real LLM, sampled)

Evals test whether the agent achieves its goal on representative inputs. They use the real LLM and cost real API budget - run them in CI on every PR, but sample inputs rather than running the full eval suite on every commit. Three eval assertion types that work for agentic outputs:

import Anthropic from '@anthropic-ai/sdk';

const client = new Anthropic();

// Eval type 1: Structural assertion - output matches required shape
async function evalOutputStructure(agentOutput: string, requiredFields: string[]): Promise {
  try {
    const parsed = JSON.parse(agentOutput);
    return requiredFields.every(field => field in parsed);
  } catch {
    return false;
  }
}

// Eval type 2: LLM-as-judge - correctness relative to a criterion
async function evalWithLLMJudge(
  task: string,
  agentOutput: string,
  criterion: string
): Promise<{ pass: boolean; score: number; reasoning: string }> {
  const judgeResponse = await client.messages.create({
    model: 'claude-sonnet-5',
    max_tokens: 512,
    system: 'You are an impartial evaluator. Score the agent output against the criterion. Be strict.',
    messages: [{
      role: 'user',
      content: `Task: ${task}

Agent output:
${agentOutput}

Criterion: ${criterion}

Respond with JSON only: { "pass": boolean, "score": 0-10, "reasoning": string }`,
    }],
  });

  const text = judgeResponse.content[0].type === 'text' ? judgeResponse.content[0].text : '{}';
  return JSON.parse(text);
}

// Eval type 3: Behavioural assertion - agent took expected actions
async function evalToolUsage(
  agentMessages: Anthropic.MessageParam[],
  requirements: {
    mustUseTools: string[];    // Tools that should have been called
    mustNotUseTools: string[]; // Tools that should NOT have been called
    minToolCalls: number;
  }
): Promise<{ pass: boolean; violations: string[] }> {
  const toolCalls = agentMessages
    .flatMap(m => Array.isArray(m.content) ? m.content : [])
    .filter((b): b is Anthropic.ToolUseBlock => (b as any).type === 'tool_use')
    .map(b => b.name);

  const violations: string[] = [];

  for (const required of requirements.mustUseTools) {
    if (!toolCalls.includes(required)) {
      violations.push(`Required tool not called: ${required}`);
    }
  }

  for (const forbidden of requirements.mustNotUseTools) {
    if (toolCalls.includes(forbidden)) {
      violations.push(`Forbidden tool was called: ${forbidden}`);
    }
  }

  if (toolCalls.length < requirements.minToolCalls) {
    violations.push(`Too few tool calls: ${toolCalls.length} < ${requirements.minToolCalls}`);
  }

  return { pass: violations.length === 0, violations };
}

Layer 4 - Regression tests with golden outputs

For agent tasks where the output should be stable (a formatting task, a classification task, a structured extraction), save the output from a known-good run as a golden file and compare future outputs against it - not with exact string matching but with a similarity threshold or field-level comparison:

import { writeFileSync, readFileSync, existsSync } from 'fs';

async function goldenTest(
  testName: string,
  agentFn: () => Promise,
  updateGolden = false
): Promise {
  const goldenPath = `./tests/golden/${testName}.json`;
  const actual = await agentFn();

  if (updateGolden || !existsSync(goldenPath)) {
    writeFileSync(goldenPath, JSON.stringify(actual, null, 2));
    console.log(`Golden file written: ${goldenPath}`);
    return;
  }

  const expected = JSON.parse(readFileSync(goldenPath'utf8'));

  // Field-level comparison - not exact string match
  for (const [key, value] of Object.entries(expected)) {
    if (!(key in (actual as object))) {
      throw new Error(`Golden regression: missing field "${key}" in actual output`);
    }
    // For string fields, check that the actual includes key terms from the golden
    if (typeof value === 'string' && typeof (actual as any)[key] === 'string') {
      const goldenTerms = value.split(/s+/).filter(w => w.length > 5);
      const missingTerms = goldenTerms.filter(
        term => !(actual as any)[key].toLowerCase().includes(term.toLowerCase())
      );
      if (missingTerms.length > goldenTerms.length * 0.3) {
        throw new Error(`Golden regression on field "${key}": output diverged significantly`);
      }
    }
  }
}

CI setup for agent tests

A practical CI pipeline for agents: Layer 1 (unit tests, mocked) runs on every commit - fast, free, no API calls. Layer 2 (loop integration tests, mocked) runs on every PR - still fast and free. Layer 3 (real LLM evals) runs on PR merges to main and on a nightly schedule - costly but authoritative. Use a separate Anthropic API key for CI with a monthly budget cap configured in the Anthropic console to prevent runaway eval costs. For tracking eval results over time and catching gradual quality regressions, see agent tracing and observability. For managing the API costs of running evals against different model tiers, see agent cost management.