AI Agents
How to test AI agents that you cannot predict - unit testing tool handlers, integration testing the agentic loop, writing LLM evals.
Testing a deterministic function is straightforward: given input X, assert output Y. Testing an AI agent is different in a fundamental way: given input X, you cannot assert a specific output Y, because the same input may produce output Y on one run and output Z on another, both of which are correct. The standard unit-test mental model breaks down. What you can test - and what actually matters for production confidence - is a different set of properties: tool handler correctness, loop termination behaviour, output structure, and goal achievement. This post covers a four-layer testing strategy that builds meaningful confidence in agent behaviour without requiring determinism.
Tool handlers are pure functions: they receive structured input, do something (call an API, query a database, write a file), and return a result. Test them like any other function - with mocked dependencies and exact assertions. This layer catches the bugs that would otherwise only surface during expensive agentic runs:
import { describe, it, expect, vi } from 'vitest';
// Mock the external dependencies - not the tool handler itself
vi.mock('../src/http', () => ({
fetch: vi.fn(),
}));
describe('webSearch tool handler', () => {
it('returns formatted results for a valid query', async () => {
const mockFetch = vi.mocked(await import('../src/http')).fetch;
mockFetch.mockResolvedValue({
ok: true,
json: async () => ({
web: {
results: [
{ title: 'Test Result', url: 'https://example.com', description: 'A test result' },
],
},
}),
});
const result = await webSearch('test query', 1);
expect(result).toContain('[1] Test Result');
expect(result).toContain('https://example.com');
});
it('throws a descriptive error when the search API returns 429', async () => {
const mockFetch = vi.mocked(await import('../src/http')).fetch;
mockFetch.mockResolvedValue({ ok: false, status: 429 });
await expect(webSearch('query', 1)).rejects.toThrow('Search API error: 429');
});
it('caps num_results at 10 regardless of input', async () => {
// Implementation should enforce the 1-10 range
const mockFetch = vi.mocked(await import('../src/http')).fetch;
mockFetch.mockResolvedValue({ ok: true, json: async () => ({ web: { results: [] } }) });
await webSearch('query', 100); // Should cap internally
const callUrl = mockFetch.mock.calls[0][0] as string;
expect(callUrl).toContain('count=10');
});
});
describe('executeTools', () => {
it('returns is_error true when a tool handler throws', async () => {
const blocks: Anthropic.ContentBlock[] = [{
type: 'tool_use',
id: 'tu_123',
name: 'web_search',
input: { query: 'test' },
}];
// Force the handler to throw
vi.spyOn(global'fetch').mockRejectedValue(new Error('Network error'));
const results = await executeTools(blocks);
expect(results[0].is_error).toBe(true);
expect(results[0].content).toContain('Network error');
});
});
Integration tests verify that the loop structure itself is correct: that tool results are fed back properly, that the loop terminates on end_turn, that the max-iteration guard fires. Mock the Anthropic client so tests run instantly without API calls or cost:
import { describe, it, expect, vi } from 'vitest';
vi.mock('@anthropic-ai/sdk');
describe('agentic loop structure', () => {
it('terminates on end_turn and returns the final text', async () => {
const mockCreate = vi.fn().mockResolvedValue({
stop_reason: 'end_turn',
content: [{ type: 'text', text: 'Task complete.' }],
});
vi.mocked(Anthropic).mockImplementation(() => ({
messages: { create: mockCreate },
} as any));
const result = await runAgentLoop('Do something', [], 10);
expect(result).toBe('Task complete.');
expect(mockCreate).toHaveBeenCalledTimes(1);
});
it('throws when max iterations are reached without end_turn', async () => {
const mockCreate = vi.fn().mockResolvedValue({
stop_reason: 'tool_use',
content: [{ type: 'tool_use', id: 'tu_1', name: 'web_search', input: { query: 'x' } }],
});
vi.mocked(Anthropic).mockImplementation(() => ({
messages: { create: mockCreate },
} as any));
await expect(runAgentLoop('Do something', [], 3)).rejects.toThrow(
'Agent did not complete within 3 iterations'
);
expect(mockCreate).toHaveBeenCalledTimes(3);
});
it('feeds tool results back as a user message', async () => {
const mockCreate = vi.fn()
.mockResolvedValueOnce({
stop_reason: 'tool_use',
content: [{ type: 'tool_use', id: 'tu_1', name: 'web_search', input: { query: 'test' } }],
})
.mockResolvedValueOnce({
stop_reason: 'end_turn',
content: [{ type: 'text', text: 'Done.' }],
});
vi.mocked(Anthropic).mockImplementation(() => ({
messages: { create: mockCreate },
} as any));
await runAgentLoop('Search for test', [webSearchTool], 5);
// Second call should include the tool result as a user message
const secondCallMessages = mockCreate.mock.calls[1][0].messages;
const lastMessage = secondCallMessages[secondCallMessages.length - 1];
expect(lastMessage.role).toBe('user');
expect(Array.isArray(lastMessage.content)).toBe(true);
expect(lastMessage.content[0].type).toBe('tool_result');
});
});
Evals test whether the agent achieves its goal on representative inputs. They use the real LLM and cost real API budget - run them in CI on every PR, but sample inputs rather than running the full eval suite on every commit. Three eval assertion types that work for agentic outputs:
import Anthropic from '@anthropic-ai/sdk';
const client = new Anthropic();
// Eval type 1: Structural assertion - output matches required shape
async function evalOutputStructure(agentOutput: string, requiredFields: string[]): Promise {
try {
const parsed = JSON.parse(agentOutput);
return requiredFields.every(field => field in parsed);
} catch {
return false;
}
}
// Eval type 2: LLM-as-judge - correctness relative to a criterion
async function evalWithLLMJudge(
task: string,
agentOutput: string,
criterion: string
): Promise<{ pass: boolean; score: number; reasoning: string }> {
const judgeResponse = await client.messages.create({
model: 'claude-sonnet-5',
max_tokens: 512,
system: 'You are an impartial evaluator. Score the agent output against the criterion. Be strict.',
messages: [{
role: 'user',
content: `Task: ${task}
Agent output:
${agentOutput}
Criterion: ${criterion}
Respond with JSON only: { "pass": boolean, "score": 0-10, "reasoning": string }`,
}],
});
const text = judgeResponse.content[0].type === 'text' ? judgeResponse.content[0].text : '{}';
return JSON.parse(text);
}
// Eval type 3: Behavioural assertion - agent took expected actions
async function evalToolUsage(
agentMessages: Anthropic.MessageParam[],
requirements: {
mustUseTools: string[]; // Tools that should have been called
mustNotUseTools: string[]; // Tools that should NOT have been called
minToolCalls: number;
}
): Promise<{ pass: boolean; violations: string[] }> {
const toolCalls = agentMessages
.flatMap(m => Array.isArray(m.content) ? m.content : [])
.filter((b): b is Anthropic.ToolUseBlock => (b as any).type === 'tool_use')
.map(b => b.name);
const violations: string[] = [];
for (const required of requirements.mustUseTools) {
if (!toolCalls.includes(required)) {
violations.push(`Required tool not called: ${required}`);
}
}
for (const forbidden of requirements.mustNotUseTools) {
if (toolCalls.includes(forbidden)) {
violations.push(`Forbidden tool was called: ${forbidden}`);
}
}
if (toolCalls.length < requirements.minToolCalls) {
violations.push(`Too few tool calls: ${toolCalls.length} < ${requirements.minToolCalls}`);
}
return { pass: violations.length === 0, violations };
}
For agent tasks where the output should be stable (a formatting task, a classification task, a structured extraction), save the output from a known-good run as a golden file and compare future outputs against it - not with exact string matching but with a similarity threshold or field-level comparison:
import { writeFileSync, readFileSync, existsSync } from 'fs';
async function goldenTest(
testName: string,
agentFn: () => Promise
A practical CI pipeline for agents: Layer 1 (unit tests, mocked) runs on every commit - fast, free, no API calls. Layer 2 (loop integration tests, mocked) runs on every PR - still fast and free. Layer 3 (real LLM evals) runs on PR merges to main and on a nightly schedule - costly but authoritative. Use a separate Anthropic API key for CI with a monthly budget cap configured in the Anthropic console to prevent runaway eval costs. For tracking eval results over time and catching gradual quality regressions, see agent tracing and observability. For managing the API costs of running evals against different model tiers, see agent cost management.