AI Agents
AI agents fail in ways that traditional software does not - stuck loops, malformed tool calls, context overflow.
An agent that fails silently is worse than one that fails loudly. In a standard API call, failure is binary - you get a 200 or you do not. In an agentic loop, failure is a spectrum: the loop might complete but return a wrong answer, loop indefinitely without completing, produce a partially correct result before crashing, or generate tool calls with malformed arguments that cascade into further errors. Each failure mode requires a different recovery strategy, and applying a generic retry to all of them makes things worse, not better.
Agent error recovery is a set of layered strategies that detect, classify, and respond to different failure types without losing user trust or producing silent corruptions. The layers are: (1) input validation before the loop starts, (2) tool-call error surfacing inside the loop, (3) retry logic with backoff for transient API failures, (4) loop depth and cost circuit breakers, and (5) graceful degradation to simpler alternatives when the agent cannot complete.
Any agent in production that takes actions with real consequences - writes files, calls external APIs, sends messages, stores data - needs error recovery. An agent that only reads and synthesises can fail more gracefully (return partial results or an error message). An agent that writes to a database, posts to an API, or triggers workflows needs recovery logic that prevents partial state corruption.
import Anthropic from '@anthropic-ai/sdk';
const client = new Anthropic();
interface AgentRunResult {
success: boolean;
output?: string;
error?: string;
iterationsUsed: number;
fallbackUsed: boolean;
}
async function runWithRecovery(
goal: string,
tools: Anthropic.Tool[],
options: {
maxIterations?: number;
maxRetries?: number;
timeoutMs?: number;
fallbackFn?: (goal: string) => Promise;
} = {}
): Promise {
const { maxIterations = 20, maxRetries = 3, timeoutMs = 60_000, fallbackFn } = options;
// Wrap the entire run in a timeout
const timeoutPromise = new Promise((_, reject) =>
setTimeout(() => reject(new Error('Agent timeout')), timeoutMs)
);
try {
const result = await Promise.race([
runAgentLoop(goal, tools, maxIterations),
timeoutPromise,
]);
return { success: true, output: result, iterationsUsed: 0, fallbackUsed: false };
} catch (err) {
const error = err as Error;
// Classify the error
const isTransient = isTransientError(error);
const isLoopError = error.message.includes('max iterations');
const isTimeout = error.message.includes('timeout');
// Transient API errors: retry with backoff
if (isTransient && maxRetries > 0) {
await backoff(3 - maxRetries);
return runWithRecovery(goal, tools, { ...options, maxRetries: maxRetries - 1 });
}
// Loop exhaustion or timeout: try fallback
if ((isLoopError || isTimeout) && fallbackFn) {
try {
const fallbackOutput = await fallbackFn(goal);
return { success: true, output: fallbackOutput, iterationsUsed: maxIterations, fallbackUsed: true };
} catch (fallbackErr) {
// Fallback also failed - return graceful degradation response
return {
success: false,
error: 'The task could not be completed automatically. Please try again or simplify your request.',
iterationsUsed: maxIterations,
fallbackUsed: true,
};
}
}
return {
success: false,
error: error.message,
iterationsUsed: maxIterations,
fallbackUsed: false,
};
}
}
function isTransientError(error: Error): boolean {
const transientPatterns = ['529', '529', 'overloaded', 'rate limit', 'ECONNRESET', 'timeout'];
return transientPatterns.some(p => error.message.toLowerCase().includes(p));
}
async function backoff(attempt: number): Promise {
const delayMs = Math.min(1000 * Math.pow(2, attempt) + Math.random() * 500, 30_000);
await new Promise(r => setTimeout(r, delayMs));
}
The most common error inside an agent loop is a failed tool call. How you return that error to the model determines whether it recovers or spirals. A tool result that silently returns an empty string on failure tells the model nothing - it will retry the same call with the same arguments. A tool result with an explicit error message and a hint about what to try instead gives the model what it needs to change course:
async function executeToolWithErrorHandling(
toolName: string,
toolInput: unknown,
executors: Record Promise>
): Promise {
const executor = executors[toolName];
if (!executor) {
return {
type: 'tool_result',
tool_use_id: '', // filled in by caller
content: `Error: No executor found for tool "${toolName}". Available tools: ${Object.keys(executors).join('')}`,
is_error: true,
};
}
try {
const result = await executor(toolInput);
const serialised = typeof result === 'string' ? result : JSON.stringify(result, null, 2);
// Truncate large results to prevent context overflow
const maxChars = 8_000;
const content = serialised.length > maxChars
? serialised.slice(0, maxChars) + `
[Result truncated at ${maxChars} characters. Use a more specific query to get a smaller result.]`
: serialised;
return { type: 'tool_result', tool_use_id: '', content };
} catch (err) {
const error = err as Error;
return {
type: 'tool_result',
tool_use_id: '',
content: `Error executing ${toolName}: ${error.message}. Consider trying a different approach or different input.`,
is_error: true,
};
}
}
Long-running agents accumulate tool results in the message history until the total context exceeds the model's limit. The model then throws a context window error, and if not handled, the entire run fails. The prevention is progressive history summarisation - when the accumulated context exceeds a threshold, replace the middle of the conversation with a summary:
async function summariseHistoryIfNeeded(
messages: Anthropic.MessageParam[],
tokenThreshold = 60_000
): Promise {
const estimatedTokens = messages.reduce((sum, m) => {
const content = typeof m.content === 'string' ? m.content : JSON.stringify(m.content);
return sum + content.length / 4; // rough: 4 chars ≈ 1 token
}, 0);
if (estimatedTokens < tokenThreshold) return messages;
// Keep the first message (original goal) and last 4 messages (recent context)
const first = messages[0];
const recent = messages.slice(-4);
const middle = messages.slice(1, -4);
if (middle.length === 0) return messages; // Already minimal, cannot summarise
const summaryResponse = await client.messages.create({
model: 'claude-haiku-4-5-20251001', // Cheap model for summarisation
max_tokens: 1024,
messages: [{
role: 'user',
content: `Summarise the following agent conversation history in 500 words or less, preserving all key findings and decisions:
${JSON.stringify(middle)}`,
}],
});
const summary = summaryResponse.content[0].type === 'text' ? summaryResponse.content[0].text : ', ';
return [
first,
{ role: 'user', content: `[History summary]: ${summary}` },
{ role: 'assistant', content: 'Understood. Continuing from where we left off.' },
...recent,
];
}
When an agent cannot complete a task - due to repeated tool failures, context overflow, or loop exhaustion - the right response is not an error modal. It is a graceful step down to a simpler alternative:
The rule: every agent that a user depends on for a real outcome needs an explicit fallback path defined before it ships to production. A fallback written under pressure after an incident is always worse than one designed in advance.
The most destructive error handling patterns in production agents:
stop_reason: 'end_turn' is not a guarantee of correctness, only of completionFor the upstream patterns that define how errors propagate across multiple agents, see the agent supervisor pattern - supervisor-level failure handling differs from in-loop recovery. For tracing which iteration produced an error in production, see agent tracing and observability.