AI Agents

Agent Error Recovery: Retry, Fallback, and Graceful Degradation

AI agents fail in ways that traditional software does not - stuck loops, malformed tool calls, context overflow.

An agent that fails silently is worse than one that fails loudly. In a standard API call, failure is binary - you get a 200 or you do not. In an agentic loop, failure is a spectrum: the loop might complete but return a wrong answer, loop indefinitely without completing, produce a partially correct result before crashing, or generate tool calls with malformed arguments that cascade into further errors. Each failure mode requires a different recovery strategy, and applying a generic retry to all of them makes things worse, not better.

Agent error recovery
Agent error recovery is a set of layered strategies applied within an agentic loop to detect, classify, and respond to failure types specific to AI agents - including stuck loops, malformed tool calls, context window overflow, and transient API errors - without producing silent corruptions or unbounded retries.

The agent error recovery pattern - what it is

Agent error recovery is a set of layered strategies that detect, classify, and respond to different failure types without losing user trust or producing silent corruptions. The layers are: (1) input validation before the loop starts, (2) tool-call error surfacing inside the loop, (3) retry logic with backoff for transient API failures, (4) loop depth and cost circuit breakers, and (5) graceful degradation to simpler alternatives when the agent cannot complete.

When to apply it

Any agent in production that takes actions with real consequences - writes files, calls external APIs, sends messages, stores data - needs error recovery. An agent that only reads and synthesises can fail more gracefully (return partial results or an error message). An agent that writes to a database, posts to an API, or triggers workflows needs recovery logic that prevents partial state corruption.

Implementation: the error-aware agent loop

import Anthropic from '@anthropic-ai/sdk';

const client = new Anthropic();

interface AgentRunResult {
  success: boolean;
  output?: string;
  error?: string;
  iterationsUsed: number;
  fallbackUsed: boolean;
}

async function runWithRecovery(
  goal: string,
  tools: Anthropic.Tool[],
  options: {
    maxIterations?: number;
    maxRetries?: number;
    timeoutMs?: number;
    fallbackFn?: (goal: string) => Promise;
  } = {}
): Promise {
  const { maxIterations = 20, maxRetries = 3, timeoutMs = 60_000, fallbackFn } = options;

  // Wrap the entire run in a timeout
  const timeoutPromise = new Promise((_, reject) =>
    setTimeout(() => reject(new Error('Agent timeout')), timeoutMs)
  );

  try {
    const result = await Promise.race([
      runAgentLoop(goal, tools, maxIterations),
      timeoutPromise,
    ]);
    return { success: true, output: result, iterationsUsed: 0, fallbackUsed: false };
  } catch (err) {
    const error = err as Error;

    // Classify the error
    const isTransient = isTransientError(error);
    const isLoopError = error.message.includes('max iterations');
    const isTimeout = error.message.includes('timeout');

    // Transient API errors: retry with backoff
    if (isTransient && maxRetries > 0) {
      await backoff(3 - maxRetries);
      return runWithRecovery(goal, tools, { ...options, maxRetries: maxRetries - 1 });
    }

    // Loop exhaustion or timeout: try fallback
    if ((isLoopError || isTimeout) && fallbackFn) {
      try {
        const fallbackOutput = await fallbackFn(goal);
        return { success: true, output: fallbackOutput, iterationsUsed: maxIterations, fallbackUsed: true };
      } catch (fallbackErr) {
        // Fallback also failed - return graceful degradation response
        return {
          success: false,
          error: 'The task could not be completed automatically. Please try again or simplify your request.',
          iterationsUsed: maxIterations,
          fallbackUsed: true,
        };
      }
    }

    return {
      success: false,
      error: error.message,
      iterationsUsed: maxIterations,
      fallbackUsed: false,
    };
  }
}

function isTransientError(error: Error): boolean {
  const transientPatterns = ['529', '529', 'overloaded', 'rate limit', 'ECONNRESET', 'timeout'];
  return transientPatterns.some(p => error.message.toLowerCase().includes(p));
}

async function backoff(attempt: number): Promise {
  const delayMs = Math.min(1000 * Math.pow(2, attempt) + Math.random() * 500, 30_000);
  await new Promise(r => setTimeout(r, delayMs));
}

Surfacing tool errors correctly inside the loop

The most common error inside an agent loop is a failed tool call. How you return that error to the model determines whether it recovers or spirals. A tool result that silently returns an empty string on failure tells the model nothing - it will retry the same call with the same arguments. A tool result with an explicit error message and a hint about what to try instead gives the model what it needs to change course:

async function executeToolWithErrorHandling(
  toolName: string,
  toolInput: unknown,
  executors: Record Promise>
): Promise {
  const executor = executors[toolName];

  if (!executor) {
    return {
      type: 'tool_result',
      tool_use_id: '', // filled in by caller
      content: `Error: No executor found for tool "${toolName}". Available tools: ${Object.keys(executors).join('')}`,
      is_error: true,
    };
  }

  try {
    const result = await executor(toolInput);
    const serialised = typeof result === 'string' ? result : JSON.stringify(result, null, 2);

    // Truncate large results to prevent context overflow
    const maxChars = 8_000;
    const content = serialised.length > maxChars
      ? serialised.slice(0, maxChars) + `

[Result truncated at ${maxChars} characters. Use a more specific query to get a smaller result.]`
      : serialised;

    return { type: 'tool_result', tool_use_id: '', content };
  } catch (err) {
    const error = err as Error;
    return {
      type: 'tool_result',
      tool_use_id: '',
      content: `Error executing ${toolName}: ${error.message}. Consider trying a different approach or different input.`,
      is_error: true,
    };
  }
}

Context overflow recovery

Long-running agents accumulate tool results in the message history until the total context exceeds the model's limit. The model then throws a context window error, and if not handled, the entire run fails. The prevention is progressive history summarisation - when the accumulated context exceeds a threshold, replace the middle of the conversation with a summary:

async function summariseHistoryIfNeeded(
  messages: Anthropic.MessageParam[],
  tokenThreshold = 60_000
): Promise {
  const estimatedTokens = messages.reduce((sum, m) => {
    const content = typeof m.content === 'string' ? m.content : JSON.stringify(m.content);
    return sum + content.length / 4; // rough: 4 chars ≈ 1 token
  }, 0);

  if (estimatedTokens < tokenThreshold) return messages;

  // Keep the first message (original goal) and last 4 messages (recent context)
  const first = messages[0];
  const recent = messages.slice(-4);
  const middle = messages.slice(1, -4);

  if (middle.length === 0) return messages; // Already minimal, cannot summarise

  const summaryResponse = await client.messages.create({
    model: 'claude-haiku-4-5-20251001', // Cheap model for summarisation
    max_tokens: 1024,
    messages: [{
      role: 'user',
      content: `Summarise the following agent conversation history in 500 words or less, preserving all key findings and decisions:

${JSON.stringify(middle)}`,
    }],
  });

  const summary = summaryResponse.content[0].type === 'text' ? summaryResponse.content[0].text : ', ';

  return [
    first,
    { role: 'user', content: `[History summary]: ${summary}` },
    { role: 'assistant', content: 'Understood. Continuing from where we left off.' },
    ...recent,
  ];
}

Graceful degradation: the fallback strategy

When an agent cannot complete a task - due to repeated tool failures, context overflow, or loop exhaustion - the right response is not an error modal. It is a graceful step down to a simpler alternative:

  • Complex agent → simple agent: If the multi-tool agentic loop fails, try a single-call response with only the information available in the context (no tools).
  • Full agent → cached response: For agents that answer questions, a relevant cached answer from a previous run is better than an error.
  • Agent → human escalation: For agents taking consequential actions (financial, legal, medical contexts), failed runs should escalate to a human rather than degrade silently.

The rule: every agent that a user depends on for a real outcome needs an explicit fallback path defined before it ships to production. A fallback written under pressure after an incident is always worse than one designed in advance.

What NOT to do

The most destructive error handling patterns in production agents:

  • Swallowing errors silently and returning empty results - the model gets no signal that something went wrong
  • Applying a blanket retry to all error types - retrying a context overflow error makes it worse
  • No maximum iteration limit - a stuck agent runs indefinitely and incurs unbounded cost
  • No timeout - a hanging tool call (network request to a slow API) blocks the entire loop
  • Treating all agent outputs as correct - a completed loop with stop_reason: 'end_turn' is not a guarantee of correctness, only of completion

For the upstream patterns that define how errors propagate across multiple agents, see the agent supervisor pattern - supervisor-level failure handling differs from in-loop recovery. For tracing which iteration produced an error in production, see agent tracing and observability.