AI Agents

Tool Error Handling for AI Agents: Errors Claude Can Fix

How to report tool failures with is_error so an agent recovers instead of looping: error message design, retryable vs fatal errors and loop guards.

Your agent calls get_invoice with a mistyped ID. Your code throws and the run crashes, or returns null and the model invents an invoice, or returns "Error" and the model repeats the same call eleven times. None of these is a model problem. In tool result error handling, the error message is the only information the model gets about what went wrong, and in all three cases it was the wrong information.

Tool result error
A tool result error is a tool_result content block, marked with is_error set to true and linked to the failed call through its tool_use_id, that tells the model a tool call did not succeed and gives it enough information to retry with corrected input, choose another tool, or report the failure to the user.

How tool errors should reach the model: the direct answer

Catch every tool failure inside your loop and return it as a tool_result with is_error: true and the matching tool_use_id. The message should say what failed, include the offending input, state whether retrying the same input can succeed, and name the next step. Never skip the result and never let the exception escape the loop.

Anatomy of an error message the model can use

An error result is a prompt, and the same rules apply: be specific, be short, and say what to do. Compare three versions of the same failure:

  • Useless: Error. The model knows something failed and nothing else, so it usually retries the identical call.
  • Harmful: a 40-line stack trace. It costs 800 tokens, leaks file paths, and still does not say what to do.
  • Useful: Invoice INV-2291 not found. Invoice IDs have the format INV- followed by 5 digits. Call search_invoices with the customer email to list valid IDs.

The useful version follows a checklist that works for almost every tool. It matches Anthropic's tool use documentation, which recommends setting is_error with an informative message:

  1. State what failed, in one sentence, including the input that caused it.
  2. Say whether retrying the same input can succeed (a timeout) or cannot (a missing record).
  3. Point to the next action: a corrected format, another tool, or asking the user.
  4. Keep it under about 60 words, and keep secrets, tokens and internal hostnames out of it.

A tool executor that classifies errors

The cleanest way to get consistent messages is to stop formatting errors inside each tool. Tools throw a typed error that carries the facts, and one executor turns every failure into a well-formed result. The executor below also validates input with Zod before running anything, runs parallel calls concurrently, applies a timeout, and returns all results in a single user message:

import Anthropic from '@anthropic-ai/sdk';

export class ToolError extends Error {
  constructor(message: string, readonly retryable: boolean, readonly hint?: string) {
    super(message);
  }
}

type Handler = {
  schema: z.ZodType;
  run: (input: any, signal: AbortSignal) => Promise<string>;
};

function formatError(err: unknown): string {
  if (err instanceof ToolError) {
    const retry = err.retryable
      ? 'Retrying the same call may succeed.'
      : 'Retrying with the same input will fail again.';
    return [err.message, retry, err.hint].filter(Boolean).join(' ');
  }
  if (err instanceof Error && err.name === 'TimeoutError') {
    return 'The tool timed out after 15 seconds. Retry once; if it times out again, continue without this data and tell the user.';
  }
  // Unknown failure: log it on our side, give the model a safe instruction.
  console.error('tool crashed', err);
  return 'Internal error in this tool. Do not retry it. Continue with other tools or tell the user this step failed.';
}

export async function runToolCalls(
  content: Anthropic.ContentBlock[],
  handlers: Record<string, Handler>,
): Promise<Anthropic.ToolResultBlockParam[]> {
  const calls = content.filter((b): b is Anthropic.ToolUseBlock => b.type === 'tool_use');

  const settled = await Promise.allSettled(
    calls.map(async (call) => {
      const handler = handlers[call.name];
      if (!handler) {
        throw new ToolError('Unknown tool "' + call.name + '".', false, 'Use only the tools you were given.');
      }
      const parsed = handler.schema.safeParse(call.input);
      if (!parsed.success) {
        const issues = parsed.error.issues.map((i) => i.path.join('.') + ': ' + i.message).join('; ');
        throw new ToolError('Invalid input: ' + issues + '.', false, 'Fix those fields and call the tool again.');
      }
      return handler.run(parsed.data, AbortSignal.timeout(15_000));
    }),
  );

  // One result per call, in order, all returned together in one user message.
  return calls.map((call, i): Anthropic.ToolResultBlockParam => {
    const outcome = settled[i];
    return outcome.status === 'fulfilled'
      ? { type: 'tool_result', tool_use_id: call.id, content: outcome.value }
      : { type: 'tool_result', tool_use_id: call.id, content: formatError(outcome.reason), is_error: true };
  });
}

Inside a tool, a failure now reads like this: throw new ToolError('Invoice ' + id + ' not found.', false, 'Call search_invoices with the customer email to list valid IDs.'). The tool author supplies the facts; the executor guarantees the shape. For strict schema guarantees at generation time, you can also set strict: true on a tool definition, which makes the API constrain tool_use.input to your JSON Schema. Keep the Zod check anyway: it also covers business rules a schema cannot express.

Guarding the loop against repeat failures

Good messages prevent most retry loops, but not all. Add two guards in the loop itself: a per-call counter that escalates the wording after identical failures, and a total error budget for the run.

const failures = new Map<string, number>();
const MAX_ERRORS_PER_RUN = 6;
let totalErrors = 0;

function guard(call: Anthropic.ToolUseBlock, result: Anthropic.ToolResultBlockParam): Anthropic.ToolResultBlockParam {
  const key = call.name + ':' + JSON.stringify(call.input);
  if (!result.is_error) {
    failures.delete(key);
    return result;
  }
  totalErrors++;
  if (totalErrors > MAX_ERRORS_PER_RUN) {
    throw new Error('Error budget exhausted: ' + totalErrors + ' tool errors in one run');
  }
  const count = (failures.get(key) ?? 0) + 1;
  failures.set(key, count);
  if (count < 2) return result;
  return {
    ...result,
    content: result.content + ' This exact call has now failed ' + count + ' times. Do not call it again with this input.',
  };
}

When the budget runs out, the loop stops and your code decides what the user sees: a partial answer, a handoff to a person, or a clear failure message. That decision belongs to your application, not to the model. The broader retry and fallback strategies are covered in agent error recovery.

Hardening an LLM integration?

Download a ready-made Claude Code subagent that reviews tool definitions, error paths and retry logic in LLM features.

Get the LLM integration engineer subagent

Errors that should not go back to the model

Not every failure is the model's to handle. Route these elsewhere:

  • Errors from the Claude API itself. A 429 rate limit or 529 overload happens before the model sees anything. The SDK retries these automatically (maxRetries defaults to 2); if retries run out, back off or re-queue the run. See rate limiting AI agents.
  • Failures the model cannot fix. An expired service credential or a database that is down will not improve with a better query. Return a short is_error result so the conversation stays valid, then stop the run and alert whoever owns the dependency.
  • Permission denials. Return them, but make them final: "Not permitted: this agent cannot issue refunds above 100 EUR. Tell the user a person will review it." Ambiguous wording invites the model to try a different route to the same action.
  • Content from untrusted sources. An error message that echoes a web page or email body can carry instructions. Treat those strings like any other tool output, as described in prompt injection defenses for agents.

Failure modes checklist

  • Dropped results. Skipping a failed call leaves a tool_use without a tool_result, and the next request fails with a 400. Answer every id, including calls you chose not to run.
  • Results split across messages. Returning parallel results in separate user messages works, but it teaches the model to stop making parallel calls. Put them all in one message.
  • Empty success. Returning an empty string for "no rows found" looks like success with no content. Say "No invoices match that email" explicitly, as a normal result.
  • Error flag on partial success. If 9 of 10 records loaded, return them as a success with a note about the missing one, not as an error that hides the 9.
  • Retry storms. A timeout message that says "retry" with no limit, combined with an agent that takes it literally. Always say "retry once".

Testing error paths

Error handling that has never run does not work. For every tool, write tests that make it throw each kind of ToolError, a timeout and an unknown exception, then run the agent loop against those stubs and check three things: the loop does not crash, the final answer mentions the failure instead of inventing data, and the number of retries stays within your guard. Agent testing strategies shows how to wire those stubs into an eval suite, and advanced tool use patterns covers the parallel and dependent calls that make error handling harder.

FAQ

How do I tell Claude that a tool call failed?

Return a tool_result block with the matching tool_use_id, is_error set to true, and a short message that says what failed, whether retrying the same input can work, and what to try instead. Do not throw out of the loop and do not drop the result: every tool_use block needs exactly one tool_result.

What happens if I leave out a tool_result?

The next request fails with a 400 invalid_request_error, because the API requires a tool_result for every tool_use id in the previous assistant turn. Even when a call is skipped or cancelled, return an error result for it so the conversation stays valid.

Should tool errors include stack traces?

No. Stack traces waste tokens, can leak internal paths or secrets, and rarely tell the model anything it can act on. Log the full error on your side and return one or two sentences describing the failure and the next step the agent should take.

How do I stop an agent from retrying a failing tool forever?

Track failures per tool and input in your loop. After two identical failures, tell the model explicitly not to repeat that call, and set a total error budget for the run, after which your code stops the loop and returns a partial result or escalates to a human.