AI Agents

Advanced Tool Use Patterns for AI Agents (2026 Guide)

Beyond basic tool use: parallel tool calls, dependent tool chains, forced tool selection, tool result caching.

Most tool use tutorials show a single tool, called once, with the result fed back to the model. That works for demos. Production agents call multiple tools simultaneously, chain dependent calls across iterations, force specific tools for critical operations, and cache expensive results to avoid redundant calls. Each of these patterns reduces latency, cuts cost, or improves reliability - often all three. This guide covers them in the order you are most likely to need them.

Parallel tool use (AI agents)
Parallel tool use is when a language model returns multiple tool_use content blocks in a single response, allowing the calling code to execute all requested tools simultaneously via Promise.all and return all results in one follow-up message - eliminating the round-trip overhead of sequential single-tool calls.

Parallel tool calls: the most impactful optimisation

By default, Claude will call one tool, wait for the result, then decide whether to call another. For independent operations - looking up multiple records, fetching several URLs, querying different data sources - this is unnecessary serialisation. Claude supports parallel tool calls: the model can request multiple tool calls in a single response, and your code executes them simultaneously.

import Anthropic from '@anthropic-ai/sdk';

const client = new Anthropic();

// The model returns multiple tool_use blocks in one response
const response = await client.messages.create({
  model: 'claude-sonnet-5',
  max_tokens: 4096,
  tools: [searchTool, fetchTool, dbQueryTool],
  messages: [{
    role: 'user',
    content: 'Compare the pricing pages of competitor-a.com and competitor-b.com',
  }],
});

// Response may contain multiple tool_use blocks
const toolCalls = response.content.filter(
  (b): b is Anthropic.ToolUseBlock => b.type === 'tool_use'
);

// Execute all in parallel - not sequentially
const results = await Promise.all(
  toolCalls.map(async (call) => ({
    type: 'tool_result' as const,
    tool_use_id: call.id,
    content: await executeToolSafely(call.name, call.input),
  }))
);

// All results returned in one message
const followUp = await client.messages.create({
  model: 'claude-sonnet-5',
  max_tokens: 4096,
  tools: [searchTool, fetchTool, dbQueryTool],
  messages: [
    { role: 'user', content: 'Compare the pricing pages of competitor-a.com and competitor-b.com' },
    { role: 'assistant', content: response.content },
    { role: 'user', content: results },
  ],
});

The parallel execution happens in your code - Claude does not know whether your tool calls run sequentially or in parallel. The key is processing all tool_use blocks in one Promise.all and returning all tool_result blocks in a single user message. Returning results one at a time breaks the parallelism.

Forced tool selection: guaranteeing a specific output format

By default, Claude decides whether to use a tool or answer in plain text. For operations where you require structured output - a classification decision, an extraction schema, a yes/no judgment - you can force the model to use a specific tool using tool_choice:

const classificationTool: Anthropic.Tool = {
  name: 'classify_intent',
  description: 'Classify the user intent into exactly one category.',
  input_schema: {
    type: 'object' as const,
    properties: {
      category: {
        type: 'string',
        enum: ['bug_report', 'feature_request', 'question', 'complaint', 'praise'],
        description: 'The primary intent category',
      },
      confidence: {
        type: 'number',
        minimum: 0,
        maximum: 1,
        description: 'Confidence score from 0 to 1',
      },
      reasoning: {
        type: 'string',
        description: 'One sentence explaining the classification',
      },
    },
    required: ['category', 'confidence', 'reasoning'],
  },
};

const response = await client.messages.create({
  model: 'claude-haiku-4-5-20251001',
  max_tokens: 256,
  tools: [classificationTool],
  tool_choice: { type: 'tool', name: 'classify_intent' }, // Force this specific tool
  messages: [{ role: 'user', content: userMessage }],
});

// Response is guaranteed to contain the tool call
const toolCall = response.content.find(b => b.type === 'tool_use') as Anthropic.ToolUseBlock;
const { category, confidence, reasoning } = toolCall.input as {
  category: string;
  confidence: number;
  reasoning: string;
};

tool_choice: { type: 'tool', name: '...' } forces a specific tool. tool_choice: { type: 'any' } forces the model to call at least one tool (from whatever tools are available). tool_choice: { type: 'auto' } is the default - let the model decide.

Dependent tool chains: passing outputs between calls

Some operations require sequential tool calls where the output of one is the input to the next. The naive approach - one tool call per model round-trip - works but is slow. A better pattern is to give the model a high-level tool that internally runs the chain, returning only the final result:

// Instead of: search → model → fetch → model → extract → model
// Do: define a compound tool that runs the chain internally

const researchTool: Anthropic.Tool = {
  name: 'deep_research',
  description: 'Search for a topic, fetch the top results, and extract key facts. Use when a simple search is not enough and you need to read source content.',
  input_schema: {
    type: 'object' as const,
    properties: {
      query: { type: 'string', description: 'The research query' },
      factsNeeded: { type: 'array', items: { type: 'string' }, description: 'Specific facts to extract' },
    },
    required: ['query', 'factsNeeded'],
  },
};

// Your tool executor runs the full chain internally
async function executeDeepResearch(input: { query: string; factsNeeded: string[] }): Promise {
  const searchResults = await webSearch(input.query);
  const topUrls = searchResults.slice(0, 3).map(r => r.url);

  const pages = await Promise.all(topUrls.map(fetchPage));
  const combined = pages.join('

---

').slice(0, 20_000);

  // Use a cheap model for extraction - not another agentic loop
  const extractResponse = await client.messages.create({
    model: 'claude-haiku-4-5-20251001',
    max_tokens: 1024,
    messages: [{
      role: 'user',
      content: `From the following content, extract: ${input.factsNeeded.join('')}

Content:
${combined}`,
    }],
  });

  return extractResponse.content[0].type === 'text' ? extractResponse.content[0].text : ', ';
}

Compound tools that internalize multi-step chains reduce the number of model round-trips from N to 1, and let you use a cheaper model (Haiku) for intermediate steps without exposing that complexity to the orchestrating model.

Tool result caching: avoiding redundant calls

In long-running agent loops, the same tool is often called with the same arguments multiple times - the model fetches the same page twice, queries the same database row repeatedly. A simple in-memory cache keyed on tool name + serialised arguments prevents these redundant calls:

class CachedToolExecutor {
  private cache = new Map();
  private ttlMs: number;

  constructor(ttlMs = 5 * 60 * 1000) { // 5-minute TTL
    this.ttlMs = ttlMs;
  }

  async execute(
    toolName: string,
    toolInput: unknown,
    executors: Record Promise>
  ): Promise {
    const cacheKey = `${toolName}:${JSON.stringify(toolInput)}`;
    const cached = this.cache.get(cacheKey);

    if (cached && Date.now() - cached.timestamp < this.ttlMs) {
      return cached.result; // Cache hit - no API call
    }

    const executor = executors[toolName];
    if (!executor) throw new Error(`Unknown tool: ${toolName}`);

    const result = await executor(toolInput);
    this.cache.set(cacheKey, { result, timestamp: Date.now() });
    return result;
  }

  invalidate(toolName: string, toolInput?: unknown): void {
    if (toolInput) {
      this.cache.delete(`${toolName}:${JSON.stringify(toolInput)}`);
    } else {
      // Invalidate all entries for this tool
      for (const key of this.cache.keys()) {
        if (key.startsWith(`${toolName}:`)) this.cache.delete(key);
      }
    }
  }
}

Use the cache for read-only tools (search, fetch, database queries). Do not cache write tools - caching a file write means a subsequent call with the same arguments silently skips the write, which is almost always a bug.

Schema design: what produces reliable structured outputs

Tool schemas that produce reliable, parseable outputs from Claude share a few consistent traits:

  • Use enum for fields with a finite valid set - Rather than type: 'string', description: 'one of: low, medium, high', use type: 'string', enum: ['low', 'medium', 'high']. The model respects enum constraints far more reliably than description-based constraints.
  • Mark all required fields in required - Optional fields get omitted unpredictably. If you need the field, require it. Set a sensible default in the description rather than making it optional.
  • Describe each field from the model's perspective - "The primary search query to use, based on the user's intent (not a verbatim copy)" gives the model more useful guidance than "Search query".
  • Keep array items simple - Arrays of complex objects are less reliable than arrays of strings. If you need complex items, use a separate required object property per item and impose a max count.

Failure modes to watch for

  • Model hallucinates tool names: The model calls a tool that does not exist. Caused by listing too many tools or having overlapping descriptions. Fix: reduce the tool set and differentiate descriptions.
  • Required fields populated with placeholders: The model fills required fields with "N/A" or empty strings when it does not have the information. Fix: add explicit instructions in the tool description: "If unknown, use null - do not use placeholder text."
  • Parallel calls when sequential is needed: The model calls tools in parallel that have a dependency (calls fetch before search returns). Fix: structure the task as separate turns or use a compound tool that internalises the dependency.

For the decision between using tool use for structured output vs alternative approaches like JSON mode and Zod, see tool use vs structured output. For tool call strategy inside a full agentic loop, see the agentic loop explained.