AI Agents
Beyond basic tool use: parallel tool calls, dependent tool chains, forced tool selection, tool result caching.
Most tool use tutorials show a single tool, called once, with the result fed back to the model. That works for demos. Production agents call multiple tools simultaneously, chain dependent calls across iterations, force specific tools for critical operations, and cache expensive results to avoid redundant calls. Each of these patterns reduces latency, cuts cost, or improves reliability - often all three. This guide covers them in the order you are most likely to need them.
tool_use content blocks in a single response, allowing the calling code to execute all requested tools simultaneously via Promise.all and return all results in one follow-up message - eliminating the round-trip overhead of sequential single-tool calls.By default, Claude will call one tool, wait for the result, then decide whether to call another. For independent operations - looking up multiple records, fetching several URLs, querying different data sources - this is unnecessary serialisation. Claude supports parallel tool calls: the model can request multiple tool calls in a single response, and your code executes them simultaneously.
import Anthropic from '@anthropic-ai/sdk';
const client = new Anthropic();
// The model returns multiple tool_use blocks in one response
const response = await client.messages.create({
model: 'claude-sonnet-5',
max_tokens: 4096,
tools: [searchTool, fetchTool, dbQueryTool],
messages: [{
role: 'user',
content: 'Compare the pricing pages of competitor-a.com and competitor-b.com',
}],
});
// Response may contain multiple tool_use blocks
const toolCalls = response.content.filter(
(b): b is Anthropic.ToolUseBlock => b.type === 'tool_use'
);
// Execute all in parallel - not sequentially
const results = await Promise.all(
toolCalls.map(async (call) => ({
type: 'tool_result' as const,
tool_use_id: call.id,
content: await executeToolSafely(call.name, call.input),
}))
);
// All results returned in one message
const followUp = await client.messages.create({
model: 'claude-sonnet-5',
max_tokens: 4096,
tools: [searchTool, fetchTool, dbQueryTool],
messages: [
{ role: 'user', content: 'Compare the pricing pages of competitor-a.com and competitor-b.com' },
{ role: 'assistant', content: response.content },
{ role: 'user', content: results },
],
});
The parallel execution happens in your code - Claude does not know whether your tool calls run sequentially or in parallel. The key is processing all tool_use blocks in one Promise.all and returning all tool_result blocks in a single user message. Returning results one at a time breaks the parallelism.
By default, Claude decides whether to use a tool or answer in plain text. For operations where you require structured output - a classification decision, an extraction schema, a yes/no judgment - you can force the model to use a specific tool using tool_choice:
const classificationTool: Anthropic.Tool = {
name: 'classify_intent',
description: 'Classify the user intent into exactly one category.',
input_schema: {
type: 'object' as const,
properties: {
category: {
type: 'string',
enum: ['bug_report', 'feature_request', 'question', 'complaint', 'praise'],
description: 'The primary intent category',
},
confidence: {
type: 'number',
minimum: 0,
maximum: 1,
description: 'Confidence score from 0 to 1',
},
reasoning: {
type: 'string',
description: 'One sentence explaining the classification',
},
},
required: ['category', 'confidence', 'reasoning'],
},
};
const response = await client.messages.create({
model: 'claude-haiku-4-5-20251001',
max_tokens: 256,
tools: [classificationTool],
tool_choice: { type: 'tool', name: 'classify_intent' }, // Force this specific tool
messages: [{ role: 'user', content: userMessage }],
});
// Response is guaranteed to contain the tool call
const toolCall = response.content.find(b => b.type === 'tool_use') as Anthropic.ToolUseBlock;
const { category, confidence, reasoning } = toolCall.input as {
category: string;
confidence: number;
reasoning: string;
};
tool_choice: { type: 'tool', name: '...' } forces a specific tool. tool_choice: { type: 'any' } forces the model to call at least one tool (from whatever tools are available). tool_choice: { type: 'auto' } is the default - let the model decide.
Some operations require sequential tool calls where the output of one is the input to the next. The naive approach - one tool call per model round-trip - works but is slow. A better pattern is to give the model a high-level tool that internally runs the chain, returning only the final result:
// Instead of: search → model → fetch → model → extract → model
// Do: define a compound tool that runs the chain internally
const researchTool: Anthropic.Tool = {
name: 'deep_research',
description: 'Search for a topic, fetch the top results, and extract key facts. Use when a simple search is not enough and you need to read source content.',
input_schema: {
type: 'object' as const,
properties: {
query: { type: 'string', description: 'The research query' },
factsNeeded: { type: 'array', items: { type: 'string' }, description: 'Specific facts to extract' },
},
required: ['query', 'factsNeeded'],
},
};
// Your tool executor runs the full chain internally
async function executeDeepResearch(input: { query: string; factsNeeded: string[] }): Promise {
const searchResults = await webSearch(input.query);
const topUrls = searchResults.slice(0, 3).map(r => r.url);
const pages = await Promise.all(topUrls.map(fetchPage));
const combined = pages.join('
---
').slice(0, 20_000);
// Use a cheap model for extraction - not another agentic loop
const extractResponse = await client.messages.create({
model: 'claude-haiku-4-5-20251001',
max_tokens: 1024,
messages: [{
role: 'user',
content: `From the following content, extract: ${input.factsNeeded.join('')}
Content:
${combined}`,
}],
});
return extractResponse.content[0].type === 'text' ? extractResponse.content[0].text : ', ';
}
Compound tools that internalize multi-step chains reduce the number of model round-trips from N to 1, and let you use a cheaper model (Haiku) for intermediate steps without exposing that complexity to the orchestrating model.
In long-running agent loops, the same tool is often called with the same arguments multiple times - the model fetches the same page twice, queries the same database row repeatedly. A simple in-memory cache keyed on tool name + serialised arguments prevents these redundant calls:
class CachedToolExecutor {
private cache = new Map();
private ttlMs: number;
constructor(ttlMs = 5 * 60 * 1000) { // 5-minute TTL
this.ttlMs = ttlMs;
}
async execute(
toolName: string,
toolInput: unknown,
executors: Record Promise>
): Promise {
const cacheKey = `${toolName}:${JSON.stringify(toolInput)}`;
const cached = this.cache.get(cacheKey);
if (cached && Date.now() - cached.timestamp < this.ttlMs) {
return cached.result; // Cache hit - no API call
}
const executor = executors[toolName];
if (!executor) throw new Error(`Unknown tool: ${toolName}`);
const result = await executor(toolInput);
this.cache.set(cacheKey, { result, timestamp: Date.now() });
return result;
}
invalidate(toolName: string, toolInput?: unknown): void {
if (toolInput) {
this.cache.delete(`${toolName}:${JSON.stringify(toolInput)}`);
} else {
// Invalidate all entries for this tool
for (const key of this.cache.keys()) {
if (key.startsWith(`${toolName}:`)) this.cache.delete(key);
}
}
}
}
Use the cache for read-only tools (search, fetch, database queries). Do not cache write tools - caching a file write means a subsequent call with the same arguments silently skips the write, which is almost always a bug.
Tool schemas that produce reliable, parseable outputs from Claude share a few consistent traits:
enum for fields with a finite valid set - Rather than type: 'string', description: 'one of: low, medium, high', use type: 'string', enum: ['low', 'medium', 'high']. The model respects enum constraints far more reliably than description-based constraints.required - Optional fields get omitted unpredictably. If you need the field, require it. Set a sensible default in the description rather than making it optional.object property per item and impose a max count.required fields with "N/A" or empty strings when it does not have the information. Fix: add explicit instructions in the tool description: "If unknown, use null - do not use placeholder text."For the decision between using tool use for structured output vs alternative approaches like JSON mode and Zod, see tool use vs structured output. For tool call strategy inside a full agentic loop, see the agentic loop explained.