AI Agents
Three ways to get typed data from an LLM: tool use (forced schema), structured output (JSON mode), and free-form with Zod validation.
Every production AI feature that needs structured data from a model eventually faces the same choice: force the model to use a tool, use JSON structured output mode, or parse free-form text with a validation layer. Each approach works. Each has failure modes the others do not. The wrong choice produces an AI feature that occasionally returns malformed data in production, requires retry logic that adds latency, or is harder to schema-version than it needs to be. The right choice is usually obvious once you know the trade-offs.
Use tool use when you need maximum reliability, the schema is complex, or the model must take an action (not just return data). Use JSON structured output (where available) when you want simplicity and the schema is flat. Use free-form + Zod validation when you want the model to reason through a problem before producing structured output, or when you are on a provider that does not support native structured output. Most production applications end up using all three - for different features, for different reasons.
| Dimension | Tool Use | JSON Structured Output | Free-form + Zod |
|---|---|---|---|
| Schema enforcement | Strong - model constrained at generation time | Strong - JSON guaranteed, schema varies by provider | None at generation; enforced at parse time |
| Failure rate (production) | <0.5% with good schemas | 1 to 3% schema violations | 3 to 8% parse failures (lower with retry) |
| Schema complexity | High - nested objects, arrays, enums, conditionals | Medium - flat schemas best | Any - Zod handles arbitrary complexity |
| Model reasoning | Limited - model outputs schema, not reasoning prose | Limited - same as tool use | Full - model reasons before structuring |
| Latency | Standard | Standard | Higher with retries |
| Anthropic support | Full - native SDK support | Partial - via tool use workaround | Full - always works |
Tool use is the right choice when you need the lowest possible failure rate for a structured output. The model is constrained to produce arguments that match your JSON Schema at generation time - not parsed after the fact. Use it when:
import Anthropic from '@anthropic-ai/sdk';
const client = new Anthropic();
// Tool use: strong schema enforcement, low failure rate
const extractionTool: Anthropic.Tool = {
name: 'extract_invoice_data',
description: 'Extract structured invoice data from the provided text.',
input_schema: {
type: 'object' as const,
properties: {
invoice_number: { type: 'string' },
vendor: {
type: 'object',
properties: {
name: { type: 'string' },
address: { type: 'string' },
tax_id: { type: 'string' },
},
required: ['name'],
},
line_items: {
type: 'array',
items: {
type: 'object',
properties: {
description: { type: 'string' },
quantity: { type: 'number' },
unit_price: { type: 'number' },
total: { type: 'number' },
},
required: ['description', 'quantity', 'unit_price', 'total'],
},
},
total_amount: { type: 'number' },
currency: { type: 'string', enum: ['USD', 'EUR', 'GBP', 'JPY'] },
due_date: { type: 'string', description: 'ISO date format YYYY-MM-DD' },
},
required: ['invoice_number', 'vendor', 'line_items', 'total_amount', 'currency'],
},
};
const response = await client.messages.create({
model: 'claude-sonnet-5',
max_tokens: 1024,
tools: [extractionTool],
tool_choice: { type: 'tool', name: 'extract_invoice_data' }, // Force it
messages: [{ role: 'user', content: invoiceText }],
});
const toolCall = response.content.find(b => b.type === 'tool_use') as Anthropic.ToolUseBlock;
const invoiceData = toolCall.input; // Fully typed, validated at generation
Free-form text with validation is the right choice when the model needs to reason at length before arriving at a structured answer. Chain-of-thought reasoning - where the model thinks through a problem in prose before producing a conclusion - is incompatible with tool use (the tool call is the output, not a step toward it). Use free-form + Zod when:
import { z } from 'zod';
const SentimentSchema = z.object({
sentiment: z.enum(['positive', 'negative', 'neutral', 'mixed']),
score: z.number().min(-1).max(1),
reasoning: z.string().min(10),
key_phrases: z.array(z.string()).min(1).max(5),
});
type Sentiment = z.infer<typeof SentimentSchema>;
async function analyseSentiment(text: string, maxRetries = 2): Promise<Sentiment> {
for (let attempt = 0; attempt <= maxRetries; attempt++) {
const response = await client.messages.create({
model: 'claude-haiku-4-5-20251001',
max_tokens: 512,
system: `Analyse sentiment. Return ONLY a JSON object with fields:
- sentiment: "positive" | "negative" | "neutral" | "mixed"
- score: number from -1.0 (very negative) to 1.0 (very positive)
- reasoning: one sentence explanation
- key_phrases: array of 1-5 phrases that most influenced the sentiment`,
messages: [{ role: 'user', content: text }],
});
const raw = response.content[0].type === 'text' ? response.content[0].text : ', ';
try {
// Extract JSON from potential prose wrapper
const jsonMatch = raw.match(/{[sS]*}/);
if (!jsonMatch) throw new Error('No JSON found in response');
return SentimentSchema.parse(JSON.parse(jsonMatch[0]));
} catch (err) {
if (attempt === maxRetries) throw err;
// On retry, the next call is identical - the model will produce a different output
}
}
throw new Error('Sentiment analysis failed after retries');
}
For complex analytical tasks that need both reasoning and structured output, combine the two: let the model reason in free-form for one call, then extract the structured output in a second forced tool use call:
async function analyseWithReasoning(problem: string): Promise<{ reasoning: string; result: AnalysisResult }> {
// Step 1: Let the model reason freely
const reasoningResponse = await client.messages.create({
model: 'claude-sonnet-5',
max_tokens: 2048,
system: 'Think through this problem carefully. Show your reasoning.',
messages: [{ role: 'user', content: problem }],
});
const reasoning = reasoningResponse.content[0].type === 'text'
? reasoningResponse.content[0].text
: ', ';
// Step 2: Extract structured result from the reasoning
const extractionResponse = await client.messages.create({
model: 'claude-haiku-4-5-20251001', // Cheap model for extraction
max_tokens: 512,
tools: [analysisExtractionTool],
tool_choice: { type: 'tool', name: 'submit_analysis' },
messages: [{
role: 'user',
content: `Extract the structured findings from this analysis:
${reasoning}`,
}],
});
const toolCall = extractionResponse.content.find(b => b.type === 'tool_use') as Anthropic.ToolUseBlock;
return { reasoning, result: toolCall.input as AnalysisResult };
}
This pattern gets the best of both approaches: the model reasons with full prose flexibility (free-form), and the structured output is extracted with maximum reliability (tool use forced on a simple extraction task with Haiku). The two-call overhead is small and the reliability improvement is significant for complex analytical tasks.
For the advanced tool use mechanics - parallel calls, forced selection, compound tools - that make tool use more powerful than basic forced-schema calls, see advanced tool use patterns. For LLM output validation with Zod at the TypeScript level, see LLM output validation in TypeScript.