AI Agents

Tool Use vs Structured Output vs Free-Form: When to Use Each

Three ways to get typed data from an LLM: tool use (forced schema), structured output (JSON mode), and free-form with Zod validation.

Every production AI feature that needs structured data from a model eventually faces the same choice: force the model to use a tool, use JSON structured output mode, or parse free-form text with a validation layer. Each approach works. Each has failure modes the others do not. The wrong choice produces an AI feature that occasionally returns malformed data in production, requires retry logic that adds latency, or is harder to schema-version than it needs to be. The right choice is usually obvious once you know the trade-offs.

LLM structured output
LLM structured output is any technique that produces machine-parseable typed data from a language model - including tool use (schema-constrained at generation time), JSON mode (guaranteed JSON syntax, schema varies), and free-form text with post-hoc validation via Zod - each with different reliability rates and schema complexity limits.

The short answer: which to use

Use tool use when you need maximum reliability, the schema is complex, or the model must take an action (not just return data). Use JSON structured output (where available) when you want simplicity and the schema is flat. Use free-form + Zod validation when you want the model to reason through a problem before producing structured output, or when you are on a provider that does not support native structured output. Most production applications end up using all three - for different features, for different reasons.

Head-to-head comparison

Dimension Tool Use JSON Structured Output Free-form + Zod
Schema enforcement Strong - model constrained at generation time Strong - JSON guaranteed, schema varies by provider None at generation; enforced at parse time
Failure rate (production) <0.5% with good schemas 1 to 3% schema violations 3 to 8% parse failures (lower with retry)
Schema complexity High - nested objects, arrays, enums, conditionals Medium - flat schemas best Any - Zod handles arbitrary complexity
Model reasoning Limited - model outputs schema, not reasoning prose Limited - same as tool use Full - model reasons before structuring
Latency Standard Standard Higher with retries
Anthropic support Full - native SDK support Partial - via tool use workaround Full - always works

When to use tool use

Tool use is the right choice when you need the lowest possible failure rate for a structured output. The model is constrained to produce arguments that match your JSON Schema at generation time - not parsed after the fact. Use it when:

  • The schema is complex (nested objects, arrays of objects, conditional fields)
  • The feature is in a critical path where a validation failure would require a full retry (adding latency the user feels)
  • The model needs to decide when to produce structured output vs give a prose response (tool use lets the model skip the tool call if inappropriate)
  • You are building an agent where the structured output drives an action (write to database, call an API) - not just a display
import Anthropic from '@anthropic-ai/sdk';

const client = new Anthropic();

// Tool use: strong schema enforcement, low failure rate
const extractionTool: Anthropic.Tool = {
  name: 'extract_invoice_data',
  description: 'Extract structured invoice data from the provided text.',
  input_schema: {
    type: 'object' as const,
    properties: {
      invoice_number: { type: 'string' },
      vendor: {
        type: 'object',
        properties: {
          name: { type: 'string' },
          address: { type: 'string' },
          tax_id: { type: 'string' },
        },
        required: ['name'],
      },
      line_items: {
        type: 'array',
        items: {
          type: 'object',
          properties: {
            description: { type: 'string' },
            quantity: { type: 'number' },
            unit_price: { type: 'number' },
            total: { type: 'number' },
          },
          required: ['description', 'quantity', 'unit_price', 'total'],
        },
      },
      total_amount: { type: 'number' },
      currency: { type: 'string', enum: ['USD', 'EUR', 'GBP', 'JPY'] },
      due_date: { type: 'string', description: 'ISO date format YYYY-MM-DD' },
    },
    required: ['invoice_number', 'vendor', 'line_items', 'total_amount', 'currency'],
  },
};

const response = await client.messages.create({
  model: 'claude-sonnet-5',
  max_tokens: 1024,
  tools: [extractionTool],
  tool_choice: { type: 'tool', name: 'extract_invoice_data' }, // Force it
  messages: [{ role: 'user', content: invoiceText }],
});

const toolCall = response.content.find(b => b.type === 'tool_use') as Anthropic.ToolUseBlock;
const invoiceData = toolCall.input; // Fully typed, validated at generation

When to use free-form + Zod

Free-form text with validation is the right choice when the model needs to reason at length before arriving at a structured answer. Chain-of-thought reasoning - where the model thinks through a problem in prose before producing a conclusion - is incompatible with tool use (the tool call is the output, not a step toward it). Use free-form + Zod when:

  • The model needs to explain its reasoning as part of the output
  • The schema is simple enough that Zod validation failure rate is acceptable (typically flat objects with 3 to 5 fields)
  • You want the flexibility to use the prose output in some cases and the structured data in others
import { z } from 'zod';

const SentimentSchema = z.object({
  sentiment: z.enum(['positive', 'negative', 'neutral', 'mixed']),
  score: z.number().min(-1).max(1),
  reasoning: z.string().min(10),
  key_phrases: z.array(z.string()).min(1).max(5),
});

type Sentiment = z.infer<typeof SentimentSchema>;

async function analyseSentiment(text: string, maxRetries = 2): Promise<Sentiment> {
  for (let attempt = 0; attempt <= maxRetries; attempt++) {
    const response = await client.messages.create({
      model: 'claude-haiku-4-5-20251001',
      max_tokens: 512,
      system: `Analyse sentiment. Return ONLY a JSON object with fields:
- sentiment: "positive" | "negative" | "neutral" | "mixed"
- score: number from -1.0 (very negative) to 1.0 (very positive)
- reasoning: one sentence explanation
- key_phrases: array of 1-5 phrases that most influenced the sentiment`,
      messages: [{ role: 'user', content: text }],
    });

    const raw = response.content[0].type === 'text' ? response.content[0].text : ', ';

    try {
      // Extract JSON from potential prose wrapper
      const jsonMatch = raw.match(/{[sS]*}/);
      if (!jsonMatch) throw new Error('No JSON found in response');
      return SentimentSchema.parse(JSON.parse(jsonMatch[0]));
    } catch (err) {
      if (attempt === maxRetries) throw err;
      // On retry, the next call is identical - the model will produce a different output
    }
  }
  throw new Error('Sentiment analysis failed after retries');
}

The chain-of-thought + tool use hybrid

For complex analytical tasks that need both reasoning and structured output, combine the two: let the model reason in free-form for one call, then extract the structured output in a second forced tool use call:

async function analyseWithReasoning(problem: string): Promise<{ reasoning: string; result: AnalysisResult }> {
  // Step 1: Let the model reason freely
  const reasoningResponse = await client.messages.create({
    model: 'claude-sonnet-5',
    max_tokens: 2048,
    system: 'Think through this problem carefully. Show your reasoning.',
    messages: [{ role: 'user', content: problem }],
  });

  const reasoning = reasoningResponse.content[0].type === 'text'
    ? reasoningResponse.content[0].text
    : ', ';

  // Step 2: Extract structured result from the reasoning
  const extractionResponse = await client.messages.create({
    model: 'claude-haiku-4-5-20251001', // Cheap model for extraction
    max_tokens: 512,
    tools: [analysisExtractionTool],
    tool_choice: { type: 'tool', name: 'submit_analysis' },
    messages: [{
      role: 'user',
      content: `Extract the structured findings from this analysis:

${reasoning}`,
    }],
  });

  const toolCall = extractionResponse.content.find(b => b.type === 'tool_use') as Anthropic.ToolUseBlock;
  return { reasoning, result: toolCall.input as AnalysisResult };
}

This pattern gets the best of both approaches: the model reasons with full prose flexibility (free-form), and the structured output is extracted with maximum reliability (tool use forced on a simple extraction task with Haiku). The two-call overhead is small and the reliability improvement is significant for complex analytical tasks.

Decision guide

  • Critical path, complex schema, or action-triggering → Tool use with forced selection
  • Needs chain-of-thought reasoning → Free-form + Zod (or hybrid: reasoning call → tool extraction call)
  • Simple schema, non-critical, or prototype → Free-form + Zod with retry
  • High frequency, cost-sensitive → Tool use on Haiku for flat schemas; Zod on Haiku for simple classification

For the advanced tool use mechanics - parallel calls, forced selection, compound tools - that make tool use more powerful than basic forced-schema calls, see advanced tool use patterns. For LLM output validation with Zod at the TypeScript level, see LLM output validation in TypeScript.