AI Agents

Build a Data Extraction Agent with Claude: Text to Typed Schema

Build a production data extraction agent that converts unstructured text to typed schemas using Claude tool use - with confidence scoring.

You have 500 invoices in PDF, 1,000 support tickets in plain text, and 200 contract documents in Word format. Each contains structured information - amounts, dates, names, addresses, line items - but it is buried in unstructured prose. Manually extracting this data takes weeks. An extraction agent that reads each document and populates a typed schema takes hours, and once the schema is defined correctly, it scales to any volume. This tutorial builds an extraction agent that handles the real-world complexities: fields that are sometimes absent, values that appear in multiple formats, and documents where the extraction confidence is low enough to route to human review.

Data extraction agent
A data extraction agent is an AI agent that processes unstructured text documents and maps their content to a predefined typed schema using tool use - returning not just the extracted values but confidence scores and missing field indicators, enabling downstream systems to differentiate high-confidence automatic extractions from uncertain results that require human verification.

What we're building

Unstructured documents (PDFs, emails, text)
        │
        ▼
┌──────────────────────────────────┐
│  Extraction Agent                │
│  (claude-sonnet-5 or haiku)      │
│                                  │
│  Tool: extract_fields            │
│  (forced via tool_choice)        │
└──────────────┬───────────────────┘
               │
               ▼
┌──────────────────────────────────┐
│  Typed extraction result         │
│  {                               │
│    fields: { name, date... },  │
│    confidence: { per-field },    │
│    missing: string[],            │
│    review_required: boolean      │
│  }                               │
└──────────────────────────────────┘

Step 1 - Define the extraction schema as a tool

The key architectural choice: use forced tool selection (tool_choice: { type: "tool", name: "extract_fields" }) rather than asking the model to return JSON in its text response. Forced tool use guarantees the output matches the schema - no regex parsing, no JSON.parse failures from prose leaking into the output:

import Anthropic from '@anthropic-ai/sdk';

const client = new Anthropic({ apiKey: process.env.ANTHROPIC_API_KEY });

// Define your extraction schema as a tool
function buildExtractionTool(schema: ExtractionSchema): Anthropic.Tool {
  return {
    name: 'extract_fields',
    description: `Extract the specified fields from the document. For each field, provide the extracted value and a confidence score (0.0-1.0). If a field is not present in the document, omit it from the output rather than guessing.`,
    input_schema: {
      type: 'object' as const,
      properties: {
        fields: {
          type: 'object',
          description: 'Extracted field values.',
          properties: Object.fromEntries(
            Object.entries(schema.fields).map(([key, def]) => [
              key,
              { type: def.type, description: def.description },
            ])
          ),
        },
        confidence: {
          type: 'object',
          description: 'Confidence score (0.0-1.0) for each extracted field.',
          properties: Object.fromEntries(
            Object.keys(schema.fields).map(key => [key, { type: 'number', minimum: 0, maximum: 1 }])
          ),
        },
        missing_fields: {
          type: 'array',
          items: { type: 'string' },
          description: 'Required fields that could not be found in the document.',
        },
        extraction_notes: {
          type: 'string',
          description: 'Any ambiguities, multiple values found, or extraction decisions made.',
        },
      },
      required: ['fields', 'confidence', 'missing_fields'],
    },
  };
}

// Example invoice schema
const invoiceSchema: ExtractionSchema = {
  fields: {
    invoice_number: { type: 'string', description: 'Invoice or reference number.', required: true },
    invoice_date: { type: 'string', description: 'Invoice date in ISO 8601 format (YYYY-MM-DD).', required: true },
    due_date: { type: 'string', description: 'Payment due date in ISO 8601 format.', required: false },
    vendor_name: { type: 'string', description: 'Name of the vendor or supplier.', required: true },
    vendor_address: { type: 'string', description: 'Full vendor mailing address.', required: false },
    total_amount: { type: 'number', description: 'Total invoice amount as a decimal number.', required: true },
    currency: { type: 'string', description: 'Three-letter currency code (e.g., USD, EUR, GBP).', required: true },
    tax_amount: { type: 'number', description: 'Tax amount as a decimal number.', required: false },
    line_items: {
      type: 'array',
      description: 'Array of line items: [{ description, quantity, unit_price, total }]',
      required: false,
    },
  },
};

Step 2 - The extraction call (single document)

interface ExtractionResult {
  fields: Record;
  confidence: Record;
  missing_fields: string[];
  extraction_notes?: string;
  review_required: boolean;
  low_confidence_fields: string[];
}

const CONFIDENCE_THRESHOLD = 0.7; // Fields below this score are flagged for review

async function extractFromDocument(
  documentText: string,
  schema: ExtractionSchema,
  options: {
    model?: string;
    documentType?: string;
  } = {}
): Promise {
  const model = options.model ?? 'claude-haiku-4-5-20251001'; // Haiku for cost efficiency at scale
  const docType = options.documentType ?? 'document';
  const tool = buildExtractionTool(schema);

  const response = await client.messages.create({
    model,
    max_tokens: 2048,
    system: `You are a precise data extraction specialist. Extract exactly what is written in the ${docType} - do not infer, estimate, or fill in values that are not explicitly present. When a date is written as "March 15, 2026", convert it to ISO format "2026-03-15". When a currency symbol appears without a code (e.g., "$"), infer the code from context (USD for US documents). If you are uncertain, assign a lower confidence score rather than omitting the field.`,
    tools: [tool],
    tool_choice: { type: 'tool', name: 'extract_fields' },
    messages: [{
      role: 'user',
      content: `Extract the required fields from this ${docType}:

${documentText}`,
    }],
  });

  const toolCall = response.content.find(
    (b): b is Anthropic.ToolUseBlock => b.type === 'tool_use' && b.name === 'extract_fields'
  );

  if (!toolCall) throw new Error('Extraction tool was not called - check tool_choice configuration');

  const raw = toolCall.input as {
    fields: Record;
    confidence: Record;
    missing_fields: string[];
    extraction_notes?: string;
  };

  // Identify fields below the confidence threshold
  const lowConfidenceFields = Object.entries(raw.confidence)
    .filter(([, score]) => score < CONFIDENCE_THRESHOLD)
    .map(([field]) => field);

  // Identify required fields that are missing
  const requiredMissing = raw.missing_fields.filter(
    field => schema.fields[field]?.required
  );

  const reviewRequired = lowConfidenceFields.length > 0 || requiredMissing.length > 0;

  return {
    ...raw,
    review_required: reviewRequired,
    low_confidence_fields: lowConfidenceFields,
  };
}

Step 3 - Batch extraction with concurrency control

interface BatchExtractionResult {
  documentId: string;
  result?: ExtractionResult;
  error?: string;
  processingTimeMs: number;
}

async function batchExtract(
  documents: Array<{ id: string; text: string }>,
  schema: ExtractionSchema,
  options: {
    concurrency?: number;
    model?: string;
    documentType?: string;
  } = {}
): Promise {
  const concurrency = options.concurrency ?? 10;
  const results: BatchExtractionResult[] = [];
  const queue = [...documents];
  let processed = 0;

  async function worker(): Promise {
    while (queue.length > 0) {
      const doc = queue.shift();
      if (!doc) break;

      const start = Date.now();
      try {
        const result = await extractFromDocument(doc.text, schema, {
          model: options.model,
          documentType: options.documentType,
        });
        results.push({
          documentId: doc.id,
          result,
          processingTimeMs: Date.now() - start,
        });
      } catch (err) {
        results.push({
          documentId: doc.id,
          error: (err as Error).message,
          processingTimeMs: Date.now() - start,
        });
      }

      processed++;
      if (processed % 10 === 0) {
        console.log(`Extracted ${processed}/${documents.length} documents`);
      }
    }
  }

  await Promise.all(Array.from({ length: concurrency }, () => worker()));
  return results;
}

// Downstream: route results to review queue vs auto-approve
function classifyResults(results: BatchExtractionResult[]) {
  const autoApproved = results.filter(r => r.result && !r.result.review_required);
  const requiresReview = results.filter(r => r.result?.review_required);
  const failed = results.filter(r => r.error);

  console.log(`Auto-approved: ${autoApproved.length}`);
  console.log(`Requires human review: ${requiresReview.length}`);
  console.log(`Failed (retry): ${failed.length}`);

  return { autoApproved, requiresReview, failed };
}

Model selection by volume and accuracy requirements

For extraction tasks at scale, model choice significantly affects cost:

  • claude-haiku-4-5-20251001 - ~5× cheaper than Sonnet. Suitable for well-structured documents (invoices, forms) where fields appear in predictable positions. Confidence scores tend to be accurately calibrated.
  • claude-sonnet-5 - better at extracting from dense prose, tables, and ambiguous formats. Use for legal documents, contracts, and anything where field boundaries are unclear.
  • Hybrid approach - run Haiku first; route documents where review_required: true or where any confidence score is below 0.5 to a Sonnet re-run. In practice, 80-90% of documents are handled by Haiku and only 10-20% need the more capable model.

Failure modes in extraction agents

  • Hallucinated field values - the model extracts a value that is not in the document. Mitigate by: instructing the model explicitly to omit rather than guess, using low confidence scores as a signal, and spot-checking a sample of auto-approved extractions against the source documents.
  • Format inconsistency - dates extracted as "15/03/2026" on some documents and "2026-03-15" on others. Include format requirements in the field description and add a post-processing normalisation step rather than relying on the model to be consistent.
  • Multi-value fields - a document may contain two invoice numbers (an amended invoice). The extraction notes field captures this; your post-processing code should flag these for human review rather than silently taking the first value.

For the forced tool use pattern that guarantees structured output from this agent, see tool use vs structured output. For running the extraction agent across large document batches with cost control, see agent cost management.