AI Agents
Build a production data extraction agent that converts unstructured text to typed schemas using Claude tool use - with confidence scoring.
You have 500 invoices in PDF, 1,000 support tickets in plain text, and 200 contract documents in Word format. Each contains structured information - amounts, dates, names, addresses, line items - but it is buried in unstructured prose. Manually extracting this data takes weeks. An extraction agent that reads each document and populates a typed schema takes hours, and once the schema is defined correctly, it scales to any volume. This tutorial builds an extraction agent that handles the real-world complexities: fields that are sometimes absent, values that appear in multiple formats, and documents where the extraction confidence is low enough to route to human review.
Unstructured documents (PDFs, emails, text)
│
▼
┌──────────────────────────────────┐
│ Extraction Agent │
│ (claude-sonnet-5 or haiku) │
│ │
│ Tool: extract_fields │
│ (forced via tool_choice) │
└──────────────┬───────────────────┘
│
▼
┌──────────────────────────────────┐
│ Typed extraction result │
│ { │
│ fields: { name, date... }, │
│ confidence: { per-field }, │
│ missing: string[], │
│ review_required: boolean │
│ } │
└──────────────────────────────────┘
The key architectural choice: use forced tool selection (tool_choice: { type: "tool", name: "extract_fields" }) rather than asking the model to return JSON in its text response. Forced tool use guarantees the output matches the schema - no regex parsing, no JSON.parse failures from prose leaking into the output:
import Anthropic from '@anthropic-ai/sdk';
const client = new Anthropic({ apiKey: process.env.ANTHROPIC_API_KEY });
// Define your extraction schema as a tool
function buildExtractionTool(schema: ExtractionSchema): Anthropic.Tool {
return {
name: 'extract_fields',
description: `Extract the specified fields from the document. For each field, provide the extracted value and a confidence score (0.0-1.0). If a field is not present in the document, omit it from the output rather than guessing.`,
input_schema: {
type: 'object' as const,
properties: {
fields: {
type: 'object',
description: 'Extracted field values.',
properties: Object.fromEntries(
Object.entries(schema.fields).map(([key, def]) => [
key,
{ type: def.type, description: def.description },
])
),
},
confidence: {
type: 'object',
description: 'Confidence score (0.0-1.0) for each extracted field.',
properties: Object.fromEntries(
Object.keys(schema.fields).map(key => [key, { type: 'number', minimum: 0, maximum: 1 }])
),
},
missing_fields: {
type: 'array',
items: { type: 'string' },
description: 'Required fields that could not be found in the document.',
},
extraction_notes: {
type: 'string',
description: 'Any ambiguities, multiple values found, or extraction decisions made.',
},
},
required: ['fields', 'confidence', 'missing_fields'],
},
};
}
// Example invoice schema
const invoiceSchema: ExtractionSchema = {
fields: {
invoice_number: { type: 'string', description: 'Invoice or reference number.', required: true },
invoice_date: { type: 'string', description: 'Invoice date in ISO 8601 format (YYYY-MM-DD).', required: true },
due_date: { type: 'string', description: 'Payment due date in ISO 8601 format.', required: false },
vendor_name: { type: 'string', description: 'Name of the vendor or supplier.', required: true },
vendor_address: { type: 'string', description: 'Full vendor mailing address.', required: false },
total_amount: { type: 'number', description: 'Total invoice amount as a decimal number.', required: true },
currency: { type: 'string', description: 'Three-letter currency code (e.g., USD, EUR, GBP).', required: true },
tax_amount: { type: 'number', description: 'Tax amount as a decimal number.', required: false },
line_items: {
type: 'array',
description: 'Array of line items: [{ description, quantity, unit_price, total }]',
required: false,
},
},
};
interface ExtractionResult {
fields: Record;
confidence: Record;
missing_fields: string[];
extraction_notes?: string;
review_required: boolean;
low_confidence_fields: string[];
}
const CONFIDENCE_THRESHOLD = 0.7; // Fields below this score are flagged for review
async function extractFromDocument(
documentText: string,
schema: ExtractionSchema,
options: {
model?: string;
documentType?: string;
} = {}
): Promise {
const model = options.model ?? 'claude-haiku-4-5-20251001'; // Haiku for cost efficiency at scale
const docType = options.documentType ?? 'document';
const tool = buildExtractionTool(schema);
const response = await client.messages.create({
model,
max_tokens: 2048,
system: `You are a precise data extraction specialist. Extract exactly what is written in the ${docType} - do not infer, estimate, or fill in values that are not explicitly present. When a date is written as "March 15, 2026", convert it to ISO format "2026-03-15". When a currency symbol appears without a code (e.g., "$"), infer the code from context (USD for US documents). If you are uncertain, assign a lower confidence score rather than omitting the field.`,
tools: [tool],
tool_choice: { type: 'tool', name: 'extract_fields' },
messages: [{
role: 'user',
content: `Extract the required fields from this ${docType}:
${documentText}`,
}],
});
const toolCall = response.content.find(
(b): b is Anthropic.ToolUseBlock => b.type === 'tool_use' && b.name === 'extract_fields'
);
if (!toolCall) throw new Error('Extraction tool was not called - check tool_choice configuration');
const raw = toolCall.input as {
fields: Record;
confidence: Record;
missing_fields: string[];
extraction_notes?: string;
};
// Identify fields below the confidence threshold
const lowConfidenceFields = Object.entries(raw.confidence)
.filter(([, score]) => score < CONFIDENCE_THRESHOLD)
.map(([field]) => field);
// Identify required fields that are missing
const requiredMissing = raw.missing_fields.filter(
field => schema.fields[field]?.required
);
const reviewRequired = lowConfidenceFields.length > 0 || requiredMissing.length > 0;
return {
...raw,
review_required: reviewRequired,
low_confidence_fields: lowConfidenceFields,
};
}
interface BatchExtractionResult {
documentId: string;
result?: ExtractionResult;
error?: string;
processingTimeMs: number;
}
async function batchExtract(
documents: Array<{ id: string; text: string }>,
schema: ExtractionSchema,
options: {
concurrency?: number;
model?: string;
documentType?: string;
} = {}
): Promise {
const concurrency = options.concurrency ?? 10;
const results: BatchExtractionResult[] = [];
const queue = [...documents];
let processed = 0;
async function worker(): Promise {
while (queue.length > 0) {
const doc = queue.shift();
if (!doc) break;
const start = Date.now();
try {
const result = await extractFromDocument(doc.text, schema, {
model: options.model,
documentType: options.documentType,
});
results.push({
documentId: doc.id,
result,
processingTimeMs: Date.now() - start,
});
} catch (err) {
results.push({
documentId: doc.id,
error: (err as Error).message,
processingTimeMs: Date.now() - start,
});
}
processed++;
if (processed % 10 === 0) {
console.log(`Extracted ${processed}/${documents.length} documents`);
}
}
}
await Promise.all(Array.from({ length: concurrency }, () => worker()));
return results;
}
// Downstream: route results to review queue vs auto-approve
function classifyResults(results: BatchExtractionResult[]) {
const autoApproved = results.filter(r => r.result && !r.result.review_required);
const requiresReview = results.filter(r => r.result?.review_required);
const failed = results.filter(r => r.error);
console.log(`Auto-approved: ${autoApproved.length}`);
console.log(`Requires human review: ${requiresReview.length}`);
console.log(`Failed (retry): ${failed.length}`);
return { autoApproved, requiresReview, failed };
}
For extraction tasks at scale, model choice significantly affects cost:
claude-haiku-4-5-20251001 - ~5× cheaper than Sonnet. Suitable for well-structured documents (invoices, forms) where fields appear in predictable positions. Confidence scores tend to be accurately calibrated.claude-sonnet-5 - better at extracting from dense prose, tables, and ambiguous formats. Use for legal documents, contracts, and anything where field boundaries are unclear.review_required: true or where any confidence score is below 0.5 to a Sonnet re-run. In practice, 80-90% of documents are handled by Haiku and only 10-20% need the more capable model.For the forced tool use pattern that guarantees structured output from this agent, see tool use vs structured output. For running the extraction agent across large document batches with cost control, see agent cost management.