AI Agents
When to hard-code a sequence of LLM calls and when to let a classifier route: prompt chaining vs routing on cost, latency and reliability, in TypeScript.
One prompt extracts facts from a sales call, drafts a follow-up email and formats a CRM note. It works on 85% of calls; on the rest it skips a commitment or leaves a placeholder in the email, and you cannot tell which part failed. Adding instructions rarely fixes that. Changing the control flow does: prompt chaining splits the work into steps your code sequences, and routing sends each input to a handler built for it.
Prompt chaining means your code, not the model, owns the order of operations. Each call does one narrow job, each handoff is checked, and a failure is visible at the step where it happened. Use it when a task has a known, repeatable structure that is too much for one prompt, and when you care more about reliability than about shaving off a round trip.
Anthropic's December 2024 essay Building effective agents draws a line that is still the most useful framing: workflows are systems where LLMs and tools are orchestrated through predefined code paths, while agents are systems where the model directs its own process and tool usage. Chaining and routing are the two most common workflows:
Chain: input -> [extract] -> gate -> [draft] -> gate -> output
Router: input -> [classify] -+-> refund handler
+-> bug report handler
+-> account handler
+-> general handler (fallback)
Agent: input -> [model <-> tools] repeated until done -> output
Moving right on this spectrum buys flexibility and costs predictability. The mistake teams make most often is starting at the agent end for a process that could have been drawn as a flowchart on day one.
A good chain has three properties: each step has one job, each handoff is structured, and code checks the handoff before the next step runs. Structured outputs (output_config.format) are the natural glue, because the next step receives a typed object instead of prose it has to re-parse.
import Anthropic from '@anthropic-ai/sdk';
const client = new Anthropic();
const CallFacts = z.object({
customer: z.string(),
problems: z.array(z.string()),
commitments: z.array(z.string()),
next_meeting: z.string().nullable(),
});
type CallFacts = z.infer<typeof CallFacts>;
// Step 1: extraction. Small, cheap model, schema-constrained output.
async function extractFacts(transcript: string): Promise<CallFacts> {
const res = await client.messages.parse({
model: 'claude-haiku-4-5',
max_tokens: 2048,
messages: [{
role: 'user',
content: [
{ type: 'text', text: 'Extract the facts from this sales call transcript. Only include commitments someone explicitly made.' },
{ type: 'text', text: transcript },
],
}],
output_config: { format: zodOutputFormat(CallFacts) },
});
if (!res.parsed_output) throw new Error('Extraction failed, stop_reason=' + res.stop_reason);
return res.parsed_output;
}
// Gate: plain code, no model call. Cheap checks catch most bad handoffs.
function checkFacts(facts: CallFacts): string[] {
const issues: string[] = [];
if (!facts.customer.trim()) issues.push('customer name missing');
if (facts.commitments.length === 0) issues.push('no commitments captured');
return issues;
}
// Step 2: drafting. Stronger model, receives only the validated facts.
async function draftFollowUp(facts: CallFacts): Promise<string> {
const res = await client.messages.create({
model: 'claude-sonnet-5',
max_tokens: 2048,
system: 'You write short follow-up emails after sales calls. Plain text, under 150 words, no placeholders.',
messages: [{ role: 'user', content: 'Write the follow-up email for these call facts: ' + JSON.stringify(facts) }],
});
if (res.stop_reason !== 'end_turn') throw new Error('Draft incomplete: ' + res.stop_reason);
return res.content.flatMap((b) => (b.type === 'text' ? [b.text] : [])).join('');
}
export async function processCall(transcript: string) {
const facts = await extractFacts(transcript);
const issues = checkFacts(facts);
if (issues.length > 0) return { status: 'needs_review' as const, issues, facts };
const email = await draftFollowUp(facts);
// Second gate, on the output: catch placeholders and runaway length.
if (email.includes('[') || email.split(' ').length > 180) {
return { status: 'needs_review' as const, issues: ['draft failed output checks'], facts, email };
}
return { status: 'ready' as const, facts, email };
}
Three details carry most of the value. The extraction step runs on claude-haiku-4-5, because pulling facts out of a transcript does not need the strongest model. The gates are plain code, so they cost nothing and never hallucinate. And a failed gate returns needs_review with the reason attached, so a human sees exactly which step went wrong instead of a vaguely bad email.
A router solves a different problem. Your inputs fall into categories that need different instructions, different context, or different models, and one prompt that covers all of them is long, expensive and mediocre at each. The router spends one cheap call on classification, then sends the input to a handler tuned for its category:
const Route = z.object({
category: z.enum(['refund', 'bug_report', 'account', 'other']),
confidence: z.enum(['high', 'low']),
});
const HANDLERS = {
refund: { model: 'claude-sonnet-5', system: REFUND_POLICY_PROMPT },
bug_report: { model: 'claude-sonnet-5', system: BUG_TRIAGE_PROMPT },
account: { model: 'claude-haiku-4-5', system: ACCOUNT_HELP_PROMPT },
other: { model: 'claude-sonnet-5', system: GENERAL_SUPPORT_PROMPT },
} as const;
async function classify(message: string): Promise<keyof typeof HANDLERS> {
const res = await client.messages.parse({
model: 'claude-haiku-4-5',
max_tokens: 256,
system: 'Classify the customer message. Use confidence "low" if it fits more than one category or none.',
messages: [{ role: 'user', content: message }],
output_config: { format: zodOutputFormat(Route) },
});
const route = res.parsed_output;
// Anything uncertain goes to the general handler, never to a specialist.
return route && route.confidence === 'high' ? route.category : 'other';
}
export async function answer(message: string): Promise<string> {
const handler = HANDLERS[await classify(message)];
const res = await client.messages.create({
model: handler.model,
max_tokens: 2048,
system: handler.system,
messages: [{ role: 'user', content: message }],
});
return res.content.flatMap((b) => (b.type === 'text' ? [b.text] : [])).join('');
}
The fallback route matters more than any specialist. A refund prompt that receives a bug report will confidently apply the refund policy to it; a general handler that receives a refund request will at worst give a less specific answer. Always route uncertainty to the handler with the least damaging failure mode.
Download a ready-made Claude Code subagent that reviews and tightens prompts, schemas and handoffs between LLM calls.
Get the prompt engineer subagent| Dimension | Chain | Router | Agent loop |
|---|---|---|---|
| Who decides the next step | Your code, always the same | One classification call | The model, every turn |
| Latency | Sum of all steps | Classifier plus one handler | Unbounded without a turn cap |
| Cost per input | Predictable | Predictable | Varies by an order of magnitude |
| Debugging | Failure is pinned to a step | Check the route, then the handler | Needs full traces |
| Handles unexpected input | Poorly | Through the fallback route | Well |
Chaining is the wrong tool when the path depends on what the run discovers. A research task, a debugging session or a data investigation cannot be drawn as a flowchart in advance, and forcing it into fixed steps produces a chain that either skips what matters or grows a branch for every case. That is where an agent loop earns its extra cost. Chaining is also unnecessary when a single well-structured prompt already passes your evals: measure first, then split.
In larger systems the patterns nest. A supervisor agent can call a chain as one of its tools, and a router can send only its hardest category to a full agent. For keeping the per-step model choices affordable, see agent cost management, and for where each step's state lives, stateful vs stateless agents.
Prompt chaining splits a task into a fixed sequence of LLM calls, where the output of each call becomes the input of the next. Your code decides the order, and programmatic checks between steps catch bad output before it propagates. It trades a little latency for accuracy and much easier debugging.
A chain runs every step in the same order for every input. A router runs one classification call first, then sends the input to one of several specialised prompts or models. Chains break a hard task into easy steps; routers send different kinds of input to handlers tuned for them.
Use an agent loop when the next step depends on what the previous step discovered and you cannot list the possible paths in advance, as in open-ended research or debugging. If you can draw the flowchart before the run starts, a chain or router is cheaper, faster and easier to test.
It adds calls, but often lowers total cost, because each step can use the cheapest model that handles it, for example Claude Haiku 4.5 for extraction and Claude Sonnet 5 only for drafting. The main cost is latency: every step in a chain is a sequential model round trip.