AI & Development
LLM outputs are not type-safe - until you make them so. How to use Zod schemas, structured outputs, and retry logic to get reliable typed data from any LLM in TypeScript.
The most common runtime error in AI applications is not an API failure or a network timeout. It is a shape mismatch: the LLM returned something that was almost the object you expected - a field with the wrong type, a missing required property, a number serialised as a string - and the application crashed silently or surfaced a confusing error to the user. This is not a reliability problem with the LLM. It is an architectural problem: the application trusted unvalidated data as if it were typed. Zod-based validation solves this at the boundary, before the data reaches your application logic.
Even when you use structured output features (Anthropic's tool use, OpenAI's JSON mode, Gemini's response schema), the guarantee is weaker than it appears. The model is constrained to produce syntactically valid JSON - but within that JSON, field names can vary, numeric values can appear as strings, optional fields can be absent, and arrays can be empty when you expected at least one item. The schema enforcement is syntactic, not semantic.
The more fundamental issue: LLM output depends on the exact prompt, the model version, the temperature, and the specific input. A prompt that has produced correct outputs for 10,000 requests can produce a malformed output on request 10,001 for an input you did not anticipate. Without validation, this failure is invisible until a user reports it - or until the bug propagates to a database write, an API call, or a payment flow.
Define your expected output schema as a Zod schema, then parse the LLM's response through it. The parse either succeeds and returns a correctly typed object, or throws a ZodError with details about what was wrong:
import { z } from 'zod';
const ArticleMetaSchema = z.object({
title: z.string().min(10).max(80),
summary: z.string().min(50).max(200),
tags: z.array(z.string()).min(2).max(6),
readingTimeMinutes: z.number().int().min(1).max(60),
difficulty: z.enum(['beginner', 'intermediate', 'advanced']),
});
type ArticleMeta = z.infer<typeof ArticleMetaSchema>;
async function getArticleMeta(articleText: string): Promise<ArticleMeta> {
const response = await callLLM(prompt(articleText)); // your API call
const parsed = JSON.parse(response);
return ArticleMetaSchema.parse(parsed); // throws ZodError if invalid
}
The critical step is ArticleMetaSchema.parse(parsed). After this line, result is fully typed - TypeScript knows the exact shape, and the runtime has verified it. Before this line, parsed is unknown and nothing about it can be trusted.
A validation failure on the first attempt does not mean the feature is broken - it means the LLM produced a response that needs correction. The right response is to retry with the validation error included in the follow-up prompt. Models are generally good at self-correction when given precise error information:
async function getValidatedOutput<T>(
schema: z.ZodType<T>,
userPrompt: string,
maxRetries = 3
): Promise<T> {
let lastError: string | null = null;
for (let attempt = 0; attempt < maxRetries; attempt++) {
const systemPrompt = lastError
? `You are a JSON generator. Your previous response failed validation: ${lastError}. Fix the issues and return only valid JSON.`
: 'You are a JSON generator. Return only valid JSON with no explanation or markdown.';
const raw = await callLLM({ systemPrompt, userPrompt });
try {
const parsed = JSON.parse(extractJSON(raw));
return schema.parse(parsed);
} catch (err) {
if (err instanceof z.ZodError) {
lastError = err.errors.map(e => `${e.path.join('.')}: ${e.message}`).join('; ');
} else {
lastError = 'Response was not valid JSON';
}
}
}
throw new Error(`Failed to get valid output after ${maxRetries} attempts. Last error: ${lastError}`);
}
Including the Zod error details in the retry prompt is key. A retry that says "your response was invalid, try again" gives the model no actionable information. A retry that says "readingTimeMinutes: Expected number, received string; difficulty: Invalid enum value, expected 'beginner' | 'intermediate' | 'advanced'" gives the model exactly what it needs to correct the output.
Not every LLM call needs to throw on validation failure. For features where partial data is better than no data - enrichment, suggestions, metadata - use safeParse to handle validation failures gracefully:
const result = ArticleMetaSchema.safeParse(parsed);
if (result.success) {
// result.data is ArticleMeta, fully typed
return result.data;
} else {
// Partial recovery: extract what we can, fall back for the rest
console.warn('LLM output validation failed:', result.error.issues);
return {
title: parsed?.title ?? 'Untitled',
summary: parsed?.summary ?? '',
tags: [],
readingTimeMinutes: 5,
difficulty: 'intermediate',
};
}
The choice between parse (throws) and safeParse (returns success/error) maps to the feature's criticality. For a checkout flow that depends on a price extraction, throw. For a tag suggestion that enriches a blog post, fall back gracefully.
When the model supports tool use (function calling), structured outputs are substantially more reliable than JSON-in-prose. Instead of asking the model to "return JSON in this format," you define a tool schema and force the model to call it. The model has to produce arguments that match the schema - the constraint is applied at generation time, not after.
In practice, tool use reduces validation failure rates from ~5% to under 0.5% for most structured output tasks. The tradeoff is slightly more verbose API call setup. For any production feature that depends on structured LLM output, tool use is the right default - Zod validation on top of it is defence in depth, not your primary reliability mechanism.
Zod offers both strict parsing and coercion. z.number() rejects a value of "5". z.coerce.number() accepts it and converts it. The choice matters for LLM output:
A practical approach: use z.coerce.number() and z.coerce.boolean() for scalar types where the semantic is clear regardless of serialisation format, but keep structural validators strict so shape errors surface clearly in logs.
The runtime boundary between LLM output and application state is the most important validation point in an AI application. Everything before it is uncertain; everything after it should be as typed and verified as any other data your application handles. Zod at that boundary is not defensive programming - it is the basic discipline that makes AI features reliable enough to put in production.