AI & Development

Structured Outputs from LLMs: JSON, Schemas, and Reliable Parsing

Getting reliable structured JSON from LLMs is a core challenge for AI applications. Here is how structured output modes, JSON schemas, and validation

The majority of LLM use cases in production require structured data output. You are not asking the model to write an essay - you are asking it to extract entities from a document, classify an input into one of several categories, or populate a data object that gets inserted into a database. For these applications, receiving free-form prose from the model and then trying to parse it is fragile, unreliable, and an unnecessary source of failure. Structured output modes solve this problem.

Structured output mode

The major LLM providers now offer a structured output mode where you provide a JSON Schema along with your prompt, and the model is constrained to generate a response that exactly matches that schema. This is not best-effort - it uses constrained decoding to guarantee that every required field is present, every type matches, and every enum value is one of the allowed options.

Structured output mode is the right default for any application that needs to parse model outputs programmatically. It eliminates an entire class of production errors: invalid JSON, missing fields, wrong types, values outside enum constraints. The validation happens at the model, not in your parsing code.

Designing schemas for reliability

The schema you provide shapes not just the output format but also the model's reasoning about the task. A schema with clear, well-named fields and descriptive field descriptions produces more accurate outputs than an identical schema with cryptic field names. "customer_sentiment: One of 'positive', 'neutral', 'negative' - the overall sentiment of the customer message" gives the model more signal than "sentiment: string".

Enum constraints are particularly valuable. Wherever the model's output should be one of a finite set of values - classification labels, status codes, category names - defining an enum in the schema constrains the output to those values and eliminates the model's tendency to invent variations ("Very Positive" instead of "positive", or "neg" instead of "negative").

Deeply nested schemas can reduce model accuracy on the deeper fields. For complex data models, flatter schemas with explicit descriptions of the relationships between fields outperform deeply nested ones. If you need nested output, consider breaking it into multiple calls where each call populates a portion of the schema.

Tool use as structured output

An alternative approach to structured output mode is using the function calling feature with a tool whose parameters define your desired schema. The model generates a tool call with the extracted or classified data as arguments. This works reliably and has the advantage of being supported by more models and in more API configurations than structured output mode.

The tool-based approach is particularly well-suited to extraction tasks: define a tool named "extract_invoice_data" with parameters for vendor_name, invoice_number, date, line_items, and total_amount, and the model reliably populates those parameters from the invoice text. The tool call output is structured by construction and can be parsed directly.

Validation and error recovery

Even with structured output mode, validation in your application code is good practice. JSON Schema validation libraries (Zod in TypeScript, Pydantic in Python) let you validate the parsed output against your schema and surface specific field-level errors when something does not match expectations. Combined with structured output guarantees, this gives you multiple layers of protection.

For cases where the model produces output that fails validation despite your best efforts, a retry with a more targeted prompt is more reliable than trying to parse and fix malformed output. "The previous output was missing the required field X. Generate the output again, making sure to include X" is effective for fixing specific validation failures.

Streaming structured output

For applications where response latency is important, structured output can be streamed - the model sends tokens as it generates them, and your application can begin processing partial JSON before the full response is complete. Streaming partial JSON requires an incremental JSON parser (IJSON in Python, streaming-json-parser in Node.js) rather than waiting for the complete response to parse. For long structured outputs like lists of extracted items or multi-field objects, streaming can significantly improve perceived performance.