AI & Development
How to build AI features in React apps: streaming responses, state management for LLM outputs, tool use patterns.
React applications in 2026 are increasingly expected to do things that traditional frontend code cannot: respond intelligently to freeform user input, generate content on demand, summarise complex data, and adapt to context in ways that no switch statement or template can capture. The LLM has become a standard backend dependency, as familiar in a React project as a REST API or a database. But wiring up an LLM properly - handling streaming, managing async state, designing UX for non-deterministic outputs - is not the same as calling a REST endpoint. This guide covers the patterns that are working in production React apps right now.
The first decision is whether the LLM call happens in a React Server Component (RSC), a server action, an API route, or a client-side fetch. The right answer depends on your use case, but the constraint is clear: your API key cannot be in client-side code. This rules out direct client-side LLM calls unless you are proxying through a secure backend.
For most React apps in 2026, the practical options are:
useActionState hook. No separate API route needed.Non-streamed LLM responses create a blank wait. The spinner spins for 2 to 5 seconds, then the full response appears at once. This is cognitively jarring - the human brain interprets a sudden content appearance differently from content that arrives gradually. Streamed responses, where tokens appear as the model generates them, feel dramatically faster even when the total time to completion is identical.
Implementing streaming in a React app requires three things:
ReadableStream (supported natively in Next.js App Router and Hono, requires configuration in Express)The cleanest client-side pattern for reading a stream in React:
const [content, setContent] = useState('');
async function fetchStream(prompt) {
const res = await fetch('/api/ai', {
method: 'POST',
body: JSON.stringify({ prompt }),
});
const reader = res.body.getReader();
const decoder = new TextDecoder();
let accumulated = ', ';
while (true) {
const { done, value } = await reader.read();
if (done) break;
accumulated += decoder.decode(value, { stream: true });
setContent(accumulated);
}
}
For SSE (server-sent events) streams from Anthropic's API, you parse the event data differently, but the React state update pattern is the same: accumulate and set on each chunk.
AI features introduce state that standard React patterns do not handle cleanly: loading states that span seconds, partial content that is not yet meaningful, error states from external APIs with different failure modes than REST calls, and results that need to be cached to avoid re-calling on re-render.
A custom hook solves all of these in one place. The minimum viable AI hook:
function useAI() {
const [status, setStatus] = useState('idle'); // idle | loading | streaming | complete | error
const [content, setContent] = useState('');
const [error, setError] = useState(null);
const abortRef = useRef(null);
async function run(prompt) {
abortRef.current?.abort();
const controller = new AbortController();
abortRef.current = controller;
setStatus('loading');
setContent('');
setError(null);
try {
const res = await fetch('/api/ai', {
method: 'POST',
body: JSON.stringify({ prompt }),
signal: controller.signal,
});
setStatus('streaming');
// ... stream reading logic
setStatus('complete');
} catch (err) {
if (err.name !== 'AbortError') {
setError(err.message);
setStatus('error');
}
}
}
return { status, content, error, run, abort: () => abortRef.current?.abort() };
}
The abort controller is not optional - it prevents stale responses from overwriting current content when users trigger multiple requests before the first one completes.
Free-form text generation is one LLM use case. But many React features need structured data from the model - a categorisation decision, a set of extracted entities, a list of recommendations in a specific format. Parsing free-form text to extract structured data is fragile. Tool use (function calling) solves this by requiring the model to output data in a schema you define.
In a React context, the pattern is: define the tool schema on the server, call the API with the tool definition, receive tool call results as JSON, return the JSON to the client, update React state directly from the typed object. No parsing, no fallback logic for malformed responses. The model either calls the tool with valid arguments or it does not - and you handle both cases cleanly.
The most common failure in AI feature UX is treating the LLM output as if it were a deterministic database query. It is not. The model may produce different outputs for identical inputs, may occasionally produce outputs that are slightly off, and will sometimes refuse to answer in expected ways. Your UI needs to be designed for this reality:
LLM API errors have different characteristics from REST API errors. The most important categories to handle:
React developers are often responsible for features without visibility into their API cost impact. A few numbers to have in mind: at 2026 pricing, a typical user query with a 500-token input and 800-token output on claude-haiku-4-5 costs approximately $0.0008. On claude-sonnet-5, the same query costs approximately $0.009. At 10,000 daily active users making two AI queries per session, the difference between Haiku and Sonnet for a feature is approximately $100/day vs $1,000/day. Model selection is a product decision that belongs in a conversation between engineers and product managers, not a default choice.
The patterns covered here - streaming, custom state hooks, tool use, deliberate error handling - are not advanced optimisations. They are the baseline for AI features that feel professional rather than bolted on. An AI feature that loads quickly, handles errors gracefully, gives users control, and produces reliably structured output is one that earns the premium experience users expect from AI-native apps in 2026.