AI & Development

AI-Powered React Development in 2026: Patterns and Tools

How to build AI features in React apps: streaming responses, state management for LLM outputs, tool use patterns.

React applications in 2026 are increasingly expected to do things that traditional frontend code cannot: respond intelligently to freeform user input, generate content on demand, summarise complex data, and adapt to context in ways that no switch statement or template can capture. The LLM has become a standard backend dependency, as familiar in a React project as a REST API or a database. But wiring up an LLM properly - handling streaming, managing async state, designing UX for non-deterministic outputs - is not the same as calling a REST endpoint. This guide covers the patterns that are working in production React apps right now.

The foundational architecture decision: where does the LLM call live?

The first decision is whether the LLM call happens in a React Server Component (RSC), a server action, an API route, or a client-side fetch. The right answer depends on your use case, but the constraint is clear: your API key cannot be in client-side code. This rules out direct client-side LLM calls unless you are proxying through a secure backend.

For most React apps in 2026, the practical options are:

  • Server Actions (Next.js 15+) - Cleanest pattern for server-rendered apps. The action runs server-side, the API key stays on the server, and the result can be streamed back to the client with the useActionState hook. No separate API route needed.
  • API routes (/api/...) - Still the right choice for CRA apps, Vite apps without RSC, and any case where you want the LLM call available as a proper HTTP endpoint (e.g., for mobile clients sharing the same backend).
  • Dedicated backend service - For production apps with complex rate limiting, per-user token budgets, or multi-tenant requirements. Worth the overhead once you have more than a few AI features or significant scale.

Streaming responses: the most impactful UX improvement you can make

Non-streamed LLM responses create a blank wait. The spinner spins for 2 to 5 seconds, then the full response appears at once. This is cognitively jarring - the human brain interprets a sudden content appearance differently from content that arrives gradually. Streamed responses, where tokens appear as the model generates them, feel dramatically faster even when the total time to completion is identical.

Implementing streaming in a React app requires three things:

  1. An API route that returns a ReadableStream (supported natively in Next.js App Router and Hono, requires configuration in Express)
  2. Client-side code to read the stream incrementally and update state on each chunk
  3. React state that handles partial content gracefully - text that is still being written, punctuation that appears mid-word, sentence fragments

The cleanest client-side pattern for reading a stream in React:

const [content, setContent] = useState('');

async function fetchStream(prompt) {
  const res = await fetch('/api/ai', {
    method: 'POST',
    body: JSON.stringify({ prompt }),
  });
  const reader = res.body.getReader();
  const decoder = new TextDecoder();
  let accumulated = ', ';

  while (true) {
    const { done, value } = await reader.read();
    if (done) break;
    accumulated += decoder.decode(value, { stream: true });
    setContent(accumulated);
  }
}

For SSE (server-sent events) streams from Anthropic's API, you parse the event data differently, but the React state update pattern is the same: accumulate and set on each chunk.

Managing AI state in React: what custom hooks solve

AI features introduce state that standard React patterns do not handle cleanly: loading states that span seconds, partial content that is not yet meaningful, error states from external APIs with different failure modes than REST calls, and results that need to be cached to avoid re-calling on re-render.

A custom hook solves all of these in one place. The minimum viable AI hook:

function useAI() {
  const [status, setStatus] = useState('idle'); // idle | loading | streaming | complete | error
  const [content, setContent] = useState('');
  const [error, setError] = useState(null);
  const abortRef = useRef(null);

  async function run(prompt) {
    abortRef.current?.abort();
    const controller = new AbortController();
    abortRef.current = controller;

    setStatus('loading');
    setContent('');
    setError(null);

    try {
      const res = await fetch('/api/ai', {
        method: 'POST',
        body: JSON.stringify({ prompt }),
        signal: controller.signal,
      });
      setStatus('streaming');
      // ... stream reading logic
      setStatus('complete');
    } catch (err) {
      if (err.name !== 'AbortError') {
        setError(err.message);
        setStatus('error');
      }
    }
  }

  return { status, content, error, run, abort: () => abortRef.current?.abort() };
}

The abort controller is not optional - it prevents stale responses from overwriting current content when users trigger multiple requests before the first one completes.

Tool use in React: structured outputs without the parsing headache

Free-form text generation is one LLM use case. But many React features need structured data from the model - a categorisation decision, a set of extracted entities, a list of recommendations in a specific format. Parsing free-form text to extract structured data is fragile. Tool use (function calling) solves this by requiring the model to output data in a schema you define.

In a React context, the pattern is: define the tool schema on the server, call the API with the tool definition, receive tool call results as JSON, return the JSON to the client, update React state directly from the typed object. No parsing, no fallback logic for malformed responses. The model either calls the tool with valid arguments or it does not - and you handle both cases cleanly.

UX design for non-deterministic outputs

The most common failure in AI feature UX is treating the LLM output as if it were a deterministic database query. It is not. The model may produce different outputs for identical inputs, may occasionally produce outputs that are slightly off, and will sometimes refuse to answer in expected ways. Your UI needs to be designed for this reality:

  • Always provide a regenerate button - Users who get an unsatisfying response need a one-click way to try again. Without it, they leave.
  • Show generation state, not just loading state - A streaming interface with visible token-by-token output signals "the AI is thinking" in a way that a spinner cannot. Users are more patient with visible progress than invisible waiting.
  • Make AI output editable - For any content the user will use (drafts, summaries, plans), allow in-place editing. Position the AI as a collaborator, not a final authority.
  • Scope the feature clearly - Do not let users believe the AI can do things it cannot. A clear label ("Summarise this document" rather than "Ask me anything") prevents the frustration of mismatched expectations.

Error handling patterns for LLM API calls

LLM API errors have different characteristics from REST API errors. The most important categories to handle:

  • Rate limit errors (429) - Implement exponential backoff with jitter. Retry automatically up to three times before surfacing an error to the user.
  • Context window errors - The user's input exceeds the model's context limit. Truncate on the server before calling the API, not after the call fails.
  • Safety filter refusals - The model declines to answer. These should not reach the user as raw API errors. Detect the refusal response and surface a user-friendly message.
  • Stream interruptions - Network drops mid-stream are common on mobile browsers. Save the partial content and offer a "continue" option rather than losing everything generated so far.

Cost awareness in the frontend

React developers are often responsible for features without visibility into their API cost impact. A few numbers to have in mind: at 2026 pricing, a typical user query with a 500-token input and 800-token output on claude-haiku-4-5 costs approximately $0.0008. On claude-sonnet-5, the same query costs approximately $0.009. At 10,000 daily active users making two AI queries per session, the difference between Haiku and Sonnet for a feature is approximately $100/day vs $1,000/day. Model selection is a product decision that belongs in a conversation between engineers and product managers, not a default choice.

The patterns covered here - streaming, custom state hooks, tool use, deliberate error handling - are not advanced optimisations. They are the baseline for AI features that feel professional rather than bolted on. An AI feature that loads quickly, handles errors gracefully, gives users control, and produces reliably structured output is one that earns the premium experience users expect from AI-native apps in 2026.