AI Agents

Stateful vs Stateless AI Agents: Which Design to Choose

Stateless agents rebuild context on every request; stateful agents keep it between turns. When each wins on scaling, cost and failure recovery, with code.

An agent that works on your laptop starts forgetting things in production: the second message lands on another container, the in-memory history is gone, and the agent greets the user as a stranger. The opposite mistake is just as common: persisting every run of a one-shot summarizer to a database it never needed. Choosing between stateful and stateless agents is an architecture decision, and it is cheaper to make it on purpose.

Stateful vs stateless AI agent
A stateless AI agent receives everything it needs in a single request and keeps nothing afterwards, so any server instance can handle any request. A stateful AI agent persists its conversation history, task progress or memory between requests, so later turns build on earlier ones without the client resending that context.

Stateful vs stateless agents: the short answer

Use a stateless agent when the whole task fits inside one request: classification, extraction, or a bounded tool loop that finishes in seconds. Use a stateful agent when work spans several user turns, outlives a single process, or must survive failures. Most production systems end up with both: stateless workers that load and save state in an external store.

What "state" means for an agent

The Claude Messages API is itself stateless. Every call to client.messages.create() carries the full messages array, and the API remembers nothing between calls. So when people call an agent stateful, they mean their application keeps some of the following somewhere:

  • Conversation history: the Anthropic.MessageParam[] array, including every tool_use and tool_result block.
  • Task progress: which sub-tasks are done, the current iteration count, partial outputs.
  • Working artifacts: files written, records fetched, drafts in progress.
  • Long-term memory: facts about the user or project that should outlive one conversation. See agent memory and context persistence for that layer.
  • Side effects: emails sent or rows written. These live in other systems, but the agent needs to know they already happened.

The real question is not whether your agent has state (every multi-step agent does) but how long that state must live, and who holds it between requests: the process, the client, or a store.

Head-to-head comparison

DimensionStateless agentStateful agent
Horizontal scalingTrivial: any instance serves any requestNeeds a shared store or sticky sessions
Failure recoveryRetry the whole request from scratchResume from the last saved step
Per-turn overheadNo store round tripOne read and one write per turn, typically 2 to 20 ms each
Cost over timeFlat per requestHistory grows every turn unless you compact it
DebuggingRequest and response tell the whole storyYou must inspect stored state to reproduce a bug
Security surfaceHistory supplied by the client is untrustedHistory stays server-side
Best forExtraction, classification, one-shot tool loopsChat, long-running tasks, human-in-the-loop flows

When a stateless agent is the right choice

A stateless design wins whenever the task completes within one invocation:

  • The input arrives whole (a ticket, a document, a diff) and the output leaves whole.
  • The tool loop is bounded, typically under 10 iterations and under 60 seconds, so it fits inside one HTTP request or serverless invocation.
  • A failed run can simply be retried, because the tools are read-only or idempotent.
  • You deploy to autoscaling or serverless infrastructure and do not want to run a session database.

The whole agent loop runs inside the handler. Nothing is read from or written to a session store, so any instance can serve the request:

import Anthropic from '@anthropic-ai/sdk';

const client = new Anthropic();
const MAX_TURNS = 8;

// Stateless: the task starts and finishes inside this call.
export async function triageTicket(ticket: string): Promise<string> {
  const messages: Anthropic.MessageParam[] = [
    {
      role: 'user',
      content: [
        { type: 'text', text: 'Triage this support ticket. Look up the customer first, then reply with a priority (P1 to P4) and a one-line reason.' },
        { type: 'text', text: ticket },
      ],
    },
  ];

  for (let turn = 0; turn < MAX_TURNS; turn++) {
    const response = await client.messages.create({
      model: 'claude-sonnet-5',
      max_tokens: 4096,
      tools,
      messages,
    });
    messages.push({ role: 'assistant', content: response.content });

    if (response.stop_reason === 'end_turn') {
      return response.content.flatMap((b) => (b.type === 'text' ? [b.text] : [])).join('');
    }
    if (response.stop_reason !== 'tool_use') {
      throw new Error('Triage stopped early: ' + response.stop_reason);
    }

    // Every tool_use block gets exactly one tool_result, errors included.
    const results: Anthropic.ToolResultBlockParam[] = [];
    for (const block of response.content) {
      if (block.type !== 'tool_use') continue;
      try {
        results.push({ type: 'tool_result', tool_use_id: block.id, content: await runTool(block.name, block.input) });
      } catch (err) {
        results.push({ type: 'tool_result', tool_use_id: block.id, content: String(err), is_error: true });
      }
    }
    messages.push({ role: 'user', content: results });
  }
  throw new Error('Triage did not finish within ' + MAX_TURNS + ' turns');
}

If the process dies mid-run, the caller retries and the agent starts again from the ticket. For a task that costs a few cents and takes ten seconds, that is the right trade: the retry is cheaper than the infrastructure needed to avoid it.

When you need a stateful agent

Move to a stateful design when any of these is true:

  • The user talks to the agent over several turns, and each turn depends on the previous ones.
  • A run takes minutes or hours, so a crash near the end would waste real money. See agent checkpointing for saving progress mid-run.
  • The agent pauses for a human approval and resumes later, possibly on another machine.
  • Several requests can touch the same session at once, for example a user with two browser tabs open.

The core of a stateful agent is a load, append, run, save cycle. The detail most first implementations miss is concurrency: two requests for the same session both load version 4, both append a turn, and the second save silently erases the first. An optimistic version check turns that silent data loss into an error you can retry:

import Anthropic from '@anthropic-ai/sdk';

const client = new Anthropic();
const db = new Pool({ connectionString: process.env.DATABASE_URL });

type Session = { id: string; messages: Anthropic.MessageParam[]; version: number };

export class SessionConflictError extends Error {}

async function loadSession(id: string): Promise<Session> {
  const { rows } = await db.query('SELECT messages, version FROM agent_sessions WHERE id = $1', [id]);
  return rows.length
    ? { id, messages: rows[0].messages, version: rows[0].version }
    : { id, messages: [], version: 0 };
}

// The write only lands if nobody else saved this session since we loaded it.
async function saveSession(s: Session): Promise<void> {
  const { rowCount } = await db.query(
    `INSERT INTO agent_sessions (id, messages, version, updated_at)
     VALUES ($1, $2, 1, now())
     ON CONFLICT (id) DO UPDATE
       SET messages = $2, version = agent_sessions.version + 1, updated_at = now()
       WHERE agent_sessions.version = $3`,
    [s.id, JSON.stringify(s.messages), s.version],
  );
  if (rowCount === 0) throw new SessionConflictError('Session ' + s.id + ' changed concurrently');
}

export async function chatTurn(sessionId: string, userText: string): Promise<string> {
  const session = await loadSession(sessionId);
  session.messages.push({ role: 'user', content: userText });

  const response = await client.messages.create({
    model: 'claude-sonnet-5',
    max_tokens: 4096,
    system: 'You are the support assistant for Acme. Keep answers under 120 words.',
    messages: session.messages,
  });
  if (response.stop_reason === 'refusal') throw new Error('Request declined by the model');

  session.messages.push({ role: 'assistant', content: response.content });
  await saveSession(session); // on SessionConflictError the caller reloads and retries
  return response.content.flatMap((b) => (b.type === 'text' ? [b.text] : [])).join('');
}

Notice that the handler itself is still stateless. It holds the session only for the duration of one request, which is why this pattern scales like a stateless service while behaving like a stateful one.

Designing an agent architecture?

Download a ready-made Claude Code subagent that reviews agent designs for state, memory and failure handling.

Get the AI agent architect subagent

Failure modes on both sides

Each design fails in its own characteristic way, and most production incidents trace back to one of these.

Stateless pitfalls

  • Client-held history is forgeable. If the browser sends the full message history back on every turn, a user can edit it: insert a fake tool_result saying their refund was approved, or a fake assistant turn agreeing to something. Keep history server-side, or sign it with an HMAC and verify the signature before use.
  • Payload growth. A 40-turn conversation with tool results can reach hundreds of kilobytes per request, which hurts mobile clients and runs into request size limits.
  • Timeouts. A loop that usually finishes in 20 seconds occasionally takes 90, and your serverless platform kills it at 60. Measure the p99, not the average.

Stateful pitfalls

  • Unbounded history. Every turn resends the whole transcript, so input tokens grow with each turn. Summarize old turns, clear stale tool results, or use server-side compaction (beta header compact-2026-01-12) for very long sessions. Prompt caching also helps, because the stable prefix of the history is billed at the cache read rate.
  • Orphaned tool calls. A crash between saving an assistant turn that contains tool_use and saving its tool_result leaves a history the API rejects with a 400. Repair it on load by adding error results for any unanswered calls.
  • Stale sessions. Without a TTL, abandoned sessions accumulate forever. Expire them, and decide explicitly whether any facts should be promoted to long-term memory first.
  • Schema drift. Stored histories outlive code releases. Version your session format so a deploy does not break every conversation in flight.

The hybrid most teams end up with

The common production shape is stateless compute with externalized state: request handlers and queue workers that hold nothing between requests, plus a store (Postgres, Redis, DynamoDB) that holds the session. It scales horizontally, survives deploys, and keeps history out of the client's hands.

If you would rather not run that store yourself, there are managed options. Claude Managed Agents (beta) keep sessions and their workspace on Anthropic's side, and the Claude Agent SDK can resume an earlier session by ID. Both move the state problem rather than remove it, so the questions about expiry, size and concurrency still apply.

For how stored history flows between cooperating agents, see agent handoff patterns. For keeping long sessions affordable, see agent cost management, and for the loop that sits inside every turn, the agent loop explained.

Decision guide

  1. Can the task finish inside one request, reliably under your platform's timeout? Start stateless.
  2. Does the user come back for another turn that depends on this one? Store the history server-side.
  3. Would a crash near the end waste more than a database write per step costs? Add checkpoints.
  4. Can two requests touch the same session at once? Add a version check before you ship, not after the first support ticket.
  5. Is the history growing past what one request should carry? Add compaction or summarization.

FAQ

Is the Claude API stateful?

No. The Messages API is stateless: each request must include the full conversation history in the messages array, and nothing is remembered between calls. Any state an agent has comes from your application storing and resending that history, or from a managed layer such as Claude Managed Agents.

Should I store agent conversation history in the client or on the server?

On the server for anything that matters. History held by the client can be edited, including fake tool results or fake assistant turns, so decisions based on it are unsafe. If you must keep it client-side, sign it and verify the signature on every request.

How do I scale a stateful agent horizontally?

Keep the request handlers stateless and put session state in a shared store keyed by session ID. Each request loads the session, runs one turn and saves it with an optimistic version check, so concurrent requests cannot overwrite each other. Any instance can then serve any turn.

When does a stateless agent stop being enough?

When a task spans several user turns, runs longer than your request timeout, needs to pause for human approval, or would waste significant cost if it restarted from scratch after a failure. Any one of these is a signal to persist state between requests.