AI Agents
Stateless agents rebuild context on every request; stateful agents keep it between turns. When each wins on scaling, cost and failure recovery, with code.
An agent that works on your laptop starts forgetting things in production: the second message lands on another container, the in-memory history is gone, and the agent greets the user as a stranger. The opposite mistake is just as common: persisting every run of a one-shot summarizer to a database it never needed. Choosing between stateful and stateless agents is an architecture decision, and it is cheaper to make it on purpose.
Use a stateless agent when the whole task fits inside one request: classification, extraction, or a bounded tool loop that finishes in seconds. Use a stateful agent when work spans several user turns, outlives a single process, or must survive failures. Most production systems end up with both: stateless workers that load and save state in an external store.
The Claude Messages API is itself stateless. Every call to client.messages.create() carries the full messages array, and the API remembers nothing between calls. So when people call an agent stateful, they mean their application keeps some of the following somewhere:
Anthropic.MessageParam[] array, including every tool_use and tool_result block.The real question is not whether your agent has state (every multi-step agent does) but how long that state must live, and who holds it between requests: the process, the client, or a store.
| Dimension | Stateless agent | Stateful agent |
|---|---|---|
| Horizontal scaling | Trivial: any instance serves any request | Needs a shared store or sticky sessions |
| Failure recovery | Retry the whole request from scratch | Resume from the last saved step |
| Per-turn overhead | No store round trip | One read and one write per turn, typically 2 to 20 ms each |
| Cost over time | Flat per request | History grows every turn unless you compact it |
| Debugging | Request and response tell the whole story | You must inspect stored state to reproduce a bug |
| Security surface | History supplied by the client is untrusted | History stays server-side |
| Best for | Extraction, classification, one-shot tool loops | Chat, long-running tasks, human-in-the-loop flows |
A stateless design wins whenever the task completes within one invocation:
The whole agent loop runs inside the handler. Nothing is read from or written to a session store, so any instance can serve the request:
import Anthropic from '@anthropic-ai/sdk';
const client = new Anthropic();
const MAX_TURNS = 8;
// Stateless: the task starts and finishes inside this call.
export async function triageTicket(ticket: string): Promise<string> {
const messages: Anthropic.MessageParam[] = [
{
role: 'user',
content: [
{ type: 'text', text: 'Triage this support ticket. Look up the customer first, then reply with a priority (P1 to P4) and a one-line reason.' },
{ type: 'text', text: ticket },
],
},
];
for (let turn = 0; turn < MAX_TURNS; turn++) {
const response = await client.messages.create({
model: 'claude-sonnet-5',
max_tokens: 4096,
tools,
messages,
});
messages.push({ role: 'assistant', content: response.content });
if (response.stop_reason === 'end_turn') {
return response.content.flatMap((b) => (b.type === 'text' ? [b.text] : [])).join('');
}
if (response.stop_reason !== 'tool_use') {
throw new Error('Triage stopped early: ' + response.stop_reason);
}
// Every tool_use block gets exactly one tool_result, errors included.
const results: Anthropic.ToolResultBlockParam[] = [];
for (const block of response.content) {
if (block.type !== 'tool_use') continue;
try {
results.push({ type: 'tool_result', tool_use_id: block.id, content: await runTool(block.name, block.input) });
} catch (err) {
results.push({ type: 'tool_result', tool_use_id: block.id, content: String(err), is_error: true });
}
}
messages.push({ role: 'user', content: results });
}
throw new Error('Triage did not finish within ' + MAX_TURNS + ' turns');
}
If the process dies mid-run, the caller retries and the agent starts again from the ticket. For a task that costs a few cents and takes ten seconds, that is the right trade: the retry is cheaper than the infrastructure needed to avoid it.
Move to a stateful design when any of these is true:
The core of a stateful agent is a load, append, run, save cycle. The detail most first implementations miss is concurrency: two requests for the same session both load version 4, both append a turn, and the second save silently erases the first. An optimistic version check turns that silent data loss into an error you can retry:
import Anthropic from '@anthropic-ai/sdk';
const client = new Anthropic();
const db = new Pool({ connectionString: process.env.DATABASE_URL });
type Session = { id: string; messages: Anthropic.MessageParam[]; version: number };
export class SessionConflictError extends Error {}
async function loadSession(id: string): Promise<Session> {
const { rows } = await db.query('SELECT messages, version FROM agent_sessions WHERE id = $1', [id]);
return rows.length
? { id, messages: rows[0].messages, version: rows[0].version }
: { id, messages: [], version: 0 };
}
// The write only lands if nobody else saved this session since we loaded it.
async function saveSession(s: Session): Promise<void> {
const { rowCount } = await db.query(
`INSERT INTO agent_sessions (id, messages, version, updated_at)
VALUES ($1, $2, 1, now())
ON CONFLICT (id) DO UPDATE
SET messages = $2, version = agent_sessions.version + 1, updated_at = now()
WHERE agent_sessions.version = $3`,
[s.id, JSON.stringify(s.messages), s.version],
);
if (rowCount === 0) throw new SessionConflictError('Session ' + s.id + ' changed concurrently');
}
export async function chatTurn(sessionId: string, userText: string): Promise<string> {
const session = await loadSession(sessionId);
session.messages.push({ role: 'user', content: userText });
const response = await client.messages.create({
model: 'claude-sonnet-5',
max_tokens: 4096,
system: 'You are the support assistant for Acme. Keep answers under 120 words.',
messages: session.messages,
});
if (response.stop_reason === 'refusal') throw new Error('Request declined by the model');
session.messages.push({ role: 'assistant', content: response.content });
await saveSession(session); // on SessionConflictError the caller reloads and retries
return response.content.flatMap((b) => (b.type === 'text' ? [b.text] : [])).join('');
}
Notice that the handler itself is still stateless. It holds the session only for the duration of one request, which is why this pattern scales like a stateless service while behaving like a stateful one.
Download a ready-made Claude Code subagent that reviews agent designs for state, memory and failure handling.
Get the AI agent architect subagentEach design fails in its own characteristic way, and most production incidents trace back to one of these.
tool_result saying their refund was approved, or a fake assistant turn agreeing to something. Keep history server-side, or sign it with an HMAC and verify the signature before use.compact-2026-01-12) for very long sessions. Prompt caching also helps, because the stable prefix of the history is billed at the cache read rate.tool_use and saving its tool_result leaves a history the API rejects with a 400. Repair it on load by adding error results for any unanswered calls.The common production shape is stateless compute with externalized state: request handlers and queue workers that hold nothing between requests, plus a store (Postgres, Redis, DynamoDB) that holds the session. It scales horizontally, survives deploys, and keeps history out of the client's hands.
If you would rather not run that store yourself, there are managed options. Claude Managed Agents (beta) keep sessions and their workspace on Anthropic's side, and the Claude Agent SDK can resume an earlier session by ID. Both move the state problem rather than remove it, so the questions about expiry, size and concurrency still apply.
For how stored history flows between cooperating agents, see agent handoff patterns. For keeping long sessions affordable, see agent cost management, and for the loop that sits inside every turn, the agent loop explained.
No. The Messages API is stateless: each request must include the full conversation history in the messages array, and nothing is remembered between calls. Any state an agent has comes from your application storing and resending that history, or from a managed layer such as Claude Managed Agents.
On the server for anything that matters. History held by the client can be edited, including fake tool results or fake assistant turns, so decisions based on it are unsafe. If you must keep it client-side, sign it and verify the signature on every request.
Keep the request handlers stateless and put session state in a shared store keyed by session ID. Each request loads the session, runs one turn and saves it with an optimistic version check, so concurrent requests cannot overwrite each other. Any instance can then serve any turn.
When a task spans several user turns, runs longer than your request timeout, needs to pause for human approval, or would waste significant cost if it restarted from scratch after a failure. Any one of these is a signal to persist state between requests.