AI Agents
How to checkpoint AI agent state so long-running agents survive failures and can resume mid-task - what to persist, storage backend options.
A long-running agent - one that searches dozens of web pages, processes hundreds of records, or orchestrates five sub-tasks over 20 minutes - has a meaningful probability of failure before it finishes. Network timeouts, API rate limits, context window overflow, and infrastructure outages all increase with task duration. Without checkpointing, a failure at iteration 47 of 50 costs you exactly as much as a failure at iteration 1: you restart from zero. Checkpointing converts that total loss into a partial recovery, and for production agents running costly tasks, the difference is significant.
Agent checkpointing is the practice of serialising an agent's execution state to durable storage at defined points during a run, then loading that snapshot on restart so the agent continues from the saved milestone rather than from the task's beginning. A checkpoint captures the message history (the agent's working memory), any intermediate results, the current goal and sub-task status, and sufficient context for the agent to resume decision-making correctly.
Not all agent state has equal value. The checkpoint should be selective:
Anthropic.MessageParam[] array. This is the agent's full working memory and is the most critical thing to preserve. Without it, the resumed agent has no knowledge of what it has already done.Do not checkpoint: ephemeral tool state (open file handles, in-progress HTTP requests), model-internal reasoning that is already reflected in the message history, or redundant copies of data already in tool results.
The checkpoint write should happen after every meaningful milestone - typically after a tool result is appended to the message history, or after a sub-task completes. Checkpointing after every single API call is excessive for most tasks; checkpointing only at the end defeats the purpose. A heuristic: checkpoint every 5-10 iterations, and always checkpoint immediately after completing any irreversible action (a write, a send, a create):
import Anthropic from '@anthropic-ai/sdk';
const client = new Anthropic();
const supabase = createClient(process.env.SUPABASE_URL!, process.env.SUPABASE_KEY!);
interface AgentCheckpoint {
task_id: string;
goal: string;
messages: Anthropic.MessageParam[];
completed_tasks: string[];
iteration: number;
created_at: string;
updated_at: string;
}
async function saveCheckpoint(taskId: string, state: Omit): Promise {
const now = new Date().toISOString();
const { error } = await supabase
.from('agent_checkpoints')
.upsert({
task_id: taskId,
...state,
updated_at: now,
}, { onConflict: 'task_id' });
if (error) {
// Log but don't throw - a checkpoint failure should not abort the agent
console.error('Checkpoint save failed:', error.message);
}
}
async function loadCheckpoint(taskId: string): Promise {
const { data, error } = await supabase
.from('agent_checkpoints')
.select('*')
.eq('task_id', taskId)
.single();
if (error || !data) return null;
return data as AgentCheckpoint;
}
async function runCheckpointedAgent(taskId: string, goal: string): Promise {
// Attempt to resume from checkpoint
const checkpoint = await loadCheckpoint(taskId);
let messages: Anthropic.MessageParam[] = checkpoint
? checkpoint.messages
: [{ role: 'user', content: goal }];
let iteration = checkpoint ? checkpoint.iteration : 0;
const completedTasks = checkpoint ? new Set(checkpoint.completed_tasks) : new Set();
const MAX_ITERATIONS = 30;
console.log(checkpoint
? `Resuming from checkpoint at iteration ${iteration}`
: 'Starting fresh agent run'
);
for (; iteration < MAX_ITERATIONS; iteration++) {
const response = await client.messages.create({
model: 'claude-sonnet-5',
max_tokens: 4096,
tools: agentTools,
messages,
});
messages.push({ role: 'assistant', content: response.content });
if (response.stop_reason === 'end_turn') {
// Clean up checkpoint on success
await supabase.from('agent_checkpoints').delete().eq('task_id', taskId);
return extractFinalAnswer(response.content);
}
if (response.stop_reason === 'tool_use') {
const toolResults = await executeTools(response.content, completedTasks);
messages.push({ role: 'user', content: toolResults });
// Checkpoint every 5 iterations and after tool execution
if (iteration % 5 === 0) {
await saveCheckpoint(taskId, {
goal,
messages,
completed_tasks: Array.from(completedTasks),
iteration: iteration + 1,
});
}
}
}
throw new Error(`Agent did not complete within ${MAX_ITERATIONS} iterations`);
}
The checkpoint must outlive the process. Three practical options in order of increasing durability and complexity:
task_id primary key. The messages column is JSONB. Fast, transactional, queryable. Recommended for most production agents.The resume logic must handle one non-obvious case: the agent may have been mid-tool-execution when it failed. The message history will contain a tool_use block from the assistant but no corresponding tool_result in the following user message. The Anthropic API will reject a message sequence that ends with an unanswered tool_use. Fix: on resume, check whether the last assistant message contains unmatched tool calls, and inject error tool_result blocks before continuing:
function repairIncompleteToolCalls(messages: Anthropic.MessageParam[]): Anthropic.MessageParam[] {
const last = messages[messages.length - 1];
if (last?.role !== 'assistant') return messages;
const content = last.content as Anthropic.ContentBlock[];
const pendingToolCalls = content.filter(
(b): b is Anthropic.ToolUseBlock => b.type === 'tool_use'
);
if (pendingToolCalls.length === 0) return messages;
// Check if the next message provides tool results
// (shouldn't exist yet if we're repairing - this handles edge cases)
const repairMessage: Anthropic.MessageParam = {
role: 'user',
content: pendingToolCalls.map((call) => ({
type: 'tool_result' as const,
tool_use_id: call.id,
content: 'Tool execution interrupted - the agent was restarted. Please retry this tool call.',
is_error: true,
})),
};
return [...messages, repairMessage];
}
Checkpointing adds storage writes, read latency on startup, and code complexity. It pays off when: the task runs for more than 60 seconds or 10 tool-call iterations, the cost of a full restart exceeds the cost of a checkpoint (true for any agent that consumes expensive API calls or processes large datasets), or the task is user-facing and restart-from-zero produces a noticeably degraded experience. For short agents (under 5 iterations, under 10 seconds), skip checkpointing - the overhead outweighs the recovery benefit. For those tasks, focus instead on error handling within the loop. See agent error recovery strategies for the retry and fallback patterns that complement checkpointing. For tracking which checkpoint was active during a given run, see agent tracing and observability.