AI Agents

Agent Checkpointing: Save and Resume Agent State in Production

How to checkpoint AI agent state so long-running agents survive failures and can resume mid-task - what to persist, storage backend options.

A long-running agent - one that searches dozens of web pages, processes hundreds of records, or orchestrates five sub-tasks over 20 minutes - has a meaningful probability of failure before it finishes. Network timeouts, API rate limits, context window overflow, and infrastructure outages all increase with task duration. Without checkpointing, a failure at iteration 47 of 50 costs you exactly as much as a failure at iteration 1: you restart from zero. Checkpointing converts that total loss into a partial recovery, and for production agents running costly tasks, the difference is significant.

Agent checkpoint
An agent checkpoint is a persisted snapshot of an agent's execution state - its message history, tool results, completed sub-tasks, and any accumulated artifacts - saved at a defined milestone so the agent can resume from that point after a failure rather than restarting from the beginning of the task.

What agent checkpointing is - the direct answer

Agent checkpointing is the practice of serialising an agent's execution state to durable storage at defined points during a run, then loading that snapshot on restart so the agent continues from the saved milestone rather than from the task's beginning. A checkpoint captures the message history (the agent's working memory), any intermediate results, the current goal and sub-task status, and sufficient context for the agent to resume decision-making correctly.

What state to checkpoint

Not all agent state has equal value. The checkpoint should be selective:

  • Message history - the accumulated Anthropic.MessageParam[] array. This is the agent's full working memory and is the most critical thing to preserve. Without it, the resumed agent has no knowledge of what it has already done.
  • Completed sub-task results - if the agent is orchestrating multiple sub-tasks, record which are done and their outputs. The resume logic can skip completed tasks.
  • Accumulated artifacts - files written, records processed, URLs fetched. These allow the resume to avoid duplicate work.
  • Current iteration count - so the resumed agent's max-iteration budget accounts for iterations already spent.
  • Metadata - task ID, user ID, created-at, checkpoint-at timestamps, the original goal. These support querying and debugging.

Do not checkpoint: ephemeral tool state (open file handles, in-progress HTTP requests), model-internal reasoning that is already reflected in the message history, or redundant copies of data already in tool results.

Implementing checkpoints

The checkpoint write should happen after every meaningful milestone - typically after a tool result is appended to the message history, or after a sub-task completes. Checkpointing after every single API call is excessive for most tasks; checkpointing only at the end defeats the purpose. A heuristic: checkpoint every 5-10 iterations, and always checkpoint immediately after completing any irreversible action (a write, a send, a create):

import Anthropic from '@anthropic-ai/sdk';

const client = new Anthropic();
const supabase = createClient(process.env.SUPABASE_URL!, process.env.SUPABASE_KEY!);

interface AgentCheckpoint {
  task_id: string;
  goal: string;
  messages: Anthropic.MessageParam[];
  completed_tasks: string[];
  iteration: number;
  created_at: string;
  updated_at: string;
}

async function saveCheckpoint(taskId: string, state: Omit): Promise {
  const now = new Date().toISOString();
  const { error } = await supabase
    .from('agent_checkpoints')
    .upsert({
      task_id: taskId,
      ...state,
      updated_at: now,
    }, { onConflict: 'task_id' });

  if (error) {
    // Log but don't throw - a checkpoint failure should not abort the agent
    console.error('Checkpoint save failed:', error.message);
  }
}

async function loadCheckpoint(taskId: string): Promise {
  const { data, error } = await supabase
    .from('agent_checkpoints')
    .select('*')
    .eq('task_id', taskId)
    .single();

  if (error || !data) return null;
  return data as AgentCheckpoint;
}

async function runCheckpointedAgent(taskId: string, goal: string): Promise {
  // Attempt to resume from checkpoint
  const checkpoint = await loadCheckpoint(taskId);

  let messages: Anthropic.MessageParam[] = checkpoint
    ? checkpoint.messages
    : [{ role: 'user', content: goal }];

  let iteration = checkpoint ? checkpoint.iteration : 0;
  const completedTasks = checkpoint ? new Set(checkpoint.completed_tasks) : new Set();
  const MAX_ITERATIONS = 30;

  console.log(checkpoint
    ? `Resuming from checkpoint at iteration ${iteration}`
    : 'Starting fresh agent run'
  );

  for (; iteration < MAX_ITERATIONS; iteration++) {
    const response = await client.messages.create({
      model: 'claude-sonnet-5',
      max_tokens: 4096,
      tools: agentTools,
      messages,
    });

    messages.push({ role: 'assistant', content: response.content });

    if (response.stop_reason === 'end_turn') {
      // Clean up checkpoint on success
      await supabase.from('agent_checkpoints').delete().eq('task_id', taskId);
      return extractFinalAnswer(response.content);
    }

    if (response.stop_reason === 'tool_use') {
      const toolResults = await executeTools(response.content, completedTasks);
      messages.push({ role: 'user', content: toolResults });

      // Checkpoint every 5 iterations and after tool execution
      if (iteration % 5 === 0) {
        await saveCheckpoint(taskId, {
          goal,
          messages,
          completed_tasks: Array.from(completedTasks),
          iteration: iteration + 1,
        });
      }
    }
  }

  throw new Error(`Agent did not complete within ${MAX_ITERATIONS} iterations`);
}

Storage backend options

The checkpoint must outlive the process. Three practical options in order of increasing durability and complexity:

  • PostgreSQL / Supabase - Simple upsert on a task_id primary key. The messages column is JSONB. Fast, transactional, queryable. Recommended for most production agents.
  • Redis with persistence enabled - Sub-millisecond reads and writes, ideal for frequent checkpoints (every iteration). Requires AOF or RDB persistence configured; in-memory-only Redis loses the checkpoint on restart.
  • Object storage (S3 / GCS) - Good for very large checkpoints (agents with gigabytes of intermediate artifacts). Higher read/write latency (~50-200ms) makes this unsuitable for every-iteration checkpoints.

Resuming correctly

The resume logic must handle one non-obvious case: the agent may have been mid-tool-execution when it failed. The message history will contain a tool_use block from the assistant but no corresponding tool_result in the following user message. The Anthropic API will reject a message sequence that ends with an unanswered tool_use. Fix: on resume, check whether the last assistant message contains unmatched tool calls, and inject error tool_result blocks before continuing:

function repairIncompleteToolCalls(messages: Anthropic.MessageParam[]): Anthropic.MessageParam[] {
  const last = messages[messages.length - 1];
  if (last?.role !== 'assistant') return messages;

  const content = last.content as Anthropic.ContentBlock[];
  const pendingToolCalls = content.filter(
    (b): b is Anthropic.ToolUseBlock => b.type === 'tool_use'
  );

  if (pendingToolCalls.length === 0) return messages;

  // Check if the next message provides tool results
  // (shouldn't exist yet if we're repairing - this handles edge cases)
  const repairMessage: Anthropic.MessageParam = {
    role: 'user',
    content: pendingToolCalls.map((call) => ({
      type: 'tool_result' as const,
      tool_use_id: call.id,
      content: 'Tool execution interrupted - the agent was restarted. Please retry this tool call.',
      is_error: true,
    })),
  };

  return [...messages, repairMessage];
}

When checkpointing is worth the overhead

Checkpointing adds storage writes, read latency on startup, and code complexity. It pays off when: the task runs for more than 60 seconds or 10 tool-call iterations, the cost of a full restart exceeds the cost of a checkpoint (true for any agent that consumes expensive API calls or processes large datasets), or the task is user-facing and restart-from-zero produces a noticeably degraded experience. For short agents (under 5 iterations, under 10 seconds), skip checkpointing - the overhead outweighs the recovery benefit. For those tasks, focus instead on error handling within the loop. See agent error recovery strategies for the retry and fallback patterns that complement checkpointing. For tracking which checkpoint was active during a given run, see agent tracing and observability.