AI Agents

Agent Parallelisation Patterns: Fan-Out, Batch, Parallel Tool Calls

How to parallelise AI agent work with fan-out/fan-in, batch agents, and parallel tool calls - the patterns that collapse multi-step task latency from minutes.

Sequential agent pipelines are slow by design: agent A completes, then agent B runs, then agent C, each waiting for the previous to finish. If each step takes 5 seconds and there are ten steps, the total wall-clock time is 50 seconds - regardless of how fast each individual agent is. The fix is not faster models or lower latency - it is running independent work in parallel. This post covers the three parallelisation patterns that matter in practice: fan-out/fan-in, batch processing, and parallel tool calls within a single agent loop.

Agent fan-out
Agent fan-out is a parallelisation pattern where an orchestrator agent distributes independent sub-tasks to multiple worker agents that run concurrently, then collects and synthesises all worker outputs - reducing total wall-clock time from O(n × task_time) to O(max_task_time + overhead) for n independent tasks.

Pattern 1 - Fan-out / fan-in

Fan-out is the most impactful parallelisation pattern: the orchestrator decomposes a goal into independent sub-tasks and fires all workers simultaneously. The fan-in collects results when all (or enough) workers complete. Total time is bounded by the slowest worker, not the sum of all workers.

import Anthropic from '@anthropic-ai/sdk';

const client = new Anthropic();

interface WorkerTask {
  id: string;
  instruction: string;
  workerType: 'researcher' | 'analyst' | 'writer';
}

interface WorkerResult {
  id: string;
  output: string;
  error?: string;
}

async function runWorker(task: WorkerTask): Promise {
  const systemPrompts: Record = {
    researcher: 'You are a research specialist. Return well-sourced, specific findings.',
    analyst: 'You are a data analyst. Return structured analysis with key metrics.',
    writer: 'You are a writing specialist. Return polished, publication-ready prose.',
  };

  try {
    const response = await client.messages.create({
      model: 'claude-sonnet-5',
      max_tokens: 2048,
      system: systemPrompts[task.workerType],
      messages: [{ role: 'user', content: task.instruction }],
    });

    const text = response.content.find(b => b.type === 'text');
    return {
      id: task.id,
      output: text?.text ?? '',
    };
  } catch (err) {
    return {
      id: task.id,
      output: '',
      error: (err as Error).message,
    };
  }
}

async function fanOut(tasks: WorkerTask[]): Promise {
  // All tasks fire simultaneously - total time ≈ slowest single task
  const settled = await Promise.allSettled(tasks.map(runWorker));

  return settled.map((outcome, i) => {
    if (outcome.status === 'fulfilled') return outcome.value;
    return {
      id: tasks[i].id,
      output: '',
      error: `Worker failed: ${outcome.reason}`,
    };
  });
}

async function fanOutPipeline(goal: string): Promise {
  // Phase 1: orchestrator decomposes goal into parallel tasks
  const planResponse = await client.messages.create({
    model: 'claude-sonnet-5',
    max_tokens: 1024,
    system: 'Decompose the goal into 3-5 independent research tasks that can run in parallel. Return JSON: { "tasks": [{ "id": string, "instruction": string, "workerType": string }] }',
    messages: [{ role: 'user', content: goal }],
    tool_choice: { type: 'auto' },
    tools: [decomposeTool],
  });

  const tasks = extractTasks(planResponse);

  // Phase 2: fan-out - all workers run concurrently
  console.log(`Fanning out to ${tasks.length} workers...`);
  const startTime = Date.now();
  const results = await fanOut(tasks);
  console.log(`Fan-in complete in ${Date.now() - startTime}ms`);

  // Phase 3: synthesise all results
  const synthesisResponse = await client.messages.create({
    model: 'claude-sonnet-5',
    max_tokens: 4096,
    system: 'Synthesise the worker outputs into a coherent response.',
    messages: [{
      role: 'user',
      content: `Goal: ${goal}

Worker outputs:
${results.map(r => `[${r.id}]: ${r.output}`).join('

')}`,
    }],
  });

  return synthesisResponse.content[0].type === 'text'
    ? synthesisResponse.content[0].text : ', ';
}

Pattern 2 - Batch processing

Batch processing applies the same agent logic to many items in parallel, with a concurrency cap to avoid rate limiting. This is the right pattern when you have N items and need to process each independently - analysing 100 documents, classifying 500 support tickets, extracting data from 200 web pages:

async function processBatch(
  items: T[],
  processFn: (item: T, index: number) => Promise,
  options: {
    concurrency: number;      // max parallel workers (typically 5-15)
    retries: number;          // retries per item on failure
    onProgress?: (done: number, total: number) => void;
  }
): Promise> {
  const results: Array<{ item: T; result?: R; error?: string }> = [];
  const queue = [...items.entries()];
  let done = 0;

  async function worker(): Promise {
    while (queue.length > 0) {
      const entry = queue.shift();
      if (!entry) break;
      const [index, item] = entry;

      let lastError: string | undefined;
      for (let attempt = 0; attempt <= options.retries; attempt++) {
        try {
          const result = await processFn(item, index);
          results.push({ item, result });
          lastError = undefined;
          break;
        } catch (err) {
          lastError = (err as Error).message;
          if (attempt < options.retries) {
            // Exponential backoff: 1s, 2s, 4s
            await new Promise(res => setTimeout(res, 1000 * Math.pow(2, attempt)));
          }
        }
      }

      if (lastError) results.push({ item, error: lastError });
      done++;
      options.onProgress?.(done, items.length);
    }
  }

  // Start concurrency-limited workers
  await Promise.all(
    Array.from({ length: options.concurrency }, () => worker())
  );

  return results;
}

// Example: classify 200 support tickets in parallel with 10 concurrent agents
async function classifyTickets(tickets: SupportTicket[]): Promise {
  const results = await processBatch(tickets, async (ticket) => {
    const response = await client.messages.create({
      model: 'claude-haiku-4-5-20251001', // Fast model for classification
      max_tokens: 256,
      system: 'Classify the support ticket. Return JSON: { "category": string, "priority": "high"|"medium"|"low", "sentiment": "positive"|"neutral"|"negative" }',
      messages: [{ role: 'user', content: ticket.body }],
    });
    return JSON.parse(response.content[0].type === 'text' ? response.content[0].text : '{}');
  }, {
    concurrency: 10,
    retries: 2,
    onProgress: (done, total) => console.log(`${done}/${total} tickets classified`),
  });

  return results.filter(r => r.result).map(r => ({ ...r.item...r.result! }));
}

Pattern 3 - Parallel tool calls within a single agent

Within a single agent loop iteration, Claude can request multiple tool calls simultaneously when they are independent. Your tool executor must then run those calls in parallel. Most developers execute tool calls sequentially, missing this built-in parallelism:

async function executeToolsInParallel(
  contentBlocks: Anthropic.ContentBlock[]
): Promise {
  const toolCalls = contentBlocks.filter(
    (b): b is Anthropic.ToolUseBlock => b.type === 'tool_use'
  );

  if (toolCalls.length === 0) return [];

  console.log(`Executing ${toolCalls.length} tool calls in parallel`);

  // All tool calls from one agent iteration run concurrently
  const settled = await Promise.allSettled(
    toolCalls.map(async (call) => {
      const result = await dispatchTool(call.name, call.input);
      return {
        type: 'tool_result' as const,
        tool_use_id: call.id,
        content: JSON.stringify(result),
      };
    })
  );

  return settled.map((outcome, i): Anthropic.ToolResultBlockParam => {
    if (outcome.status === 'fulfilled') return outcome.value;
    return {
      type: 'tool_result',
      tool_use_id: toolCalls[i].id,
      content: `Error: ${(outcome.reason as Error).message}`,
      is_error: true,
    };
  });
}

Rate limiting and concurrency caps

Anthropic's API has per-minute token and request rate limits that vary by tier. Running 50 agents simultaneously on a Tier 1 account will produce rate limit errors. Practical concurrency caps by tier:

  • Tier 1 - cap at 3-5 concurrent agents; use claude-haiku-4-5-20251001 for batch tasks
  • Tier 2 - cap at 10-15 concurrent agents
  • Tier 3+ - cap at 25-50 depending on model; monitor the x-ratelimit-remaining-requests response header and back off when it approaches zero

When a rate limit error (HTTP 429) occurs, exponential backoff with jitter is the correct response - not a hard stop. The batch processor pattern above includes retry logic for this reason.

When NOT to parallelise

Parallelism adds complexity: error handling becomes more nuanced, debugging requires correlating logs across concurrent runs, and cost spikes are instant rather than gradual. Do not parallelise when: tasks are strictly sequential (each depends on the previous), the total task time is under 10 seconds (overhead negates gains), or the concurrency would exhaust your API rate limit tier. For sequential dependencies between workers, the supervisor pattern's dependency resolution handles this correctly. For understanding the underlying single-agent execution model that fan-out extends, see the agentic loop explained.