AI Agents

How to Rate Limit AI Agents in Production (Claude API)

Keep AI agents inside Claude API rate limits and your own budgets: 429 handling, retry-after, cache-aware ITPM, per-user token budgets and queues.

A chat feature makes one API call per message. An agent makes one per loop iteration, and a coding agent can run 30 iterations for one task. Let a power user queue 20 runs and your input token budget is gone in a minute, leaving every other user with 429 errors mid-run. Rate limiting AI agents happens at two layers: respecting the provider's limits, and enforcing your own so one user cannot starve the rest.

Agent rate limiting
Agent rate limiting is the practice of controlling how fast AI agents consume model API capacity, combining respect for the provider's request and token limits with application-level token budgets, concurrency caps and queues, so that each user and each agent gets a fair share and overload turns into delay rather than failure.

How to rate limit an AI agent: the direct answer

Rate limit agents on tokens, not requests, and at the run level, not only the call level. Read the rate limit headers on every response, let the SDK retry 429s, cap concurrent runs per user, charge each user the tokens every call actually used, and push runs through a queue so excess demand waits instead of failing mid-task.

Know the upstream limits

The Claude Messages API enforces rate limits per organization and per model class, measured in three units: requests per minute (RPM), input tokens per minute (ITPM) and output tokens per minute (OTPM). A few properties change how you should design around them:

  • Token bucket, not a fixed window. Capacity refills continuously. A limit of 60 RPM can behave like 1 request per second, so a burst of 30 parallel calls can trip it even when your per-minute total is fine.
  • Cache-aware ITPM. For most current models, cache_read_input_tokens do not count toward ITPM; only input_tokens and cache_creation_input_tokens do. Agents resend a long, stable prefix (system prompt, tools, early history) on every iteration, so prompt caching can multiply effective throughput several times over. See prompt caching explained.
  • Separate buckets per model. Claude Opus 5, Claude Sonnet 5 and Claude Haiku 4.5 each have their own limits, so a cheap classifier on Haiku does not eat into your Sonnet capacity.
  • Rate limit 429s carry retry-after, the number of seconds to wait. The exception is the monthly spend cap: that 429 has no retry-after, and error.details.error_code is enforced_spend_limit_reached. Retrying that one is pointless until the cap resets or rises.
  • 529 is not a quota problem. overloaded_error means the service is temporarily saturated. Back off, but do not count it against a user.

Read your headroom from response headers

Every response includes headers such as anthropic-ratelimit-requests-remaining, anthropic-ratelimit-input-tokens-remaining, anthropic-ratelimit-output-tokens-remaining and matching -limit and -reset headers (the reset values are RFC 3339 timestamps). The TypeScript SDK exposes them through .withResponse(), which lets a scheduler slow down before the API starts refusing:

import Anthropic from '@anthropic-ai/sdk';
import { metrics } from './telemetry.js'; // your StatsD or OpenTelemetry client

// The SDK retries 408, 409, 429 and 5xx responses with backoff. Default is 2 retries.
const client = new Anthropic({ maxRetries: 4 });

export async function createWithHeadroom(params: Anthropic.MessageCreateParamsNonStreaming) {
  const { data, response } = await client.messages.create(params).withResponse();

  const inputLeft = Number(response.headers.get('anthropic-ratelimit-input-tokens-remaining'));
  const inputLimit = Number(response.headers.get('anthropic-ratelimit-input-tokens-limit'));
  const requestsLeft = Number(response.headers.get('anthropic-ratelimit-requests-remaining'));

  // Publish headroom so the queue can throttle before the API says no.
  metrics.gauge('claude.itpm_headroom_ratio', inputLeft / inputLimit);
  metrics.gauge('claude.requests_remaining', requestsLeft);
  return data;
}

A headroom ratio that sits below 0.2 for several minutes is your signal to lower worker concurrency, improve caching, or request a higher tier, well before users see errors.

Per-user token budgets and concurrency caps

Upstream limits protect the provider. Your own limits protect your other users and your bill. Two controls cover most cases: a daily token budget per user, charged from the usage object on every response, and a cap on how many runs one user can have in flight. Counting requests does not work for agents, because one run can be 3 calls or 60.

import Anthropic from '@anthropic-ai/sdk';

const redis = new Redis(process.env.REDIS_URL!);
const DAILY_BUDGET = 2_000_000;   // weighted tokens per user per day
const MAX_CONCURRENT_RUNS = 2;    // agent runs in flight per user

export class BudgetExceededError extends Error {}

const today = () => new Date().toISOString().slice(0, 10);

export async function assertBudget(userId: string): Promise<void> {
  const used = Number((await redis.get('tokens:' + userId + ':' + today())) ?? 0);
  if (used >= DAILY_BUDGET) throw new BudgetExceededError('Daily agent budget used for ' + userId);
}

// Charge what each call actually used, weighted roughly by price.
export async function chargeUsage(userId: string, usage: Anthropic.Usage): Promise<void> {
  const weighted =
    usage.input_tokens +
    (usage.cache_creation_input_tokens ?? 0) * 1.25 +
    (usage.cache_read_input_tokens ?? 0) * 0.1 +
    usage.output_tokens * 5;
  const key = 'tokens:' + userId + ':' + today();
  await redis.multi().incrby(key, Math.ceil(weighted)).expire(key, 172_800).exec();
}

export async function acquireRunSlot(userId: string): Promise<() => Promise<void>> {
  const key = 'runs:' + userId;
  const running = await redis.incr(key);
  await redis.expire(key, 900); // self-heals if a worker dies holding a slot
  if (running > MAX_CONCURRENT_RUNS) {
    await redis.decr(key);
    throw new BudgetExceededError('Too many concurrent agent runs for ' + userId);
  }
  return async () => { await redis.decr(key); };
}

The weights mirror relative prices: cache writes cost 1.25 times base input, cache reads 0.1 times, and output tokens cost 5 times input on Claude Sonnet 5. Call assertBudget before every model call inside the loop, not just at the start of a run, so a run that crosses the line stops at the next iteration instead of finishing on someone else's capacity.

Designing rate limits for an API?

The skills library covers token buckets, sliding windows, quota design and 429 handling in depth.

Read the rate limiting skill

Queue runs instead of rejecting them

Budgets decide who may run. A queue decides when. Putting every agent run on a job queue gives you three things at once: a global concurrency limit, a start-rate limit, and retries with backoff when a run fails on a 429 that outlived the SDK's own retries. With BullMQ on Redis:

import { Queue, Worker, UnrecoverableError } from 'bullmq';

const connection = { host: process.env.REDIS_HOST!, port: 6379 };
export const agentRuns = new Queue('agent-runs', { connection });

export function enqueueRun(userId: string, task: string) {
  return agentRuns.add('run', { userId, task }, {
    attempts: 5,
    backoff: { type: 'exponential', delay: 15_000 }, // 15 s, 30 s, 60 s, 120 s
    removeOnComplete: 1000,
  });
}

new Worker('agent-runs', async (job) => {
  try {
    return await runAgentForUser(job.data.userId, job.data.task);
  } catch (err) {
    // Retrying cannot fix an exhausted budget, so fail the job for good.
    if (err instanceof BudgetExceededError) throw new UnrecoverableError(err.message);
    throw err; // rate limit errors that outlived SDK retries: BullMQ backs off and retries
  }
}, {
  connection,
  concurrency: 8,                         // runs in flight per worker process
  limiter: { max: 60, duration: 60_000 }, // runs started per minute across workers
});

Note what the limiter counts: runs started, not API calls. If an average run makes 12 calls, 60 runs per minute means roughly 720 calls per minute, so size it from your own traces. Pair queue retries with agent checkpointing, otherwise a retried run pays again for every iteration that had already succeeded.

Failure modes

  • Retry storms. Forty workers hit a 429 at the same moment and retry at the same moment. Exponential backoff with jitter spreads them out. The SDK adds jitter to its own retries; hand-written loops and queue retries usually need it added explicitly, as a small random delay on top of the backoff.
  • Budget races. Two runs for the same user both pass assertBudget just under the line and both finish. The overshoot is bounded by your per-call size, which is usually acceptable. If it is not, reserve an estimate before the call and settle it afterwards.
  • Counting only output tokens. Agents are input-heavy. A run that resends a 30,000 token history 20 times has used 600,000 input tokens and perhaps 8,000 output tokens.
  • Treating the spend cap like a rate limit. Retrying enforced_spend_limit_reached burns attempts for nothing. Detect it and page a human.
  • No acceleration ramp. A sudden jump in traffic can trigger acceleration limits even under your nominal quota. Ramp new features up gradually.

Rate limits and cost controls overlap heavily. For model selection per role, caching strategy and budget alerts, see agent cost management; for how the loop should report failed tool calls rather than failed API calls, see tool result error handling.

FAQ

What are the Claude API rate limits measured in?

The Messages API limits requests per minute (RPM), input tokens per minute (ITPM) and output tokens per minute (OTPM), separately for each model class, at the organization level. Limits use a token bucket, so capacity refills continuously rather than resetting at the top of each minute.

Do cached tokens count toward Claude rate limits?

For most current Claude models, tokens read from the prompt cache do not count toward the input tokens per minute limit; only uncached input tokens and tokens written to the cache do. Caching a long system prompt and tool list therefore raises your effective throughput as well as lowering cost.

How should an agent handle a 429 error?

Let the SDK retry first, since it retries 429 responses with backoff. If the error survives those retries, do not spin in the loop: save the agent state and re-queue the run with a delay based on the retry-after header. A 429 without retry-after may mean the monthly spend cap was reached, which retrying cannot fix.

Should I rate limit users by requests or by tokens?

By tokens, or better by weighted cost. Agent runs vary by two orders of magnitude in size, so a request counter lets one heavy user consume most of your capacity. Charge each user the actual usage reported by every API response, and cap concurrent runs per user as well.