AI Agents
AI agents can cost 10-50x more than a single LLM call for the same task.
An agent that runs 15 iterations with large tool results accumulates 50,000+ tokens of context by the final iteration. At those token counts, a 20-iteration run on claude-sonnet-5 costs roughly $0.15, $0.50 per task. At 10,000 agent runs per day, that is $1,500, $5,000 daily before accounting for tool API costs. Agent cost management is not a nice-to-have - it is the difference between a feature that is economically viable and one that is not. Most of the cost reduction comes from four levers: model tiering, token budget enforcement, prompt caching, and tool result compression.
Not every step in an agent loop requires the same model. A capable model (Sonnet or Opus) is needed for complex reasoning and planning. A fast, cheap model (Haiku) is sufficient for tool result parsing, simple classification, content truncation, and sub-task routing. Tiering models across agent roles can cut total run cost by 60 to 80% with minimal quality impact:
// Model costs (per million tokens, approximate 2026 pricing)
const MODEL_COSTS = {
'claude-opus-5': { input: 15.00, output: 75.00 },
'claude-sonnet-5': { input: 3.00, output: 15.00 },
'claude-haiku-4-5-20251001': { input: 0.80, output: 4.00 },
} as const;
// Example: research agent with tiered models
async function tieredResearchAgent(goal: string): Promise {
// Planning: use capable model - decisions here affect all downstream steps
const plan = await runSingleCall('claude-sonnet-5', PLANNING_PROMPT, goal);
// Subtask routing: use cheap model - it's a simple classification
const subtasks = await runSingleCall(
'claude-haiku-4-5-20251001',
'Parse the plan into a JSON array of subtask strings. Return only the JSON array.',
plan
);
// Research execution: use capable model - quality matters here
const researchResults = await Promise.all(
JSON.parse(subtasks).map((task: string) =>
runAgentLoop('claude-sonnet-5', task, researchTools)
)
);
// Synthesis: use capable model - final output quality
return runSingleCall('claude-sonnet-5', SYNTHESIS_PROMPT, researchResults.join('
'));
}
The rule: use the cheapest model that produces acceptable quality for each specific sub-task. Run internal evaluations comparing Haiku and Sonnet outputs for your specific sub-tasks - the quality gap is often smaller than expected for structured and classification tasks.
Without budget enforcement, a misbehaving agent can run 50 iterations, accumulate 200,000 tokens of context, and generate a $5 run for what should have been a $0.10 task. Token budgets enforce a ceiling on per-run spend before it occurs:
class TokenBudgetedLoop {
private totalInputTokens = 0;
private totalOutputTokens = 0;
private readonly maxInputTokens: number;
private readonly maxOutputTokens: number;
constructor(maxInputTokens = 50_000, maxOutputTokens = 10_000) {
this.maxInputTokens = maxInputTokens;
this.maxOutputTokens = maxOutputTokens;
}
get estimatedCostUsd(): number {
return (this.totalInputTokens / 1_000_000 * 3.00) +
(this.totalOutputTokens / 1_000_000 * 15.00); // Sonnet pricing
}
async call(
client: Anthropic,
params: Anthropic.MessageCreateParamsNonStreaming
): Promise {
// Pre-check: estimate if this call will blow the budget
const estimatedInputTokens = estimateTokens(params.messages);
if (this.totalInputTokens + estimatedInputTokens > this.maxInputTokens) {
throw new Error(
`Token budget exceeded: ${this.totalInputTokens + estimatedInputTokens} > ${this.maxInputTokens} input tokens`
);
}
const response = await client.messages.create(params);
this.totalInputTokens += response.usage.input_tokens;
this.totalOutputTokens += response.usage.output_tokens;
// Post-check
if (this.totalOutputTokens > this.maxOutputTokens) {
throw new Error(`Output token budget exceeded: ${this.totalOutputTokens} tokens used`);
}
return response;
}
}
// Usage
const budget = new TokenBudgetedLoop(40_000, 8_000);
try {
const result = await runAgentLoopWithBudget(goal, tools, budget);
} catch (err) {
if ((err as Error).message.includes('budget exceeded')) {
// Return partial result or trigger fallback
return partialResultOrFallback(goal);
}
throw err;
}
If your agent uses a long system prompt (tool descriptions, knowledge base, persona) that is the same across many runs, prompt caching can reduce costs by 90% on cached tokens. The cache is populated on the first call and reused for subsequent calls within the TTL (5 minutes standard, 1 hour extended):
import Anthropic from '@anthropic-ai/sdk';
const client = new Anthropic();
// Long system prompt with cache control
const cachedSystemPrompt: Anthropic.TextBlockParam = {
type: 'text',
text: LONG_SYSTEM_PROMPT_WITH_KNOWLEDGE_BASE, // 2,000+ tokens
cache_control: { type: 'ephemeral' }, // Mark for caching
};
async function callWithCaching(userMessage: string): Promise {
const response = await client.messages.create({
model: 'claude-sonnet-5',
max_tokens: 2048,
system: [cachedSystemPrompt], // Array form with cache_control
messages: [{ role: 'user', content: userMessage }],
});
// Check cache usage
const usage = response.usage as any;
console.log('Cache read tokens:', usage.cache_read_input_tokens ?? 0);
console.log('Cache creation tokens:', usage.cache_creation_input_tokens ?? 0);
return response.content[0].type === 'text' ? response.content[0].text : ', ';
}
At 2026 pricing, prompt caching charges 10% of the standard input price for cache reads. A 4,000-token system prompt on Sonnet costs $0.012 without caching. With caching after the first call, it costs $0.0012 per read. At 1,000 agent runs per day, that is $12/day vs $1.20/day for just the system prompt - a $3,900/year saving from one annotation.
Every tool result appended to the message history persists for all subsequent iterations. A 10,000-character web page fetched in iteration 1 is re-sent as context in iterations 2 to 15. Compressing tool results before they enter the history is among the most effective cost reductions available:
async function compressToolResult(
toolName: string,
rawResult: string,
goal: string,
maxChars = 1500
): Promise {
if (rawResult.length <= maxChars) return rawResult; // No compression needed
// Use Haiku to extract only the goal-relevant portion
const compressionResponse = await client.messages.create({
model: 'claude-haiku-4-5-20251001',
max_tokens: 512,
messages: [{
role: 'user',
content: `Extract the portions of this ${toolName} result that are relevant to: "${goal}"
Limit your extraction to ${maxChars} characters.
Result:
${rawResult.slice(0, 20_000)}`,
}],
});
return compressionResponse.content[0].type === 'text'
? compressionResponse.content[0].text
: rawResult.slice(0, maxChars);
}
Compressing a 15,000-character page to 1,500 goal-relevant characters costs approximately $0.0008 in Haiku calls but saves approximately $0.03, $0.15 in accumulated context costs across 10 to 15 subsequent iterations on Sonnet. The compression is almost always profitable after 3+ iterations.
Track cost per run, per agent type, and per user. Set budget alerts at 80% of daily ceiling - not 100%, because you need time to investigate before hitting the limit:
// Per-run cost tracking
async function runWithCostTracking(
goal: string,
tools: Anthropic.Tool[],
userId: string
): Promise<{ output: string; costUsd: number }> {
let totalInputTokens = 0;
let totalOutputTokens = 0;
const result = await runAgentLoopTracked(goal, tools, (response) => {
totalInputTokens += response.usage.input_tokens;
totalOutputTokens += response.usage.output_tokens;
});
const costUsd =
(totalInputTokens / 1_000_000 * 3.00) +
(totalOutputTokens / 1_000_000 * 15.00);
// Record to your analytics backend
await recordAgentCost({ userId, costUsd, tokens: { input: totalInputTokens, output: totalOutputTokens } });
return { output: result, costUsd };
}
The four levers above - model tiering, token budgets, caching, result compression - occasionally conflict with output quality. The right trade-off depends on the consequences of a low-quality output. For an agent that generates a draft that a human reviews, aggressive cost reduction is appropriate. For an agent that takes autonomous actions (sends emails, writes to databases, makes API calls), prioritise quality and reserve cost optimisations for the parts of the run where they have the least impact on decision quality.
Cost tracking pairs directly with agent tracing and observability - the span structure captures per-iteration token counts that feed cost dashboards. For reducing cost by right-sizing the structured output approach, see tool use vs structured output.