AI & Development

Prompt Caching Explained: How to Reduce LLM Costs by Up to 90%

Prompt caching lets providers reuse the processed representation of repeated prompt prefixes. Here is how it works, how to design for it, and how much it

Every LLM API call involves the model processing all input tokens to compute the key-value representations it needs for attention. For applications where a large portion of the input is the same across many calls - a long system prompt, a reference document, a set of few-shot examples - this repeated processing is wasteful. Prompt caching addresses this by allowing the provider to store and reuse the processed representation of repeated prompt prefixes, so subsequent calls with the same prefix pay a dramatically lower price for those tokens.

How caching works at the provider level

When a caching-enabled API call is made, the provider computes the key-value cache for the full input and stores it server-side with a time-to-live (typically 5 minutes on Anthropic, varying by provider). If a subsequent call arrives within the TTL with the same prefix, the provider retrieves the cached KV state and uses it directly rather than recomputing it. This eliminates the processing cost for the cached tokens on the subsequent call.

The pricing model reflects this: cache write tokens (the first call that populates the cache) are priced slightly higher than uncached input - typically 25% more. Cache read tokens (subsequent calls that hit the cache) are priced much lower - typically 10-25% of uncached input token price. At any volume above a handful of calls per 5 minutes, the math strongly favors caching for long, stable prompt prefixes.

Designing prompts for cache efficiency

The cache is keyed on the prefix of the prompt - the initial, unchanging portion. Everything after the first varying token breaks the cache. This means the order of content in your prompt critically affects caching efficiency: static content (system prompt, reference documents, few-shot examples) should come first, and dynamic content (current user message, per-call context variables) should come last.

A common mistake is inserting dynamic content - timestamps, user names, session IDs - into the system prompt early in the conversation structure. Each unique value creates a unique prompt that cannot hit any cached prefix. Moving these dynamic values to the end of the system prompt, or injecting them only in the user-turn messages, restores caching for the static portion of the system prompt.

Cache hit rates and their impact

The economic impact of prompt caching scales with cache hit rate - the fraction of calls that hit the cache rather than triggering a cache write. For a high-volume application making 1,000 calls per hour with a 5-minute cache TTL, most calls will hit the cache; each 5-minute window has many calls that reuse the same prefix. For a low-volume application making 10 calls per hour, calls are often spaced more than 5 minutes apart, and the cache expires between calls - no benefit is realized.

The relationship is: prompt caching delivers the most value to high-volume applications with long, stable system prompts. It delivers less value to low-volume applications or applications with highly variable system prompts. Evaluating whether caching applies to your specific traffic pattern before investing in cache-friendly prompt restructuring is worthwhile.

Anthropic's extended cache TTL

Anthropic offers an extended cache TTL option for the first 1,024 tokens cached in a session - these are kept alive for the lifetime of the session rather than 5 minutes. For long-running conversations where the system prompt is stable but the session lasts hours, this prevents cache expiry in the middle of a conversation. Extended cache TTL is automatically applied to the longest eligible prefix in a cached call, with no configuration required beyond enabling caching.

When caching does not help

Caching provides no benefit when the cache TTL expires between calls (low-volume applications), when system prompts are highly dynamic (per-user or per-request system prompts that vary significantly), or when the application structure puts dynamic content early in the prompt. For these cases, other cost optimizations - model tier selection, output length control, batch API - are more impactful than caching. Measure your cache hit rate after enabling caching to verify that it is delivering the expected savings rather than assuming it is.