AI & Development

LLM Token Economics: Understanding the Real Cost of AI in Production

Understanding how LLM pricing works - per token, input vs output, caching, batch - is essential for building AI products with sustainable unit economics.

Token pricing is the fundamental economic model of LLM APIs. You pay for every token you send (input) and every token the model generates (output). The price per token varies by model, provider, and in some cases by the type of input (cached vs. uncached). Understanding these pricing structures - and designing your application to work favorably within them - is a fundamental product decision, not an afterthought.

Input vs. output token pricing

Output tokens are consistently more expensive than input tokens across all major providers - typically 3x to 5x more per token. This asymmetry has important design implications. Tasks that require long outputs - generating complete documents, writing extensive code, producing detailed analysis - cost more per call than tasks that require extracting a short answer from a long input. Designing your prompts to minimize output tokens, where quality allows, is meaningful cost optimization.

Concrete implications: asking the model to "write a comprehensive guide to X" produces far more output tokens than "write a concise explanation of X in under 200 words." For many use cases, the concise version is actually more useful. The disciplined use of output length constraints ("limit your response to N sentences," "output only the JSON, no explanation") reduces output token cost and often improves response quality simultaneously.

Model tier pricing

The major providers offer several model tiers at different price points. Frontier models (GPT-4o, Claude Opus, Gemini Ultra) are the most capable and the most expensive - typically $5-$25 per million input tokens and $15-$75 per million output tokens. Mid-tier models (GPT-4o-mini, Claude Sonnet, Gemini Flash) offer 80-95% of frontier capability on most tasks at 10-20% of the price. Small/fast models (Claude Haiku, Gemini Flash 8B) are extremely cheap - fractions of a cent per 1,000 tokens - and still competitive on simple tasks.

The single highest-impact cost optimization for most applications is using the smallest model that achieves acceptable quality on each task. This requires evaluating models against your evaluation set rather than assuming frontier model quality is required everywhere. For many tasks in production applications, a mid-tier or small model achieves acceptable quality at dramatically lower cost.

Prompt caching economics

Prompt caching allows providers to charge less for re-processing a prefix that appeared in a recent call. The economics are substantial: cache write tokens (the first call that populates the cache) are priced at 25% higher than uncached input, but cache read tokens (subsequent calls that hit the cache) are priced at 10-25% of uncached input. For a long system prompt repeated on every call, the amortized cost after the first call drops by 75-90%.

Cache effectiveness depends on the cache TTL (typically 5 minutes on Anthropic) and the volume of calls. If you make 1 call per 10 minutes, the cache expires between calls and you always pay the full uncached price. If you make 100 calls per minute, the cache stays warm and you pay the cached price on 99 out of every 100 calls. High-volume, stateless applications benefit most from caching; low-volume or stateful applications benefit less.

Batch API pricing

Batch APIs - where you submit a large number of requests to be processed asynchronously - are priced at 50% of synchronous pricing across most providers. For tasks that do not require real-time responses - nightly content generation, processing a backlog of documents, evaluating a test set - batch APIs halve the cost without any quality trade-off. The constraint is latency: batch jobs complete in minutes to hours, not milliseconds. For latency-tolerant workloads, the 50% discount makes batch the default choice.

Unit economics modeling

Before scaling an AI feature, model its unit economics. The key metric is AI cost per unit of user value delivered - per message handled, per document processed, per task completed. If the AI cost per unit does not allow for a profitable business at your target price point, the feature's economic model is broken regardless of how good the AI quality is. Building cost awareness into product decisions from the start - rather than discovering the economics problem after scaling - is how AI products end up with sustainable unit economics.