AI & Development
The context window is the fundamental constraint of every LLM application. Understand what it is, how token limits affect your design, and practical
A context window is the maximum amount of text a language model can process in a single inference call. Everything that the model can "see" when generating a response - the system prompt, conversation history, retrieved documents, tool outputs, and the current user message - must fit within this window. Anything outside the window is invisible to the model, as if it does not exist.
Context windows are measured in tokens. A token is not the same as a word or a character - it is a unit of the model's tokenization scheme. In English, a token is roughly 3-4 characters, so 100 tokens is roughly 75 words. Code is denser; structured data like JSON is denser still because brackets, commas, and whitespace each consume tokens.
The first GPT-3 models had a context window of 2,048 tokens - about 1,500 words. This severely constrained what you could build. Today's frontier models offer context windows of 128,000 tokens (roughly a 100,000-word novel) and in some configurations up to 1 million tokens. This is a 500x increase in capacity in a few years, and it has fundamentally changed what LLM applications can do.
But larger context windows do not eliminate the context constraint - they shift it. Applications that once had to work around 4,000-token limits now have new headroom, but they also have new costs: larger context windows cost more per call (pricing is typically per token for both input and output), and the "lost in the middle" problem - where models attend poorly to information buried in the middle of a long context - means that raw capacity does not translate directly into reliable performance on long-context tasks.
In a typical production LLM application, the context window is consumed by four things: the system prompt, retrieved documents (for RAG applications), conversation history, and the current message. System prompts that include detailed instructions, examples, and tool definitions can easily consume 2,000-5,000 tokens before the user has said anything. RAG-retrieved documents add thousands more. In a long conversation, history accumulates.
The practical consequence is that context window management - deciding what to include, what to truncate, and what to summarize - is a core engineering concern, not a detail. Applications that do not manage context actively will eventually hit limits, and when they do, the failure mode is often silent: the model silently ignores content that was truncated, producing responses that appear coherent but are missing critical context.
The most common strategy is conversation summarization: rather than including all conversation history verbatim, you periodically summarize older turns and replace them with the summary. The model retains the key facts from earlier in the conversation without the raw token cost. This works well for many conversational applications, but it loses exact quotes and specific details that may be important for some use cases.
Selective history is a simpler approach for some applications: include only the most recent N turns of conversation, discarding older ones. This is appropriate when each exchange is relatively self-contained and does not depend heavily on earlier context. Combined with a good system prompt that provides persistent context, selective history keeps token costs predictable.
For RAG applications, aggressive chunk filtering is important: retrieve more candidates than you will actually use, then select only the most relevant ones for the context. Including every potentially relevant chunk without filtering is a common source of context bloat and "lost in the middle" failures.
Most leading LLM providers now support prompt caching - a feature where the token processing cost of a repeated prefix (such as a long system prompt or a large document that appears in every call) is amortized after the first call. The cached portion is stored server-side and reused for subsequent calls with the same prefix. For applications with long system prompts that are static across many calls, prompt caching can reduce input token costs by 70-90%. Designing prompts to maximize cache hit rates is a meaningful optimization at scale.
Context window capacity is shared between input and output tokens, but output tokens are typically more expensive per token than input tokens and are also strictly bounded by the model's maximum output token limit. This asymmetry matters for application design: tasks that require very long outputs - generating a full article, a long piece of code, or a detailed analysis - consume output capacity that limits how much input context you can effectively use. Splitting long-output tasks into multiple shorter calls is often more reliable than trying to generate everything in a single long call.