AI & Development

LLM Inference Optimization: Serving Language Models Fast

Serving LLMs at production scale requires specialized inference infrastructure. Learn the key optimization techniques

Running a language model for inference is not the same as running a traditional web API. LLMs generate text autoregressively - one token at a time, with each token's generation depending on all previous tokens. This sequential generation pattern, combined with the massive memory requirements of large models, creates unique performance characteristics that standard web serving infrastructure is not designed to handle. Teams self-hosting LLMs for production inference need to understand these characteristics to serve efficiently.

The KV cache: the most important optimization

During autoregressive generation, the model computes key-value (KV) tensors for each token in the input to attend over when generating subsequent tokens. Without caching, these computations would be repeated for every new token, making inference quadratically expensive. The KV cache stores these tensors so they only need to be computed once per input token, reducing subsequent token generation to a linear operation.

KV cache size grows with batch size and sequence length, and is one of the primary constraints on how many concurrent requests an LLM server can handle. A 70B parameter model might require 80GB of GPU memory for weights alone, leaving limited space for the KV cache. Managing this trade-off - between the number of concurrent requests and the length of sequences each request can use - is a central challenge in production inference.

PagedAttention and vLLM

PagedAttention is the core innovation in vLLM that makes it the most widely-used open-source LLM inference server. Traditional KV caches allocate a contiguous block of GPU memory for each request's cache - the size of the maximum sequence length, even if the actual sequence is much shorter. This wastes memory and limits how many requests can be served simultaneously.

PagedAttention divides the KV cache into non-contiguous pages, similar to virtual memory in operating systems. Each request only uses as many pages as it needs for its actual sequence length, and pages can be shared between requests with identical prefixes (enabling prompt caching at the inference level). The result is significantly better memory utilization and dramatically higher throughput for the same GPU hardware. vLLM implements PagedAttention with a clean OpenAI-compatible API, making it the standard choice for self-hosted LLM serving.

Continuous batching

Naive batching groups multiple requests together and processes them in a fixed batch. This wastes GPU cycles when some requests in the batch finish early - the entire batch must wait for the longest request to complete before the GPU can accept new work. Continuous batching (also called iteration-level scheduling) inserts new requests into the batch at every generation step, keeping GPU utilization high regardless of varying sequence lengths. This technique alone can increase throughput by 3-10x compared to naive batching on typical production workloads.

Speculative decoding

Speculative decoding uses a small, fast "draft" model to propose multiple tokens ahead of the main model, then uses the main model to verify those tokens in parallel. When the draft model's predictions are correct (which is often, for predictable outputs), the main model can accept multiple tokens per verification step rather than generating one token at a time. This reduces the number of forward passes required and can produce 2-3x speedup on latency for tasks with predictable outputs, at no quality cost.

Speculative decoding is most effective when the draft model's predictions are accurate, which is application-dependent. For structured outputs, code generation, and repetitive tasks, the draft model achieves high accuracy and the speedup is substantial. For highly unpredictable creative generation, the speedup is smaller because the draft predictions are more often rejected.

Quantization for inference

Weight quantization - representing model parameters in 8-bit or 4-bit instead of 16-bit floating point - reduces memory requirements and often increases inference throughput on hardware with integer compute units. INT8 quantization is nearly lossless for most tasks; 4-bit quantization (GPTQ, AWQ, GGUF) shows more quality degradation but enables running much larger models on constrained hardware. For production serving where hardware cost is a primary concern, quantization is often the right trade-off. For applications requiring the highest possible quality, full precision or FP16 is preferred.