AI & Development

LLM Temperature and Sampling: What the Parameters Actually Do

Temperature, top-p, and top-k control how LLMs sample from probability distributions. Learn what these parameters mean, how they interact, and what values

When a language model generates text, it does not simply pick the most likely next word at each step. Instead, it produces a probability distribution over all possible next tokens and then samples from that distribution. The sampling parameters - temperature, top-p, top-k, frequency penalty, and presence penalty - control how that sampling happens. Understanding these parameters is the difference between using default settings and deliberately tuning your model's behavior for your use case.

Temperature

Temperature is the most important sampling parameter. It controls the "sharpness" of the probability distribution before sampling. At temperature 0, the model always picks the most probable token - the output is deterministic and reproducible. At temperature 1, the model samples from the unmodified probability distribution. At temperature values above 1, the distribution is flattened, making lower-probability tokens more likely to be selected.

In practice: temperature 0 is for tasks where you want maximum consistency and factual accuracy - data extraction, classification, code generation with tests. Temperature 0.3-0.7 is for most conversational and instructional applications - good quality with some natural variation. Temperature 0.8-1.0 is for creative tasks where diversity of outputs is desirable - brainstorming, creative writing, generating alternatives.

Values above 1 are rarely useful in production applications. Very high temperatures produce incoherent text as low-probability tokens become as likely as high-probability ones. If you find yourself wanting very high temperature outputs, the more likely explanation is that the prompt needs better specification, not more randomness.

Top-p (nucleus sampling)

Top-p, also called nucleus sampling, restricts token selection to the smallest set of tokens whose cumulative probability exceeds a threshold p. At top-p = 0.9, the model only considers tokens that together account for 90% of the probability mass. This has a similar effect to temperature in reducing diversity but works differently: rather than scaling the entire distribution, top-p cuts off the tail of low-probability tokens entirely.

Top-p and temperature are often used together. Most providers recommend either using temperature or top-p, not both simultaneously, since they interact in ways that can produce unexpected results. The common practical guidance: set top-p to 1 when using temperature as the control, or set temperature to 1 when using top-p. Choose one as your primary control.

Frequency penalty and presence penalty

Frequency penalty reduces the probability of tokens that have already appeared in the output, weighted by how many times they have appeared. This discourages repetition and is useful for tasks where the model tends to get stuck repeating phrases or ideas. Presence penalty is simpler: it reduces the probability of any token that has appeared at all, regardless of frequency, encouraging the model to introduce new concepts.

Both penalties are useful for applications where output diversity is important - creative writing, generating lists of alternatives, brainstorming. For factual question answering or structured extraction, these penalties should typically be left at 0 (no penalty), since factual accuracy may require repeating specific terms.

Practical settings by task type

For factual question answering and data extraction: temperature 0, top-p 1, frequency penalty 0. This maximizes consistency and accuracy. For general-purpose conversational assistance: temperature 0.7, top-p 1, frequency penalty 0.1. This produces natural-sounding responses with appropriate variation. For creative writing or ideation: temperature 0.9, top-p 0.95, frequency penalty 0.3. This encourages diverse, creative outputs while still producing coherent text.

These are starting points, not universal prescriptions. The right settings depend on your specific task, model, and content. A/B testing different temperature settings on your evaluation set with your specific prompts and model is the only reliable way to determine what works best for your application.

The interaction with model capabilities

More capable models are generally more robust to high temperatures - they maintain coherence at higher randomness levels than less capable models. A temperature setting that produces good creative outputs from a frontier model may produce incoherent outputs from a smaller model on the same task. When switching models, re-evaluating your sampling parameters rather than carrying over the same settings is good practice.