AI & Development
Few-shot prompting uses examples to shape LLM behavior more reliably than descriptions alone. Learn how to select examples, structure them, and use them
Language models learn from patterns. In their training, they were exposed to countless examples of input-output pairs - questions and answers, prompts and completions, code and documentation. At inference time, you can exploit this pattern-learning capability by providing a small number of examples (a "few shots") of exactly the input-output pattern you want. The model uses these examples to infer and apply the pattern to your actual query.
This is called few-shot prompting, and it is one of the most reliable techniques for getting consistent, correctly formatted outputs from a language model - especially when the desired format is unusual, highly specific, or difficult to describe precisely in natural language.
Natural language descriptions of format requirements are often ambiguous in ways that examples are not. "Write in a formal but approachable tone" is difficult to operationalize consistently. Three examples of responses in that tone, placed in the prompt, give the model a much clearer signal. "Extract the vendor name, invoice date, and total amount" is more precisely communicated by showing an example invoice with the correct extractions than by describing the extraction rules in prose.
Few-shot prompting is particularly valuable for: output format enforcement (JSON structure, Markdown layout, specific response templates), style and tone consistency, classification tasks with non-obvious category definitions, and tasks where the distinction between correct and incorrect outputs is subtle and hard to express as a rule.
The examples you choose strongly influence the output pattern the model applies. Examples should be: representative of the distribution of inputs the model will encounter, unambiguous (each example should clearly demonstrate the intended pattern), and diverse enough to show the model how to handle the range of inputs you expect. Showing three variations of the same type of input does not help the model generalize to other input types.
The number of examples matters up to a point. One example is often enough to establish a format. Three to five examples are sufficient for most tasks. More than ten examples in a prompt typically produces diminishing returns and increases context usage significantly. Quality over quantity applies strongly here - five well-chosen examples outperform twenty mediocre ones.
In fixed few-shot prompting, the same examples appear in every call. In dynamic few-shot selection, the examples are chosen at query time based on the similarity of the current query to a library of labeled examples. The intuition is that examples similar to the current input are more instructive than generic examples.
Dynamic few-shot requires maintaining a library of labeled examples and a retrieval system that returns the most similar ones for each query. For classification and extraction tasks, this approach consistently outperforms fixed examples. The implementation overhead is modest - the same embedding-based retrieval used for RAG applies here - and the quality improvement for diverse input distributions is meaningful.
The order in which examples are presented affects model performance. A consistent finding in research is that the most recent examples (closest to the query in the prompt) have the strongest influence. Placing the most representative or most important example last - immediately before the query - gives it the highest weight. Placing an unusual or edge-case example in this position can bias the model toward that pattern even for unrelated queries.
For classification tasks with multiple classes, interleaving examples from different classes rather than grouping by class prevents the model from developing a recency bias toward the last class shown. For format examples, ordering from simpler to more complex gives the model a gradual build-up of the pattern rather than starting with the most complex case.
Clearly delineating the boundary between the input and the output in each example makes the pattern clearer to the model. XML-style tags (<input>...</input><output>...</output>) or labeled sections (Input:...Output:...) make the structure explicit. Clear delimiters also reduce the risk of the model treating part of the example as the actual query, which can happen with ambiguous formatting. Consistency in the delimiter style across all examples and the actual query position is important for reliable results.