AI & Development
Model distillation trains a smaller model to mimic a larger teacher model, achieving similar performance at lower cost.
Model distillation is a training technique where a small "student" model is trained to replicate the outputs of a larger "teacher" model. The fundamental insight is that a large model, despite being expensive to run, has learned rich representations of the problem - and much of that knowledge can be transferred to a smaller, faster, cheaper model through a training process that uses the large model's outputs as supervision signals rather than raw labels.
The result is a student model that is significantly smaller than the teacher but often approaches the teacher's performance on the target task - without requiring the computational resources to run the teacher at inference time.
When a large model generates a probability distribution over possible next tokens or output classes, it encodes information beyond just the most likely answer. The distribution reflects the model's uncertainty, the relationships between similar classes, and the structure of the output space. Training a student model to match these distributions - rather than just training on one-hot labels - transfers this "soft" knowledge to the student.
This is distinct from training the student on the teacher's final predictions (greedy or sampled outputs). Using the full probability distribution as a training signal, often called "soft targets," is what makes distillation more effective than simply curating a dataset of the teacher's outputs and fine-tuning the student on it. The soft targets carry more information than discrete predictions.
For large language models, distillation takes several forms. Response distillation - the most common and simplest form - generates a dataset of prompt-response pairs using the teacher model and fine-tunes the student model on those pairs. This is essentially fine-tuning on teacher-generated data, and it transfers the teacher's style and knowledge at the cost of discarding the soft target information. It is practical because generating a large dataset from a capable teacher model is straightforward with an API.
Token-level distillation matches the student's next-token probability distribution to the teacher's at each step in the generation. This requires access to the teacher model's logits (probability scores), which is possible when you control the teacher model but not when using a closed API where only the final output is returned. Token-level distillation is more powerful than response distillation but requires infrastructure to capture and use teacher logits.
Distillation is most valuable when you have a specific, well-defined task that a large model solves well, and you need to run that capability at scale at lower cost. A customer intent classification task that a large model handles with 95% accuracy can often be distilled into a small model that achieves 93% accuracy at one-tenth the inference cost. For high-volume tasks, this trade-off is highly favorable.
Distillation is less suitable for tasks that require the full generalization capability of a large model - tasks with many diverse sub-types, tasks where the distribution shifts frequently, or tasks where performance differences between 90% and 95% accuracy are consequential. In these cases, the distilled model's reduced capacity shows up as fragility on out-of-distribution inputs.
The distillation data quality is the primary determinant of the student model's quality. If the teacher model's outputs are good, the student model will learn good behaviors. If the teacher is prompted poorly or the tasks selected for distillation are not representative of production inputs, the student will learn those deficiencies as well. The same data quality practices that apply to fine-tuning apply equally to distillation.
Evaluating the distilled model comprehensively before deployment is critical, especially on edge cases and inputs that are underrepresented in the distillation dataset. Student models often show the same failure modes as the teacher on common inputs but diverge more significantly on unusual ones. Running the student model's outputs through the same evaluation pipeline used for the teacher is the right practice for catching these divergences.