What is distillation?
Distillation is a training method where a small "student" model learns to imitate a larger "teacher" model, so the student delivers part of the teacher's quality in a form that is cheaper to serve. The cost saving comes from size: a student with fewer parameters needs less VRAM and less compute per token than the teacher.
The original idea
Hinton, Vinyals and Dean introduced the technique for neural networks in Distilling the Knowledge in a Neural Network. Their motivation is that predicting with a whole ensemble of models is cumbersome and may be too expensive to deploy, so they compress the ensemble's knowledge into a single model. The key device is soft targets: instead of training only on hard labels, the student trains on the probability distributions the teacher outputs, which carry more information per example than a single correct label. A temperature setting softens those distributions so that the relative probabilities of wrong answers become visible to the student. The paper also describes specialist models that focus on fine-grained distinctions the full models confuse.
Distillation for language models
Language models are distilled in more than one way. The classic approach matches the teacher's output distribution. Another, widely used for reasoning models, is to fine-tune the student on text the teacher generated. The DeepSeek-R1 distills are an example of the second: their model cards state that they are fine-tuned on samples generated by DeepSeek-R1, starting from existing Qwen and Llama base models. The Qwen-based distills are described as fine-tuned with 800k samples curated with DeepSeek-R1. The student is a normal checkpoint with the base model's architecture, so it runs in the same serving software as any other model of that size.
Example models in the catalog
- DeepSeek R1 Distill Qwen 32B: Qwen2.5-32B as the base, with DeepSeek-R1 as the teacher. The card notes the Qwen2.5 base carries an Apache 2.0 license.
- DeepSeek R1 Distill Llama 8B: Llama-3.1-8B as the base; the card says it is originally licensed under the Llama 3.1 license.
- DeepSeek R1 Distill Qwen 7B and DeepSeek R1 Distill Qwen 1.5B: the same recipe at smaller sizes.
- DeepSeek R1 Distill Llama 70B FP8 dynamic: a third-party FP8 quantization of a distill, showing that distilling and quantizing stack.
How it differs from related techniques
- Quantization keeps the model's size in parameters and reduces the bits per weight. Distillation reduces the parameter count. The two combine freely: quantization of a distilled model is common.
- Fine-tuning with LoRA adapts a model to a task with small added weights (LoRA, QLoRA). Distillation can use LoRA as its training method, but its goal is to transfer behavior from a teacher.
- Model merging combines the weights of existing models without training (model merging).
What it means when you pick a GPU
Distillation shifts the choice from "which big model" to "how small a student is acceptable". A distilled model is sized like its base: an 8B student needs the memory of any 8B model and fits cards a 70B teacher cannot. Use the VRAM calculator with the student's parameter count, and add room for the KV cache, which grows with context length. Whether the student is good enough is a question to test on your own prompts. Aquanode rents GPUs by the hour, so you can try a student and its larger sibling on the same task.
Building on GPUs? Aquanode runs the workload.
Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.
See also
Abliteration
Abliteration edits a language model's weights to remove its refusal direction, with no training run. How it works and what it means for GPU memory.
Model merging
Model merging combines the weights of several fine-tuned models into one, with no extra training and no extra inference cost. Methods and GPU needs.
Quantization
Quantization stores a model's weights, and sometimes activations, in lower-precision formats like INT8 or 4-bit, cutting VRAM use for a small accuracy cost.
VRAM
VRAM is the memory attached to a GPU that holds the data it works on, and it caps which AI models fit. VRAM vs RAM, how to check yours, and how much AI needs.
LoRA
LoRA fine-tunes an LLM by freezing its weights and training small low-rank matrices added to its layers, cutting trainable parameters and GPU memory.