What is distillation?

Distillation is a training method where a small "student" model learns to imitate a larger "teacher" model, so the student delivers part of the teacher's quality in a form that is cheaper to serve. The cost saving comes from size: a student with fewer parameters needs less VRAM and less compute per token than the teacher.

The original idea

Hinton, Vinyals and Dean introduced the technique for neural networks in Distilling the Knowledge in a Neural Network. Their motivation is that predicting with a whole ensemble of models is cumbersome and may be too expensive to deploy, so they compress the ensemble's knowledge into a single model. The key device is soft targets: instead of training only on hard labels, the student trains on the probability distributions the teacher outputs, which carry more information per example than a single correct label. A temperature setting softens those distributions so that the relative probabilities of wrong answers become visible to the student. The paper also describes specialist models that focus on fine-grained distinctions the full models confuse.

Distillation for language models

Language models are distilled in more than one way. The classic approach matches the teacher's output distribution. Another, widely used for reasoning models, is to fine-tune the student on text the teacher generated. The DeepSeek-R1 distills are an example of the second: their model cards state that they are fine-tuned on samples generated by DeepSeek-R1, starting from existing Qwen and Llama base models. The Qwen-based distills are described as fine-tuned with 800k samples curated with DeepSeek-R1. The student is a normal checkpoint with the base model's architecture, so it runs in the same serving software as any other model of that size.

Example models in the catalog

How it differs from related techniques

  • Quantization keeps the model's size in parameters and reduces the bits per weight. Distillation reduces the parameter count. The two combine freely: quantization of a distilled model is common.
  • Fine-tuning with LoRA adapts a model to a task with small added weights (LoRA, QLoRA). Distillation can use LoRA as its training method, but its goal is to transfer behavior from a teacher.
  • Model merging combines the weights of existing models without training (model merging).

What it means when you pick a GPU

Distillation shifts the choice from "which big model" to "how small a student is acceptable". A distilled model is sized like its base: an 8B student needs the memory of any 8B model and fits cards a 70B teacher cannot. Use the VRAM calculator with the student's parameter count, and add room for the KV cache, which grows with context length. Whether the student is good enough is a question to test on your own prompts. Aquanode rents GPUs by the hour, so you can try a student and its larger sibling on the same task.

Building on GPUs? Aquanode runs the workload.

Deploy on H100, H200, B200, A100 and MI300X across a multi-provider marketplace, without racking your own hardware or committing to one cloud's spec sheet.

See also

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.