The Curator

Model Distillation: A Beginner’s Guide to Teaching Smaller AI Systems

Last updated: 9/28/2026

Back to blog
Hana Berg avatarHana Berg 7 min read
Cover image for Model Distillation: A Beginner’s Guide to Teaching Smaller AI Systems
AI-assisted, human-reviewed. Drafted with AI research tools from public sources and edited by our team. How we build these →

Model distillation begins with a practical observation: the most capable AI model is not always the right model to deploy. A large system may solve a task well, yet be too slow, costly, private, or operationally demanding for the product surrounding it. Distillation offers another path. Instead of asking a smaller model to learn only from original training data, you let it study the behavior of a more capable model.

The result is not a miniature copy. It is a specialized student shaped by what the teacher demonstrates and what the training process rewards. That distinction is the key to understanding both the promise and the limits of distillation.

The Core Mental Model: Compress Behavior, Not Files

Distillation is sometimes described as model compression, but the phrase can mislead. You are not placing a large model into an archive and extracting a smaller equivalent. You are training a separate model to reproduce selected aspects of the larger model’s behavior.

Imagine a teacher that can classify support messages, draft replies, explain policy, translate text, and write code. If your product only needs to route support messages, the student does not need the teacher’s full repertoire. It needs a reliable approximation of one bounded capability.

A typical process looks like this:

  1. Define the behavior the deployed system must perform.
  2. Collect representative inputs for that behavior.
  3. Ask a capable teacher model to produce labels, answers, probabilities, or reasoning traces.
  4. Train a smaller student on those demonstrations.
  5. Evaluate the student on cases that were withheld from training.

The teacher therefore acts as a generator of learning signals. The student’s architecture, training data, and objective determine how much of that signal survives.

The Vocabulary That Makes Distillation Legible

TermMeaningWhy it matters
TeacherThe model whose outputs guide training.Its strengths, errors, and biases can all be transferred.
StudentThe smaller or otherwise deployable model being trained.Its capacity places a ceiling on what can be learned.
Hard labelA single target answer, such as “billing.”Simple to store and train on, but discards uncertainty.
Soft targetA probability distribution over possible answers.Reveals relationships among alternatives when model access permits it.
LogitsRaw scores produced before probabilities are calculated.They contain richer information than the winning output alone.
Sequence distillationTraining on complete text sequences generated by a teacher.Common when the teacher is accessible only through generated text.
TemperatureA setting that changes how concentrated or varied output probabilities are.It can expose alternatives or diversify generated examples.
Fine-tuningUpdating a pretrained model using task-specific examples.Distillation is often implemented as fine-tuning with teacher-produced targets.

These terms describe different layers of the process. Fine-tuning is the training operation; distillation describes where the supervision comes from and what behavior it is meant to transfer.

What the Student Actually Learns

Consider a routing system with three destinations: billing, technical support, and account security. A hard training label tells the student that “I cannot access my account after changing phones” belongs to account security. A soft target might indicate that account security is strongly preferred while technical support remains plausible.

That second signal contains structure. It teaches the student that some categories are easily confused while others are remote. Traditional distillation can train the student to match the teacher’s softened probability distribution, often alongside the verified label.

For generative systems, internal probabilities may be unavailable. The teacher instead creates outputs: summaries, classifications, tool arguments, rewritten queries, or answers following a required format. The student learns through ordinary next-token prediction on those sequences.

There are several useful forms of supervision:

  • Answer supervision: Train on the teacher’s final response.
  • Rationale supervision: Train on an explanation or intermediate decomposition, provided it is useful and safe to expose.
  • Critique supervision: Have the teacher identify faults in candidate responses.
  • Preference supervision: Ask the teacher to rank alternatives, then train the student to prefer the stronger one.
  • Tool-use supervision: Capture which tool was selected and the arguments supplied.

More detail is not automatically better. Long rationales can add noise, teach stylistic imitation rather than correct decisions, and increase training cost. The target should expose only the structure the student needs.

Where Distillation Creates an Advantage

The strongest use cases are narrow enough to evaluate and frequent enough for deployment economics to matter. Classification, extraction, query rewriting, policy-constrained drafting, and recurring tool selection are natural candidates.

Suppose an application receives customer messages and must emit a category, urgency level, and destination queue. A general model can produce this record, but much of its capability is irrelevant. A distilled student can be trained on representative messages paired with teacher-generated records, then checked against human-reviewed examples. If it meets the required quality threshold, it may offer lower latency, simpler capacity planning, or deployment within a controlled environment.

Distillation can also create architectural options. A local student might handle routine requests while uncertain cases escalate to the teacher. The smaller model is no longer merely a cheaper substitute; it becomes the first stage of a routed system.

The Trade-offs Hidden by the Word “Smaller”

A student generally loses breadth before it loses competence on its designated task. This is desirable when the boundary is explicit and dangerous when users can silently cross it.

Capability becomes conditional

Performance may hold on familiar phrasing but collapse on new domains, languages, adversarial inputs, or unusually long context. Average accuracy can conceal these edges.

Teacher errors become training data

A confident teacher can generate plausible but incorrect labels. Repetition turns those mistakes into stable student behavior. Human review is especially important for rare, consequential, or ambiguous cases.

Data coverage matters more than volume alone

Thousands of near-duplicate easy examples may teach less than a deliberately balanced set containing borderline cases, missing information, malformed inputs, and situations where the correct action is abstention.

Operational savings are not guaranteed

Generating demonstrations, cleaning them, training the student, serving it, and monitoring drift all carry costs. Distillation makes sense when the resulting system has a durable role, not merely because a smaller checkpoint exists.

A First Distillation Experiment

Begin with a task whose result can be judged without subjective debate. A structured classifier is more instructive than an open-ended assistant.

  1. Write the contract. Define the allowed inputs, output schema, classes, abstention behavior, and prohibited actions.
  2. Build a human-reviewed evaluation set. Include ordinary cases, boundaries between classes, incomplete requests, and inputs outside the task.
  3. Establish the teacher baseline. Run the teacher against the evaluation set using a fixed prompt. Record outputs, latency characteristics, and failure types.
  4. Create training examples. Use real, appropriately governed inputs where possible. Add synthetic variations only to cover known gaps, not to disguise missing evidence.
  5. Generate targets. Request structured outputs and validate them mechanically. Reject malformed records rather than teaching them.
  6. Train the student. Fine-tune a suitable pretrained model on the accepted input-output pairs.
  7. Compare failure slices. Evaluate by class, input source, language, ambiguity, and consequence—not only by one aggregate score.
  8. Design escalation. Route low-confidence, invalid, or out-of-scope cases to a stronger system or a person.

A useful worked pattern is invoice-email routing. Inputs are subject lines and message bodies; outputs are payment status, invoice request, dispute, suspected fraud, or other. The teacher labels a governed corpus, reviewers inspect ambiguous and high-risk examples, and the student learns the mapping. At runtime, suspected fraud and uncertain outputs always escalate. This preserves the student’s efficiency without pretending it should own every decision.

What to Ignore for Now

Do not begin by comparing every student architecture, reproducing research benchmarks, or transferring hidden representations between models. Those questions matter later, particularly when you control both teacher and student internals.

Also ignore the ambition to distill a general assistant in your first attempt. Generality makes coverage elusive and evaluation ambiguous. A bounded task will teach you more about data design, error transfer, and deployment thresholds.

Finally, do not treat resemblance to the teacher as the ultimate objective. The deployed student must satisfy the product contract. A student that occasionally disagrees with its teacher may be better if human-reviewed evidence supports it. The teacher is a source of supervision, not an oracle.

The Strategic Insight

Distillation changes the role of a frontier model. It need not answer every production request directly. It can operate upstream—as demonstrator, critic, labeler, and curriculum designer—while smaller systems perform stable work downstream.

This suggests a disciplined sequence for AI products: discover the behavior with a capable model, make the behavior measurable, distill the repeatable portion, and reserve expensive intelligence for uncertainty. The durable advantage is not simply a smaller model. It is knowing precisely which capability deserves to become infrastructure.

This post was drafted with AI assistance and reviewed against our editorial policy before publication. Corrections are made at the source, on the page, with the date shown.

model distillationsmall language modelsAI inferencemachine learningAI product design

From our own rounds

Measured on The Curator, from real sessions people played on this site — not a third-party dataset.

Rounds played here
161
Questions per round
1.7
Play a round and add to these numbers
Share this post

Rate this article

No ratings yet

Discussion

Comments are moderated. Read our editorial policy.