
Distillation means training a small AI model on the answers of a big one. The idea was published in 2015 by Geoffrey Hinton, Oriol Vinyals and Jeff Dean at Google. The small "student" learns from the big "teacher's" answers and from how confident it was about each, which teaches far more per example than a plain right-or-wrong label. Meta used it in 2025 to rescue value from Llama 4 Behemoth, a nearly 2-trillion-parameter model that the research firm SemiAnalysis says suffered from flawed design choices and messy data: Meta says it distilled its Maverick model from Behemoth, and SemiAnalysis reports Scout was distilled too, adding that distillation is far more efficient than reinforcement learning (trial-and-error training) for small models. In January 2025 DeepSeek released six small models distilled from its R1 reasoning model, and found that distilling beat training small models with reinforcement learning directly. Distillation is legitimate and everywhere, but it has a dark side: in February 2026 Anthropic accused DeepSeek, Moonshot and MiniMax of using about 24,000 fake accounts and over 16 million conversations to distil its Claude models without permission. The catch: a student is rarely better than its teacher.
What happened
In April 2025 Meta announced Llama 4 Behemoth, its biggest AI model ever: nearly 2 trillion parameters (the adjustable numbers a model learns during training). It never shipped to the public. According to SemiAnalysis, a research firm that tracks the AI industry, the training run went wrong in several ways: Meta switched how the model routes work between its internal specialists halfway through, which left it meaningfully worse, and it switched to a new, poorly cleaned source of web data mid-run.
Yet Meta did not throw it away. It used Behemoth as a teacher. Meta says it "codistilled" its smaller Llama 4 Maverick model from Behemoth, and SemiAnalysis reports that Scout, the smallest, was distilled from it too. In SemiAnalysis's words, "distillation is far more efficient than reinforcement learning for smaller models".
How distillation works
Normally an AI learns from examples with a single right answer: this photo is a cat, the next word is "Paris". Distillation adds a teacher. You ask a big, expensive model a huge number of questions, record its answers, and train a small model to copy them.
The clever part, described by Geoffrey Hinton, Oriol Vinyals and Jeff Dean at Google in a 2015 paper, is to copy more than the final answer. A big model does not just say "cat". It says something like 90% cat, 9% fox, 1% car. That spread is information. It tells the student that cats and foxes look alike and that cars do not, something a plain "correct answer: cat" never teaches. So the student learns far more from each example and needs much less training to get good.
The reward is speed and cost. A small model can run on a laptop or phone, answer in a fraction of a second, and cost a sliver of what the teacher costs to run. Many of the cheap "mini" and "flash" models you use every day are built this way.
The everyday version
Picture learning to cook. You could spend years experimenting alone, burning dishes until you work out what goes wrong. That is roughly reinforcement learning, trial and error with a score at the end. Or you could stand beside a master chef for a month, watching what they do and hearing them mutter "nearly there, a touch more salt, definitely not yet". You will not become as good as the chef, but you will get surprisingly close, surprisingly fast. The muttering, the teacher's confidence, is what makes distillation work.
A student rarely beats its teacher. But a cheap student that is nearly as good is often all you need.
Is this actually new?
No. The core idea is more than ten years old, and it went mainstream in January 2025 when the Chinese lab DeepSeek released its R1 reasoning model along with six small versions, from 1.5 to 70 billion parameters, built on Alibaba's Qwen and Meta's Llama models and trained on R1's answers. DeepSeek reported that distilling worked better than training those small models with reinforcement learning directly, and that its 32-billion-parameter student beat OpenAI's o1-mini on several benchmarks, by DeepSeek's own measurements.
What is newer is the fight. Labs happily distil their own models, but a teacher's answers can also be harvested by someone else. In February 2026 Anthropic said three Chinese labs, DeepSeek, Moonshot and MiniMax, had used about 24,000 fake accounts to hold more than 16 million conversations with Claude, in breach of its terms, to train their own models. MiniMax alone accounted for over 13 million, according to Anthropic. These are Anthropic's allegations, and the labs involved did not confirm them. Anthropic itself calls distillation "a widely used and legitimate training method"; the dispute is about whose teacher you are allowed to learn from.
What it means
Distillation explains two things that look puzzling from the outside. First, why companies spend fortunes on giant models they barely release: the giant is often the teacher, and the money is made on its smaller, cheaper students. Second, why a lab that falls behind can catch up so fast: if you can learn from a leader's answers, you skip much of the expensive trial and error.
There is a limit, and SemiAnalysis names it. Meta's students were "still bound by the limitations of their source" and were not best in class for their size. A copy inherits its teacher's blind spots. That is why the race at the very top is still about building better teachers, and why the companies that own the best ones are now guarding them so carefully.
Curious about AI? Come build with us.
Oslo Vibe Coding runs free, beginner-friendly drop-ins where we build real things with AI. No one codes alone.