Distillation (Part 1)
Neil Haddley • September 13, 2026
How DeepSeek stands accused of breaking OpenAI's terms of service.
"Distillation" usually means one specific thing in the research literature and a looser, more common thing in practice. This post is about the common thing: training a smaller model on a bigger model's answers, not its internal probabilities.
Two kinds of distillation
Same idea — a small model learning from a large one — two different training signals, with two different requirements.
Soft-label (logit) distillation — already built on this blog — trains a student to match a teacher's whole probability distribution over the next token.
Hard-label (response) distillation is cruder: ask the teacher a question, keep the one answer it wrote, and train the student on that text with ordinary cross-entropy — the same loss every model uses for plain next-token prediction. All you need is the text, which is why it works with a closed API just as well as an open one.
How Microsoft does this — and the unproven accusation against DeepSeek
Phi (Microsoft) — hard-label, not soft-label. Starting with phi-1 in 2023, Microsoft trained small models substantially on synthetic data — plain text generated by GPT-3.5, then GPT-4, then GPT-4o as the series progressed — rather than on web text alone. Training on that generated text still just means ordinary next-token cross-entropy over words GPT-4 wrote, the hard-label approach from earlier in this post; the OpenAI API Microsoft called does not expose GPT-4's internal probability distribution in a form soft-label distillation could even use, so the logit-matching approach was never on the table here. The pitch was that a small model taught on carefully generated, textbook-quality examples learns faster than one taught on the internet's raw average. Microsoft has direct, licensed access to OpenAI's models, so this is uncontroversial, done at large scale, disclosed openly in their own technical reports.
DeepSeek — an unproven accusation. Around the time DeepSeek released R1, OpenAI told the Financial Times it had seen evidence suggesting accounts linked to DeepSeek had queried its API at a scale consistent with harvesting outputs to train a competing model — a possible violation of OpenAI's terms of service, never publicly confirmed with hard evidence. Same technique as Phi's either way; the real difference, if the accusation is true, is that the answers being trained on were not used with permission.
Trying it myself
I wanted to see the mechanism directly rather than just read about it, using a teacher I could actually run: Qwen2.5-32B-Instruct (4-bit, via mlx-community), local, free, via MLX. I generated fresh answers to 8,000 Alpaca prompts, then trained a small model (a few million parameters) on them with the same masked cross-entropy loss described above.
One thing became clear immediately: this only works on top of ordinary pretraining. A model with no prior exposure to raw text, trained purely on a few thousand Q&A pairs, learns the shape of a good answer — numbered lists, section headers — long before it learns enough English to fill that shape with anything coherent. Every real system that uses this technique, including both examples above, starts from an already-pretrained model. So I did too: pretraining first on WikiText-103, then response-distilling on top.
Pretraining alone only ever teaches next-token prediction, so the plain pretrained checkpoint has no reason to treat a prompt as a question to answer — it was never trained to do that. Prompted with one of the same 8,000 questions before any response-distillation at all, it just continues the text as if it were the opening of a Wikipedia article:
"Does the word 'malfunctioning' have any synonyms?" "...is not a source of the other, a means that a great power might have been used by the GIS's book. The book's name is also believed to have been made by the novel..."
That is exactly what the response-distillation stage exists to fix: it is the step that teaches the model to answer at all, not just to write fluent text. This checkpoint is the shared starting point both students below are fine-tuned from; everything that follows is about what happens after this point.
Worth being precise about what that stage does and does not produce: every training example is one prompt and one response, so both students end up as single-turn instruction-followers, not chat models. They use the same <|user|>/<|assistant|> token scaffolding a chat model's training data uses, but nothing in that data spans more than one exchange — there is no prior turn to condition on, no system prompt, no notion of a conversation. One raw Alpaca pair from the training set, wrapped exactly the way tokenize_qa.py wraps every example before it reaches the model:
CODE
1<|user|>Edit this sentence so it is in the form of a questions. 2 3I love ice cream.<|assistant|>Do you love ice cream?<|endoftext|>
That is the entire training example — one <|user|> block, one <|assistant|> block, then <|endoftext|>. There is no second <|user|> block anywhere in the format for a follow-up to occupy. Ask one of these models a follow-up that depends on its previous answer and it has no mechanism to know what the follow-up refers to, because it was never trained on anything but isolated, one-shot exchanges like this one.
Base, Instruct, Chat — what those suffixes actually promise
Released model families name these three stages fairly consistently, and the naming maps directly onto what I built:
- Base (Llama-3-8B, Qwen2.5-32B with no suffix) is the raw pretrained checkpoint — next-token prediction only, no notion of a question or a task. My own equivalent is the plain pretrained checkpoint shown above, before either fine-tuning arm touches it.
- -Instruct (Llama-3-Instruct, Gemma-it) marks a model fine-tuned on top of a base to follow single instructions well — precise, direct task execution, not necessarily sustained dialogue. distilled and baseline are Instruct models in this narrow sense.
- -Chat (Llama-2-Chat) generally implies the same instruction-following ability plus training on multi-turn conversations, often with a further preference-tuning stage (RLHF or DPO) layered on top of the SFT step to shape how it converses, not just what it answers.
In practice labs use "-Instruct" and "-Chat" close to interchangeably, and the suffix alone does not reliably tell you which stages a given release actually went through — the only way to know for certain is to read the model card. Neither distilled nor baseline earns a "-Chat" label: turning either into one would need a different training set entirely, built from multi-turn conversations with real dialogue history rather than Alpaca's isolated prompt/response pairs, and very likely a preference-tuning stage this project never attempted.
The test that actually mattered
Getting a coherent small model running was the easy part. The real question — does the teacher's quality actually matter, or would any answers do? — needed a proper controlled comparison, which I built once the pretraining stage was in place:
Of the 8,000 Alpaca prompts, I split off 200 for validation and 200 for testing, leaving 7,600 for fine-tuning — the same 7,600 prompts, and the same split boundaries, for both arms:
- distilled: the pretrained model, fine-tuned on Qwen2.5-32B-Instruct's answers to those 7,600 prompts.
- baseline: the same pretrained model, fine-tuned on Alpaca's own original 2023 answers to the identical 7,600 prompts (written by the much older text-davinci-003, and swapped out for fresh Qwen answers everywhere else in this project).
Neither model ever saw the 200 held-out test prompts during training — that reserved set is what the judge scored below.
Two of the 7,600 training prompts, answered by both:
"Answer the given question in yes or no. Question: Does social media have a negative effect?" Alpaca (2023): "Answer: Yes" Qwen (2025): "Yes, social media can have negative effects, though these can vary widely depending on usage and individual circumstances."
"Does the word 'malfunctioning' have any synonyms?" Alpaca (2023): "Yes, the word 'malfunctioning' has synonyms such as failing, faltering, defective, impaired, and deficient." Qwen (2025): "Yes, the word 'malfunctioning' does have several synonyms. Some of these include: Faulty, Defective, Broken, Not working, Out of order, Malfunctioned, Dysfunctional, Inoperative..."
Alpaca's answers are not wrong — both examples above are perfectly correct. They are just thinner: a bare "Yes," five synonyms instead of eight with more natural phrasing around them. That gap, repeated across all 7,600 training prompts, is the entire independent variable in the comparison below.
Comparing perplexity between the two would be rigged — each model would simply score best on its own training source's writing style. Grading my own two models' answers would be no better. So I built the test to remove every source of bias I could think of:
1. Held out 100 prompts neither model had trained on — half of the 200-prompt test split set aside before fine-tuning even started.
2. Generated a fresh answer from each model to every one of those 100 prompts.
3. Randomly swapped which answer was labelled "A" and which was "B" for each prompt, so consistently favouring "A" or "B" could not manufacture a result.
4. Sent both answers, unlabelled as to source, to a third model to judge — Llama-3.1-70B-Instruct (4-bit, via mlx-community) — deliberately not Qwen, since a judge from the same family as one of the training sources would likely rate that source's style more favourably.
5. Told the judge to privately weigh what a genuinely good answer would contain before comparing A and B — not to write one out itself, just to use it silently as the yardstick so the choice is closer-to-correct rather than merely closer-to-the-other-flawed-answer, and explicitly not to favour whichever answer just sounded longer or more confident.
6. Tallied the picks across all 100 prompts and checked the result against a one-sided binomial test, to rule out the win margin being noise from a 100-prompt sample.

100 held-out prompts, neither model trained on.
distilled won 60, baseline won 36, 4 ties — a 62.5% win rate excluding ties, with a one-sided binomial test putting this at p ≈ 0.009. Modest, not dramatic, but real: on identical prompts, identical architecture, and identical pretraining, the model trained on the better teacher's answers won more often than not, and not by chance.
What I took from it
- A smaller model can absorb real value from the work poured into a bigger one, and this experiment shows why that is not just a Phi/DeepSeek-scale phenomenon. All it takes is the bigger model's answers — no logits, no special access, nothing an API-only teacher would not hand over voluntarily. That is exactly what makes the practice usable with a teacher's permission (Phi, licensed OpenAI output) and, allegedly, without it (the accusation against DeepSeek) — the technique itself does not care which.
- Catching my own judge bias mattered more than any other decision here. A biased evaluator does not announce its bias in its output; the only way to find it was to notice the conflict of interest before trusting the number.
- A real effect can be small, and reporting it as small is the honest thing to do. 62.5% is not a rout. It is a genuine, statistically real edge from using a better teacher, at a scale where "genuine but modest" is exactly what you'd expect.
The same test, at a much larger scale
62.5% is the effect size a 2.8M-parameter model, 8,000 prompts, and 100M pretraining tokens can show. So I reran the identical one-variable comparison much bigger: an 18.9M-parameter student (about 7x the parameters), pretrained on 807M tokens of raw Wikipedia text (about 8x the corpus), then fine-tuned on 19,400 prompts (2.5x as many) — same two arms, same judge, same held-out methodology, just scaled up across every dimension at once.
The result was not just bigger — it was clearer. Judged on 300 held-out prompts this time (up from 100): distilled won 186, baseline won 100, 14 ties — a 65.0% win rate excluding ties, with a one-sided binomial test putting this at p ≈ 2 × 10⁻⁷. The effect got stronger, not weaker, once the setup was less data-starved — exactly the direction you would hope for if the original 62.5% reflected a real, modest teacher-quality effect rather than noise from a small model and a small test set.
Two of the judged prompts, unedited, both judged clear distilled wins:
"In not more than two sentences, summarize the purpose of Machine Learning."
baseline: "Inputing Machine Learning is an example of a machine learning field that is a field of Machine Learning. It is a machine learning field that allows for manual learning to be done through a machine learning algorithm. Machine Learning is a field of machine learning that uses algorithms to learn from..."distilled: "Machine Learning is a subset of artificial intelligence that focuses on the development of algorithms and statistical models that enable computers to make predictions or decisions without explicit instructions or relying instead of using explicit instructions to maintain the accuracy of a program."
"Give an example of a way to investigate the cause of air pollution."
baseline: "An example of a way to investigate the cause of air pollution is to try to reduce air pollution."distilled: "An example of a way to investigate the cause of air pollution is the burning of fossil fuels (coal, oil, and natural gas) for energy production and industrial processes. This process is straightforward and can be quite different than traditional fossil fuels. Here's an example: 1. Increase Energy..."
Neither is polished — this is still an 18.9M-parameter model, and distilled's second answer visibly trails off into a listicle that does not really go anywhere. But the pattern across both is the same one the judge picked up 186 times: baseline tends to circle a topic using its own words back at itself ("a field of Machine Learning... a machine learning field", "investigate... reduce", answering a different question than the one asked), while distilled states something substantive and correct before it runs out of road. That gap, not either model's absolute polish, is what the win rate is measuring.
What's next
The mechanism holds up at both scales tested so far. The next thing worth varying is not scale again, but structure: everything in this post — both arms, both scales — was single-turn instruction-following, built on Alpaca's one-prompt-one-answer format. Part 2 picks that up: a real pretrained base model, LoRA fine-tuning, genuine multi-turn chat behaviour, and a second teacher (DeepSeek) instead of a bigger run with the same one.
Try it yourself
The code is in github.com/Haddley/distillation: part1/ (teacher-generated answers), part2/ (the two-stage recipe at small scale), part3/ (the fair comparison above), and part4/ (the same comparison rerun at the larger scale below). See the repo's README for the exact commands and how the pieces fit together.
Requires Apple Silicon for MLX. The judge alone needs about 40GB free for Llama-3.1-70B-Instruct-4bit.