Distillation (Part 2)
Neil Haddley • September 15, 2026
Turning a pretrained base model into a chat model with LoRA on a 64GB Mac Studio, then testing whether a second teacher, DeepSeek, produces a better one than GPT-3.5 did
Part 1 trained a small model from scratch and showed that a better teacher's answers produce a measurably better student, at two different scales. Everything in it was single-turn instruction-following, and the student was a custom architecture trained from random initialisation. This post changes both of those things: a real pretrained base model, Llama-3.1-8B (8 billion parameters), fine-tuned into a genuine multi-turn chat model with LoRA — and then the same one-variable question Part 1 asked, run again with a second teacher, DeepSeek, instead of GPT-3.5.
The base model does not know where its turn ends
Llama-3.1-8B-base has never seen a chat template. Prompted with a real two-turn conversation from UltraChat-200k — someone asking for a product name, then a tagline — it answers the immediate question reasonably:
User: "I like the name PureSprout Baby Formula, but could you add a tagline that emphasizes the all-natural, non-GMO aspect of the product?" Base model: "Sure! How about 'PureSprout Baby Formula: The All-Natural, Non-GMO Choice for Your Little One'?"
Then it keeps going — inventing a new User: turn, describing a logo that does not exist, offering a fake image link, and starting a packaging-design pitch nobody asked for. It was never taught that a reply ends; it was only ever taught to predict the next plausible token, and a plausible continuation of a chat transcript is more chat transcript. This is the same lesson Part 1's plain pretrained checkpoint taught with single-turn prompts, showing up again in a more specific way: fluency is not the same skill as knowing when to stop.
Why LoRA, not full fine-tuning
This machine is a 64GB M1 Max Mac Studio. Just loading the model takes about 16GB. Fully fine-tuning it needs far more than that on top: training has to track how every one of the model's 8 billion numbers should change, plus a running history of those changes to keep the process stable — in effect, several extra full-size copies of the entire model held in memory at once, each more precise (and so larger) than the copy you'd use simply to run it. Add all of that up and it comes to well over 100GB. Not close to fitting.
LoRA leaves the original model's numbers untouched and trains a small set of new ones alongside it instead — 10.5M trainable parameters out of 8.03 billion, 0.131% of the model. Because the original numbers never change, they can stay in the same compact 16-bit format (called bf16) used simply to run the model, with no need for an extra, more precise copy set aside for training. So in practice, memory only has to hold that unchanged model, the small set of new numbers actually being trained, and the temporary intermediate values the model produces while working out how to adjust them — landing around 25–35GB in total, comfortably inside the budget. That also gave a real quality improvement over QLoRA's 4-bit-quantised alternative, without the memory pressure that would have forced that trade-off.
Every turn needs its own training example
Here is a real, two-turn UltraChat conversation, used exactly as it appears in the dataset:
User 1: "Experiment with creating your own homemade natural body wash." Assistant (A1): "I do not have a physical body. However, I can suggest a recipe for a homemade natural body wash: ingredients — 1/4 cup liquid castile soap, 1/4 cup honey, 1/4 cup sweet almond oil..." User 2: "Can you suggest some essential oils that are good for sensitive skin?" Assistant (A2): "Yes, here are some essential oils that are good for sensitive skin: 1. Lavender oil 2. Chamomile oil 3. Sandalwood oil..."
UltraChat conversations average 3.2 assistant turns each, and mlx-lm's chat-data format only ever scores the final message in a {"messages": [...]} example against the loss — pass it a whole multi-turn conversation as one example and every earlier assistant turn is silently masked out along with the prompt. Supervising every turn meant expanding each conversation into one training example per assistant turn, each one a growing prefix of the last:
- Example 1: User 1 → A1, with A1 scored.
- Example 2: User 1 → A1 → User 2 → A2, with only A2 scored — A1 is now unscored context, exactly like the User turns.
One two-turn conversation becomes two training examples, each with a different final turn scored.
A 3-turn conversation becomes three examples the same way; this is why the counts elsewhere in this post talk about training examples in the tens of thousands, built from a training set of only 7,600 conversations.
That also fixes the context-length question in advance: since the scored answer is always the last thing in one of these examples, any example longer than the model's context window has to be filtered out before training, not truncated during it — truncating from the front, which is what happens by default, would cut the answer off and score nothing.
How far the context window can safely go
So the real question was how long a context window this hardware could actually support, since UltraChat's conversations run well past the 512-token window Part 1 used throughout. The answer was worth measuring directly rather than assuming: attention memory scales roughly with the square of sequence length, not linearly. Testing the actual longest training example at each candidate limit — not a random sample, which can miss the tail entirely — a 2048-token sequence peaked at a stable 31GB; a 3072-token sequence (1.5x the length) peaked at 69GB, already past this machine's 64GB of physical memory, surviving only by swapping to disk. 4096 tokens crashed outright, and inconsistently — sometimes completing, sometimes not, depending on what else the allocator was doing at that moment. 2048 is not an arbitrary round number here; it is the actual, empirically validated ceiling for an 8B model on this machine.
Building the second arm's data the same way
The distilled arm needs the identical structure — the same fixed User 1 / User 2 turns — but DeepSeek writing the assistant side instead of GPT-3.5. Critically, DeepSeek answers User 2 using its own answer to User 1 as context, never GPT-3.5's:
Both User boxes are identical text in both chains. Only who writes the answers differs — and each chain only ever builds on its own prior answer, never the other chain's.
DeepSeek's version of the same conversation, for comparison:
User 1: "Experiment with creating your own homemade natural body wash." DeepSeek (D1): "Here's a solid DIY natural body wash recipe, plus some variations and tips so you can tweak it to your skin type. Basic Natural Body Wash — You'll need: 1/4 cup liquid Castile soap... 1 tsp vitamin E oil (natural preservative and skin nourisher)..." — followed by scent/skin-type variations, shelf-life notes, and a "quick shortcut version," none of which GPT-3.5's answer included. User 2: "Can you suggest some essential oils that are good for sensitive skin?" DeepSeek (D2): "Great question — sensitive skin needs gentle, low-irritation oils. Here are the safest bets, plus a few to avoid. Lavender — The gold standard for sensitive skin... Chamomile — Excellent for redness, irritation, and reactive skin..." — followed by a "use with caution" list of oils to avoid and a suggested starter blend.
Once both chains exist, each gets turn-expanded exactly as shown earlier — baseline's two examples from GPT-3.5's A1/A2, distilled's two examples from DeepSeek's D1/D2 — before the length filtering and population-matching described next.
Training on UltraChat's own answers
The first arm — baseline — fine-tunes Llama-3.1-8B-base on UltraChat-200k's own conversations: 8,000 sampled from its ~208k total (7,600 train, 200 validation, 200 held-out test, matching Part 1's scale). UltraChat itself is worth being precise about: both sides of every conversation, not just the assistant, were generated by GPT-3.5-turbo playing two roles against itself, then cleaned up by Hugging Face's H4 team for supervised fine-tuning (SFT). Training this arm the same way Part 1 taught response-distillation more generally — masked cross-entropy, response tokens only — took the raw completion behaviour shown above and turned it into something that answers once and stops:
"PureSprout Baby Formula: All-Natural, Non-GMO Nutrition for Your Little One. The tagline emphasizes the product's natural and organic ingredients, which are free from harmful chemicals and genetically modified organisms..."
No fabricated follow-up turns, no invented logo. It is also considerably more repetitive than the single crisp sentence UltraChat's own GPT-3.5 answer gave for the same prompt — a real, honest limitation of roughly one epoch of LoRA on 14,000 examples, not something to hide.
A second teacher: DeepSeek
The interesting question Part 1 already answered once — does the teacher's quality actually matter? — deserved testing again with a genuinely different teacher, not just a bigger training run. DeepSeek (deepseek-flash, called through its API) is that second teacher, regenerating the assistant side of the same conversations as shown above.
Two real mechanical problems came up building this arm, both worth recording plainly:deepseek-flash is a reasoning model. Its reasoning_content field can consume an entire token budget before writing a single word of the actual answer — one real prompt from this dataset produced finish_reason: "length" with 100% of a 600-token budget spent "thinking" and zero characters of actual response. This task is straightforward instruction-following, not the kind of problem reasoning models exist for, so the fix was to disable it outright with reasoning_effort: "none" — which also cut cost substantially, since reasoning tokens are billed the same as any other output token.
DeepSeek's answers are considerably longer than GPT-3.5's. Matching each conversation's turn count to whatever baseline's cutoff already established (so both arms target the same population, not whichever teacher happened to be more concise) is necessary but not sufficient: DeepSeek's own, longer text at those same positions still individually exceeded the 2048-token safety limit for 40% of the candidate turns, even though UltraChat's shorter version of the identical position fit fine. Generation for the full 8,000-conversation set — parallelised across 40 concurrent workers, since this is a remote network call rather than a local GPU job — cost roughly $10–20 in total.
Making it a fair comparison, properly
Simply training distilled on whatever survived that length filter would have left it with fewer training examples than baseline purely because DeepSeek writes more — conflating "better teacher" with "teacher happened to cost less context per turn." Fixing this meant computing the true intersection: for every (conversation, turn) position, keeping it only if both UltraChat's own text and DeepSeek's own text fit under the limit, and training both arms exclusively on that matched set. That meant retraining baseline a second time — its first checkpoint, trained on its own full 23,351-example population, was discarded — so that the only variable left between the two arms is whose answer fills an identical set of 14,078 positions.
The result
Both arms — baseline on UltraChat's answers, distilled on DeepSeek's, same matched training positions, same LoRA recipe, same base checkpoint — were judged on 200 held-out conversations neither had trained on. The judge is Llama-3.1-70B-Instruct (4-bit, via mlx-community) again — a different lineage from both Llama-based students and both teachers, so it has no stake in either answer's style — comparing each arm's own generated reply to the same final turn, position-swapped per conversation exactly as Part 1's judge worked.

200 held-out conversations, neither arm trained on.
distilled won 118, baseline won 81, 1 tie — a 59.3% win rate excluding ties, one-sided binomial p ≈ 0.0053. Not a landslide, and the losses are as informative as the wins.
"¿Podría decirme el tamaño de la casa?" ("Could you tell me the size of the house?")
baseline(UltraChat/GPT-3.5): "...Esto significa que la casa tiene un área total de 1,200 pies cuadrados. Esto incluye el espacio de la casa... Esto significa que la casa tiene una sola planta y no tiene escaleras..." — the same "this means that" construction repeated three times, circling the answer rather than moving past it.distilled(DeepSeek): "Claro, el tamaño de la casa es de 1,200 pies cuadrados (aproximadamente 111 metros cuadrados). La descripción completa es: 'Starting out or ready to downsize, this cozy ranch style home offers you 1 floor living...'" — states the number, adds a genuine unit conversion, then moves on to substantive detail instead of restating itself.
"Wow, I didn't realize Australia beat Namibia by such a huge margin! That must have been a crazy game to watch."
baseline: "...Australia, who scored 14 tries in the match... played in front of a crowd of over 60,000 people at the Stade de France in Paris..."distilled: "...Namibia actually led 7-0 early on, but Australia scored 14 tries... It's worth noting that Namibia has since improved significantly and has become a much more competitive team..." Judge: "distilled is closer to the standard as it provides more specific details about the game and acknowledges Namibia's improvement since then."
A loss, to keep this honest — a request for error-handling code that should keep re-prompting rather than exit:
"Can you add some error handling in case the user enters negative values or zeros... I don't want the program to crash if that happens."
baseline: used awhileloop that keeps re-prompting the user until a valid positive integer is entered.distilled: produced code that validated the input but exited on the first invalid entry instead of looping. Judge: "baseline provides more robust error handling... whereas distilled simply exits the program if invalid input is detected." — the correct call. Not every comparison favours the better teacher; that is what a 59% win rate actually looks like, not 100%.
What I took from it
- The effect replicates with a real teacher swap, not just more scale. Part 1 varied scale twice (2.8M and 18.9M parameters) with one teacher pair (Qwen vs Alpaca's original answers) and got a consistent, modest edge both times. This post held scale and architecture fixed and swapped the teacher entirely — a different technique for reaching the same underlying claim: teacher quality, not just teacher size or student size, is what response distillation is actually transferring.
- A memory ceiling is worth measuring, not assuming. The 2048-token limit was not a guess dressed up as an engineering constraint — it came from directly testing the actual longest training examples and finding that 1.5x the context length cost 2.2x the memory, consistent with attention's quadratic scaling, and that the next reasonable-looking value (3072) already exceeded this machine's physical RAM.
- Matching the training population mattered as much as matching the recipe. A teacher that happens to write more isn't automatically a better teacher by this test's design — it needed to be, and discovering that the first version of this experiment quietly favoured brevity was worth retraining baseline a second time to fix.
Try it yourself
The code is in github.com/Haddley/chat-distillation, in part1/. Requires Apple Silicon for MLX, a Hugging Face account with no special access needed for the base model (mlx-community/Meta-Llama-3.1-8B-bf16 is openly available), and a DeepSeek API key for the second teacher arm.