MiniGPT (Part 2)
Neil Haddley โข October 5, 2026
How a GPT is grown: retraining my MiniGPT exhibit model from random numbers on a Mac, one guess at a time, and watching every embedding and parameter from the first post appear
In the first post, I took a trained MiniGPT apart while it was running: the token and position embeddings, attention, the MLPs, and lm_head with its 65 rows. Every number in it came from one model, my exhibit. This post answers the obvious next question: where did those numbers come from? Nobody typed them in. They were grown, and here I grow them again from scratch, on my Mac, and end up with exactly the same numbers.
In Part 1, writing one letter took five steps: letters to numbers, embeddings, the blocks, chances, and spin the wheel. Growing the machine uses exactly the same steps, with step 5 changed and one step added. So if Part 1 made sense, most of this post will already feel familiar.
One promise, the same as last time: no magic. Every number in this post is either worked out in front of you, or comes from a run on my own Mac Studio.
| This post's machines | The exhibit | The bigger machine |
|---|---|---|
| What changed | grown from random numbers | bigger, in the notebook's part 3 |
| Text | Tiny Shakespeare | Tiny Shakespeare |
| Pieces | letters: 65 | letters: 65 |
| Blocks | 4, each 4 heads | 6, each 6 heads |
| Vector size | 128 | 384 |
| Positions | 128, position embeddings | 256, position embeddings |
| Engine | PyTorch | PyTorch |
| Size | 826,433 numbers | 10.77 million, token embeddings doubling as lm_head's rows |
| Score | 1.70 surprise per letter | 1.47 |
Growing a GPT, in plain English
Starting from nothing
Build Part 1's machine, but put random numbers in every embedding and every table of weights, and run its five steps. Step 4 then gives an almost even wheel: every slice about the same size, 1 in 65. So what it writes is pure noise. I asked the untrained MiniGPT model from the notebook to carry on from ROMEO:, and this is what it wrote:
CODE
1ROMEO:dHrNUBK.!oWffHGBTyysS:hvnRyuZtONMQhv,dAw$xBM.VQu.!nhulODzyrBM;sta'fy,dlAvIzRlwMl
Every letter in that line was a near-random spin. The untrained model's chances were all between about 1% and 3%, close to an even 1 in 65.
That is the whole secret of training. The machine is never taught a rule like "a capital letter starts a name" or "q is followed by u". Training only reshapes the wheel, one tiny nudge at a time, until the slices for letters that tend to come next in Shakespeare are bigger than the rest. Shape it well enough, and the spins start to look like writing.
Keeping score: the surprise score
How do you tell a good forecaster from a bad one? You wait to see what the weather actually does, and then you look at the chance they gave it. A forecaster who said "90% chance of rain" on a day it rained did well. One who said "5% chance of rain" on a day it poured did badly. A good forecaster should not be surprised.
The machine is marked in exactly the same way. Once the real next letter is revealed, I look up the chance the machine gave that one letter, ignore all the others, and turn it into a surprise score: a number for how surprised the machine should be by what actually happened. A high chance gives a low surprise score, and a low chance gives a high one.
How the surprise score is worked out: the halving rule. The surprise score follows one simple rule:
- If the machine gave the right letter 100%, the surprise score is 0. It was certain, and it was right, so it is not surprised at all.
- Every time that chance is cut in half, add the same amount to the surprise score: 0.69.
So the surprise score climbs one step of 0.69 for every halving:
| Chance the machine gave the right letter | Halvings from 100% | Surprise score |
|---|---|---|
| 100% | 0 | 0 |
| 50% | 1 | 0.69 |
| 25% | 2 | 1.39 |
| 12.5% | 3 | 2.08 |
| 6.25% | 4 | 2.77 |
| 3.1% | 5 | 3.47 |
| 1.6% (1 in 64) | 6 | 4.16 |
Chances that fall between two rows get a surprise score between those two rows. Take my trained model's real chances after good m, as in good my lord or good madam: y 40.8%, e 26.9%, and a 15.2%.
- If the next letter turns out to be
y, the machine gave it 40.8%. That is a little less than 50%, so the surprise score is a little more than one step: 0.90. - If it turns out to be
a, the machine gave it only 15.2%. That falls between 25% and 12.5%, so the surprise score lands between 1.39 and 2.08: 1.88. A bigger surprise.
Why add 0.69 for each halving, and not a round 1? There is no deep reason. Adding 1 for each halving would work just as well: every surprise score would come out about 1.44 times bigger, and the machine would learn exactly the same thing. The maths in the notebook happens to use 0.69, so I use it too, which means the surprise scores in this post match the numbers the notebook prints.
A score of 0 means the machine gave the right answer 100%: it was certain, and it was right. It cannot do better than that. In the other direction there is no limit: every further halving adds another 0.69, so give almost 0% to the letter that actually turns up, and the surprise score shoots up. That is how the game punishes a machine for being confidently wrong.
This one number is what the whole project is about. Training means making the machine's average surprise score, over millions of guesses, as small as possible.
A machine that knows nothing at all gives the same 1-in-65 chance to every letter. 1 in 65 is almost exactly 1 in 64, the bottom row of the table: six halvings from 100%. So whatever comes next, its surprise score is about 6 ร 0.69, or 4.17 to be exact. That is almost exactly where my exhibit started: 4.18 on its very first step.
Under the hood: Where the surprise numbers come from
The halving rule is exactly what a logarithm does. The surprise score is the natural logarithm of one over the chance, written ln(1 รท chance), and every scientific calculator has an ln button. The step of 0.69 is ln(2): the cost of one halving.
- A chance of 40.8%: ln(1 รท 0.408) = 0.90
- A chance of 15.2%: ln(1 รท 0.152) = 1.88
- A chance of 1 in 65: ln(65) = 4.17
- A chance of 100%: ln(1) = 0
To add 1 for each halving instead, swap ln for logโ, the logarithm that counts halvings directly. Then the blind guess scores logโ(65) = 6.02: just over the six halvings in the table.
There is a friendlier way to read the surprise score. Turn it back into a sentence (in the maths, raise e, about 2.718, to the power of the surprise score): "the machine was as unsure as if it were choosing between N equally likely letters". At the start, N is 65. After my small model had trained for 90 seconds, N was about 5.5. After the bigger model, it was about 4.3. That is the whole scoreboard for this project: from 65 down to 4.
Growing it: the same five steps, then check and nudge
Everything Part 1 called fixed, the token embeddings, the position embeddings, every table of weights in the blocks, the stretch and shift in every normalisation, and the rows of lm_head, comes down to numbers, the parameters. I picture each one as a dial. The small model has 826,433 parameters, and the bigger model has 10.77 million. Growing the machine means finding good settings for all of them, and it uses the five steps from Part 1, with one change and one addition:
| Step | When the machine writes (Part 1) | When the machine grows |
|---|---|---|
| 1. Letters to numbers | the text so far | 32 random snippets of practice text, each 128 letters long: exactly enough to fill every position. The 32 snippets for one step are called a batch |
| 2. Embeddings | a token embedding plus a position embedding, for each letter | exactly the same, except that the embeddings start as random numbers |
| 3. The blocks | attention, then the MLP, four times | exactly the same |
| 4. Chances | a wheel for the last position only | a wheel for every position: 32 ร 128 = 4,096 wheels, because in practice text, every position's next letter is already known |
| 5. | spin the wheel, and write whatever comes up | check the answer: look up the chance each wheel gave the real next letter, turn it into a surprise score, and average all 4,096 |
| 6. | nudge every one of the 826,433 fixed numbers a tiny amount, in whichever direction would make that average surprise smaller |
Each pass through all six steps is one training step. Then the machine does it all again with 32 new snippets: 3,000 times in all.
Here is the very first step of growing my exhibit, with the real numbers:
- Steps 1 to 4. One of the 32 snippets begins " guess who caused yo". Its first hidden state has seen only the space, and its wheel gives the real next letter,
g, a chance of 0.87%. - Step 5. That one guess has a surprise score of 4.74. Averaged over all 4,096 guesses, the surprise is 4.18: almost exactly the 4.17 of an even wheel, because the machine knows nothing yet.
- Step 6. Almost every one of the 826,433 numbers moves by 0.0003. That distance is a setting I chose, called the learning rate: how far each nudge goes. The
gtoken embedding's first number goes from โ0.04793 to โ0.04823, the position 1 embedding's from 0.02622 to 0.02592, and the first number oflm_head'sdrow from 0.01135 to 0.01105. (On the very first step, every number with a slope moves almost exactly the same distance, because the training method, AdamW, starts with equal-sized steps. Later on, it sizes each number's step separately. The exceptions are 5 token embeddings, which you will meet below.)
Watching it grow
To see what those steps do, I grew the exhibit model from the first post again, and stopped it five times along the way to write from the same prompt. Nothing changes between the frames except the parameters:
This is what I meant at the start by a model being grown rather than built. Nobody wrote anything into the parameters between those frames. The guessing game did it.
The payoff: growing the exhibit, exactly
Every number in the first post came from one trained model, my exhibit. Here is where they came from. The exhibit is exactly the training loop above: 3,000 steps of 32 snippets, with the notebook's settings. I ran it on my Mac's CPU rather than its GPU, with the random choices fixed in advance, because a CPU does its arithmetic in exactly the same order every time. That makes the whole run repeatable to the last digit: run it again, and you get the same machine. It takes about 6 minutes.
Here is one token embedding, and one guess, growing:
| Training step | The g token embedding begins | Chance of d after goo |
|---|---|---|
| 0, all random | โ0.048, 0.010, โ0.012, โ0.008, โฆ | 2.0% |
| 100 | โ0.044, 0.013, โ0.002, 0.008, โฆ | 3.8% |
| 300 | โ0.043, 0.015, 0.000, 0.010, โฆ | 3.4% |
| 1,000 | โ0.034, 0.005, โ0.021, 0.010, โฆ | 39.5% |
| 3,000, finished | โ0.049, 0.032, 0.014, 0.030, โฆ | 96.6% |
At step 0, the chance of d after goo is 2.0%, little better than a blind 1-in-65 guess. By step 3,000, it is the 96.6% from the first post's guessing game, and the g token embedding is the one from its step 2. I checked the whole machine, not just these two examples: every one of the 826,433 numbers matches the exhibit exactly. Nobody typed any of them in. They grew.
You can grow it yourself. My follow-along workbook runs this exact training loop, reproduces the table above, and then compares every number it grew with my published exhibit: open it in Colab, or download it. On my Mac Studio's CPU it takes about 6 minutes and matches the exhibit exactly. On a different computer, which does some of its arithmetic in a slightly different order, expect a machine that is very close but not identical to the last digit.
How does it know which way to nudge?
826,433 numbers, and every one of them could be turned up or down. It sounds overwhelming: how could the training code possibly know which ones to turn, and which way?
The answer is that it never searches or guesses. For each number, it asks just one question: if this number went up a tiny bit, would the average surprise go up or down, and how fast? That answer is the number's slope. Then step 6 moves every number a tiny step against its own slope: down if raising it would raise the surprise, up if raising it would lower the surprise.
What makes this possible is that every one of steps 2 to 5 is a chain of small, simple calculations: adding a token embedding to a position embedding, multiplying a vector by a table of weights, sharing out attention, scoring against a row of lm_head, and taking the surprise. Each small calculation comes with a simple rule for passing blame backwards. Take a multiplication, score = a ร b: if the score needs to go up, then a gets blamed in proportion to b, and b in proportion to a.
So the code starts at the surprise score and walks back through the chain: from step 5 to lm_head, back through the four blocks, and finally to the token and position embeddings, passing blame along at each calculation. This is called backpropagation, and it is why PyTorch quietly records every calculation the machine makes while it trains: the records are the chain it walks back along.
One trip back gives every one of the 826,433 numbers its slope. On my Mac's CPU, for a batch of 32 snippets, the trip forward took 83 milliseconds and the trip back took 48. Trying each number up and down, one at a time, would take 826,433 trips.
What does the whole set of slopes look like? It is simply a second copy of the machine: 826,433 numbers, one slope for every parameter, in exactly the same shape. There is a slope vector for every token embedding, a slope grid for every table of weights, and a slope row for every row of lm_head. Step 6 takes the machine and moves each number a little way along its slope.
One way to picture it is a hilly landscape, where your position is set by the 826,433 numbers and the height is the average surprise. The slopes say which way is downhill from exactly where you stand, and each training step takes one small step that way. Nobody can draw a landscape with 826,433 directions, but the idea is the same as walking downhill in fog: you cannot see the valley, but you can always feel which way the ground slopes under your feet.
Following the blame back, with real numbers
Here is the last link of the chain, worked by hand, using my exhibit and one guess from Part 1: the letter after good m. The real next letter is y, and the wheel gave it 40.81%, so the surprise score is ln(1 / 0.4081) = 0.896. Three rules carry the blame back from that surprise score to the rows of lm_head.
Rule 1, the wheel and the surprise score: the blame on each letter's score is its chance, minus 1 for the real letter.
| Letter | Score | Chance | Blame on its score |
|---|---|---|---|
y (the real next letter) | 8.438 | 40.81% | 0.4081 โ 1 = โ0.592 |
e | 8.021 | 26.90% | +0.269 |
a | 7.451 | 15.21% | +0.152 |
o | 6.679 | 7.03% | +0.070 |
z | โ3.737 | 0.00% | +0.000 |
A negative blame means that raising the score would lower the surprise. So y's score should go up, and every other letter's should go down, in proportion to the chance it took. A letter with no chance, like z, gets no blame at all.
Rule 2, multiplying: the blame on one number is scaled by the other number. Each score is a row of lm_head multiplied by the final hidden state, number by number, and added up. The hidden state's first number is โ1.081, so:
- the first number of the
yrow gets โ0.592 ร โ1.081 = +0.640. A positive slope means step 6 turns it down, from โ0.046 towards โ1.081: theyrow is pulled towards the hidden state. - the first number of the
erow gets +0.269 ร โ1.081 = โ0.291, so it is turned up, from โ0.094, away from the hidden state.
Rule 3, adding: the blame passes on unchanged. Each row's bias is simply added to its score, so the y bias gets โ0.592, and is turned up.
The blame keeps going. The hidden state's numbers were multiplied by all 65 rows, so by rule 2 its first number collects each row's first number times that row's blame, added up: โ0.014. From there, the blame passes back through the final normalisation, and then through the four blocks. Every block added its results onto the hidden states, so by rule 3 the blame passes straight through each addition, both into the block and past it: the clear route back that Part 1's "add, never replace" promised.
To check my arithmetic, I worked out these three rules in Python for all 65 rows of lm_head, and compared the results with PyTorch's own loss.backward():
PYTHON
1p = F.softmax(scores, dim=-1) # the wheel: 65 chances 2blame_scores = p.clone() 3blame_scores[target] -= 1 # rule 1: chance, minus 1 for the real letter 4blame_rows = torch.outer(blame_scores, hidden) # rule 2: each row's blame, scaled by the hidden state 5blame_biases = blame_scores # rule 3: added, so passed on unchanged 6blame_hidden = model.lm_head.weight.T @ blame_scores # rule 2 again: back to the hidden state 7 8loss.backward() # PyTorch walks the whole chain 9print((blame_rows - model.lm_head.weight.grad).abs().max())
The biggest difference, across all 8,385 numbers in lm_head, was 0.00000003: rounding. These four lines are the key moves of backpropagation, and loss.backward() repeats them link by link, all the way back to the token embeddings. It knows the chain because PyTorch recorded every calculation on the way forward: 229 records for this one guess, plus one more for each of the 70 named sets of parameters, where their slopes are collected.
Nobody wrote these backward steps for MiniGPT. Every operation the model code uses, such as +, @, F.softmax, and F.cross_entropy, comes with its own blame rule built into PyTorch, so writing the trip forward is all it takes to get the trip back. That is also why Part 1 switched the records off with requires_grad_(False): a machine that is only writing never needs to walk back.
Under the hood: The records PyTorch walks back along
Following the main line back from the surprise score, the records are:
| Record | What it is |
|---|---|
NllLossBackward0, LogSoftmaxBackward0 | step 5 and the wheel: rule 1 |
AddmmBackward0 | lm_head: multiply and add the bias, rules 2 and 3 |
NativeLayerNormBackward0 | the final normalisation |
AddBackward0, 8 times | the 8 additions: attention and the MLP, in each of 4 blocks |
AddBackward0 | the token embedding plus the position embedding |
EmbeddingBackward0 | the g, o, o, d, space, and m token embeddings |
Who gets nudged, and when
Every number gets a slope on every step, but not every slope is the same, and some are exactly zero. Here is every parameter the machine has, all 826,433 of them, with what each one does and what happened on the very first step of growing my exhibit:
| Parameters | What they do | Numbers | Given a direction at step 0 |
|---|---|---|---|
| token embeddings | a vector of 128 numbers for each of the 65 letters | 8,320 | 60 of the 65 |
| position embeddings | a vector for each of the 128 positions | 16,384 | all 128 |
| query, key, and value weights | in each block, three grids, plus biases, that make the queries, keys, and values | 198,144 | all |
| mixing the heads | in each block, one grid, plus biases, that combines what the four heads found | 66,048 | all |
| the MLPs | in each block, two grids, plus biases, that rework each hidden state on its own | 526,848 | all |
| normalising | a stretch and a shift for each of the 128 numbers, before every attention step and MLP, and once at the end | 2,304 | all |
lm_head | a row of 128 numbers, plus a bias, for each of the 65 letters | 8,385 | all 65 rows |
- A token embedding only gets a slope when its letter is in the batch. The 32 snippets in the first step happened to contain no
$,&,3,X, orz, so those 5 token embeddings played no part in any calculation, and their slope was exactly 0. (That does not mean they stood still on later steps: see the watch-it below.) - Every position embedding gets a slope on every step, because every snippet fills all 128 positions.
- Every table of weights, and every normalising parameter, gets a slope on every step, because every hidden state in every snippet passes through every block. How big those slopes are is another matter: the query and key weights' slopes start out tiny, for reasons explained below.
- Every row of
lm_headis nudged on every step, because every guess scores all 65 of them. As the worked example showed, the real next letter's row is pulled towards the hidden state, and every other row is pushed away, harder the bigger the chance it took.
Why the query and key weights wake up late
The query and key weights are nudged on every step, but for the first few hundred steps, hardly at all. Nobody planned that. The training code does exactly the same thing on every step, to every number: loss.backward() works out every slope by the same rules, and optimizer.step() follows them. The late start falls out of the arithmetic. To see how, I followed the blame into attention, measuring the slopes on the same 32 snippets as the exhibit grew.
Attention scores a match by multiplying a query by a key, number by number, and adding up. Multiplying is the rule from the pencil exercise above: when a score is multiplied, the blame for one number is scaled by the other number. So a query number's slope comes down to two things multiplied together:
- How much the surprise cares about the shares of attention. If moving attention from one earlier position to another would not change the guess, nothing gets blamed.
- The size of the key it is matched against. Every query number's slope is scaled by a key number, and every key number's slope by a query number.
At the start, both are tiny:
- The keys and queries start small. Every table of weights starts as tiny random numbers, so every number in a key or query is about 0.23. Every match scores close to 0, so the shares of attention are almost even.
- The easy wins need no looking back. On the first steps, the quickest way to lower the surprise is to learn which letters are common at all. The token embeddings,
lm_head, and the biases can do that without attention, so the surprise fell from 4.17 to 3.51 in 10 steps while the blame on the shares shrank to a fifth. Only when the easy wins run out is looking at the right earlier letters the best way left to lower the surprise, and the blame on the shares grows.
Then it snowballs. As the key weights grow the keys, the query weights' slopes grow; as the query weights grow the queries, the key weights' slopes grow. Each one waits for the other, and then each one speeds the other up:
| Step | Average surprise | Query weights' slope | Key weights' slope | Size of a key number | Biggest share of attention |
|---|---|---|---|---|---|
| 0 | 4.17 | 0.011 | 0.011 | 0.23 | 2% (an even share would be 1.8%) |
| 10 | 3.51 | 0.0025 | 0.0024 | 0.23 | 2% |
| 100 | 2.64 | 0.016 | 0.012 | 0.36 | 3% |
| 300 | 2.46 | 0.060 | 0.082 | 0.54 | 6% |
| 1,000 | 2.04 | 0.080 | 0.161 | 0.95 | 22% |
To check the snowball, I grew the machine again with the key weights frozen at their random start. The keys stayed small, about 0.23 to 0.27, and at step 1,000 the query weights' slope was 0.049 instead of 0.080. Attention's biggest share reached only 11% instead of 22%, and the surprise was 2.20 instead of 2.04.
The wider picture
The student who memorised the textbook
Before training starts, the notebook locks away the last 10% of the Shakespeare text. The machine never practises on it. Every so often, it sits an exam on that locked-away text instead. If its practice scores keep improving while its exam scores get worse, it has started learning the textbook by heart instead of learning how to write. The notebook's bigger machine does exactly that, in section 3.10 below, and the fix is simple: keep a copy of the machine whenever it sets a new best exam score, and use the best copy at the end.
Expensive for computers, cheap for people
Notice who is missing from the learning loop: a teacher. The text marks its own homework. Every time the machine guesses, the answer is simply the next letter of Shakespeare, already sitting there on the page. Nobody has to write questions, check answers, or label anything. All it takes is a pile of text and a computer to grind through it.
The grinding is the expensive part, and it is the learning that takes the time, not the guessing. My small model's 12.3 million practice guesses took 90 seconds on my Mac Studio's GPU; the bigger model's 82 million took 47 minutes. Once a model is trained, writing one new letter is a single guess, which takes a tiny fraction of a second. The models behind today's chatbots play the same learning game on a large slice of the internet, on thousands of GPUs, for weeks or months. That bill is paid in electricity and hardware, not in people's time, which is why it can be scaled up so far. This first stage of learning is the "Pre-trained" in Generative Pre-trained Transformer.
A machine trained only this way is not a chatbot, though. It is a document completer: give it the start of any text, and it writes the most likely continuation. Ask it a question and it may simply carry on writing more questions, because carrying on the text is all it knows how to do. The simplest trick needs no extra training at all: start the text with a pretend conversation, a line beginning "User:" and then a line beginning "Assistant:", so that the most likely way to carry on is to write the assistant's reply. That works, after a fashion. Turning it into something you can properly chat with takes a second stage, and that stage needs people. They write example conversations showing how a helpful assistant should reply, and the machine is trained to copy them. That is the same guessing game, played on a much smaller pile of much more carefully chosen text.
Copying examples only goes so far, so there is usually a third and fourth stage, and they contain the cleverest trick in the whole process. People are shown two of the machine's answers to the same question and asked which is better. Their choices are used to train a second model, a judge, whose only job is to predict which answer people would prefer. Then the chatbot practises: it writes answers, the judge scores them, and the parameters are nudged towards answers the judge scores highly. People compare thousands of answers, and the judge then scores millions, so it stretches their effort a very long way.
Even so, the human work in stages 2 and 3 is slow and costly for every example. These stages use far less data than pre-training, but far more human effort.
There is also a shortcut: let an existing chat model write the example conversations, and train the new model on its answers instead of on people's. This is called distillation, and I wrote about it in Distillation (Part 1) and Distillation (Part 2).
A different kind of distillation, where a small model learns from a bigger model's chances for whatever comes next during pre-training, is the subject of MiniGPT (Part 6). MiniGPT itself stops at pre-training, so everything in this post is the cheap-for-people, expensive-for-computers half.
The jargon decoder
The terms for the whole series are collected in one table, the series glossary.
Here are the comparisons for growing the machine, next to the names the notebook uses. The ones for the machine itself are in the first post.
| What I called it | What the experts call it |
|---|---|
| learning from the text itself, with no people marking answers | self-supervised learning, or pre-training |
| people writing example conversations for the machine to copy | supervised fine-tuning (SFT) |
| people comparing answers, to train a judge | the reward model |
| practising against the judge | reinforcement learning from human feedback (RLHF), often using a method called PPO |
| letting another model do some of the comparing | reinforcement learning from AI feedback (RLAIF) |
| letting an existing model write the examples, or teach its chances | distillation |
| the surprise score | the loss (cross-entropy loss) |
| "as unsure as choosing between N letters" | perplexity |
| working out which way to turn every parameter | backpropagation |
| the 32 snippets for one step | a batch (batch size 32) |
| the slopes of all 826,433 parameters, together | the gradient |
| nudging every number a little in its direction | an optimiser step (here, with AdamW) |
| how far each nudge goes | the learning rate |
| AdamW's running average of recent slopes | momentum |
| shrinking every number very slightly on every step | weight decay |
| the locked-away exam text | the validation set |
| memorising the textbook | overfitting |
| keeping the best copy | checkpoint selection |
| walking downhill on the surprise-score landscape | gradient descent |
Opening the notebook
The setup is in the first post. The notebook's part 1 builds the machine; its parts 2 and 3, below, grow it.
Using the Mac's GPU
Training on a CPU is slow, so for my notebook runs I made one change, so that it would use the Mac Studio's GPU. The notebook's device line only checks for CUDA, NVIDIA's GPU platform, so on a Mac it falls straight through to the CPU. I replaced it with a version that tries Metal Performance Shaders (MPS), Apple's GPU backend for PyTorch, first:
PYTHON
1device = ( 2 "cuda" if torch.cuda.is_available() 3 else "mps" if torch.backends.mps.is_available() 4 else "cpu" 5)
Nothing else needed changing: the notebook's mixed-precision code, which saves time on NVIDIA GPUs by doing some of the arithmetic with fewer digits, only switches on for CUDA, so on MPS it trains in full precision. (The exhibit itself was grown on the CPU, because only the CPU repeats its arithmetic exactly.)
Notebook part 2: the training pipeline
The notebook is in parts of its own, and its section numbers below are its own too. Its part 2 loads the text, turns it into numbers, and trains the model from its part 1 at baseline settings. Loading the text and turning letters into numbers are the first post, so I skip sections 2.1 and 2.2.
2.3 Convert Text to Token Tensor and Split into Train/Validation Sets
The whole text becomes one long list of letter IDs. I keep the first 90% for practice, and the last 10% becomes the locked-away exam text.
PYTHON
1data = torch.tensor(encode(text), dtype=torch.long) 2n = int(0.9 * len(data)) 3train_data, val_data = data[:n], data[n:] # 90 / 10 split
2.5 Create Minibatches for Next-Token Prediction
Each snippet is block_size letters long. The answers are the same snippet shifted one position to the right, so every position has its real next letter.
PYTHON
1def get_batch(split): 2 d = train_data if split == "train" else val_data 3 ix = torch.randint(len(d) - block_size, (batch_size,)) 4 x = torch.stack([d[i : i + block_size] for i in ix]) 5 y = torch.stack([d[i + 1 : i + block_size + 1] for i in ix]) 6 return x.to(device), y.to(device)
If the snippet is goo, the answers are ood.
This is where the four lessons from the first post's "Putting it together" come from: every position in the window is checked against the letter that really came next.
2.9 Training Loop (baseline)
This is the training step from the introduction, in code. Inside the loop, with the progress checks taken out, the six steps are five lines:
PYTHON
1# step 1: 32 snippets of 128 letter IDs, and the real next letter at every position 2xb, yb = get_batch("train") 3 4# steps 2 to 5: embeddings, the blocks, a wheel for every position, and the average surprise score 5logits, loss = model(xb, yb) 6 7# step 6: work out which way to turn every number, then turn it a little 8optimizer.zero_grad(set_to_none=True) 9loss.backward() 10optimizer.step()
get_batch(section 2.5) picks 32 random starting points in the practice text, and takes 128 letter IDs from each asxb.ybis the same 128 letters shifted one place along, so that it holds the real next letter for every position.model(xb, yb)runs exactly the code from Part 1's walk-through: the token and position embeddings, the four blocks, andlm_head, for every position at once. Because it is givenyb, it also does step 5, in the lines Part 1 left out:F.cross_entropylooks up the chance each wheel gave the real next letter and averages the 4,096 surprise scores into one number,loss.loss.backward()is the maths that works backwards from the surprise score, through every calculation, and works out which way to turn each of the 826,433 numbers.optimizer.zero_gradfirst clears the directions left over from the last step.optimizer.step()is the nudge: AdamW turns every number a little in its direction.
The rest of the loop only records progress: every few hundred steps, estimate_loss (section 2.8) scores the machine on both the practice text and the locked-away exam text, without nudging anything.
The first configuration is deliberately tiny: the exhibit's settings.
| Setting | Baseline |
|---|---|
| Layers / heads / embedding dim | 4 / 4 / 128 |
| Context length | 128 |
| Parameters | 826,433 |
| Batch size | 32 |
| Optimiser | AdamW, fixed learning rate 3 ร 10โปโด |
| Iterations | 3,000 |
| Checkpoint | final step |
Both surprise scores, on the practice text and on the exam text, start near 4.20, roughly an even wheel's 4.17, and then fall steadily.
2.10 Plot Training and Validation Loss (baseline)
By step 3,000, my exam score was 1.71: as unsure as choosing between about 5.5 letters. The run took 89 seconds on my Mac Studio's GPU. The output is not good Shakespeare yet, but every part of the pipeline works.
Notebook part 3: a bigger machine
The notebook's part 3 keeps the same text and model code, but grows a much bigger machine, keeps the best copy, and judges the result by reading what it writes.
3.3 Stronger Hyperparameter Configuration
The second configuration is close to nanoGPT's small Shakespeare setup.
| Setting | Stronger |
|---|---|
| Layers / heads / embedding dim | 6 / 6 / 384 |
| Context length | 256 |
| Parameters | 10.77M |
| Batch size | 64 |
| Dropout | 0.2 |
| Optimiser | AdamW, betas (0.9, 0.99), weight decay 0.1 on 2โD tensors only |
| Learning rate | 100 warmup steps to 10โปยณ, cosine decay to 10โปโด over 5,000 steps |
| Also | gradient clipping at 1.0, and weight tying: the token embeddings double as the rows of lm_head |
| Checkpoint | best validation loss |
In plain words, the new settings are:
- Bigger steps that change over time. The learning rate starts small and grows over the first 100 steps (warmup), so that the random starting numbers are not knocked about too hard, then shrinks gradually along a curve (cosine decay) for fine adjustments at the end.
- A cap on the slopes. If the slopes on one step are unusually big, they are scaled down before the nudge (gradient clipping), so that one odd batch cannot throw the machine off course.
- Stronger weight decay, ten times stronger than the small model's, and only on the tables of weights and embeddings, not on the biases and normalising parameters.
betasset how long AdamW's momentum remembers earlier slopes.- More dropout: during training, a fifth of the numbers are switched off at random on every step, to make the machine harder to memorise with.
3.9 Training Loop with Best Checkpoint
The exam score starts at 4.29 and drops fast. It was best at step 1,500, at 1.47: as unsure as choosing between about 4.3 letters. The whole run took 47 minutes on my Mac Studio's GPU.
3.10 Plot Training Loss Curves
This is the student who memorised the textbook, caught in the act.
What happens after step 1,500 is the most useful part of the experiment. The practice score keeps improving, down to 0.61 by step 5,000, but the exam score climbs back to 1.71, no better than the tiny baseline's. The machine is memorising the practice text. The last copy is not the best one, which is why the notebook keeps the copy with the best exam score.
3.13 Generate High-Quality Samples
Prompted with ROMEO:, at the notebook's temperature of 0.8, my best copy wrote this (shortened):
CODE
1ROMEO: 2And this true sluck of voices, and kills thee 3To all the sun of the city the People: 4Your blood standing am I am about to have. 5 6JULIET: 7These grapes of you, noble like my brother did not 8To give the butcher of the noble gentleman: if you live 9But be his mind own so time, if he receive your brother's death. 10 11ROMEO: 12Ay, if with him, he conceal'd me.
It is not coherent, but the shape is unmistakable: speaker names in capitals, colons, line breaks, verse-like line lengths, plausible Shakespearean vocabulary. A machine that only ever sees one letter at a time picked all of that up from 1 MB of text, in under an hour.
My runs against the paper's
| The paper (Colab A100) | My runs (Mac Studio, MPS) | |
|---|---|---|
| Baseline, step 3,000: practice and exam scores | 1.5304 and 1.7236 | 1.5306 and 1.7126 |
| Baseline time | 50.79 seconds | 89.35 seconds |
| Bigger machine: best step | 1,750 | 1,500 |
| Bigger machine: best exam score | 1.4780 | 1.4679 |
| Bigger machine: whole run | 4.76 minutes | 46.98 minutes |
What I took from it
Running MiniGPT end to end took under an hour on my own machine and cost nothing. A few things stuck with me:
- Growing is writing, plus two steps. Check the answer, and nudge every number. Everything else is the machine from Part 1.
- Nobody decides anything. Every number follows its own slope, and behaviour like attention waking up late emerges from the arithmetic.
- Overfitting is not an edge case. The stronger model's best validation loss arrived at step 1,500 out of 5,000 on my run (step 1,750 in the paper). Without validation-based checkpoint selection the notebook would have shipped a worse model that scored well on its own training text.
- Scale is doing the heavy lifting elsewhere. MiniGPT is honest that it is a reproducibility study, not a competitive model. The same training idea, with far more data, compute, and parameters, is what produces the models I use every day.
To get a feel for "far more", here is my small model next to GPT-3, the 2020 model whose family the first ChatGPT grew out of, using the figures from its paper:
| My small MiniGPT | GPT-3 (2020) | How much bigger | |
|---|---|---|---|
| Parameters | 826,433 | 175 billion | about 200,000 times |
| Practice text | 1 million letters | about 300 billion pieces of words, roughly 1.2 trillion letters | about 1 million times |
| Blocks | 4 | 96 | 24 times |
| Numbers in each vector | 128 | 12,288 | 96 times |
If all of Tiny Shakespeare were one book on a shelf, GPT-3's practice text would fill a shelf tens of kilometres long. And the models behind today's chatbots are bigger again, though most companies no longer publish their sizes.
The value of a paper like this is not a benchmark number. It is that the path from raw text to generated samples is now something I have run rather than something I have read about.
Try it yourself
- My follow-along workbook: open it in Colab. It grows the exhibit from random numbers and reproduces this post's numbers
- Running the finished model: the first post's workbook
- The notebook: github.com/jibin10/MiniGPT




