MiniGPT (Part 2)

Neil Haddley โ€ข October 5, 2026

How a GPT is grown: retraining my MiniGPT exhibit model from random numbers on a Mac, one guess at a time, and watching every embedding and parameter from the first post appear

AIgpttransformerspytorchnanogptmachine-learning

In the first post, I took a trained MiniGPT apart while it was running: the token and position embeddings, attention, the MLPs, and lm_head with its 65 rows. Every number in it came from one model, my exhibit. This post answers the obvious next question: where did those numbers come from? Nobody typed them in. They were grown, and here I grow them again from scratch, on my Mac, and end up with exactly the same numbers.

In Part 1, writing one letter took five steps: letters to numbers, embeddings, the blocks, chances, and spin the wheel. Growing the machine uses exactly the same steps, with step 5 changed and one step added. So if Part 1 made sense, most of this post will already feel familiar.

One promise, the same as last time: no magic. Every number in this post is either worked out in front of you, or comes from a run on my own Mac Studio.

This post's machinesThe exhibitThe bigger machine
What changedgrown from random numbersbigger, in the notebook's part 3
TextTiny ShakespeareTiny Shakespeare
Piecesletters: 65letters: 65
Blocks4, each 4 heads6, each 6 heads
Vector size128384
Positions128, position embeddings256, position embeddings
EnginePyTorchPyTorch
Size826,433 numbers10.77 million, token embeddings doubling as lm_head's rows
Score1.70 surprise per letter1.47

Growing a GPT, in plain English

Starting from nothing

Build Part 1's machine, but put random numbers in every embedding and every table of weights, and run its five steps. Step 4 then gives an almost even wheel: every slice about the same size, 1 in 65. So what it writes is pure noise. I asked the untrained MiniGPT model from the notebook to carry on from ROMEO:, and this is what it wrote:

CODE
1ROMEO:dHrNUBK.!oWffHGBTyysS:hvnRyuZtONMQhv,dAw$xBM.VQu.!nhulODzyrBM;sta'fy,dlAvIzRlwMl

Every letter in that line was a near-random spin. The untrained model's chances were all between about 1% and 3%, close to an even 1 in 65.

That is the whole secret of training. The machine is never taught a rule like "a capital letter starts a name" or "q is followed by u". Training only reshapes the wheel, one tiny nudge at a time, until the slices for letters that tend to come next in Shakespeare are bigger than the rest. Shape it well enough, and the spins start to look like writing.

Keeping score: the surprise score

How do you tell a good forecaster from a bad one? You wait to see what the weather actually does, and then you look at the chance they gave it. A forecaster who said "90% chance of rain" on a day it rained did well. One who said "5% chance of rain" on a day it poured did badly. A good forecaster should not be surprised.

The machine is marked in exactly the same way. Once the real next letter is revealed, I look up the chance the machine gave that one letter, ignore all the others, and turn it into a surprise score: a number for how surprised the machine should be by what actually happened. A high chance gives a low surprise score, and a low chance gives a high one.

How the surprise score is worked out: the halving rule. The surprise score follows one simple rule:

  • If the machine gave the right letter 100%, the surprise score is 0. It was certain, and it was right, so it is not surprised at all.
  • Every time that chance is cut in half, add the same amount to the surprise score: 0.69.

So the surprise score climbs one step of 0.69 for every halving:

Chance the machine gave the right letterHalvings from 100%Surprise score
100%00
50%10.69
25%21.39
12.5%32.08
6.25%42.77
3.1%53.47
1.6% (1 in 64)64.16

Chances that fall between two rows get a surprise score between those two rows. Take my trained model's real chances after good m, as in good my lord or good madam: y 40.8%, e 26.9%, and a 15.2%.

  • If the next letter turns out to be y, the machine gave it 40.8%. That is a little less than 50%, so the surprise score is a little more than one step: 0.90.
  • If it turns out to be a, the machine gave it only 15.2%. That falls between 25% and 12.5%, so the surprise score lands between 1.39 and 2.08: 1.88. A bigger surprise.
The halving rule drawn as a curve. Near 100% the surprise score barely moves; near 0% it shoots upScroll sideways, or tap the diagram to open it full size

The halving rule drawn as a curve. Near 100% the surprise score barely moves; near 0% it shoots up

Why add 0.69 for each halving, and not a round 1? There is no deep reason. Adding 1 for each halving would work just as well: every surprise score would come out about 1.44 times bigger, and the machine would learn exactly the same thing. The maths in the notebook happens to use 0.69, so I use it too, which means the surprise scores in this post match the numbers the notebook prints.

A score of 0 means the machine gave the right answer 100%: it was certain, and it was right. It cannot do better than that. In the other direction there is no limit: every further halving adds another 0.69, so give almost 0% to the letter that actually turns up, and the surprise score shoots up. That is how the game punishes a machine for being confidently wrong.

This one number is what the whole project is about. Training means making the machine's average surprise score, over millions of guesses, as small as possible.

A machine that knows nothing at all gives the same 1-in-65 chance to every letter. 1 in 65 is almost exactly 1 in 64, the bottom row of the table: six halvings from 100%. So whatever comes next, its surprise score is about 6 ร— 0.69, or 4.17 to be exact. That is almost exactly where my exhibit started: 4.18 on its very first step.

Under the hood: Where the surprise numbers come from

The halving rule is exactly what a logarithm does. The surprise score is the natural logarithm of one over the chance, written ln(1 รท chance), and every scientific calculator has an ln button. The step of 0.69 is ln(2): the cost of one halving.

  • A chance of 40.8%: ln(1 รท 0.408) = 0.90
  • A chance of 15.2%: ln(1 รท 0.152) = 1.88
  • A chance of 1 in 65: ln(65) = 4.17
  • A chance of 100%: ln(1) = 0

To add 1 for each halving instead, swap ln for logโ‚‚, the logarithm that counts halvings directly. Then the blind guess scores logโ‚‚(65) = 6.02: just over the six halvings in the table.

There is a friendlier way to read the surprise score. Turn it back into a sentence (in the maths, raise e, about 2.718, to the power of the surprise score): "the machine was as unsure as if it were choosing between N equally likely letters". At the start, N is 65. After my small model had trained for 90 seconds, N was about 5.5. After the bigger model, it was about 4.3. That is the whole scoreboard for this project: from 65 down to 4.

Growing it: the same five steps, then check and nudge

Everything Part 1 called fixed, the token embeddings, the position embeddings, every table of weights in the blocks, the stretch and shift in every normalisation, and the rows of lm_head, comes down to numbers, the parameters. I picture each one as a dial. The small model has 826,433 parameters, and the bigger model has 10.77 million. Growing the machine means finding good settings for all of them, and it uses the five steps from Part 1, with one change and one addition:

One training step. Steps 1 to 4 are exactly the steps the machine takes when it writes; step 5 checks instead of spinning, and step 6 is newScroll sideways, or tap the diagram to open it full size

One training step. Steps 1 to 4 are exactly the steps the machine takes when it writes; step 5 checks instead of spinning, and step 6 is new

StepWhen the machine writes (Part 1)When the machine grows
1. Letters to numbersthe text so far32 random snippets of practice text, each 128 letters long: exactly enough to fill every position. The 32 snippets for one step are called a batch
2. Embeddingsa token embedding plus a position embedding, for each letterexactly the same, except that the embeddings start as random numbers
3. The blocksattention, then the MLP, four timesexactly the same
4. Chancesa wheel for the last position onlya wheel for every position: 32 ร— 128 = 4,096 wheels, because in practice text, every position's next letter is already known
5.spin the wheel, and write whatever comes upcheck the answer: look up the chance each wheel gave the real next letter, turn it into a surprise score, and average all 4,096
6.nudge every one of the 826,433 fixed numbers a tiny amount, in whichever direction would make that average surprise smaller

Each pass through all six steps is one training step. Then the machine does it all again with 32 new snippets: 3,000 times in all.

Here is the very first step of growing my exhibit, with the real numbers:

  • Steps 1 to 4. One of the 32 snippets begins " guess who caused yo". Its first hidden state has seen only the space, and its wheel gives the real next letter, g, a chance of 0.87%.
  • Step 5. That one guess has a surprise score of 4.74. Averaged over all 4,096 guesses, the surprise is 4.18: almost exactly the 4.17 of an even wheel, because the machine knows nothing yet.
  • Step 6. Almost every one of the 826,433 numbers moves by 0.0003. That distance is a setting I chose, called the learning rate: how far each nudge goes. The g token embedding's first number goes from โˆ’0.04793 to โˆ’0.04823, the position 1 embedding's from 0.02622 to 0.02592, and the first number of lm_head's d row from 0.01135 to 0.01105. (On the very first step, every number with a slope moves almost exactly the same distance, because the training method, AdamW, starts with equal-sized steps. Later on, it sizes each number's step separately. The exceptions are 5 token embeddings, which you will meet below.)
Watching it grow

To see what those steps do, I grew the exhibit model from the first post again, and stopped it five times along the way to write from the same prompt. Nothing changes between the frames except the parameters:

From babbling to Shakespeare in one training run. After 100 steps it has found spaces and line breaks; after 1,000, short words and a speaker's name; after 3,000, lines that look like ShakespeareScroll sideways, or tap the diagram to open it full size

From babbling to Shakespeare in one training run. After 100 steps it has found spaces and line breaks; after 1,000, short words and a speaker's name; after 3,000, lines that look like Shakespeare

This is what I meant at the start by a model being grown rather than built. Nobody wrote anything into the parameters between those frames. The guessing game did it.

The payoff: growing the exhibit, exactly

Every number in the first post came from one trained model, my exhibit. Here is where they came from. The exhibit is exactly the training loop above: 3,000 steps of 32 snippets, with the notebook's settings. I ran it on my Mac's CPU rather than its GPU, with the random choices fixed in advance, because a CPU does its arithmetic in exactly the same order every time. That makes the whole run repeatable to the last digit: run it again, and you get the same machine. It takes about 6 minutes.

Here is one token embedding, and one guess, growing:

Training stepThe g token embedding beginsChance of d after goo
0, all randomโˆ’0.048, 0.010, โˆ’0.012, โˆ’0.008, โ€ฆ2.0%
100โˆ’0.044, 0.013, โˆ’0.002, 0.008, โ€ฆ3.8%
300โˆ’0.043, 0.015, 0.000, 0.010, โ€ฆ3.4%
1,000โˆ’0.034, 0.005, โˆ’0.021, 0.010, โ€ฆ39.5%
3,000, finishedโˆ’0.049, 0.032, 0.014, 0.030, โ€ฆ96.6%

At step 0, the chance of d after goo is 2.0%, little better than a blind 1-in-65 guess. By step 3,000, it is the 96.6% from the first post's guessing game, and the g token embedding is the one from its step 2. I checked the whole machine, not just these two examples: every one of the 826,433 numbers matches the exhibit exactly. Nobody typed any of them in. They grew.

You can grow it yourself. My follow-along workbook runs this exact training loop, reproduces the table above, and then compares every number it grew with my published exhibit: open it in Colab, or download it. On my Mac Studio's CPU it takes about 6 minutes and matches the exhibit exactly. On a different computer, which does some of its arithmetic in a slightly different order, expect a machine that is very close but not identical to the last digit.

How does it know which way to nudge?

826,433 numbers, and every one of them could be turned up or down. It sounds overwhelming: how could the training code possibly know which ones to turn, and which way?

The answer is that it never searches or guesses. For each number, it asks just one question: if this number went up a tiny bit, would the average surprise go up or down, and how fast? That answer is the number's slope. Then step 6 moves every number a tiny step against its own slope: down if raising it would raise the surprise, up if raising it would lower the surprise.

What makes this possible is that every one of steps 2 to 5 is a chain of small, simple calculations: adding a token embedding to a position embedding, multiplying a vector by a table of weights, sharing out attention, scoring against a row of lm_head, and taking the surprise. Each small calculation comes with a simple rule for passing blame backwards. Take a multiplication, score = a ร— b: if the score needs to go up, then a gets blamed in proportion to b, and b in proportion to a.

So the code starts at the surprise score and walks back through the chain: from step 5 to lm_head, back through the four blocks, and finally to the token and position embeddings, passing blame along at each calculation. This is called backpropagation, and it is why PyTorch quietly records every calculation the machine makes while it trains: the records are the chain it walks back along.

One trip back gives every one of the 826,433 numbers its slope. On my Mac's CPU, for a batch of 32 snippets, the trip forward took 83 milliseconds and the trip back took 48. Trying each number up and down, one at a time, would take 826,433 trips.

What does the whole set of slopes look like? It is simply a second copy of the machine: 826,433 numbers, one slope for every parameter, in exactly the same shape. There is a slope vector for every token embedding, a slope grid for every table of weights, and a slope row for every row of lm_head. Step 6 takes the machine and moves each number a little way along its slope.

One way to picture it is a hilly landscape, where your position is set by the 826,433 numbers and the height is the average surprise. The slopes say which way is downhill from exactly where you stand, and each training step takes one small step that way. Nobody can draw a landscape with 826,433 directions, but the idea is the same as walking downhill in fog: you cannot see the valley, but you can always feel which way the ground slopes under your feet.

Following the blame back, with real numbers

Here is the last link of the chain, worked by hand, using my exhibit and one guess from Part 1: the letter after good m. The real next letter is y, and the wheel gave it 40.81%, so the surprise score is ln(1 / 0.4081) = 0.896. Three rules carry the blame back from that surprise score to the rows of lm_head.

The whole trip in one picture, with the numbers worked out below. The blame lights up from the surprise score back to the token embeddings, and then the picture replaysScroll sideways, or tap the diagram to open it full size

The whole trip in one picture, with the numbers worked out below. The blame lights up from the surprise score back to the token embeddings, and then the picture replays

Rule 1, the wheel and the surprise score: the blame on each letter's score is its chance, minus 1 for the real letter.

LetterScoreChanceBlame on its score
y (the real next letter)8.43840.81%0.4081 โˆ’ 1 = โˆ’0.592
e8.02126.90%+0.269
a7.45115.21%+0.152
o6.6797.03%+0.070
zโˆ’3.7370.00%+0.000

A negative blame means that raising the score would lower the surprise. So y's score should go up, and every other letter's should go down, in proportion to the chance it took. A letter with no chance, like z, gets no blame at all.

Rule 2, multiplying: the blame on one number is scaled by the other number. Each score is a row of lm_head multiplied by the final hidden state, number by number, and added up. The hidden state's first number is โˆ’1.081, so:

  • the first number of the y row gets โˆ’0.592 ร— โˆ’1.081 = +0.640. A positive slope means step 6 turns it down, from โˆ’0.046 towards โˆ’1.081: the y row is pulled towards the hidden state.
  • the first number of the e row gets +0.269 ร— โˆ’1.081 = โˆ’0.291, so it is turned up, from โˆ’0.094, away from the hidden state.

Rule 3, adding: the blame passes on unchanged. Each row's bias is simply added to its score, so the y bias gets โˆ’0.592, and is turned up.

The blame keeps going. The hidden state's numbers were multiplied by all 65 rows, so by rule 2 its first number collects each row's first number times that row's blame, added up: โˆ’0.014. From there, the blame passes back through the final normalisation, and then through the four blocks. Every block added its results onto the hidden states, so by rule 3 the blame passes straight through each addition, both into the block and past it: the clear route back that Part 1's "add, never replace" promised.

To check my arithmetic, I worked out these three rules in Python for all 65 rows of lm_head, and compared the results with PyTorch's own loss.backward():

PYTHON
1p = F.softmax(scores, dim=-1)                       # the wheel: 65 chances
2blame_scores = p.clone()
3blame_scores[target] -= 1                           # rule 1: chance, minus 1 for the real letter
4blame_rows = torch.outer(blame_scores, hidden)      # rule 2: each row's blame, scaled by the hidden state
5blame_biases = blame_scores                         # rule 3: added, so passed on unchanged
6blame_hidden = model.lm_head.weight.T @ blame_scores  # rule 2 again: back to the hidden state
7
8loss.backward()                                     # PyTorch walks the whole chain
9print((blame_rows - model.lm_head.weight.grad).abs().max())

The biggest difference, across all 8,385 numbers in lm_head, was 0.00000003: rounding. These four lines are the key moves of backpropagation, and loss.backward() repeats them link by link, all the way back to the token embeddings. It knows the chain because PyTorch recorded every calculation on the way forward: 229 records for this one guess, plus one more for each of the 70 named sets of parameters, where their slopes are collected.

Nobody wrote these backward steps for MiniGPT. Every operation the model code uses, such as +, @, F.softmax, and F.cross_entropy, comes with its own blame rule built into PyTorch, so writing the trip forward is all it takes to get the trip back. That is also why Part 1 switched the records off with requires_grad_(False): a machine that is only writing never needs to walk back.

Under the hood: The records PyTorch walks back along

Following the main line back from the surprise score, the records are:

RecordWhat it is
NllLossBackward0, LogSoftmaxBackward0step 5 and the wheel: rule 1
AddmmBackward0lm_head: multiply and add the bias, rules 2 and 3
NativeLayerNormBackward0the final normalisation
AddBackward0, 8 timesthe 8 additions: attention and the MLP, in each of 4 blocks
AddBackward0the token embedding plus the position embedding
EmbeddingBackward0the g, o, o, d, space, and m token embeddings
Who gets nudged, and when

Every number gets a slope on every step, but not every slope is the same, and some are exactly zero. Here is every parameter the machine has, all 826,433 of them, with what each one does and what happened on the very first step of growing my exhibit:

ParametersWhat they doNumbersGiven a direction at step 0
token embeddingsa vector of 128 numbers for each of the 65 letters8,32060 of the 65
position embeddingsa vector for each of the 128 positions16,384all 128
query, key, and value weightsin each block, three grids, plus biases, that make the queries, keys, and values198,144all
mixing the headsin each block, one grid, plus biases, that combines what the four heads found66,048all
the MLPsin each block, two grids, plus biases, that rework each hidden state on its own526,848all
normalisinga stretch and a shift for each of the 128 numbers, before every attention step and MLP, and once at the end2,304all
lm_heada row of 128 numbers, plus a bias, for each of the 65 letters8,385all 65 rows
  • A token embedding only gets a slope when its letter is in the batch. The 32 snippets in the first step happened to contain no $, &, 3, X, or z, so those 5 token embeddings played no part in any calculation, and their slope was exactly 0. (That does not mean they stood still on later steps: see the watch-it below.)
  • Every position embedding gets a slope on every step, because every snippet fills all 128 positions.
  • Every table of weights, and every normalising parameter, gets a slope on every step, because every hidden state in every snippet passes through every block. How big those slopes are is another matter: the query and key weights' slopes start out tiny, for reasons explained below.
  • Every row of lm_head is nudged on every step, because every guess scores all 65 of them. As the worked example showed, the real next letter's row is pulled towards the hidden state, and every other row is pushed away, harder the bigger the chance it took.
Why the query and key weights wake up late

The query and key weights are nudged on every step, but for the first few hundred steps, hardly at all. Nobody planned that. The training code does exactly the same thing on every step, to every number: loss.backward() works out every slope by the same rules, and optimizer.step() follows them. The late start falls out of the arithmetic. To see how, I followed the blame into attention, measuring the slopes on the same 32 snippets as the exhibit grew.

Attention scores a match by multiplying a query by a key, number by number, and adding up. Multiplying is the rule from the pencil exercise above: when a score is multiplied, the blame for one number is scaled by the other number. So a query number's slope comes down to two things multiplied together:

  1. How much the surprise cares about the shares of attention. If moving attention from one earlier position to another would not change the guess, nothing gets blamed.
  2. The size of the key it is matched against. Every query number's slope is scaled by a key number, and every key number's slope by a query number.

At the start, both are tiny:

  • The keys and queries start small. Every table of weights starts as tiny random numbers, so every number in a key or query is about 0.23. Every match scores close to 0, so the shares of attention are almost even.
  • The easy wins need no looking back. On the first steps, the quickest way to lower the surprise is to learn which letters are common at all. The token embeddings, lm_head, and the biases can do that without attention, so the surprise fell from 4.17 to 3.51 in 10 steps while the blame on the shares shrank to a fifth. Only when the easy wins run out is looking at the right earlier letters the best way left to lower the surprise, and the blame on the shares grows.

Then it snowballs. As the key weights grow the keys, the query weights' slopes grow; as the query weights grow the queries, the key weights' slopes grow. Each one waits for the other, and then each one speeds the other up:

StepAverage surpriseQuery weights' slopeKey weights' slopeSize of a key numberBiggest share of attention
04.170.0110.0110.232% (an even share would be 1.8%)
103.510.00250.00240.232%
1002.640.0160.0120.363%
3002.460.0600.0820.546%
1,0002.040.0800.1610.9522%

To check the snowball, I grew the machine again with the key weights frozen at their random start. The keys stayed small, about 0.23 to 0.27, and at step 1,000 the query weights' slope was 0.049 instead of 0.080. Attention's biggest share reached only 11% instead of 22%, and the surprise was 2.20 instead of 2.04.

The wider picture

The student who memorised the textbook

Before training starts, the notebook locks away the last 10% of the Shakespeare text. The machine never practises on it. Every so often, it sits an exam on that locked-away text instead. If its practice scores keep improving while its exam scores get worse, it has started learning the textbook by heart instead of learning how to write. The notebook's bigger machine does exactly that, in section 3.10 below, and the fix is simple: keep a copy of the machine whenever it sets a new best exam score, and use the best copy at the end.

Expensive for computers, cheap for people

Notice who is missing from the learning loop: a teacher. The text marks its own homework. Every time the machine guesses, the answer is simply the next letter of Shakespeare, already sitting there on the page. Nobody has to write questions, check answers, or label anything. All it takes is a pile of text and a computer to grind through it.

The grinding is the expensive part, and it is the learning that takes the time, not the guessing. My small model's 12.3 million practice guesses took 90 seconds on my Mac Studio's GPU; the bigger model's 82 million took 47 minutes. Once a model is trained, writing one new letter is a single guess, which takes a tiny fraction of a second. The models behind today's chatbots play the same learning game on a large slice of the internet, on thousands of GPUs, for weeks or months. That bill is paid in electricity and hardware, not in people's time, which is why it can be scaled up so far. This first stage of learning is the "Pre-trained" in Generative Pre-trained Transformer.

A machine trained only this way is not a chatbot, though. It is a document completer: give it the start of any text, and it writes the most likely continuation. Ask it a question and it may simply carry on writing more questions, because carrying on the text is all it knows how to do. The simplest trick needs no extra training at all: start the text with a pretend conversation, a line beginning "User:" and then a line beginning "Assistant:", so that the most likely way to carry on is to write the assistant's reply. That works, after a fashion. Turning it into something you can properly chat with takes a second stage, and that stage needs people. They write example conversations showing how a helpful assistant should reply, and the machine is trained to copy them. That is the same guessing game, played on a much smaller pile of much more carefully chosen text.

Copying examples only goes so far, so there is usually a third and fourth stage, and they contain the cleverest trick in the whole process. People are shown two of the machine's answers to the same question and asked which is better. Their choices are used to train a second model, a judge, whose only job is to predict which answer people would prefer. Then the chatbot practises: it writes answers, the judge scores them, and the parameters are nudged towards answers the judge scores highly. People compare thousands of answers, and the judge then scores millions, so it stretches their effort a very long way.

The four stages from a document completer to an assistant. MiniGPT only does stage 1. The purple boxes are the shortcuts, where another model stands in for peopleScroll sideways, or tap the diagram to open it full size

The four stages from a document completer to an assistant. MiniGPT only does stage 1. The purple boxes are the shortcuts, where another model stands in for people

Even so, the human work in stages 2 and 3 is slow and costly for every example. These stages use far less data than pre-training, but far more human effort.

There is also a shortcut: let an existing chat model write the example conversations, and train the new model on its answers instead of on people's. This is called distillation, and I wrote about it in Distillation (Part 1) and Distillation (Part 2).

A different kind of distillation, where a small model learns from a bigger model's chances for whatever comes next during pre-training, is the subject of MiniGPT (Part 6). MiniGPT itself stops at pre-training, so everything in this post is the cheap-for-people, expensive-for-computers half.

The jargon decoder

The terms for the whole series are collected in one table, the series glossary.

Here are the comparisons for growing the machine, next to the names the notebook uses. The ones for the machine itself are in the first post.

What I called itWhat the experts call it
learning from the text itself, with no people marking answersself-supervised learning, or pre-training
people writing example conversations for the machine to copysupervised fine-tuning (SFT)
people comparing answers, to train a judgethe reward model
practising against the judgereinforcement learning from human feedback (RLHF), often using a method called PPO
letting another model do some of the comparingreinforcement learning from AI feedback (RLAIF)
letting an existing model write the examples, or teach its chancesdistillation
the surprise scorethe loss (cross-entropy loss)
"as unsure as choosing between N letters"perplexity
working out which way to turn every parameterbackpropagation
the 32 snippets for one stepa batch (batch size 32)
the slopes of all 826,433 parameters, togetherthe gradient
nudging every number a little in its directionan optimiser step (here, with AdamW)
how far each nudge goesthe learning rate
AdamW's running average of recent slopesmomentum
shrinking every number very slightly on every stepweight decay
the locked-away exam textthe validation set
memorising the textbookoverfitting
keeping the best copycheckpoint selection
walking downhill on the surprise-score landscapegradient descent

Opening the notebook

The setup is in the first post. The notebook's part 1 builds the machine; its parts 2 and 3, below, grow it.

Using the Mac's GPU

Training on a CPU is slow, so for my notebook runs I made one change, so that it would use the Mac Studio's GPU. The notebook's device line only checks for CUDA, NVIDIA's GPU platform, so on a Mac it falls straight through to the CPU. I replaced it with a version that tries Metal Performance Shaders (MPS), Apple's GPU backend for PyTorch, first:

PYTHON
1device = (
2    "cuda" if torch.cuda.is_available()
3    else "mps" if torch.backends.mps.is_available()
4    else "cpu"
5)

Nothing else needed changing: the notebook's mixed-precision code, which saves time on NVIDIA GPUs by doing some of the arithmetic with fewer digits, only switches on for CUDA, so on MPS it trains in full precision. (The exhibit itself was grown on the CPU, because only the CPU repeats its arithmetic exactly.)

Notebook part 2: the training pipeline

The notebook is in parts of its own, and its section numbers below are its own too. Its part 2 loads the text, turns it into numbers, and trains the model from its part 1 at baseline settings. Loading the text and turning letters into numbers are the first post, so I skip sections 2.1 and 2.2.

2.3 Convert Text to Token Tensor and Split into Train/Validation Sets

The whole text becomes one long list of letter IDs. I keep the first 90% for practice, and the last 10% becomes the locked-away exam text.

PYTHON
1data = torch.tensor(encode(text), dtype=torch.long)
2n = int(0.9 * len(data))
3train_data, val_data = data[:n], data[n:]   # 90 / 10 split
The notebook printed the 65 letters and the practice and exam splitTap the image to open it full size

The notebook printed the 65 letters and the practice and exam split

2.5 Create Minibatches for Next-Token Prediction

Each snippet is block_size letters long. The answers are the same snippet shifted one position to the right, so every position has its real next letter.

PYTHON
1def get_batch(split):
2    d = train_data if split == "train" else val_data
3    ix = torch.randint(len(d) - block_size, (batch_size,))
4    x = torch.stack([d[i : i + block_size] for i in ix])
5    y = torch.stack([d[i + 1 : i + block_size + 1] for i in ix])
6    return x.to(device), y.to(device)

If the snippet is goo, the answers are ood.

This is where the four lessons from the first post's "Putting it together" come from: every position in the window is checked against the letter that really came next.

2.9 Training Loop (baseline)

This is the training step from the introduction, in code. Inside the loop, with the progress checks taken out, the six steps are five lines:

PYTHON
1# step 1: 32 snippets of 128 letter IDs, and the real next letter at every position
2xb, yb = get_batch("train")
3
4# steps 2 to 5: embeddings, the blocks, a wheel for every position, and the average surprise score
5logits, loss = model(xb, yb)
6
7# step 6: work out which way to turn every number, then turn it a little
8optimizer.zero_grad(set_to_none=True)
9loss.backward()
10optimizer.step()
  • get_batch (section 2.5) picks 32 random starting points in the practice text, and takes 128 letter IDs from each as xb. yb is the same 128 letters shifted one place along, so that it holds the real next letter for every position.
  • model(xb, yb) runs exactly the code from Part 1's walk-through: the token and position embeddings, the four blocks, and lm_head, for every position at once. Because it is given yb, it also does step 5, in the lines Part 1 left out: F.cross_entropy looks up the chance each wheel gave the real next letter and averages the 4,096 surprise scores into one number, loss.
  • loss.backward() is the maths that works backwards from the surprise score, through every calculation, and works out which way to turn each of the 826,433 numbers. optimizer.zero_grad first clears the directions left over from the last step.
  • optimizer.step() is the nudge: AdamW turns every number a little in its direction.

The rest of the loop only records progress: every few hundred steps, estimate_loss (section 2.8) scores the machine on both the practice text and the locked-away exam text, without nudging anything.

The first configuration is deliberately tiny: the exhibit's settings.

SettingBaseline
Layers / heads / embedding dim4 / 4 / 128
Context length128
Parameters826,433
Batch size32
OptimiserAdamW, fixed learning rate 3 ร— 10โปโด
Iterations3,000
Checkpointfinal step

Both surprise scores, on the practice text and on the exam text, start near 4.20, roughly an even wheel's 4.17, and then fall steadily.

The baseline training log: the surprise score on the practice and exam text, every few hundred stepsTap the image to open it full size

The baseline training log: the surprise score on the practice and exam text, every few hundred steps

2.10 Plot Training and Validation Loss (baseline)
My reproduction of the paper's Figure 1: the practice and exam scores fall to about 1.53 and 1.71 by step 3,000, with no sign of memorisingTap the image to open it full size

My reproduction of the paper's Figure 1: the practice and exam scores fall to about 1.53 and 1.71 by step 3,000, with no sign of memorising

By step 3,000, my exam score was 1.71: as unsure as choosing between about 5.5 letters. The run took 89 seconds on my Mac Studio's GPU. The output is not good Shakespeare yet, but every part of the pipeline works.

Notebook part 3: a bigger machine

The notebook's part 3 keeps the same text and model code, but grows a much bigger machine, keeps the best copy, and judges the result by reading what it writes.

3.3 Stronger Hyperparameter Configuration

The second configuration is close to nanoGPT's small Shakespeare setup.

SettingStronger
Layers / heads / embedding dim6 / 6 / 384
Context length256
Parameters10.77M
Batch size64
Dropout0.2
OptimiserAdamW, betas (0.9, 0.99), weight decay 0.1 on 2โ€‘D tensors only
Learning rate100 warmup steps to 10โปยณ, cosine decay to 10โปโด over 5,000 steps
Alsogradient clipping at 1.0, and weight tying: the token embeddings double as the rows of lm_head
Checkpointbest validation loss

In plain words, the new settings are:

  • Bigger steps that change over time. The learning rate starts small and grows over the first 100 steps (warmup), so that the random starting numbers are not knocked about too hard, then shrinks gradually along a curve (cosine decay) for fine adjustments at the end.
  • A cap on the slopes. If the slopes on one step are unusually big, they are scaled down before the nudge (gradient clipping), so that one odd batch cannot throw the machine off course.
  • Stronger weight decay, ten times stronger than the small model's, and only on the tables of weights and embeddings, not on the biases and normalising parameters.
  • betas set how long AdamW's momentum remembers earlier slopes.
  • More dropout: during training, a fifth of the numbers are switched off at random on every step, to make the machine harder to memorise with.
3.9 Training Loop with Best Checkpoint

The exam score starts at 4.29 and drops fast. It was best at step 1,500, at 1.47: as unsure as choosing between about 4.3 letters. The whole run took 47 minutes on my Mac Studio's GPU.

3.10 Plot Training Loss Curves

This is the student who memorised the textbook, caught in the act.

My reproduction of the paper's Figure 2: the exam score is best at step 1,500 (step 1,750 in the paper), then gets worse while the practice score keeps improvingTap the image to open it full size

My reproduction of the paper's Figure 2: the exam score is best at step 1,500 (step 1,750 in the paper), then gets worse while the practice score keeps improving

What happens after step 1,500 is the most useful part of the experiment. The practice score keeps improving, down to 0.61 by step 5,000, but the exam score climbs back to 1.71, no better than the tiny baseline's. The machine is memorising the practice text. The last copy is not the best one, which is why the notebook keeps the copy with the best exam score.

3.13 Generate High-Quality Samples

Prompted with ROMEO:, at the notebook's temperature of 0.8, my best copy wrote this (shortened):

CODE
1ROMEO:
2And this true sluck of voices, and kills thee
3To all the sun of the city the People:
4Your blood standing am I am about to have.
5
6JULIET:
7These grapes of you, noble like my brother did not
8To give the butcher of the noble gentleman: if you live
9But be his mind own so time, if he receive your brother's death.
10
11ROMEO:
12Ay, if with him, he conceal'd me.

It is not coherent, but the shape is unmistakable: speaker names in capitals, colons, line breaks, verse-like line lengths, plausible Shakespearean vocabulary. A machine that only ever sees one letter at a time picked all of that up from 1 MB of text, in under an hour.

I generated 800 letters from the "ROMEO:" prompt using the best copyTap the image to open it full size

I generated 800 letters from the "ROMEO:" prompt using the best copy

My runs against the paper's
The paper (Colab A100)My runs (Mac Studio, MPS)
Baseline, step 3,000: practice and exam scores1.5304 and 1.72361.5306 and 1.7126
Baseline time50.79 seconds89.35 seconds
Bigger machine: best step1,7501,500
Bigger machine: best exam score1.47801.4679
Bigger machine: whole run4.76 minutes46.98 minutes

What I took from it

Running MiniGPT end to end took under an hour on my own machine and cost nothing. A few things stuck with me:

  • Growing is writing, plus two steps. Check the answer, and nudge every number. Everything else is the machine from Part 1.
  • Nobody decides anything. Every number follows its own slope, and behaviour like attention waking up late emerges from the arithmetic.
  • Overfitting is not an edge case. The stronger model's best validation loss arrived at step 1,500 out of 5,000 on my run (step 1,750 in the paper). Without validation-based checkpoint selection the notebook would have shipped a worse model that scored well on its own training text.
  • Scale is doing the heavy lifting elsewhere. MiniGPT is honest that it is a reproducibility study, not a competitive model. The same training idea, with far more data, compute, and parameters, is what produces the models I use every day.

To get a feel for "far more", here is my small model next to GPT-3, the 2020 model whose family the first ChatGPT grew out of, using the figures from its paper:

My small MiniGPTGPT-3 (2020)How much bigger
Parameters826,433175 billionabout 200,000 times
Practice text1 million lettersabout 300 billion pieces of words, roughly 1.2 trillion lettersabout 1 million times
Blocks49624 times
Numbers in each vector12812,28896 times

If all of Tiny Shakespeare were one book on a shelf, GPT-3's practice text would fill a shelf tens of kilometres long. And the models behind today's chatbots are bigger again, though most companies no longer publish their sizes.

The value of a paper like this is not a benchmark number. It is that the path from raw text to generated samples is now something I have run rather than something I have read about.

Try it yourself