At a glance
Four hours. One empty file. One trained model that beats the thing it was copying.
Andrej Karpathy opens an empty train_gpt2.py, writes the GPT-2 124 million parameter architecture from scratch in PyTorch, loads the released OpenAI weights into his own module tree to prove the implementation is correct, throws those weights away, and trains the thing himself on FineWeb-Edu. By the end of a run that costs about ten dollars and takes under two hours on eight rented GPUs, his from scratch model is above the published GPT-2 124M on both the validation loss and HellaSwag. The companion repository is build-nanogpt, with one git commit per step of the video.
The architecture is the easy part and takes about forty minutes. What makes this the video people come back to is the middle two hours, where every optimization is applied one at a time and measured, so you watch the step time fall from 1,000 milliseconds to 90 on a single A100: TF32, bfloat16 autocast, torch.compile, flash attention, padding an ugly vocabulary size of 50257 up to a pretty 50304, and fused AdamW. Then the full optimization recipe copied out of the GPT-3 paper, gradient accumulation to reach the intended 524,288 token batch, and eight way distributed data parallel at 1.5 million tokens per second.
It is a lecture, a benchmark log, and a debugging diary at the same time. He leaves in the bug where he forgets to move a tensor to the GPU, the generations that refuse to match Hugging Face's pipeline, the torch.compile error he cannot solve on camera, and a loss curve with a periodicity he does not fully understand at the end of four hours. Those are not blemishes. They are most of why the video is loved.
The setup: what we are reproducing, and what it costs
He is careful about the word "reproduce" in the first thirty seconds. GPT-2 is not one model, it is a miniseries, and the biggest one usually gets the name. OpenAI released four sizes in 2019, from 124 million parameters up to 1558 million, so that you can put model size on the x axis of a plot, downstream metrics like translation, summarization and question answering on the y axis, and chart the scaling laws. This video does the smallest one, 124M.
Right away a small correction to the published record: the parameter counts in the GPT-2 paper's table disagree with the numbers he uses, and the reason is that the table is wrong. The GPT-2 repository says there was an error in how they added the parameters up. The real 124M configuration is 12 layers, 12 attention heads, and 768 channels.
Two assets exist and two do not, and the whole video is shaped by that asymmetry:
- GPT-2: weights, no details. OpenAI released a blog post, a paper, the code and, unusually, the weights. But the paper is extremely vague about optimization, and the released code is inference only, with no training script and almost no hyperparameters.
- GPT-3: details, no weights. The GPT-3 paper is far more concrete about hyperparameters and optimization settings, and the architecture is barely a departure from GPT-2. The context length went from 1024 to 2048, some Transformer hyperparameters moved, and GPT-3 was trained much longer on a bigger dataset with much more thorough evaluation. The models were never released.
So the plan is to take the architecture and the target weights from GPT-2, and the training recipe from GPT-3. The cost framing he gives up front is the headline of the whole video: this was probably a fairly complicated optimization five years ago on much smaller GPUs, and today you can reproduce it in roughly an hour or less for about ten dollars of rented cloud compute.
Start at the target: opening the OpenAI checkpoint
The first move is to start at the end. Before writing a line of his own, he loads the released GPT-2 124M and takes it for a spin.
The original code is TensorFlow, which he calls "not used as much anymore", so instead he goes through Hugging Face Transformers, which has already done the work of converting the weights to PyTorch friendly tensors. One awkward detail he flags immediately: in Hugging Face, from_pretrained("gpt2") gives you the 124M model, not the famous 1.5 billion one. For that you want gpt2-xl.
He pulls the state_dict and prints every key with its shape, in a Jupyter notebook running inside VS Code, because he likes one interface for everything.
The two shapes that explain the model:
wte.weightis 50257 by 768. 50257 tokens in the GPT-2 vocabulary, each one getting a 768 dimensional embedding. The number breaks down as 50,000 byte pair encoding merges, plus the 256 byte tokens that are the leaves of the BPE tree, plus one special end of text token that delimits documents and can also start a generation.wpe.weightis 1024 by 768. GPT-2 has a maximum sequence length of 1024, so there are 1024 absolute positions, and every one of them has its own learned 768 dimensional vector.
Then he does something most tutorials skip: he plots the weights. The position embedding table, visualized as 1024 rows, has obvious structure, because those embeddings end up learning sinusoids and cosines that stand in for position. In the original Attention Is All You Need the positional encodings are fixed to sinusoids of different frequencies; in GPT-2 they are ordinary parameters trained from scratch, and they recover sinusoid like features anyway.
He pushes on the picture and gets a real diagnostic out of it. The curves are a bit jagged and noisy, and that tells you the released model was not fully trained, because a more trained model would smooth them out. He also stops to be properly impressed: in principle these curves do not have to be smooth at all. The table starts as complete random noise, and the fact that anything interpretable falls out of the optimization is already remarkable.
Looking into individual channels as a function of position, three picked at random, some channels respond to parts of the position spectrum and not others. One green channel fires for everything from about 200 up to 800, much less below, with a sharp dropoff near zero. "Who knows what these embeddings are doing and why they are the way they are." He takes the first block of 300 by 300 from the first layer's weights and sees structure there too, with an aside for anyone who likes mechanistic interpretability, which is not what this video is about.
Finally he samples from the Hugging Face pipeline with the prefix "Hello, I'm a language model," asking for five sequences of 30 tokens. It produces coherent text. One wrinkle: even with the seed fixed he does not get the generations that the documentation example produced, so presumably the code changed. The point lands anyway. The weights load, they generate, and the state dict keys tell you exactly where everything lives in the model.
Which sets up the real task. The Hugging Face modeling_gpt2.py is "medium readable but not fully readable" and about 2,000 lines. He wants his own GPT class, so he has full understanding of what is happening, and his first milestone is to load this exact checkpoint into it.
We're going to have a lot of confidence that because we can load the OpenAI model we are in the same model family and model class, and we just have to rediscover a good setting of the weights, but from scratch. Andrej Karpathy, 13:20
SECTION 1: implementing the GPT-2 nn.Module
He opens the Attention Is All You Need architecture figure and starts deleting. GPT-2 is a decoder only Transformer, so the entire encoder stack is gone, and the cross attention block that was feeding off the encoder goes with it. Everything else stays almost the same, with two changes documented in section 2.3 of the GPT-2 paper:
- The layer norms moved. Instead of sitting after the attention and after the feed forward, they move to the input of each sub block. This is the pre normalization layout.
- One extra layer norm was added right before the final classifier, after the last block.
The skeleton, named to match Hugging Face
He deliberately mirrors the Hugging Face naming scheme, because matching keys is what makes the weight port trivial later. The container is an nn.ModuleDict called transformer, which lets you index submodules by string key:
wte, annn.Embeddingfor token embeddings. He notes annn.Embeddingis "a glorified wrapper around a tensor that allows you to access its elements by indexing into the rows".wpe, annn.Embeddingfor position embeddings.h, annn.ModuleList(indexed by integer, soh.0throughh.11) holdingn_layerblocks. He thinks thehprobably stands for hidden.ln_f, the extra final layer norm the GPT-2 paper added.
Then, outside the dict, lm_head, the language model head: a linear projection from 768 up to the vocabulary size of 50257. GPT-2 uses no bias on this final projection.
He then reads his own skeleton back against the figure from the paper, element by element, which is the step that makes the names stop being arbitrary:
wteis the thing the figure labels the output embedding, which is really the token embeddings.wpeis the positional encoding in the figure. Those two pieces of information add, and the sum goes into the Transformer.his all the blocks in gray in the figure, the repeated stack.ln_fis the new layer that GPT-2 adds and the 2017 figure does not have.lm_headis the linear part at the top, before the softmax.
So the module tree is not a convention he invented, it is the figure read top to bottom with GPT-2's two edits applied, and it happens to be spelled the way Hugging Face spells it so the state dict keys line up for free.
The block: a clean residual stream
The block is where he makes his first strong opinion stick. In the original figure the normalizations are inside the residual stream, so the residual pathway has normalizations in it. That, he says, is "not very good or desirable".
The argument runs through addition. Recall from micrograd that addition simply distributes gradients to both of its branches equally during the backward pass. So if the residual path is pure addition, gradients from the top flow straight down to the input tokens unchanged, while in addition flowing through the blocks, which contribute their own corrections over time and kick in gradually. You want a single clean residual stream all the way from supervision down to the tokens.
So the block is pre normalization, and in his code it reads as two lines of forward pass:
x = x + self.attn(self.ln_1(x))x = x + self.mlp(self.ln_2(x))
And then the framing that has probably been quoted more than anything else in the video. Attention is a communication operation. There are 1024 tokens lined up in a sequence and attention is where they exchange information. It is an aggregation, a pooling, a weighted sum, a reduce. The MLP happens at every token individually with no information exchanged between tokens at all, so it is the map.
So the attention is the reduce and the MLP is the map, and what you end up with is that the Transformer just ends up being a repeated application of map reduce, if you want to think about it that way. Andrej Karpathy, 19:58
Attention is where they communicate; the MLP is where each one thinks individually about what it gathered; and every block iteratively refines the representation sitting in the residual stream.
The MLP, and a tangent about GELU that is worth the detour
The MLP is two linear projections sandwiched around a nonlinearity. Up by a factor of four, from 768 to 3072, then the nonlinearity, then back down to 768.
The nonlinearity is nn.GELU(approximate='tanh'), and this is where he spends five minutes on something nobody would have blamed him for skipping.
GELU, Gaussian Error Linear Units, looks very much like a slightly smoother ReLU, except there is no exactly flat tail at zero. It comes from a paper with "some mathematical calac reasoning", as the captions render it, that connects it to stochastic regularizers and the expectation of a modification to adaptive dropout.
PyTorch offers both the exact version and a tanh approximation. There is no real good reason to use the approximation today. So why does it exist, and why is he using it?
Because of a GitHub issue. Dan Hendrycks, the GELU author, explains in PyTorch issue 39853 that at the time he developed the nonlinearity, the error function erf that you need for the exact GELU was very slow in TensorFlow, so they built the approximation instead. That approximation then got picked up by BERT, by GPT-2, and onwards. His exact words in the thread:
I used the tanh approximation simply because the error function erf was slow in tensorflow some years ago. If the exact version is fast enough now and does not have numerical issues, I do not see a reason to use an inexact version. Dan Hendrycks, PyTorch issue 39853, quoted in the video at 22:03
So the tanh form is a historical quirk, and the only reason to keep it is fidelity. Karpathy is reproducing GPT-2 exactly, GPT-2 used the tanh approximate version, so he sticks with it.
The intuitive argument for GELU over ReLU gets its own paragraph, and it is the dead ReLU neuron problem. In the flat tail of a ReLU, any activation that lands there gets exactly zero gradient. No change, no adaptation, no development of the network for that neuron. GELU always contributes a local gradient, so there is always a change and always an adaptation, and smoothing it out ends up working better empirically. He notes that more modern networks, Llama 3 among them, have moved on again to SwiGLU and variants.
Causal self attention, as tensor gymnastics
He goes through this one faster, pointing back at the previous video in the series for the slow version. The content is the same multi headed attention; the implementation is different.
In the earlier video the heads were separate modules whose outputs were concatenated, which made it obvious that heads are "just kind of like parallel streams". Here all of that collapses into one module, and the cost is "a bunch of transpose, split, tensor gymnastics to make this very efficient in PyTorch". Fundamentally and algorithmically nothing is different.
What happens, in his order:
- Each token emits three vectors, the query, the key and the value.
c_attnis a single linear layer producing all three at once, which he then splits. - The number of heads is folded into the batch dimension, so PyTorch treats both B and n_head as batch dimensions and applies everything in parallel across both.
- Queries and keys multiply to give the attention, "how interesting they find each other", which has to be a multiplicative interaction.
- The autoregressive mask makes sure tokens only attend to tokens before them and never to the future. It is a registered buffer, not a parameter, which matters in a minute.
- Softmax normalizes the attention so it sums to one.
- The attention matrix multiplied against the values is a weighted sum of the values of the tokens each token found interesting.
- A final transpose, contiguous and view reassembles everything, and that step is what actually performs the concatenation of the heads.
c_projprojects back out.
The shapes, which are the part worth having written down if you are typing along, with B the batch, T the sequence length up to 1024, C the 768 channels, nh the 12 heads and hs the 64 dimensional head size, where nh times hs equals C:
- In:
(B, T, C). c_attngives(B, T, 3C), which splits into three tensors of(B, T, C), the queries, keys and values.- Each is viewed as
(B, T, nh, hs)and transposed to(B, nh, T, hs). This is the step he calls the gymnastics, and the point of it is thatnhnow sits besideBso PyTorch treats both as batch dimensions and runs every head in parallel. - Queries against transposed keys gives the attention,
(B, nh, T, T). At T of 1024 that is a million numbers per head per batch element, which is the number that matters when flash attention arrives two hours later. - Mask, softmax, then the attention against the values gives
(B, nh, T, hs). - Transpose,
contiguous,viewback to(B, T, C). That last reshape is the concatenation of the heads, which is the part that looks like bookkeeping and is actually the operation. c_projgives(B, T, C)back out, ready to be added into the residual stream.
He is careful with variable names throughout, so his keys follow the Hugging Face schema exactly, which is what makes the port a copy loop.
The result: Hugging Face's file is about 2,000 lines. His complete GPT-2 implementation is under 100 lines of code.
Loading the Hugging Face parameters
The config gets set to the real GPT-2 124M numbers: block_size 1024, vocab_size 50257, n_layer 12, n_head 12, n_embd 768.
Then from_pretrained, a class method that returns a GPT object given a model type. He calls the loading code "kind of dry" and "not that exciting", and walks it anyway, because the two gotchas in it are exactly the kind of thing that silently ruins a reimplementation:
- Skip the buffers.
attn.biasis not a parameter, it is the autoregressive mask. Ignore those keys. - Transpose four weights. Because the checkpoint came out of the TensorFlow repo, some weights are transposed relative to what PyTorch wants. He hardcodes the list of names that need transposing and flips them. "I'm not sure how. This is a little bit annoying."
He runs it. No crash. Weights, biases and everything else load into his nn.Module.
The forward pass
The input is idx, token indices, always of shape B by T, where T cannot exceed the block size. B independent sequences stacked in a batch for efficiency.
Inside, he creates the positions with torch.arange and is careful to put them on idx.device, which is a detail that pays off twice later: once for CPU and Apple silicon support, and once as the reason a device mismatch bug does not happen here.
Position embeddings plus token embeddings. There is broadcasting hidden in that plus, because the position embeddings are identical for every row of the batch, so a dimension gets created and the two add. Then the blocks, then ln_f, then lm_head.
What comes out is logits of shape B by T by vocab_size. At every single position, the logits for what token comes next, which is "just a softmax away from becoming probabilities".
Sampling: tokenization, the loop, and top k 50
He sets up the exact same experiment as the Hugging Face pipeline: five sequences, 30 tokens, prefix "Hello, I'm a language model,".
model.eval() first, which is good practice when you are not training. Then an honest aside: he does not actually know if it is doing anything here, because nothing in this model has training versus evaluation behavior. Dropout and batch norm do; every layer he wrote should be identical in both modes. So model.eval() may be doing nothing, but he is not sure, and PyTorch internals may be doing something clever.
Then model.to('cuda'). He is SSHed into a cloud box with eight GPUs, and this ships every tensor off to what he describes as "basically a whole separate computer that is sitting on the GPU", with its own architecture, connected to the CPU and able to communicate with it, but well catered to the parallel processing that neural networks are.
Tokenization uses tiktoken and the GPT-2 encoding. The prefix comes out as eight tokens, which he cross checks by pasting the same string into Tiktokenizer. Those eight tokens get replicated five times into a five by eight tensor, moved to the GPU, and that is the starting idx.
The sampling loop, which runs under torch.no_grad() so PyTorch does not cache intermediates for a backward pass that will never come:
- Forward to get logits, then take only the last column and throw the rest away. He names this as wasteful: correct but an inefficient implementation of sampling.
- Softmax to probabilities.
- Top k of 50, because that is the Hugging Face default. Keep the 50 highest probabilities, clamp everything below the 50th to zero, renormalize. "That way we are never sampling very rare tokens." It keeps the model on track, stops it blabbering and going off the rails, and keeps it in the vicinity of likely tokens.
- Sample, append the new column to
x, and the columns grow until the loop ends with a five by 30 tensor, which he decodes back to strings.
The generations:
Hello, I'm a language model, not a program. Andrej Karpathy's implementation, first sample, 39:55
Hello, I'm a language model, and one of the main things that bothers me when they create languages is how easy it becomes to create something that ... Andrej Karpathy's implementation, second sample, 39:55
They do not match the Hugging Face pipeline's output, and he cannot find the discrepancy on camera. "I can't find the discrepancy to be honest." He suspects there is something hiding in the pipeline in addition to the top p setting. So he does the thing you should do: he replicates the Hugging Face call path directly in the notebook, gets identical results there, and concludes the model internals are not wrong, he just does not know what the pipeline is doing.
That is the milestone. Every weight ported, this is the exact OpenAI GPT-2, and it generates sensible sequences. Now throw it away.
From random initialization, and auto detecting the device
Initializing from scratch turns out to be the easy part, because PyTorch already initializes randomly by default. Every linear layer has a default constructor, using for example the Xavier initialization covered in earlier videos. So model = GPT(GPTConfig()) is a randomly initialized 124M model, and the output is, as promised, "total garbage garbled", random token string pieces chunked up at random.
Before moving on he adds device auto detection, explicitly so that people without a GPU can follow along, at least until the multi GPU section at the end. The ladder goes by compute capability: start with CPU, which every computer has, then try CUDA, then try MPS, the Apple silicon backend, which on a fairly new MacBook gives you a GPU that is "actually fairly capable depending on which MacBook you have" and will beat CPU.
With device swapped in for the hardcoded 'cuda', forcing CPU still works, because the forward pass creates the position tensor on idx.device rather than assuming. A CPU generation takes about six seconds, without torch.compile and the rest, so following along on a laptop is viable.
For now, let's just say the device makes code go fast. Andrej Karpathy, deferring the whole topic of what PyTorch does when you call .to(device), 45:32
Let's train: data batches, loss, and crushing a single batch
Tiny Shakespeare, by the numbers
His favorite debugging dataset is tiny Shakespeare, and he gives it in full detail because the arithmetic matters later. Word count on the file: about 40,000 lines, about 200,000 words, about 1 million bytes. It is all ASCII, one byte per character, so roughly a million characters. The GPT-2 tokenizer has a compression ratio of roughly 3 to 1, so a thousand characters is about 300 tokens, and the whole file comes out at 338,000 tokens.
He tokenizes the first thousand characters and prints the first 24 token ids, and points out that if you can read GPT-2 tokens you will recognize 198 as the newline character, appearing twice in a row where the text has a blank line.
Making a (B, T) batch out of a one dimensional stream
The problem: a Transformer wants a batch of B independent sequences of up to T tokens, and what you have is one very long one dimensional sequence.
His favorite way to do it is a .view(). Take the first 24 tokens, view them as 4 by 6, and the first six tokens become the first row, the next six the second row, and so on. It stacks every six tokens as an independent row.
Then the labels. You could compute the targets inside the forward pass, since the next token is just one to the right, except that the very last token in the batch has no next token loaded, so you are one short. His fix, and the pattern worth stealing:
- Fetch a buffer of
B * T + 1tokens, one extra. xis everything up to but not including the last token, viewed as (B, T).yis the buffer starting at index 1, same view.
Now token 25's target is 198 and it sits at exactly the same position in the target tensor, and the last token 13 has its label too, because of that plus one.
Cross entropy, and the number you should expect at initialization
The forward pass grows an optional targets argument, returns (logits, loss), and calls F.cross_entropy. The reshaping in that call looks scary and is not: F.cross_entropy will not take a three dimensional B by T by vocab_size input, so the logits get flattened to two dimensions, B * T rows by vocab_size columns, and the targets get flattened to a single B * T tensor.
Then the sanity check, which is the most reusable fifteen seconds in this part of the video. At initialization you want the network to be maximally uncertain, so the probability of any arbitrary token should be roughly 1 over 50257. Cross entropy is negative log likelihood, so the loss you expect is the negative natural log of that, which is 10.82. He prints 11. Not way off, so the distribution at initialization is suitably diffuse and nothing is confidently wrong before training starts.
Overfit a single batch
With AdamW, not SGD. He explains the choice as a bug fix: "AdamW is a bug fix of Adam is what I would say." It keeps two buffers per parameter, m and v, which it calls the first and second moment, one looking a bit like momentum and one a bit like RMSProp, a normalization applied to each gradient element individually that speeds up optimization, especially for language models. He treats it as a black box here.
Two things he insists on in the loop:
- Zero the gradients first.
loss.backward()always does a plus equals on the gradients, it deposits rather than assigns, which is exactly why you must zero them. loss.item()converts the single element tensor to a float, and that float lives on the CPU. Behind the scenes PyTorch takes the one element tensor, ships it back to CPU memory, and converts it.
Then the bug, left in. The run dies with expected all tensors to be on the same device, but found at least two devices, cuda:0 and cpu. The model was moved, the data buffer was not, and the reason the fix is not obvious is a real PyTorch asymmetry:
You have to be careful, because you can't just do
buff.to(device). It's not stateful, it doesn't convert it to be a device. It instead returns a pointer to a new memory which is on the device. So you see how we can just domodel.to(device), that does not apply to tensors, you have to dobuff = buff.to(device). Andrej Karpathy, 1:00:21
Fixed, and the single batch gets crushed. Starting from 10.82 or 11, optimizing the same batch over and over with no new data, the loss goes to very very low. The Transformer is memorizing one batch, which is the point: it proves the whole forward and backward path works. Learning rate 3e-4, which he calls "a pretty good default for most optimizations that you want to run at a very early debugging stage".
DataLoaderLite
A deliberately minimal loader. Read the whole text file, tokenize it, print the total token count and the number of batches in a single epoch. Start at position zero, take chunks of B * T, always advance by exactly B * T, but always fetch B * T + 1 because of the target for the last token. Run off the end, loop back to zero.
He predicts the result before running it, which is the habit worth copying. The loss should come down but not to zero, and the reason is that of 50257 tokens many never occur in Shakespeare at all, so there are very easy gains available: drive the logits of all the never seen tokens towards negative infinity. All the crazy Unicode and the other languages. "That's probably most of the loss gain that we're going to see at this scale right now." Fifty iterations is also not enough for one epoch, which with B of 4 and T of 32 takes 2,600 batches.
The run: starts in familiar territory around 11, comes down to about 6.6. Exactly as called.
Two details from the real GPT-2: weight tying and the initialization
Weight tying, found by comparing pointers
He goes back to the state dict because he missed something while loading, and calls it a bug with respect to how GPT-2 training should happen.
The token embedding at the bottom of the Transformer and the language model head at the top are both 50257 by 768. Same shape, which could be coincidence. So he checks harder:
- Element wise equality plus
.all(): every single element is identical. .data_ptr(): the pointers are identical. These are not two tensors that happen to match, they are the same tensor.
That is weight tying, and it comes from Attention Is All You Need, which says in its embeddings and softmax section that it shares the same weight matrix between the two embedding layers and the pre softmax linear transformation, citing an earlier 2017 paper that argues for the scheme. He calls that phrasing "an awkward way to phrase that these two are shared and they're tied and they're the same matrix", and he cannot find where Hugging Face does the tying, but he can find it in the original OpenAI GPT-2 TensorFlow source, where wte is used once at the bottom as the token encoder and again at the top in the matmul that produces the logits.
The intuition he gives for why it should work:
If two tokens are very similar semantically, like maybe one of them is all lowercase and the other one is all uppercase, or it's the same token in a different language or something like that, if you have similarity between two tokens presumably you would expect that they are nearby in the token embedding space. But in the exact same way you'd expect that if you have two tokens that are similar semantically you'd expect them to get the same probabilities at the output of a Transformer. Andrej Karpathy, 1:09:04
Both positions, the very bottom and the very top, want similar tokens to have similar weights. So tie them.
In the backward pass the shared tensor collects gradient contributions from both branches, from the classifier at the top and from the embedding lookup at the bottom, and they add up on it.
The implementation is one line: point wte.weight at lm_head.weight. That copies the reference, orphans the old value, Python cleans it up, and now a single tensor is used twice in the forward pass.
And the parameter count makes it more than an elegance argument. 768 times 50257 is about 40 million parameters, in a 124 million parameter model. Roughly 30 percent of the parameters saved. He offers that as a possible reason the scheme works slightly better if you are not training long enough: fewer parameters to train makes the process more efficient, and you are putting in a real inductive bias.
Initialization, read between the lines
Neither the GPT-2 nor the GPT-3 paper is explicit about initialization, so he goes to the released code instead of the paper, which he calls "quite vague".
What model.py in the OpenAI repo actually does:
- Weights: normal distribution, standard deviation 0.02.
- Biases: zero.
- Token embeddings: 0.02.
- Position embeddings: 0.01, "for some reason".
His version uses nn.Module.apply to walk every submodule and initialize. Linear weights normal with std 0.02, biases zero, embeddings 0.02. He keeps embeddings at 0.02 rather than switching position embeddings to 0.01 because "it's about the same". He flags that zero initialization for the bias is not the PyTorch default, which is a uniform distribution. LayerNorm he leaves alone, because the PyTorch default is already scale one and offset zero, which is what you want.
Then the question of whether 0.02 is a sensible number at all. If you follow Xavier, the standard deviation would be one over the square root of the incoming feature count. Run the arithmetic on the GPT-2 sizes and 0.02 lands right in the middle of the range:
- 1 over sqrt(768) gives about 0.03.
- 1 over sqrt(1600) gives about 0.02.
- Three times 1600 gives about 0.014.
So the hardcoded 0.02 is "not completely crazy", though typically you would want something that scales with model size. He keeps it, because it is what GPT-2 did.
The residual initialization scaling, derived on screen
One caveat remains, from the GPT-2 paper itself:
A modified initialization which accounts for the accumulation on the residual path with model depth is used. We scale the weights of residual layers at initialization by a factor of one over the square root of N, where N is the number of residual layers. The GPT-2 paper, read out in the video at 1:17:15
Rather than just implementing it, he motivates it in the notebook, and this is one of the clearest two minutes in four hours.
Start with a residual stream of 768 zeros. The stream has the form x = x + something, so every block contributes some amount and it gets added. Now add a standard normal draw, zero mean and unit standard deviation, one hundred times. By the end the residual stream has a standard deviation of 10, which is the square root of 100. The variance of activations in the residual stream grows with depth, because you keep adding.
Scale each contribution by one over the square root of n, which is n ** -0.5, and the standard deviation comes back to 1. The paper's factor exactly compensates for the growth.
His implementation is, in his own words, possibly not PyTorch sanctioned but it works: set a flag attribute NANOGPT_SCALE_INIT = 1 on the modules that write into the residual stream, which is the c_proj at the end of each attention and each MLP, then in the init function check for the flag and multiply the standard deviation by (2 * config.n_layer) ** -0.5. "There must be a better way in PyTorch, right? But I don't know."
Why two times the number of layers? Because every layer in the Transformer has two blocks that add into the residual pathway, the attention and the MLP. That is where the factor of two comes from.
One awkwardness he flags and chooses not to fix. Because wte and lm_head are tied, the module walk comes around to that same tensor twice: once as an embedding, initialized to 0.02, and once as a linear, initialized to 0.02 again. Same value both times, since lm_head carries no scale flag, so it is harmless, just initialized twice identically.
Seeds get set for reproducibility, and that is the GPT-2 initialization as faithfully as the public record supports.
SECTION 2: let's make it fast
This is the center of the video, and the method is as important as the content. Every change is applied on its own, timed, and compared against the number before it. Nothing is taken on faith, and twice the speedup he gets is nothing like the speedup the hardware promised, which turns into the single most useful lesson in the section.
You always want to start with: what hardware do you have, what does it offer, and are you fully utilizing it? Andrej Karpathy, 1:22:26
The hardware, and the napkin math
nvidia-smi shows eight A100 SXM 80GB GPUs. He rents these from Lambda Labs, which he names as his favorite place to spin up a box, pay by the hour, and connect VS Code to, and discloses that they sponsor his development and his projects.
He sets the benchmark configuration: batch size 16, sequence length 1024, the real GPT-2 maximum, on tiny Shakespeare. At 16 by 1024 the model occupies 35 GB of the 80 available, and one epoch of Shakespeare is only 20 batches.
Then he breaks into the code right after the loss, prints logits.dtype, and gets torch.float32. By default everything in PyTorch, every activation and every parameter, is a 32 bit float. That is a lot of memory, and:
It turns out empirically that for deep learning as a computational workload this is way too much. Andrej Karpathy, 1:24:02
From the A100 datasheet, with his commentary:
- FP64: supported, genuinely useful for scientific computing, irrelevant here.
- FP32: 19.5 TFLOPS. That is where the code currently sits, so at most 19.5 trillion floating point operations per second, mostly multiply adds.
- TF32: 8 times more, 156 TFLOPS.
- FP16 and BF16: 16 times more, up to 312 TFLOPS.
- INT8: more again, and the wrong tool. INT8 is for inference, not training, because its spacing is uniform, and training needs a float to match the normal distributions that both activations and weights actually follow.
- The asterisked numbers NVIDIA likes to quote are "with sparsity", which he is not using and does not think is widely used in industry right now, so he reads the column without it.
Then the fact the whole section rests on. Lower precision does not just make the multiply faster. Fewer bits per number means less data to move, and that is where memory bandwidth comes in. The A100 can move about 2 terabytes per second, which is a lot, and it is still the binding constraint:
Many of the deep learning workloads for training are memory bound, and what that means is actually that the tensor cores that do all these extremely fast multiplications, most of the time they're waiting around, they're idle, because we can't feed them with data fast enough. Andrej Karpathy, 1:27:06
His rule of thumb: if you are getting 60 percent hardware utilization you are doing extremely well. Half the time, in a well tuned application, your tensor cores are not multiplying anything.
Tensor cores and TF32: 1,000 to 333 milliseconds
A tensor core is just an instruction in the A100 architecture, and what it does is a little 4 by 4 matrix multiply. There are multiple configurations for the input precisions, the internal accumulate precision and the output precision, but it is fundamentally a 4 by 4 multiply, and any matrix multiplication gets broken up into it because that is the fastest way to multiply matrices.
That matters because almost all of the computational work here is matrix multiplication, hidden inside the linear layers. There are additions in the residuals, GELU nonlinearities, layer norms, but time them and they are nothing. And at this small scale the single biggest matmul by a distance is the classifier at the top, 768 going to 50257, which dominates everything else in the network.
For the mechanism he goes to the A100 architecture whitepaper, figure 9, which he calls relatively readable if you half understand what is happening. TF32 is:
- Exactly the same bit layout as FP32 at the top: one sign bit, eight exponent bits.
- The mantissa gets cropped. 19 bits kept, the last 13 dropped.
- The accumulate is still FP32, the output is FP32, the inputs are FP32. The truncation is entirely internal to the instruction.
So nothing in your PyTorch code changes and every number still looks identical. All you do is let the tensor core crop 13 bits inside the operation, and the little matrix multiply goes 8 times faster.
The reason I like TF32 is because if you can tolerate a little bit of a precision fudge then this is free. Like none of your code sees this, it's fully internal to the operation, and the operation to you just goes 8x faster. Andrej Karpathy, 1:32:17
Before timing, he fixes a trap that catches everyone. When the CPU runs, it is only scheduling work on the GPU, so it queues kernels and races ahead. You have to call torch.cuda.synchronize() to wait for the GPU to finish everything that was scheduled before you take the time. He also notes the first iteration is often slower, because PyTorch is doing initializations, allocating tensors and buffers for the gradients, so be careful when timing.
He also adds tokens per second as the metric, on the grounds that it is the objective measure, since the batch size will change over the course of the video and milliseconds per step will not be comparable.
The float32 baseline: roughly 1,000 milliseconds per iteration, about 163,000 tokens per second.
Enabling TF32 is one line, torch.set_float32_matmul_precision('high'). The default is 'highest', which keeps everything in float32; 'high' lets matmuls use TF32 where the hardware has it, which on an Ampere A100 it does.
Promised: 8 times. Delivered: 1,000 down to about 333 milliseconds, roughly 50,000 tokens per second. About 3 times.
And the explanation is the lesson:
Even though the TF32 offers in principle a lot faster throughput, all of these numbers everywhere are still float32s, and it's float32 numbers that are being shipped all over the place through the memory system, and it's just costing us way too much time to shuttle around all this data. So even though we've made the multiply itself much faster, we are memory bound. Andrej Karpathy, 1:38:29
Still, a 3 times throughput improvement for one line of code, with every variable still float32 everywhere.
bfloat16: 333 to 300 milliseconds
If the problem is that we are still moving 32 bit numbers around, move 16 bit numbers instead. Which raises the question of which 16 bit format, and the answer is the clearest explanation of bfloat16 versus float16 in any lecture:
- bfloat16 keeps the sign bit and all eight exponent bits unchanged, and truncates the mantissa even more aggressively than TF32. The exponent sets the range you can represent; the mantissa sets the precision within that range. So bfloat16 has the identical range to float32 with fewer possibilities inside it.
- float16 touches the range. It cannot represent the full FP32 range, and that is where the trouble starts, because now you need gradient scalers, which are "kind of annoying and an additional source of state and complexity".
- Historically float16 came first, on the Volta generation, before Ampere, so everyone trained in float16 and everyone had to use gradient scaling. Then bfloat16 arrived with Ampere and made it much simpler: same range, no scalers.
He follows the PyTorch mixed precision recipe, which he recommends specifically because "there are five other copies that I would not recommend", and the guidance is narrow:
- Ignore everything about gradient scalers. Look only at
torch.autocast. - Do not call
.bfloat16()on any of your tensors yourself. - Surround only the forward pass and the loss calculation. Leave the backward and the optimizer step alone.
Since his loss is computed inside the model's forward, the context manager wraps one call. Then he breaks in to show exactly what changed, and the answer is "not everything", which is why it is called mixed precision:
logits.dtypeis nowtorch.bfloat16. The activations changed.model.transformer.wte.weight.dtypeis stilltorch.float32. The parameters did not.
What gets cast and what does not is, he admits, not super clear. Matrix multiply like operations get converted. A lot of operations stay in float32, in particular normalizations like layer norms, and softmax, log softmax and the loss calculation, because those are more susceptible to precision changes while matmuls are fairly robust to them. "Unfortunately they don't document it very well, so we're not going to go into that in too much detail."
Result: 333 down to about 300 milliseconds, about 55,000 tokens per second. "We're definitely running faster, but maybe not a lot faster." Because there are still many, many bottlenecks, and we are only getting started.
torch.compile: 300 to 129 milliseconds
torch.compile is really quite incredible infrastructure from the PyTorch team, and it's basically a compiler for neural networks. Like it's almost like GCC for C and C++ code. This is just the GCC of neural nets. Andrej Karpathy, 1:48:10
One line, model = torch.compile(model). It costs compilation time, which he spends explaining what it does, and the result when it lands is 300 down to 129 milliseconds, which he calls about a 2.3 times improvement from a single line of PyTorch. Throughput about 125,000 tokens per second. His verdict is unambiguous: there is no real good reason not to use torch.compile, you should be using it almost by default, unless you are debugging.
The line from the documentation that he says actually explains it: the speedup mainly comes from reducing Python overhead and GPU read writes. Both halves get unpacked.
Python overhead. An nn.Module is just the algorithmic description of what you want to happen. Without the compiler, PyTorch runs in eager mode: the Python interpreter walks the forward pass layer by layer, dispatching and materializing each operation as it goes, with no idea what operations come later. torch.compile sees the whole thing at once, knows exactly what you intend to run, and takes the Python interpreter out of the forward pass entirely, compiling the network as a single object.
GPU read writes, which is the bigger half, and he builds it from the GELU. He writes out the tanh GELU formula by hand, as the equivalent of calling nn.GELU, and walks what happens without the compiler:
- The interpreter hits
x ** 3, dispatches a kernel. The input lives in the GPU's high bandwidth memory. It has to travel to the cores, the caches and the registers on the chip, get cubed, and the result gets saved back to memory. - Multiply by a constant: another kernel, another trip out, another trip back.
- Add the input back: travels again, adds, written back.
Every elementwise step is a full round trip to memory, because all the tensor cores and the arithmetic units are on the chip while the data is not. With the compiler, the data makes one trip to the chip, all the elementwise operations happen while it is sitting there, and it is written back once. That is kernel fusion, and it is the major way everything gets sped up.
He then supplements with the memory hierarchy, calling it a preview of what could be its own two hour video:
- The GPU chip is where the calculations happen, and it does have some memory, but most of the memory by far is the separate high bandwidth memory chip beside it, tens of gigabytes, connected but off chip.
- On the chip there are a large number of streaming multiprocessors, 120 of them in total, each with four quadrants containing tensor cores along with separate units for FP64, FP32 and integers.
- Memory is sprinkled through the chip: an L2 cache on the chip, and L1 cache and registers on each SM. These are built differently, in SRAM rather than the transistors and capacitors of HBM.
- The critical point: on chip there is only a couple of tens of megabytes of memory in total, because memory on the chip is expensive. It is lightning fast in relative terms and there is almost none of it.
- A CPU might have a terabyte of DRAM, and for a GPU it is extremely expensive to reach, because you have to go through the CPU.
The three diagrams he walks through, in his order, are worth having in your head whenever you are reasoning about why a kernel is slow:
The two chips. The GPU chip is where almost all the calculation happens, and it does contain some memory, but most of the memory by far is on a physically separate HBM chip beside it. They are connected and they are not the same thing. Everything is super fast within the chip; going to the memory is extremely expensive and takes an extremely long amount of time.
The zoom in. On the sides of the die are the links to HBM. On the die are the streaming multiprocessors, 120 of them, and zooming into one shows four quadrants. The tensor core is where the matrix multiply work happens, and around it are separate units for FP64, FP32 and integers, because they are genuinely different hardware. The L2 cache lives on the die; L1 cache and registers live on each SM.
The capacity ladder, which is the punchline. A CPU might have a terabyte of DRAM, and for a GPU reaching it is extremely expensive because you have to go through the CPU. A typical GPU has tens of gigabytes of HBM, which is also, as he keeps saying, very expensive to access. And on the chip itself, across the L2, all the L1 caches and all the registers together, there are only a couple of tens of megabytes. The memory that is lightning fast is the memory there is almost none of, because on chip memory is expensive to build, and the implementation is different in kind: transistors and capacitors in HBM versus SRAM on the die.
So the accurate picture of a kernel is: inputs live in global memory, data streams to the chip, the calculation happens, the result streams back. Without fusion you make that round trip many, many times. With fusion a chunk of data lives on the chip, you do all the operations on it right there in an elementwise fashion, which is very cheap, and you take a single round trip back. "That gives huge savings, and that's why torch.compile ends up being a lot faster."
Flash attention: 130 to 96 milliseconds
And then the operation torch.compile cannot find.
FlashAttention came out of Stanford in 2022. It replaces four lines of his attention implementation with one, and it is a kernel fusion, but it is a kernel fusion the compiler cannot discover, because it requires an algorithmic rewrite of how attention is implemented.
The remarkable part, stated plainly:
Flash attention actually, if you just count the number of FLOPs, flash attention does more FLOPs than this attention here. But flash attention is actually significantly faster. In fact they cite 7.6 times faster potentially. Andrej Karpathy, 2:01:37, on the FlashAttention paper's own numbers
It is faster because it is mindful of the memory hierarchy he just drew. It is careful about what sits in high bandwidth memory and what sits in shared memory, and it orchestrates the computation so that there are fewer reads and writes to HBM. More FLOPs, fewer expensive loads and stores, and the loads and stores are the expensive part.
Specifically, it never materializes the T by T attention matrix. That matrix is where all the queries and keys interact, and for each head and each batch element at T of 1024 it is a million numbers. FlashAttention is designed so that matrix never exists at any point and is never read from or written to HBM.
The mechanism is the online softmax trick, with intermediate variables M and L and an update rule that lets you evaluate a softmax incrementally without ever realizing all of its inputs to do the normalization. FlashAttention-2 adds further gains on top.
And then the piece of history he clearly enjoys. The online softmax paper this is built on is Online normalizer calculation for softmax, and it came out of NVIDIA in early 2018, four years before FlashAttention. Its abstract proposes a way to compute the classical softmax with fewer memory accesses and hypothesizes that the reduction in memory accesses should improve softmax performance on actual hardware.
They are extremely correct in this hypothesis, but it's really fascinating to me that they're from NVIDIA and that they had this realization but they didn't actually take it to the actual flash attention that had to come four years later from Stanford. So I don't fully understand how this happened historically. Andrej Karpathy, 2:04:10
In PyTorch you get it by calling F.scaled_dot_product_attention, the compound operation, and PyTorch dispatches to flash attention. He finds it a little bit odd that torch.compile cannot see that his four lines should become exactly that call, and says so.
Before the change, step 49 gave a specific loss at 130 milliseconds. After, the same loss "basically identical up to a floating point fudge factor", at about 95 to 96 milliseconds. The computation is the same, the kernel is faster, and he puts the improvement at 96 over 130, roughly 27 percent.
Nice and ugly numbers: 50257 to 50304, 96.5 to 93 milliseconds
We are now getting to one of my favorite optimizations, and it is simultaneously the dumbest and the most brilliant optimization, and it's always a little bit surprising to me. Andrej Karpathy, 2:06:48
The premise: 64 is a beautiful number, 128 is even nicer, 256 is beautiful. What makes a number beautiful is having many powers of two inside it, so you can halve it repeatedly. Ugly numbers are 13, 17, primes, anything odd. And you always want nice numbers in code that deals with neural networks or CUDA, because everything in CUDA works in powers of two, kernels are written in terms of them, and blocks come in sizes like 16 and 64. Everything else gets special case handling.
So the heuristic is literally: scan your code and look for ugly numbers. His own audit, out loud:
- The
3 *inc_attnis "kind of ugly", and he is not sure it can be improved. 4 *in the MLP is nice.- 1024 is very nice.
- 12 layers is "a little bit suspicious", not too many powers of two.
- 768 is great.
- 50257 is a really really ugly number. Odd, with not many powers of two in it. Highly suspicious.
- And in the GPT-2 XL config, 25 attention heads. "That's a really ugly number, that's an odd number, and actually this did cause a lot of headaches for us recently when we were trying to optimize some kernels, and required a bunch of special case handling."
The fix for the vocabulary size is to raise it to the nearest nice number above: 50304, which divides by 8, 16, 32, 64 and even 128. One line in the config. He is explicit that this increases the amount of computation the network does, more FLOPs by napkin math, and then thinks through whether it breaks anything while the compile runs:
vocab_sizeis used in exactly two places: the embedding table at the bottom and the classifier at the top.- The embedding table is now larger, "almost like we introduced more tokens at the bottom", and those rows will never be indexed, because the GPT-2 tokenizer only goes up to 50256. A little wasted space, memory that is never accessed.
- But because of weight tying, the same tensor is the classifier. So the network is now predicting extra dimensions, probabilities for tokens that can never appear in the training set, and it has to learn to drive those logits to negative infinity.
- Which is no different from what it is already doing. Shakespeare probably uses a thousand tokens out of 50257, so most tokens are already being driven to zero probability by the optimization. A few more join them.
So functionally nothing breaks. We're using a bit more extra memory, but otherwise this is a harmless operation as far as I can tell. And we're adding calculation, but it's running faster. Andrej Karpathy, 2:12:55
Result: 96.5 down to 93 milliseconds, roughly 4 percent, by doing more arithmetic.
Why it works, in his explanation: CUDA kernels chunk your input into block tiles of nice sizes, powers of two, calculations happen in chunks of 32 or 64. When your desired calculation does not fit neatly into those tiles, boundary kernels kick in to handle the leftover. In a lot of kernels they do the nice part first, then a whole second phase comes back for the remainder, and the kernels for that can be very inefficient, spinning up extra compute for a sliver of work. So you may as well pad your input and make it fit, and empirically that runs faster.
One caveat that matters enormously if you reproduce this today, and which he gives precisely:
I also have to point out that we're using PyTorch nightly, so that's why we're only seeing 4 percent. If you're using PyTorch 2.3.1 or earlier you would actually see something like 30 percent improvement just from this change, from changing it from 50257 to 50304. Andrej Karpathy, 2:14:27
At that point the running total is about 11 times, from about 1,000 milliseconds per step to 93.
SECTION 3: the GPT-3 recipe, borrowed line by line
With the hardware wrung out, he turns to algorithmic changes, and the entire section is him reading the GPT-3 paper's appendix and implementing each sentence in turn. The reason is the asymmetry from the top of the video, restated:
GPT-2, we have the weights but no details. GPT-3, we have lots of details but no weights. Andrej Karpathy, 2:15:59
GPT-2's paper is "extremely vague as to the optimization details", and the code OpenAI released is inference only, with no training code and very few hyperparameters. So the recipe comes from GPT-3, which is fine because the two architectures are very very similar: context length went from 1024 to 2048, some Transformer hyperparameters moved, and GPT-3 is 175 billion instead of 1.6 billion and was trained far longer on more data with more thorough evaluation. Otherwise, pretty much the same model.
AdamW betas and epsilon
The GPT-3 paper: "To train all versions of GPT-3 we use Adam with beta one, beta two of 0.9 and 0.95."
PyTorch's AdamW defaults to betas of 0.9 and 0.999, so the second one changes. Epsilon is 1e-8 in the paper, which happens to match the PyTorch default, and he sets it explicitly anyway.
Gradient clipping at a global norm of 1.0
The paper: "We clip the global norm of the gradient at 1.0."
One line, torch.nn.utils.clip_grad_norm_, inserted right after loss.backward(). What it computes: take every gradient on every parameter, square it, add it all up, take the big square root. That is the norm of the parameter gradient vector, its length. Then make sure the length is no more than 1.0, and clip it if it is over.
Why people do it: sometimes you get unlucky, maybe a bad data batch, and an unlucky batch gives you a really high loss, which gives you a really high gradient, which shocks your model and shocks the optimization. Clipping upper bounds the magnitude of the shock. His honest assessment of the technique:
It's a bit of a hacky solution, it's about like a patch on top of like deeper issues, but people still do it fairly frequently. Andrej Karpathy, 2:18:31
And then the piece of advice buried in it, which is worth more than the clipping. clip_grad_norm_ returns the norm, and he always prints it, because it is useful information. A well behaved norm means things are good. A climbing norm means things are bad and destabilizing during training. A spike means there is some kind of issue or instability.
On this run the norm comes out high at the start, around 30, and stabilizes below one as training continues. That is not uncommon early on: the model is completely random, there is a ton of learning happening very early, and most of what it is learning is the biases of the output tokens, which is an unstable time. The network usually stabilizes in a very few iterations. He still looks at the actual sequence, 28 then 6 then 2 then 10, and calls it "not completely insane but just kind of a little bit funky".
The learning rate schedule: warmup then cosine decay
The paper: cosine decay for the learning rate down to 10 percent of its value over the first 260 billion tokens, then training continues at 10 percent after that, with a linear warmup over the first 375 million tokens.
He implements it himself rather than using PyTorch's schedulers, and says why: it is five lines of code and he fully understands what is happening inside them. "I don't love to use abstractions where they're kind of inscrutable and then I don't know what they're doing." Personal style, stated as such.
The shape: start at essentially zero, ramp up linearly, come down with a cosine form to a minimum learning rate that is up to you. Setting the learning rate in PyTorch he calls "a little bit gnarly", because you have to iterate over the optimizer's parameter groups and set it in a for loop, even though there is currently only one group.
The peak value comes straight out of the GPT-3 table of per model hyperparameters, and this is a real change. His debugging default was 3e-4. GPT-3 Small, which is 12 layers and 768 dimensions and therefore roughly a GPT-2 124M, uses 6e-4. So the max learning rate doubles, and the minimum is 10 percent of the max per the paper's description.
A small implementation detail with a reason: the warmup uses it + 1, so that on the zeroth iteration you are not using exactly zero, because updating with a learning rate of zero would not be useful.
And one place where he knowingly departs from the paper. GPT-3's decay horizon is shorter than its training horizon: it decays to 10 percent at 260 billion tokens and then trains the remaining 40 billion at 10 percent. In his implementation the decay time and the max steps are exactly equal. "So it's not exactly faithful, but it's okay for us and for our purposes right now. I don't think it makes too, too big of a difference honestly."
He also notes the whole question is open: cosine was popularized by GPT-2 and GPT-3, people have come up with all kinds of other schedules, and which one is most effective is an active area of research.
The batch size ramp, deliberately skipped
The paper describes a gradual, linear batch size increase: start very small and ramp up to a big batch size over time. He skips it, and gives three reasons, in order:
- It complicates the arithmetic. You are changing the number of tokens processed at every single step of the optimization, and he likes to keep that math very simple.
- It is not a major improvement.
- It is not an algorithmic improvement, it is a systems and speed improvement. And then the argument for why, which is the interesting part:
Early in the optimization the model is in a very atypical setting. Mostly what it is learning is to ignore the tokens that do not come up in the training set very often, and some very simple biases. So every single example you put through the network is basically just telling you "use these tokens and don't use these tokens", which means the gradients from every example are extremely highly correlated. They all look roughly the same.
Why are you doing batch sizes of like millions when, if you do a batch size of 32k, you're basically getting the exact same gradient early on in the training? And then later in the optimization, once you've learned all the simple stuff, that's where the actual work starts, and that's where the gradients become more decorrelated per example, and that's where they actually offer you sort of statistical power in some sense. Andrej Karpathy, 2:27:43
Sampling without replacement, already satisfied
The paper says data is sampled without replacement during training until an epoch boundary is reached. He walks through what that means: they are not drawing a sequence from a fixed pool and returning it, they are exhausting a pool, so a drawn sequence is gone until the next epoch. His loader already iterates over chunks of data in order, so there is no replacement and nothing becomes eligible again until the next pass. Nothing to implement.
Weight decay 0.1, on two dimensional parameters only
The paper: all models use a weight decay of 0.1 to provide a small amount of regularization.
PyTorch's AdamW default is 0.01, so this is ten times higher than the default. Rather than passing it in flat, he writes a configure_optimizers method on the model that returns the optimizer, because the decay needs to be applied selectively. The split:
- Decay every parameter with two or more dimensions. That is the matrices that participate in matrix multiplications, and the embeddings.
- Do not decay one dimensional tensors. The biases, and the LayerNorm scales and biases. "It doesn't really make sense to weight decay those."
The reason to decay at all, which he covered in an earlier video and recaps here: you can view it as a regularization, because pulling down all the weights forces the optimization to use more of the weights and does not allow any single weight to get way too large. It forces the network to distribute the work across more channels. "There's sort of like a pull of gravity on the weights themselves."
The counts his script prints: 50 decayed tensors, holding most of the parameters, and 98 non decayed tensors, which are mostly the biases and the layer norm parameters and amount to only about 100,000 parameters.
Fused AdamW: 93 to 90 milliseconds
The last of the pure speed wins, and it sneaks in here. torch.optim.AdamW grew a fused option in a later PyTorch version, and because it did not always exist he guards it with inspect.signature, checking whether fused is a valid keyword before passing it.
What it does: instead of iterating in a for loop over every parameter tensor and updating it, which launches a lot of kernels, all those kernels are fused into one. A single kernel call updates all the parameters, and all that launch overhead disappears. It is kernel fusion for the AdamW update specifically.
PyTorch does not default to it, because it was relatively new and they wanted to give it sufficient bake time, but it is a lot faster when it is available and you are running on CUDA. His advice: if you have it, use it. "I'm not actually 100 percent sure why they don't default to it, it seems fairly benign and harmless."
Result: 93 down to 90 milliseconds per step.
One closing honesty note on this whole section, which is easy to skip past and should not be:
The relationship between weight decay, learning rate, batch size, the Adam parameters beta one beta two, the epsilon and so on, these are very complicated mathematical relationships in the optimization literature, and for the most part in this video I'm just trying to copy paste the settings that OpenAI used. But this is a complicated topic, quite deep. Andrej Karpathy, 2:34:24
| Setting | PyTorch default | GPT-3 paper | What he ships |
|---|---|---|---|
| Optimizer | AdamW | Adam (he reads it as AdamW) | AdamW, fused=True when available |
| betas | 0.9, 0.999 | 0.9, 0.95 | 0.9, 0.95 |
| eps | 1e-8 | 1e-8 | 1e-8, set explicitly |
| Gradient clipping | none | global norm 1.0 | 1.0, and he prints the returned norm every step |
| Peak learning rate | 1e-3 | 6e-4 for GPT-3 Small, 12 layers and 768 dims | 6e-4, up from his 3e-4 debugging default |
| Minimum learning rate | n/a | 10 percent of peak | 10 percent of peak |
| Warmup | none | linear over the first 375 million tokens | 715 steps, which is 375e6 divided by 2^19. He calls it very mild and says 100 would probably do |
| Decay schedule | none | cosine to 10 percent over the first 260 billion of 300 billion tokens, then flat | cosine over the full run: decay time equals max steps, knowingly unfaithful, judged not to matter |
| Weight decay | 0.01 | 0.1 | 0.1, applied only to parameters with 2 or more dims. 50 tensors decayed, 98 not, the 98 holding about 100,000 parameters |
| Total batch size | n/a | 0.5M tokens for GPT-3 Small | 524,288 = 2^19, reached by gradient accumulation |
| Batch size ramp | n/a | linear ramp from small to large | skipped: complicates the arithmetic, and early gradients are highly correlated anyway |
| Sequence length | n/a | 2048 | 1024, GPT-2's value. He gives the exact change for fidelity: T=2048 and micro batch 32, so they still multiply to half a million |
| Data sampling | n/a | without replacement until an epoch boundary | already satisfied by iterating chunks in order |
| Training tokens | n/a | 300 billion | 10 billion for the main run, 40 billion overnight. Both beat GPT-2 124M |
Gradient accumulation, and the normalization bug that catches everybody
The GPT-3 table lists a batch size per model, and the pattern across the sizes is clear: bigger networks get slightly lower learning rates and bigger batch sizes. GPT-3 Small uses 0.5 million tokens per batch.
He does the division out loud, because the units trip people. Half a million is a count of tokens, and every row is 1024 tokens, so 0.5e6 divided by 1024 is a batch of about 488 rows. "The problem is I can't come in here and set this to 488 because my GPU would explode. This would not fit for sure."
But he still wants that batch size, and the reason is not stubbornness. The batch size is correlated with all the other optimization hyperparameters, the learning rates among them, so if you want a faithful representation of the recipe you need the batch size the recipe assumes.
The answer is gradient accumulation, which simulates in a serial way any arbitrary batch size you set. Run many forward and backward passes, let the gradients add up, then do a single update.
The arithmetic as he sets it up:
total_batch_size = 524288, which is 2 to the 19, a nice number that is roughly half a million.- The micro batch is still B of 16, T of 1024, so one forward and backward covers 16,384 tokens.
- 2^19 divided by 16,384 gives grad_accum_steps = 32.
- At about 100 milliseconds per forward and backward, 32 of them makes every optimizer step roughly 3 seconds. Napkin math, stated as such.
Then the subtlety. He writes the obvious inner loop, 32 micro steps of forward and backward before everything else, pauses, and says this is actually incorrect, and invites you to work out why before he fixes it. The demonstration is in a notebook, on a toy problem, and it is the clearest treatment of this bug anywhere:
- A tiny network taking a vector of 16 numbers and returning one number, four random examples, four targets, mean squared error loss.
loss.backward(), look at the gradient. Call that the reference.- Now do it the gradient accumulation way: four accumulation steps, one example each,
backward()on each, gradients adding up. - The gradients do not match.
The reason: MSELoss defaults to reduction='mean', so the real objective has a one quarter in front of it, averaging over the four examples. In the accumulation version each loop's objective is a single example's squared error with no one quarter, and accumulating gradients is equivalent to a sum in the loss. So the accumulated version is missing the normalizer.
The fix is one line: loss = loss / 4. That puts the one quarter back in front of every individual loss, and when they accumulate by summing, every component carries its quarter. Run it again and the gradients are now identical.
Which maps straight onto the real model, because F.cross_entropy also defaults to a mean reduction, over all B * T elements. So:
loss = loss / grad_accum_steps
In the same way exactly, we are scaling down the loss so that when we do loss.backward, which basically corresponds to a sum in the objective, we are summing up the already normalized loss, and therefore when we sum up the losses divided by grad accum steps we are recovering the additional normalizer. Andrej Karpathy, 2:44:36
Two cleanups follow. Printing needs a loss_accum variable initialized to zero and accumulated into, using .detach() so the tensor comes off the graph and he is just tracking values, because otherwise he would be printing only the final micro step's loss. And the tokens processed per step is now B * T * grad_accum_steps.
And a nice consequence he points out: once you have a total batch size and gradient accumulation, the micro batch B becomes purely a performance knob. Big GPU, set it to 32 and go a bit faster. Very small GPU, try 8 or 4. You get the exact same optimization and the same answers up to floating point error, because the accumulation handles everything serially.
Eight GPUs: distributed data parallel
Now is the time to bring out the heavy weapons. You've noticed that so far we've only been using a single GPU for training, but actually I am paying for eight GPUs here, and so we should be putting all of them to work. Andrej Karpathy, 2:46:38
The tool is PyTorch's DistributedDataParallel. He flags up front that there is also a legacy DataParallel and recommends you not use it.
The model is simple to state. Eight GPUs means eight processes, one assigned to each GPU. Each process runs the training loop exactly as built so far, as far as it is concerned nothing has changed, except that secretly there are eight of them, they are each processing slightly different parts of the data, and one new step is added at the end: the gradients get averaged across all of them.
Launching and the environment variables
You no longer run python train_gpt2.py. You run torchrun, which launches eight copies in parallel and sets environment variables so each process can look up which one it is:
torchrun --standalone --nproc_per_node=8 train_gpt2.py
Three variables matter:
WORLD_SIZE, 8 for him, the total number of processes.RANK, the index of this process, 0 through 7. Every process runs the exact same code at roughly the same time; the rank is the only difference between them, and it is how you coordinate that they do not all run on the same data.LOCAL_RANK, the rank of the GPU within a single node. Only relevant in a multi node setting, which he is not in, so for him it runs 0 to 7 as well.
The presence of RANK in the environment is also, he notes, a somewhat bad way to detect whether DDP is running. If it is not set, the script falls back to single GPU: rank zero, world size one, master process true, autodetect the device, business as normal.
local_rank sets the device to cuda:<local_rank>, so no two processes collide on the same GPU. And he creates a boolean he uses everywhere after:
master_process = ddp_rank == 0. Process zero, arbitrarily, does all the printing, logging and checkpointing. The others are thought of as compute processes that assist.
Reading code with eight interpreters in your head
The advice for the rest of the section is the practical kind:
The tricky thing with running multiple processes is you always have to imagine that there's going to be eight processes running in parallel. So as you read the code now you have to imagine there's eight Python interpreters running down these lines of code, and the only difference between them is that they have a different DDP rank. Andrej Karpathy, 2:51:41
They all come to the same lines, all pick the same seed, all build the identical model, "completely unaware of the other copies running". So every calculation that depends on how much data exists has to be adjusted for world size and rank.
What actually changes
Gradient accumulation steps. Now total_batch_size / (B * T * ddp_world_size). With 16 by 1024 on 8 GPUs that is 131,072 tokens in a single forward and backward across the box, and 524,288 divided by 131,072 gives grad_accum_steps = 4, down from 32. He checks the division comes out clean, which it does.
Printing. Eight processes hit every print and you get eight copies, so everything informational gets guarded by if master_process. He demonstrates the failure first, with a script that just prints its rank and exits, and the output is instructive: process 5 prints first "just by chance", then zero, then three and two, and because a process exits without destroy_process_group, DDP complains that the process group has not been destroyed before destruction. In a real application you want to call destroy_process_group() so you clean up properly and NCCL does not complain. The ordering is not something you can guarantee; it depends on how the operating system scheduled the processes.
The data loader. Every process must get its own chunk, so the loader takes the rank and the number of processes:
- The starting position is not zero, it is
B * T * process_rank. Rank zero starts at zero, rank one atB * T, rank two at2 * B * T, and so on, striding the processes out across the file. - Advancing is not by
B * Tany more, it is byB * T * num_processes, because the whole chunk is consumed per step and the position has to jump the entire chunk. - The wraparound check and the reset follow the same pattern. With rank zero and one process the behavior is identical to before, which is the property you want.
Wrapping the model. model = DDP(model, device_ids=[ddp_local_rank]). He notes the documentation is extensive, full of caveats, and that everything complexifies by a factor of ten when multiple processes are involved. He also notes the docs for device_ids specifically are "extremely unclear" and the comment explaining it is "roughly nonsensical", but he is pretty sure it has to be the local rank, not the rank.
What DDP does for you: the forward pass behaves identically, nothing changes there. In the backward pass, once the backward is over on each independent GPU, each GPU has gradients for all parameters, and DDP calls an all reduce, averaging across all the ranks and depositing that average back on every rank. It is a bit more involved than that, because as the backward pass moves through the layers of the Transformer it can dispatch the communication for gradients that are already done while the backward is still running, so the communication overlaps the computation. More efficient that way.
raw_model. model.configure_optimizers no longer works, because model is now a DDP wrapper. The real module is at model.module, so he keeps a raw_model reference right after wrapping and calls configure_optimizers on that.
The synchronization problem, and the naughty fix
Here is the part that matters most, and the part he is openly unhappy about.
By default DDP synchronizes gradients after every loss.backward(). But inside a gradient accumulation loop that is extremely wasteful: for the first 31 micro steps, or the first 3 on eight GPUs, you are only depositing gradients locally and you do not want to pay for an all reduce. You want to add them up locally, and all reduce exactly once, on the very last micro step.
PyTorch's sanctioned way is the no_sync() context manager, a context manager that disables gradient synchronization so gradients accumulate without communication, and then you do the final step outside it.
They are asking us to do
with ddp.no_sync(), do the gradient accumulation, accumulate grads, and then they are asking us to do DDP again with another input and backward. And I just really don't love this. I just really don't like it, the fact that you have to copy paste your code here and use a context manager. This is just super ugly. Andrej Karpathy, 3:03:24
So he reads the source, finds that entering the context manager simply toggles a variable called require_backward_grad_sync, and sets that variable directly instead, right before loss.backward(), true only when the micro step is the last one.
He is completely upfront that this is a hack:
This is a naughty thing to do, because they could probably change the DDP and this variable will go away. But for now I believe this works, and it allows me to avoid the use of context managers and code duplication. Andrej Karpathy, 3:04:25
The loss also has to be averaged
One more thing that falls out of the gradients being averaged. loss_accum lives outside the DDP container, so it is not averaged. If you print it on the master process you are printing only the loss that rank zero happened to see on its own slice of the data. Since the gradients are averaged, the loss you report should be too.
So: import torch.distributed as dist, then dist.all_reduce(loss_accum, op=dist.ReduceOp.AVG). The tensor exists on every rank, the all reduce creates the average and deposits it on all of them, and now the number the master process prints is the same number every rank holds.
And finally the token counter has to be multiplied by the world size as well, because the box really is processing that many more tokens.
The run
8 GPUs, grad accum of 4, and the throughput is 1.5 million tokens per second. "Wow, we're going really fast. These are some serious numbers." Tiny Shakespeare at 338,000 tokens is now so small that it is being looped over many times per minute, which is the signal that it is time for a real dataset.
A last consistency check he does, and it is a good one, because the numbers do not match between a single GPU doing 32 accumulation steps and eight GPUs doing 4. The reason is boring: the loader is looking for an entire page of data for all eight GPUs at once, and when that chunk exceeds the remaining tokens it loops, so the single GPU run and the eight GPU run reset at slightly different places and see slightly different batches.
To convince himself nothing is wrong he shrinks the total batch size to 32,768, which is 4 by 1024 by 8, so the single GPU does 8 accumulation steps and the eight GPU run does one each. That reduces the boundary effects of the data loader, and the numbers match.
The datasets: from Shakespeare to FineWeb-Edu
With the machinery finished, tiny Shakespeare has been outgrown, and he goes and looks at what GPT-2 and GPT-3 actually trained on.
GPT-2 used WebText, and it was never released. The paper's description: they scraped all outbound links from Reddit with at least three karma, and that was the starting point. 45 million links, collected, text extracted, ending up at 40 GB of text. There is a reproduction attempt called OpenWebText.
GPT-3 used a mixture, and it was never released either. The GPT-3 paper has a training dataset section where Common Crawl gets discussed properly. His assessment of Common Crawl by itself is blunt:
It's not a very high quality dataset all by itself, because it is extremely noisy. This is a completely random subset of the internet and it's much worse than you think. So people go into great lengths to filter Common Crawl, because there's good stuff in it but most of it is just like ad spam, random tables and numbers and stock tickers, and it's just total mess. Andrej Karpathy, 3:11:09
Which is why people train on curated data mixtures. Typically a large chunk, for example 50 percent of the tokens, will be Common Crawl, and then you add WebText2, books, Wikipedia, and whatever else you decide.
Since neither original dataset exists publicly, he names the modern stand ins:
- RedPajama, and more specifically the SlimPajama subset, which is a cleaned and deduplicated version of it.
- C4, also Common Crawl as far as he knows, processed differently.
- Plus GitHub, books, arXiv, Wikipedia and Stack Exchange, the usual components of these mixtures.
And then the one he picks. FineWeb is an attempt to collect really high quality Common Crawl data and filter it, in this case down to 15 trillion tokens. More recently Hugging Face released the FineWeb-Edu subset, 1.3 trillion tokens of educational content and 5.4 trillion of high educational content, filtering Common Crawl to very high quality educational subsets. He recommends the FineWeb write up as "really fascinating reading" if you care about data mixtures and how data gets processed at these scales.
He uses the sample-10BT subsample, 10 billion tokens, and gives the reason: in his previous experiments that is enough to get really close to GPT-2 performance, and it is simple enough to work with.
One detail about the filtering that is worth knowing, and he flags it as pretty cool: the FineWeb-Edu filters were applied automatically using Llama 3 70B. An LLM judges which content is educational, and that judgement is what makes it through the filter. He browses the dataset viewer to check: nuclear energy in France, Mexican America, some Mac PJs. "Actually it seems like their filters are working pretty well."
Pre-tokenizing into shards
A separate fineweb.py downloads the dataset, pre-processes and pre-tokenizes everything, and writes shards to local disk. He does not walk the whole script because it is "not as interesting and not as LLM centric", but the choices he does narrate are the ones you would get wrong:
- Every document's token list starts with the end of text token, id 50256. It is a special token in the GPT-2 tokenizer and, despite the name, it is the first token that begins a document.
- The tokens are stored as
np.uint16. The reason is exact: 2 to the 16 minus 1 is 65,535, and the GPT-2 maximum token id is well below that, so uint16 saves a little space. - 100 shards of exactly 100 million tokens each, which is 10 billion in total. Shard
000000is the validation shard, every other shard is training. Sharding rather than one massive file is "just kind of nicer from that perspective", because a single huge file can be hard to work with on disk. - Processing takes about 30 minutes.
And a bug, left in on camera. The script fails, and the cause is using float division in Python where it must be integer division, so a count is not an int. "Apologies for that."
With the shards on disk, the data loader grows accordingly: load the uint16 numpy file, convert to a torch.long tensor because that is what the layers up top expect, enumerate all the shards, take a split argument so it can serve train or val, and track a current shard as well as a current position. Run out of tokens in a shard, advance the shard, loop if needed, get the tokens, readjust.
The run that will actually be the run
The numbers for the real run:
- 2 to the 19 tokens per step, 10 billion tokens total, so 19,073 steps for one epoch.
- Warmup of 375 million tokens divided by 2 to the 19 gives 715 steps, which exactly matches GPT-3's warmup schedule. He thinks 715 is very mild and could be significantly more aggressive, probably even 100 would be good enough, but he leaves it to keep the GPT-3 hyperparameters exact.
- Then a last minute change: he can fit a bigger micro batch. B of 64, so 64 by 1024 per GPU across 8 GPUs multiplies out to exactly 524,288 tokens. No gradient accumulation at all.
- It fits. 330 milliseconds per iteration, 1.5 million tokens per second. 19,073 steps times 0.33 seconds is about 1.7 hours, "so one and a half hour run".
If this works then this is basically a serious pre-training run. We're not logging, we're not evaluating the validation split, we're not running any evaluations yet, so we haven't crossed our t's and dotted our i's. But if we let this run for a while we're going to actually get a pretty good model, and the model that might even be on par with or better than GPT-2 124M. Andrej Karpathy, 3:21:29
So he stops, and goes and crosses the t's.
| Constant | Value | Where it comes from |
|---|---|---|
| n_layer | 12 | GPT-2 124M. The paper's own parameter count table is wrong; the repo says the addition was in error |
| n_head | 12 | GPT-2 124M. GPT-2 XL uses 25, which he calls a really ugly number that caused real kernel headaches |
| n_embd | 768 | GPT-2 124M. 12 heads of 64 each |
| block_size | 1024 | GPT-2's maximum sequence length, so wpe is 1024 by 768. GPT-3 uses 2048 |
| vocab_size | 50257, run as 50304 | 50,000 BPE merges + 256 byte tokens + 1 end of text token. Padded up to the nearest multiple of 128 for speed |
| End of text token id | 50256 | The special GPT-2 token that, despite the name, begins every document in the shards |
| Parameters | 124 million, about 40 million of them shared | 768 times 50257 is tied between wte and lm_head, so roughly 30 percent of the model is one tensor used twice |
| Expected loss at init | 10.82, he measures 11 | -ln(1/50257). The check that says the distribution is diffuse before step one |
| Init std | 0.02, biases 0 | OpenAI's released model.py. Position embeddings are 0.01 there; he keeps 0.02. Zero biases is not the PyTorch default |
| Residual init scale | (2 * n_layer) ** -0.5 | GPT-2 paper. The 2 is because attention and the MLP each add into the residual stream |
| Tiny Shakespeare | 1 MB, ~1M chars, 338,000 tokens | About 40,000 lines and 200,000 words, at a roughly 3 to 1 character to token ratio |
| FineWeb-Edu sample | 10 billion tokens, 100 shards of 100 million | Shard 000000 is validation, the rest train. Stored as np.uint16 because the max token id is well under 65,535 |
| Micro batch, final run | B=64, T=1024 | 64 times 1024 times 8 GPUs equals the total batch exactly, so no gradient accumulation at all |
| Hardware | 8x A100 SXM 80GB | Rented from Lambda Labs. 19.5 TFLOPS fp32, 156 TF32, 312 bf16, about 2 TB/s of memory bandwidth |
| Memory used | 35 GB of 80 | At B=16, T=1024. If you hit out of memory, halve the batch and keep it a nice number |
| Good utilization | 60 percent | His rule of thumb. Half the time in a well tuned run the tensor cores are idle waiting for data |
Validation, sampling and checkpointing
The validation split
The val loader is the same class with split='val', serving the one held out shard. He also adds a reset() method to the loader, called at init and again before each evaluation, which is what makes a repeatable validation pass possible.
Every 100 steps, including step zero, and later every 250: put the model in eval mode, reset the val loader, and under no_grad accumulate the loss over 20 steps and average it. Same logic as the training loop with the backward pass removed. It is only inference, measuring the loss.
He is clear eyed about what it buys. With roughly infinite data, train and val loss should be about the same, so it tells you only a little about overfitting. But it would matter a lot if you went to multiple epochs, where a big enough model might start memorizing, and the validation split is how you would catch that. And:
In any case you would always want to have a validation split in a training run like this, so that you can make sure that you are not overfitting. Andrej Karpathy, 3:25:11
The other reason he wants it is sharper: you can initialize from the released GPT-2 124M and measure its loss on this validation split, which gives you a reference line. He flags the caveat himself. It is not a super fair comparison, because GPT-2 was trained on a very different data distribution, but it is an interesting data point and a good cross check.
Sampling, moved up and given its own RNG
The orphaned sampling code from the first hour gets deleted from the bottom of the script and moved up into the loop, so that once in a while the script validates, once in a while it samples, and it trains on every step.
One real change, and it is the kind of thing that quietly ruins reproducibility if you skip it. He creates a separate torch.Generator for sampling and passes it into torch.multinomial, specifically so that drawing samples does not touch the RNG state of the global random number generator used for training. Sampling stays completely outside the training loop. He seeds it so every rank gets a different seed.
And a problem he cannot solve on camera. torch.compile breaks the sampling and the HellaSwag evaluation, with "a really scary error from PyTorch, and I have no idea how to resolve it right now". So he turns compile off to get samples, which is why the run gets slower at this point, and then turns compile back on and gives up the samples. He says plainly that he hopes it is fixed by the time you see the code, and that he will fix the bug later.
The samples at step 1000, with the model just past the peak of the warmup:
Hello, I'm a language model, and I'm not able to get more creative. The model at step 1000, 3:25:43
Hello, I'm a language model, and languages file you're learning about here is or is the beginning of a computer. The model at step 1000, 3:25:43
His read: "this is still a garble, but we're only at iteration 1000 and we've only just barely reached maximum learning rate, so this is still learning". And then the line that is the best description of a half trained model anyone has written:
The model is still a young baby. Andrej Karpathy, 3:26:14
Logging and checkpoints
A log directory with a log.txt that records the train loss, the validation loss and the HellaSwag accuracies, opened for writing so it starts empty and then appended to. A simple text file, parsed later by a matplotlib cell in the notebook.
Checkpointing, added at the same place as the validation logging: every 5,000 steps, if you are the master process, save the model's state_dict. And then the warning about what a model checkpoint is not:
- To resume optimization you also need the optimizer state dict, because AdamW carries the
mandvbuffers per parameter. - And you have to be careful with your RNG seeds and random number generators.
- "If you wanted to exactly be able to resume optimization you have to think through the state of the training process." If you just want to save the model, the state dict is enough.
The reason to save at all, beyond resuming: you may want to evaluate the model much more carefully than he is doing here, where he is "only kind of winging the HellaSwag eval", using proper infrastructure like the EleutherAI evaluation harness and comparing against the OpenAI GPT-2 on many other tasks involving math, code or different languages.
And an aside about scope, worth stating because people do get confused about it:
Everything we've built here, this is only the pre-training step. The GPT here is, it dreams documents, it just predicts the next token. You can't talk to it like you can talk to ChatGPT. Andrej Karpathy, 3:55:17
To talk to it you fine tune into the chat format, and he says that is "not actually that complicated". Supervised fine tuning really means swapping in a dataset that is a lot more conversational with a user and assistant structure, filling in the user tokens and sampling the assistant tokens. "It's not a lot more deeper than that. Basically we swap out the dataset and continue training." But this video stops at pre-training.
HellaSwag: the evaluation, and why this one
The validation loss needs a companion that is held out, comparable and somewhat standard, and for that he uses HellaSwag, from a 2019 paper by Rowan Zellers and co authors.
What it is
A sentence completion dataset, multiple choice, four candidate endings sharing a context. The example he reads out verbatim:
A woman is outside with a bucket and a dog. The dog is running around trying to avoid a bath. She:
- A. rinses the bucket off with soap and blow dries the dog's head.
- B. uses a hose to keep it from getting soapy.
- C. gets the dog wet, then it runs away again.
- D. gets into a bathtub with the dog.
The options are constructed so that one is a natural continuation and the others are not, and some of them do not make sense at all. "Uses the hose to keep it from getting soapy, that makes no sense." Models that are not trained very well cannot tell these apart; models with a lot of world knowledge can.
The sentences come from ActivityNet and WikiHow, and the paper has a chart of the WikiHow domains, computers and electronics, home and garden and so on, giving it broad coverage of the kinds of things you need to know about the world to find the most likely completion.
The construction detail that makes it good: the incorrect options are deliberately adversarially sourced. They are not random sentences, they are generated by language models, and generated such that language models find them difficult and humans find them easy. The paper reports 95 percent human accuracy against 48 percent for the state of the art at the time.
And why it is now too easy
He does not oversell it. Five years later HellaSwag "has been totally just solved", and language models sit at 96 percent, so the last 4 percent is probably errors in the dataset or genuinely very hard questions. The dataset is, in his word, "crushed".
But it is still useful here, for a specific reason, and this is the part that generalizes to choosing any eval for a small model:
HellaSwag is a smooth eval, and it is an eval that offers quote unquote early signal. So early signal means that even small language models are going to start at the random chance of 25 percent, but they're going to slowly improve, and you're going to see 25, 26, 27 etc. And you can see slow improvement even when the models are very small and it's very early. Andrej Karpathy, 3:31:59
Smooth, early signal, and around long enough that everybody uses it. It is not used in the GPT-2 paper, but it is in the GPT-3 paper, which means published GPT-3 accuracies exist for every model size and he has a reference point.
How he actually runs it
Small models cannot do multiple choice. They do not understand the concept of associating a label with one of the options; "they don't understand that". So you have to give the task to them in a native form, which is token completion.
The construction, per example:
- Build a batch of 4 rows by T tokens. The shared context tokens are repeated across all four rows, then each row continues with one of the four options.
- The options differ in length, so T is the longest one, and the shorter rows get padded.
- You need three things out of this: the tokens, the correct label, and a mask marking which tokens are active option tokens, with zeros over the padding.
- Evaluate the cross entropy loss of predicting the next token across the option tokens of each row, average it per row, and pick the row with the lowest average loss, which is equivalently the highest average probability. That is the model's answer.
He believes this is also how GPT-3 did it, and then flags the fork in the road honestly. Other harnesses may run HellaSwag in a true multiple choice format, giving the context once followed by all four completions so the model can see the other options before it picks. That is an easier task, and models at this size cannot do it.
Our models are actually slightly handicapped in this way, that they are not going to see the other options, they're only going to see one option at a time and they just have to assign probabilities, and the correct option has to win out in this metric. Andrej Karpathy, 3:35:05
The implementation is a hellaswag.py that downloads the data, renders all 10,000 examples into the format above, and provides an evaluate function that can load a GPT-2 from Hugging Face and run the eval. He calls the code "kind of a little bit tedious honestly" and does not walk it line by line.
The reference numbers it produces, which are the numbers to beat:
- Random chance: 25 percent.
- GPT-2 124M: 29.55 percent by his script. "So we haven't gone too far."
- GPT-2 XL, 1558M: about 49 percent.
- The Eleuther harness gets slightly different numbers, and he is not 100 percent sure what the discrepancy is. His guess is that they do the multiple choice version rather than the completions, but he says he would have to look.
- Today's state of the art: more like 95 percent, "so these are definitely older models by now".
Then it goes into the training script, periodic like everything else, so he can track HellaSwag over time and see when and if the run crosses 29.55. Under DDP, each process takes only the examples where the index modulo the world size equals its rank, then the counts are packaged into tensors, all reduced with a sum, unwrapped back to integers, and the master process prints and logs the accuracy.
SECTION 4: results in the morning
He goes to bed. The next cell in the notebook parses the log file and plots it, "a lot of this is just like boring matplotlib code", and the first run is done.
The one epoch run: 10 billion tokens, about two hours
Two panels. On the left, the loss: training loss in blue, validation loss in orange, and a horizontal red line for the OpenAI GPT-2 124M checkpoint evaluated on the FineWeb-Edu validation split. On the right, HellaSwag, with the OpenAI GPT-2 124M in red and the GPT-3 124M in green.
The orange is below the red. The from scratch model surpasses the released GPT-2 124M on this validation set, with the caveat he repeats: the data distribution is very different from what GPT-2 trained on, so this is not an exactly fair comparison, but it is a good cross check.
And on HellaSwag, the one that is held out and comparable:
You see that we basically surpassed the GPT-2 124M model right here, which is really nice. Now interestingly we were able to do so with only training on 10 billion tokens, while GPT-2 was trained on 100 billion tokens. Andrej Karpathy, 3:45:00
So a 10 times learning efficiency gap, in 2024, against a model that was a serious research artifact in 2019. He does not leave that unexamined, and offers three possible explanations rather than claiming a win:
- GPT-2 was trained on a much wider data distribution. FineWeb-Edu is all English, not multilingual, and does not have much math or code. Math, code and multilingual capability "could have been stealing capacity from the original GPT-2 model".
- HellaSwag is five years old and might have leaked. It is possible that aspects of HellaSwag, in some way or even identically, made it into FineWeb's training set. "If that was the case then we are basically looking at the training curve instead of the validation curve." He then gives the mitigating fact: Hugging Face used HellaSwag as an eval when they created FineWeb-Edu, so he would hope they deduplicated against it. "But we can't be sure."
- The data is probably just better per token. The original GPT-2 dataset was WebText, and "it's possible that not a lot of care and attention went into the dataset, this was very early in LLMs, whereas now there's a lot more scrutiny on good practices around deduplication, filtering, quality filtering and so on".
That is three caveats on his own headline result, unprompted, which is most of why this video is trusted.
The loss curve that is wrong, and he says so
There is a visible problem in the plot, and he does not paper over it:
The other thing I wanted to address briefly is, look at this loss curve. This looks really wrong here. I don't actually know 100 percent what this is, and I suspect it's because the 10 billion sample of FineWeb-Edu was not properly shuffled, and there's some issue here with the data that I don't fully understand yet, and there's some weird periodicity to it. Andrej Karpathy, 3:46:33
His diagnosis: his own loader is "in a very lazy way sort of serializing all the tokens and just iterating all of them from scratch without doing any permutation or any random sampling ourselves", so it is inheriting whatever ordering the dataset has. He expects it will be fixed in the repo by the time you read it.
The overnight run: 40 billion tokens, four epochs, about eight hours
Having seen the one epoch result, he wanted to know how far it would push, so he made exactly one change, multiplied the token budget by four, and went to sleep for eight hours. Four epochs, roughly 40 billion tokens.
The result, narrated honestly in both directions:
- The periodicity is now clearly visible per epoch, confirming it is a data ordering artifact, "and that is to be determined".
- HellaSwag "actually went up by a lot", and the run almost reached the GPT-3 124M accuracy. Almost. Not quite. Final step around 76,290, final HellaSwag 33.24 percent.
- "It's too bad that I didn't sleep slightly longer. I think if this was a five epoch run we may have gotten here."
And a second efficiency claim, with the same framing as the first: the run is almost matching GPT-3 accuracy with 40 billion tokens where GPT-3 trained on 300 billion. "Again we're seeing about a 10 times improvement here with respect to learning efficiency." And again he says he does not know exactly what to attribute it to beyond the reasons already listed.
Two things he wants fixed for anyone doing multi epoch runs:
- The data loader goes through the data in exactly the same format and exactly the same order every epoch, which is suboptimal. You want to permute the documents randomly within every shard on every new epoch, and possibly permute the shards too. That would go a long way to decreasing the periodicity and is better for the optimization.
- The reason it matters is specific: in every row the documents follow each other, separated by an end of text token, so the documents are currently glued together in the exact same identical manner. "We actually want to break up the documents and shuffle them around, because the order of the documents shouldn't matter. Basically we want to break up that dependence because it's a kind of a spurious correlation."
Two parting hyperparameter notes
The learning rate is probably too low. He has seen people play with this in a related repository, and it turns out you can go about three times higher on the max learning rate:
For some reason the GPT-3 hyperparameters that we are inheriting are actually extremely conservative, and you can actually get away with a higher learning rate and it would train faster. So a lot of these hyperparameters are quite tunable, and feel free to play with them. They're probably not set precisely correctly. Andrej Karpathy, 3:51:14
And the exact change for GPT-3 fidelity, if you want it. GPT-3's sequence length is double GPT-2's, 2048 instead of 1024. So set T to 2048, and then to keep the same half million tokens per step, drop the micro batch to 32, "so they still multiply to half a mil". With that, as far as he is aware, the models would be roughly identical, because GPT-2 and GPT-3 are very very similar models.
The overnight samples
The same prompt, a model that has seen four times as much data:
Hello, I'm a language model, and I try to be as accurate as possible. The overnight model, 3:53:17
Hello, I'm a language model, not a programming language. I know how to communicate. I use Python. The overnight model, 3:53:17
Hello, I'm a language model, and I'm going to be speaking English and German. The overnight model, 3:42:59, from the halfway samples
His read on the progression: the predictions are "getting less and less random", the model "is a little bit more self-aware and using language that is a bit more specific to it being a language model", and the overnight samples are "a lot more coherent" than the 10 billion token ones if you pause and compare them side by side.
The shoutout to llm.c
Everything built in this video was building towards nanoGPT, the earlier repository. But there is a second nanoGPT implementation hiding in a more recent project: llm.c, a pure CUDA implementation of GPT-2 and GPT-3 training that uses CUDA directly and is written as CUDA.
The relationship between the two is the interesting part. The nanoGPT style train_gpt2.py inside llm.c acts as the PyTorch reference code for the C implementation, so the two are exactly matched, and the hope is that the C and CUDA version is faster. Scroll through that Python file and "you'll find a lot of things that very much look like things that we've built up in this lecture". Then train_gpt2.cu is the C and CUDA implementation, full of MPI, NCCL, GPU, CUDA and C and C++, and you have to be familiar with that.
Then he runs them side by side, one GPU each, llm.c on GPU 1 and PyTorch grabbing GPU 0 by default, and the race has a comic first act:
- llm.c compiles, allocates space, and is already stepping.
- PyTorch is still compiling, because
torch.compileis slower here than llm.c'snvccCUDA compile.
Then the steady state numbers, with the honest asterisk:
- llm.c: about 223,000 tokens per second.
- PyTorch: about 185,000 tokens per second.
- "Quite a bit slower, but I don't have full confidence that I exactly squeezed out all the juice from the PyTorch implementation."
And the verification that makes the comparison worth anything: line up the steps and the losses and the gradient norms printed by the two implementations are identical. Same computation, one runs faster.
His framing of the comparison is careful. This is a very specific implementation for GPT-2 and GPT-3, and PyTorch is a very general neural network framework, so they are not exactly comparable. But if you are only interested in training GPT-2 and GPT-3, llm.c is very fast, takes less space, is faster to start and faster per step.
Wrapping up, and what is left open
I think it's getting way longer than I anticipated, but we did cover a lot of ground, and we built everything from scratch. Andrej Karpathy, 3:59:24
His own summary: they looked at the GPT-2 and GPT-3 papers, looked at how you set up these training runs and all the considerations involved, wrote everything from scratch, and then over a two hour run or an overnight run matched the 124 million parameter checkpoints of GPT-2 and GPT-3 "to a very large extent". And in principle the code would train bigger models too, if you have the patience or the computing resources, so you could think about the bigger checkpoints as well.
Then, instead of ending on the win, he lists the open bugs:
- The loss periodicity, which he suspects is the FineWeb-Edu data sampling.
- Why
torch.compilecannot be turned on, because it currently breaks generation and HellaSwag. "What's up with that." - The data loader should permute the data when it reaches boundaries.
He expects to document those over time in build-nanogpt, and he makes a point about how that repository was built that is worth more than it sounds:
I will be releasing all this code, and actually I've been very careful about making git commits every time we add something. And so I'm going to release the entire repo that starts completely from scratch all the way to now, and so everything should be exactly documented in the git commit history. Andrej Karpathy, 3:27:45
Which is why the repo is a usable companion rather than a finished artifact: there is one commit per step of the video, so you can check out the state of the code at any point in these four hours. Questions go to the repository's discussions tab, issues or pull requests, or the Zero to Hero Discord.
Key takeaways
- Writing a correct GPT-2 takes under an hour and fits in under 100 lines. Hugging Face's equivalent is 2,000. Making it train efficiently takes the rest of the day, and that is the part that is actually a skill.
- Verify a from scratch implementation by loading the released weights into it. Match the upstream naming scheme while you write it, port the weights, generate, confirm the output is sensible, and only then throw the weights away and train. "We just have to rediscover a good setting of the weights, but from scratch."
- Check your loss at initialization against the number you can compute by hand. For a 50257 token vocabulary that is -ln(1/50257), which is 10.82. He got 11, which told him nothing was confidently wrong before a single step.
- Precision, kernel fusion and flash attention are memory bandwidth wins, not arithmetic wins. TF32 promises 8 times and gives 3, because the data being moved is still float32. Flash attention does more FLOPs than the code it replaces and is 27 percent faster, because it never writes the attention matrix to memory. The GPU is starved for data, not for math.
- Ugly tensor dimensions cost real speed. Padding the vocabulary from 50257 to 50304 adds computation and runs 4 percent faster on PyTorch nightly, or about 30 percent faster on 2.3.1 and earlier, because CUDA kernels work in power of two block tiles and a bad remainder triggers inefficient boundary kernels.
- Take the hyperparameters from the GPT-3 paper, because GPT-2's does not have them. AdamW with betas 0.9 and 0.95, eps 1e-8, global gradient norm clipped at 1.0, 6e-4 peak learning rate with a 715 step linear warmup and cosine decay to 10 percent, weight decay 0.1 on two dimensional parameters only, fused AdamW.
- Log the gradient norm, not just the loss.
clip_grad_norm_returns it. Well behaved is good, climbing is bad, a spike means the data or the schedule has a problem. - If you accumulate gradients, divide the loss by the accumulation count. Cross entropy reduces by mean and accumulation sums, so without the division your gradients are silently the wrong magnitude. He proves it on a four example toy problem rather than asserting it.
- Weight tying saves about 40 million of 124 million parameters, and he found it by comparing
.data_ptr()on two tensors in the released checkpoint rather than by reading about it. - Scale the initialization of layers that write into the residual stream by one over the square root of twice the number of layers. Twice, because attention and the MLP each add into the stream. He derives the need for it by adding 100 unit normal draws to a zero vector and watching the standard deviation reach 10.
- Better data, not a better architecture, is why a 2024 run beats the 2019 model on a tenth of the tokens. Though he lists two other possible explanations for his own result and refuses to pick one.
- The economics changed completely. Ten dollars and an hour and a half of eight rented A100s, for a model that was a serious research artifact five years earlier.
Where this sits in the LLM Learning track
This is the second half of the build portion of the track, immediately after the tokenizer video, and it is where everything the earlier videos describe abstractly turns into code you can run. The attention block, the parameter counts, the scaling relationships: they all appear here with exact shapes and exact values, in a file you can execute.
It is also the video that makes the efficiency arguments from the rest of the track concrete. Every optimization in the middle two hours is an arithmetic intensity argument in disguise, and the reason the chapter titles carry millisecond counts is that none of them is taken on faith.
If you watch one video in this track with a keyboard in front of you rather than a notebook, make it this one, and keep build-nanogpt open beside it. The repository has one commit per step of these four hours, so you can check out the code at the exact moment of any chapter below.
Chapters
All 31 entries are Karpathy's own chapter markers, reproduced exactly as written, typo in the section 3 heading and all. Note that seven of them carry a millisecond count in the title, which is the whole method of the video compressed into a table of contents.
- 0:00:00 intro: Let’s reproduce GPT-2 (124M)
- 0:03:39 exploring the GPT-2 (124M) OpenAI checkpoint
- 0:13:47 SECTION 1: implementing the GPT-2 nn.Module
- 0:28:08 loading the huggingface/GPT-2 parameters
- 0:31:00 implementing the forward pass to get logits
- 0:33:31 sampling init, prefix tokens, tokenization
- 0:37:02 sampling loop
- 0:41:47 sample, auto-detect the device
- 0:45:50 let’s train: data batches (B,T) → logits (B,T,C)
- 0:52:53 cross entropy loss
- 0:56:42 optimization loop: overfit a single batch
- 1:02:00 data loader lite
- 1:06:14 parameter sharing wte and lm_head
- 1:13:47 model initialization: std 0.02, residual init
- 1:22:18 SECTION 2: Let’s make it fast. GPUs, mixed precision, 1000ms
- 1:28:14 Tensor Cores, timing the code, TF32 precision, 333ms
- 1:39:38 float16, gradient scalers, bfloat16, 300ms
- 1:48:15 torch.compile, Python overhead, kernel fusion, 130ms
- 2:00:18 flash attention, 96ms
- 2:06:54 nice/ugly numbers. vocab size 50257 → 50304, 93ms
- 2:14:55 SECTION 3: hyperpamaters, AdamW, gradient clipping
- 2:21:06 learning rate scheduler: warmup + cosine decay
- 2:26:21 batch size schedule, weight decay, FusedAdamW, 90ms
- 2:34:09 gradient accumulation
- 2:46:52 distributed data parallel (DDP)
- 3:10:21 datasets used in GPT-2, GPT-3, FineWeb (EDU)
- 3:23:10 validation data split, validation loss, sampling revive
- 3:28:23 evaluation: HellaSwag, starting the run
- 3:43:05 SECTION 4: results in the morning! GPT-2, GPT-3 repro
- 3:56:21 shoutout to llm.c, equivalent but faster code in raw C/CUDA
- 3:59:39 summary, phew, build-nanogpt github repo
Notable quotes
When we talk about reproducing GPT-2 we have to be careful, because in particular in this video we're going to be reproducing the 124 million parameter model. Andrej Karpathy, the first caveat, 0:00
The reason my numbers the way I say them disagree with this table is that this table is wrong. Andrej Karpathy, on the parameter counts in the GPT-2 paper, 1:00
Today you can reproduce this model in roughly an hour or probably less even, and it will cost you about 10 bucks if you want to do this on the cloud. Andrej Karpathy, the economics, 2:35
You can tell, for example, that because they're a bit more jagged and they're kind of noisy, you can tell that this model was not fully trained. Andrej Karpathy, reading the released GPT-2 position embeddings as a diagnostic, 9:14
You actually prefer to have a single clean residual stream all the way from supervision all the way down to the inputs, the tokens. Andrej Karpathy, on why GPT-2 moved the layer norms, 18:26
So the attention is the reduce and the MLP is the map, and what you end up with is that the Transformer just ends up being a repeated application of map reduce. Andrej Karpathy, 19:58
Today there's no real good reason to use the approximate version. You'd prefer to just use the exact version, because my expectation is that there's no big difference anymore and this is kind of like a historical quirk. But we are trying to reproduce GPT-2 exactly, and GPT-2 used the tanh approximate version, so we prefer to stick with that. Andrej Karpathy, on the GELU approximation, 22:33
We don't have to basically use this file from Hugging Face, which is fairly long, this is 2,000 lines of code. Instead we just have a less than 100 lines of code, and this is the complete GPT-2 implementation. Andrej Karpathy, 27:44
You have to be careful, because you can't just do buff.to(device). It's not stateful, it doesn't convert it to be a device, it instead returns a pointer to a new memory which is on the device. Andrej Karpathy, on the device bug he leaves in, 1:00:21
Not only are these two separate tensors that happen to have the same shape and elements, they're actually pointing to the identical tensor. Andrej Karpathy, discovering GPT-2's weight tying by comparing data pointers, 1:07:33
Unfortunately the GPT-2 paper and the GPT-3 paper are not very explicit about initialization, so we kind of have to read between the lines. Andrej Karpathy, 1:13:40
You always want to start with: what hardware do you have, what does it offer, and are you fully utilizing it? Andrej Karpathy, opening section 2, 1:22:26
It turns out empirically that for deep learning as a computational workload this is way too much. Andrej Karpathy, on float32 being the PyTorch default, 1:24:02
Many of the deep learning workloads for training are memory bound, and what that means is actually that the tensor cores that do all these extremely fast multiplications, most of the time they're waiting around, they're idle, because we can't feed them with data fast enough. Andrej Karpathy, the thesis of the whole optimization section, 1:27:06
Typical utilizations of your hardware, if you're getting 60 percent utilization you're actually doing extremely well. So half of the time, in a well tuned application, your tensor cores are not doing multiplies because the data is not available. Andrej Karpathy, 1:27:37
The reason I like TF32 is because if you can tolerate a little bit of a precision fudge then this is free. Like none of your code sees this, it's fully internal to the operation, and the operation to you just goes 8x faster. Andrej Karpathy, 1:32:17
Even though the TF32 offers in principle a lot faster throughput, all of these numbers everywhere are still float32s, and it's float32 numbers that are being shipped all over the place through the memory system, and it's just costing us way too much time to shuttle around all this data. Andrej Karpathy, on why the 8x became 3x, 1:38:29
torch.compile is really quite incredible infrastructure from the PyTorch team, and it's basically a compiler for neural networks. Like it's almost like GCC for C and C++ code. This is just the GCC of neural nets. Andrej Karpathy, 1:48:10
There's no real good reason for you to not use torch.compile in your PyTorch. I kind of feel like you should be using it almost by default, unless you're debugging. Andrej Karpathy, 1:49:15
Flash attention actually, if you just count the number of FLOPs, flash attention does more FLOPs than this attention here. But flash attention is actually significantly faster. Andrej Karpathy, 2:01:37
Great example, I think, of being aware of memory hierarchy, the fact that FLOPs don't matter, the entire memory access pattern matters, and that torch.compile is amazing but there are many optimizations that are still available to us that potentially torch.compile cannot find. Andrej Karpathy, 2:04:40
We are now getting to one of my favorite optimizations, and it is simultaneously the dumbest and the most brilliant optimization. Andrej Karpathy, introducing the vocabulary padding, 2:06:48
Basically, scan your code and look for ugly numbers is roughly the heuristic. Andrej Karpathy, 2:07:50
So functionally nothing breaks. We're using a bit more extra memory, but otherwise this is a harmless operation as far as I can tell. And we're adding calculation, but it's running faster. Andrej Karpathy, on padding the vocabulary to 50304, 2:12:55
GPT-2, we have the weights but no details. GPT-3, we have lots of details but no weights. Andrej Karpathy, 2:15:59
Sometimes you could get a spike in the norm, and that means there's some kind of an issue or an instability. Andrej Karpathy, on why you should log the gradient norm, 2:19:03
I don't love to use abstractions where they're kind of inscrutable and then I don't know what they're doing. So, personal style. Andrej Karpathy, on writing his own learning rate scheduler, 2:23:07
Why are you doing batch sizes of like millions when, if you do a batch size of 32k, you're basically getting the exact same gradient early on in the training? Andrej Karpathy, on skipping GPT-3's batch size ramp, 2:27:43
The problem is I can't come in here and set this to 488, because my GPU would explode. Andrej Karpathy, on GPT-3's half million token batch, 2:35:56
There's like a subtle and deep issue here, and this is actually incorrect. So I invite you to think about why this is not yet sufficient. Andrej Karpathy, before the gradient accumulation normalization fix, 2:39:29
Now is the time to bring out the heavy weapons. You've noticed that so far we've only been using a single GPU for training, but actually I am paying for eight GPUs here. Andrej Karpathy, 2:46:38
As you read the code now you have to imagine there's eight Python interpreters running down these lines of code, and the only difference between them is that they have a different DDP rank. Andrej Karpathy, 2:51:41
This is a naughty thing to do, because they could probably change the DDP and this variable will go away. Andrej Karpathy, on toggling require_backward_grad_sync by hand instead of using no_sync, 3:04:25
It's not a very high quality dataset all by itself, because it is extremely noisy. This is a completely random subset of the internet and it's much worse than you think. Andrej Karpathy, on Common Crawl, 3:11:09
The filters here, by the way, were applied automatically using Llama 3 70B, I believe. And so basically LLMs are judging which content is educational, and that ends up making it through the filter. Andrej Karpathy, on FineWeb-Edu, 3:14:48
If this works then this is basically a serious pre-training run. Andrej Karpathy, launching the real run, 3:21:29
The model is still a young baby. Andrej Karpathy, on the samples at step 1000, 3:26:14
I will be releasing all this code, and actually I've been very careful about making git commits every time we add something. Andrej Karpathy, on why build-nanogpt has one commit per step, 3:27:45
HellaSwag is a smooth eval, and it is an eval that offers quote unquote early signal. Andrej Karpathy, on why this benchmark and not another, 3:31:59
Our models are actually slightly handicapped in this way, that they are not going to see the other options. Andrej Karpathy, on running HellaSwag as token completion rather than multiple choice, 3:35:05
Interestingly, we were able to do so with only training on 10 billion tokens, while GPT-2 was trained on 100 billion tokens. Andrej Karpathy, on surpassing GPT-2 124M, 3:45:00
The other thing I wanted to address briefly is, look at this loss curve. This looks really wrong here. I don't actually know 100 percent what this is. Andrej Karpathy, on the periodicity in his own results, 3:46:33
It's too bad that I didn't sleep slightly longer. I think if this was a five epoch run we may have gotten here. Andrej Karpathy, on almost reaching GPT-3 124M, 3:49:41
For some reason the GPT-3 hyperparameters that we are inheriting are actually extremely conservative, and you can actually get away with a higher learning rate and it would train faster. Andrej Karpathy, 3:51:14
Everything we've built here, this is only the pre-training step. The GPT here is, it dreams documents, it just predicts the next token. You can't talk to it like you can talk to ChatGPT. Andrej Karpathy, 3:55:17
I don't have full confidence that I exactly squeezed out all the juice from the PyTorch implementation. Andrej Karpathy, on llm.c being faster than his own PyTorch code, 3:58:54
I think it's getting way longer than I anticipated, but we did cover a lot of ground, and we built everything from scratch. Andrej Karpathy, 3:59:24
A note on names, from the captions
The caption track mangles technical names relentlessly, so a few corrections are worth stating once rather than silently fixing. Each of these is the correct spelling of what he actually says:
- Dan Hendrycks, the GELU author, appears in the captions as "Daniel Hendrix". The erf, the error function, becomes "the Earth function".
- GELU becomes "G" or "Galo", and the paper title "Gaussian Error Linear Units" becomes "Gan error linear units".
- HellaSwag appears as "H swag", "hos swag" and "helis swag".
- AdamW appears as "atom W", "addom w" and "ADD and W". all reduce appears as "alberu" and "alberon".
- torchrun appears as "torrun". NCCL appears as "nickel". nvcc appears as "nbcc".
- EleutherAI's lm-evaluation-harness appears as "Uther harness" and "Luther evaluation hardness".
- Xavier initialization appears as "Javier initialization". mantissa appears as "mantisa" and "Mena".
- Common Crawl appears as "common coll", "common craw" and "comic". RedPajama and SlimPajama survive as "red pajama" and "slim pajama".
- llm.c appears as "lm. C", "LL M.C" and "llm Doc". micrograd appears as "microad". SwiGLU appears as "swiglo". The BERT paper appears as "the bird paper".
- eager mode appears as "e mode", and cosine appears as "coign".
Two numbers in the captions are also wrong by a dropped digit, and the page uses the correct ones. The one epoch run is 19,073 steps, not "1973", which you can check against 10 billion divided by 2 to the 19 and against the "19073 times 0.33" he reads out moments later. And the overnight run's final step is around 76,290, consistent with four times 19,073.
Resources mentioned
The repositories
- build-nanogpt, the repo built in this video, with one git commit per step so you can check out the code at any point in the four hours. Discussions tab for questions. The two helper scripts he writes on camera are fineweb.py and hellaswag.py.
- nanoGPT, the cleaned up repository this was building towards.
- llm.c, the same training run in raw C and CUDA, with a nanoGPT style PyTorch file inside it acting as the reference implementation.
- micrograd, referenced for the fact that addition distributes gradients equally.
- openai/gpt-2, the original TensorFlow release, inference code and weights only.
- Hugging Face Transformers' modeling_gpt2.py, the roughly 2,000 line implementation his under 100 lines replaces.
The papers
- Language Models are Unsupervised Multitask Learners, the GPT-2 paper, and its announcement blog post. Section 2.3 is where the moved layer norms and the residual initialization scaling live.
- Language Models are Few-Shot Learners, the GPT-3 paper, source of every hyperparameter in section 3.
- Attention Is All You Need, the architecture he starts from and deletes half of, and the source of the weight tying.
- Using the Output Embedding to Improve Language Models, the 2017 paper that Attention Is All You Need cites for weight tying and that he reads the motivation out of.
- Gaussian Error Linear Units by Dan Hendrycks and Kevin Gimpel, and PyTorch issue 39853, where Hendrycks explains that the tanh approximation exists because erf was slow in TensorFlow.
- FlashAttention by Tri Dao and co authors, Stanford 2022, and FlashAttention-2.
- Online normalizer calculation for softmax, NVIDIA 2018, the online softmax trick that FlashAttention is built on, four years earlier.
- HellaSwag by Rowan Zellers and co authors, with the dataset and project page. Its sentences come from ActivityNet and WikiHow.
- BERT, named as the paper that picked up the approximate GELU.
- GLU Variants Improve Transformer, the SwiGLU paper, named as where modern networks like Llama 3 went instead.
- Understanding the difficulty of training deep feedforward neural networks, Glorot and Bengio, the Xavier initialization he compares 0.02 against.
The models and datasets
- GPT-2 124M on Hugging Face, which
from_pretrained("gpt2")actually gives you, and GPT-2 XL, the 1.5 billion parameter one. - FineWeb, 15 trillion tokens of filtered Common Crawl, and FineWeb-Edu, the educational subset he trains on. He uses the sample-10BT subsample and recommends the FineWeb write up on how the data was processed.
- Llama 3 70B, the model that applied FineWeb-Edu's educational filter.
- OpenWebText, the open reproduction of GPT-2's never released WebText, which was 45 million Reddit outbound links at 3 karma or more, 40 GB of text.
- Common Crawl, and the modern mixtures he names as stand ins: RedPajama, SlimPajama, C4, plus GitHub, books, arXiv, Wikipedia and Stack Exchange.
- tiny Shakespeare, the 1 MB debugging dataset, which is 338,000 GPT-2 tokens.
The tools
- PyTorch, and specifically: torch.compile, torch.set_float32_matmul_precision, torch.autocast and AMP with the mixed precision recipe he recommends over the five other copies, F.scaled_dot_product_attention, F.cross_entropy, torch.nn.utils.clip_grad_norm_, torch.optim.AdamW, nn.GELU, torch.cuda.synchronize, DistributedDataParallel and its no_sync context manager, torchrun, and the MPS backend for Apple silicon.
- tiktoken, the GPT-2 tokenizer, and Tiktokenizer, the web tool he cross checks token ids against.
- Lambda Labs, where he rents the eight A100 box, and which he discloses sponsors his development.
- NVIDIA A100 datasheet for the TFLOPS table, and the Ampere architecture whitepaper figure 9 for the TF32 bit layout. NCCL is the collective communication library DDP is talking to.
- EleutherAI's lm-evaluation-harness, the infrastructure he recommends for evaluating a checkpoint properly rather than "winging the HellaSwag eval".
- Hugging Face for the checkpoint, the datasets and the
datasetslibrary.
The rest of the series
- Neural Networks: Zero to Hero, the playlist this video continues.
- Let's build GPT: from scratch, in code, spelled out, the previous Transformer video he points back to for attention.
- Let's build the GPT Tokenizer, the tokenization video that explains where 50257 comes from.
- Andrej Karpathy on YouTube, and karpathy.ai.
- Mechanistic interpretability, the field he gestures at while looking at the structure in a weight matrix and then declines to go into.
An honest footnote
Four things are worth saying once, at the end, for anyone planning to run this rather than watch it.
The benchmark comparison is weaker than the headline, and he says so three times. The validation loss comparison against GPT-2 124M is on FineWeb-Edu's validation split, which is not the distribution GPT-2 was trained on, and he calls it "not an exactly fair comparison" himself. The HellaSwag comparison is the better one because it is held out and standard, and even there he raises the possibility that HellaSwag content leaked into FineWeb, in which case "we are basically looking at the training curve instead of the validation curve". The honest reading of the 10 times token efficiency result is that it is real but multiply caused: better filtered data, a narrower distribution with no multilingual or code or math to pay for, and an unknown amount of benchmark contamination.
Two of the numbers in the video will not reproduce today, in opposite directions. The vocabulary padding gives him 4 percent on PyTorch nightly and he says explicitly that on 2.3.1 and earlier it was about 30 percent, so the win you measure depends entirely on your PyTorch version. And torch.compile breaking generation and HellaSwag was a live bug he could not fix on camera, which he expected to be resolved by the time people read the repo. If you are following along on a current PyTorch, expect the step times in the chapter titles to be a rough guide rather than targets.
The learning rate he ships is knowingly too low. He says at the end that you can go roughly three times higher on the max learning rate and the GPT-3 hyperparameters he inherited are "extremely conservative". That is in the video, after the runs, so the recipe in the table above is the faithful GPT-3 recipe rather than his recommended one. If you want faster convergence, that is the first dial to turn, and he says so.
And the data loader has a known defect that shows up in his own plot. It serializes every token and iterates them in a fixed order with no permutation, which is why the loss curve has a periodicity he could not fully explain in four hours, and why the periodicity becomes obviously per epoch in the overnight run. The fix he names, and did not implement on camera, is to permute the documents within every shard on every new epoch and possibly permute the shards too. If you are planning multi epoch runs, do that before you start, not after you see the curve.


