At a glance
Andrej Karpathy had wanted to make this video for a while. The pitch is simple and the execution is not: three and a half hours, no math prerequisites, and nothing waved away, walking the entire pipeline that turns a crawl of the internet into the thing answering you inside the ChatGPT text box. His framing question is the one everybody actually has and nobody can answer. You can type anything in there and press enter. What should you be putting there, what are these words coming back, how does this work, and what exactly are you talking to?
He builds the answer in three sequential stages, and he keeps the receipts at every step. Pretraining is downloading and processing the internet, and he does it through FineWeb, the one dataset whose construction is documented in public: 44 terabytes on disk, 15 trillion tokens, filtered down from Common Crawl through URL blocklists, text extraction, a 65 percent English language classifier, deduplication and PII removal. Then tokenization, where text becomes a one dimensional sequence of 100,277 symbols and the model's blind spots are quietly created. Then the training loop itself, watched live on a rented 8x H100 node at $3 per GPU per hour, one million tokens per update, seven seconds per update, 32,000 steps. What comes out is a base model, which he is emphatic is not an assistant. It is an internet document simulator, a lossy zip file of the web, and he spends sixteen minutes playing with Llama 3.1 405B base to prove it: it reciting the zebra Wikipedia article from memory, it inventing two different parallel universes for the 2024 election, it becoming a passable assistant from nothing but a cleverly written prompt.
Supervised finetuning is where the personality arrives, and his reframing of it is the single most quoted idea in the video. You throw out the internet documents, substitute a dataset of conversations written by hired human labelers following a labeling instruction document that runs to hundreds of pages, and continue training with the exact same algorithm for about three hours. What you get is an imitation. "You're not talking to a magical AI," he says. "You're talking to an average labeler." Then reinforcement learning, the third stage, the one that is still early and not yet standard in the field, where the model stops imitating our solutions and goes looking for its own.
In between, the part people quote him for: LLM psychology. Why hallucinations are the default behavior of a system trained to imitate confident answers, and the two mitigations that actually work. Why the parameters are a vague recollection and the context window is working memory. Why models need tokens to think, and why a correct answer delivered too fast is a worse answer. Why they cannot count dots or spell ubiquitous backwards. Why 9.11 looks bigger than 9.9 to a network whose Bible verse neurons just lit up. The Swiss cheese model of capability, where the holes are real and arbitrary and you do not want to trip over them. And finally the practical close: where to keep track, where to find the models, how to run them yourself, and the one piece of advice he repeats at both the beginning and the end. Use them as tools in a toolbox, check their work, own the product of your work.
The question behind the text box (0:00:00)
He opens on the ChatGPT interface, cursor blinking, and names what he is after: mental models for thinking through what this tool is. It is obviously magical and amazing in some respects, really good at some things and not very good at others, and there are a lot of sharp edges to be aware of. The plan is the entire pipeline of how the stuff is built, kept accessible to a general audience, with the cognitive and psychological implications of the tools threaded through as they come up.
So let's build ChatGPT.
Stage one: pretraining
Downloading and processing the internet (0:01:00)
The first step of the first stage is to download and process the internet. To get a sense of what that looks like he points at FineWeb, a dataset collected, created and curated by Hugging Face, who went into a lot of detail in a blog post on how they constructed it. Every major provider, OpenAI, Anthropic, Google and the rest, will have some internal equivalent. FineWeb is the one you can actually read about.
The goal: a huge quantity of very high quality documents, and very large diversity of documents, because diversity is where the knowledge comes from. Large diversity of high quality documents, and many many of them. Achieving that is quite complicated and takes multiple stages to do well.
The number he wants you to sit with first is how small the result is. FineWeb, which he calls fairly representative of what you would see in a production grade application, ends up being about 44 terabytes of disk space. You can get a USB stick for a terabyte very easily. This could almost fit on a single hard drive today. The internet is very very large, but we are working with text, and we are filtering it aggressively.
The starting point, and the thing that contributes most of the data by the end, is Common Crawl. Common Crawl is an organization that has basically been scouring the internet since 2007, and as of 2024 it had indexed 2.7 billion web pages. The mechanism is unglamorous: start with a few seed web pages, follow all the links, keep following links, keep indexing all the information, and over time you end up with a ton of data of the internet.
That raw crawl is then filtered in many many different ways, and he walks the FineWeb diagram stage by stage:
- URL filtering. Blocklists of URLs and domains you do not want to be getting data from. Malware sites, spam sites, marketing sites, racist sites, adult sites. A ton of different types of websites just eliminated at this stage because we do not want them in our dataset.
- Text extraction. Remember that what the crawlers saved is raw HTML. He opens inspect element to show it: markup, lists, CSS, all this kind of stuff, computer code almost. What we really want is just the text of the web page, and not the navigation. So there is a lot of filtering, processing and heuristics that go into adequately filtering for just the good content.
- Language filtering. FineWeb runs a language classifier, guesses what language every single web page is in, and keeps only pages with more than 65 percent English. He flags this as a design decision rather than a technical step. What fraction of all the different types of languages are we going to include? If we filter out all the Spanish, you might imagine the model later will not be very good at Spanish, because it just never saw that much data of that language. Different companies focus on multilingual performance to a different degree. FineWeb is quite focused on English, so a language model trained on it will be very good at English and maybe not very good at other languages.
- Deduplication, and a few other filtering steps.
- PII removal. Personally identifiable information. Addresses, social security numbers and things like that. You try to detect them and filter out those kinds of web pages from the dataset as well.
Then he clicks into the finished dataset, which anyone can download from the Hugging Face page, and reads two of the actual surviving documents out loud, because the substance of the model's world is worth looking at directly. One is an article about tornadoes in 2012. The other opens with "did you know you have two little yellow 9 volt battery sized adrenal glands in your body," which he calls some kind of odd medical article. Just think of these as web pages on the internet, filtered for the text in various ways.
To build intuition about scale, he takes the first 200 web pages, concatenates all the text together, and zooms out. You get raw internet text, a ton of it even in 200 pages, a massive tapestry of text data. That text data has all these patterns, and what we want to do now is train neural networks on it so they can internalize and model how this text flows. We have this giant texture of text and we want neural nets that mimic it.
Tokenization: bits, bytes, and byte pair encoding (0:07:47)
Before text goes into a neural network we have to decide how to represent it. The technology expects a one dimensional sequence of symbols from a finite set of possible symbols. So we have to decide what the symbols are, and then represent our data as a one dimensional sequence of those symbols.
What we have is already one dimensional. The text starts here and goes to there; on a monitor it is laid out in two dimensions but it reads left to right, top to bottom. And being computers, there is an underlying representation. UTF-8 encode the text and you get the raw bits, which he displays as a long bar of ones and zeros where the very first segment is the first eight bits.
So in a sense we already have what we were looking for: exactly two possible symbols, zero and one, and a very long sequence of them. The problem is the tradeoff. Sequence length is a very finite and precious resource in a neural network. We do not want extremely long sequences of just two symbols. We want to trade off the symbol size of the vocabulary against the resulting sequence length, and land somewhere with more symbols and shorter sequences.
The naive first move: take groups of eight consecutive bits and call each group a byte. Since each bit is on or off, there turn out to be only 256 possible combinations, so the sequence becomes eight times shorter at the cost of 256 possible symbols. Every number now runs 0 to 255, and he insists you not think of these as numbers. Think of them as unique IDs, unique symbols. He replaces each one with a distinct emoji on screen to make the point stick: we now have a sequence of emojis, and there are 256 possible emojis.
In production for state of the art language models you go further, because you want to keep shrinking the sequence in return for more symbols in the vocabulary, and the way that is done is the byte pair encoding algorithm. Look for consecutive bytes or symbols that are very common. It turns out, in his example, that the sequence 116 followed by 32 is quite common and occurs very frequently. So group that pair into a new symbol, mint a symbol with ID 256, and rewrite every single 116 32 pair with the new symbol. Then iterate as many times as you wish. Each time you mint a new symbol you decrease the length and increase the symbol size.
In practice a pretty good setting of the vocabulary size turns out to be about 100,000 possible symbols, and in particular GPT-4 uses 100,277 symbols. The process of converting raw text into these symbols, which we call tokens, is tokenization.
He then goes to TikTokenizer, his preferred website for exploring token representations, picks cl100k_base from the dropdown (the GPT-4 base model tokenizer), and types into the left pane.
hello worldis exactly two tokens. The tokenhellohas ID 15339. The tokenworld, with the leading space, is 1917.- Join the two words together and you get two tokens again, but split differently.
- Put two spaces between them and the tokenization changes again, with a new token 220 appearing.
- It is case sensitive. A capital H is something else entirely.
hello worldwritten a third way comes out as three tokens instead of two.
He encourages you to play with it and build an intuitive sense of how tokens work, and he promises to loop back around to tokenization later in the video. He does, twice, and both times it is to explain a failure.
One last demonstration: he takes a single line of the FineWeb text and shows what GPT-4 will see it as. A sequence of length 62. The chunks of text correspond to these symbols, there are 100,277 possible symbols, and we now have one dimensional sequences of them.
Re-representing the whole dataset this way is what gives the headline number. FineWeb is not just 44 terabytes of disk space, it is about a 15 trillion token sequence. He shows the first few thousand tokens on screen and asks you to keep in mind that there are 15 trillion of them, and that all of these represent little text chunks, atoms of these sequences, and the numbers do not make any sense because they are just unique IDs.
What goes into the network, and what comes out (0:14:27)
Now the fun part, and the part where the heavy lifting happens computationally. What we want to model is the statistical relationships of how these tokens follow each other in the sequence.
So we come into the data and take windows of tokens, fairly randomly. The window length can range anywhere from zero tokens all the way up to some maximum size we decide on. In practice you could see windows of say 8,000 tokens. In principle arbitrary window lengths are possible, but processing very long windows gets very computationally expensive, so we decide that 8,000 is a good number, or 4,000, or 16,000, and we crop it there.
For his worked example he takes the first four tokens so everything fits nicely on screen. The four tokens spell out something like "bar view in and single" in the FineWeb text, and the real token that comes next in the sequence is 3962, which is the token post.
Those four tokens are the context. They feed into the neural network, and that is the input. The output is a prediction for what comes next. Because our vocabulary has 100,277 possible tokens, the network outputs exactly that many numbers, and all of those numbers are the probability of that token coming next in the sequence. It is making guesses about what comes next.
In the beginning the network is randomly initialized, so it is a random transformation and the probabilities at the start of training are also random. On his screen three of the hundred thousand are visible: the token Direction is at 4 percent, token 11799 is at 2 percent, and token 3962, post, which is the one that actually came next, is at 3 percent.
But we sampled this window out of our dataset, so we know the answer. 3962 is the label. And now we have a mathematical process for doing an update: we know we want the probability of 3 percent to be higher and the probabilities of all the other tokens to be lower, and we have a way of mathematically calculating how to adjust and update the network so that the correct answer gets a slightly higher probability.
After one update, feed the same four tokens in again and post is maybe 4 percent, case is maybe 1 percent, Direction has come down to 2 percent. We have a way of nudging the net to give a higher probability to the correct next token.
And then the crucial scaling point. This does not happen just for this one window. It happens at the same time for all of the tokens in the entire dataset. In practice we sample little batches of windows, and at every single one of those tokens we want to adjust the network so the probability of the correct token becomes slightly higher, all in parallel in large batches. That is the process of training: a sequence of updates so that the network's predictions match up with the statistics of what actually happens in your training set, and its probabilities become consistent with the statistical patterns of how these tokens follow each other.
Inside the network (0:20:11)
A brief look at the internals, just to give you a sense of what is in there.
The inputs are sequences of tokens, four in his example, anywhere from zero up to say 8,000 in practice. In principle this could be an infinite number of tokens; it would just be too computationally expensive, so we crop at a certain length and that becomes the maximum context length of the model.
Those inputs X get mixed up in a giant mathematical expression together with the parameters, or weights, of the network. He shows six example parameters with their settings. In practice modern networks have billions of them, and in the beginning they are completely randomly set. With a random setting of parameters you might expect random predictions, and that is exactly what you get at the start. It is through the process of iteratively updating the network, which is what we call training, that the parameters get adjusted so the outputs become consistent with the patterns seen in the training set.
His analogy: think of the parameters as knobs on a DJ set. As you twiddle the knobs you get different predictions for every possible token sequence input, and training a neural network just means discovering a setting of knobs that seems to be consistent with the statistics of the training set.
Then he shows what the giant mathematical expression actually looks like, to demystify it. Modern networks are massive expressions with probably trillions of terms, but the simple example on screen is not very scary: inputs x1 and x2 get mixed with the weights w0, w1, w2, w3 and so on, and the mixing is simple things like multiplication, addition, exponentiation and division. Designing effective expressions with convenient characteristics, expressive, optimizable, parallelizable, is the subject of neural network architecture research. But at the end of the day these are not complex expressions. They mix up the inputs with the parameters to make predictions, and we optimize the parameters so the predictions come out consistent with the training set.
For a production grade example he sends you to Brendan Bycroft's LLM visualization, where the network used in production settings has a special kind of structure. This network is called the Transformer, and this particular one has roughly 85,000 parameters.
Information flows top to bottom. The token sequences come in at the top, flow through, and come out at the bottom as the logits and softmax, which are the predictions for what token comes next. In between is a sequence of transformations and all the intermediate values the expression produces.
- First the tokens are embedded into a distributed representation. Every possible token has a vector that represents it inside the network.
- Those values flow through layer norms, matrix multiplications and softmaxes, all very simple mathematical expressions individually.
- There is the attention block of the Transformer, and then information flows into the multi layer perceptron block, and so on.
The intermediate numbers you can almost think of as the firing rates of these synthetic neurons, and here he inserts a careful caution. Do not think of it too much like neurons. These are extremely simple neurons compared to the ones in your brain. Your biological neurons are very complex dynamical processes that have memory. There is no memory in this expression. It is a fixed mathematical expression from input to output with no memory. It is stateless. You can still loosely think of it as a synthetic piece of brain tissue if you like to think about it that way.
He deliberately does not dwell on the precise mathematical details of all these transformations, saying honestly that he does not think it is that important. What is really important to understand is that this is a mathematical function, parameterized by a fixed set of parameters, say 85,000 of them, that transforms inputs into outputs, and as we twiddle the parameters we get different kinds of predictions, and we need to find a good setting so the predictions match the patterns in the training set. That is the Transformer.
Inference: flipping biased coins (0:26:01)
Training is one stage. The other major stage of working with these networks is inference, where we generate new data from the model to see what patterns it has internalized.
Generating is relatively straightforward. Start with some prefix tokens, whatever you want to start with. Say we start with token 91. Feed it in, and remember that the network gives us probabilities, a probability vector over the whole vocabulary. Now flip a biased coin: sample a token from that probability distribution, so tokens the model gives high probability to are more likely to be sampled.
- Sample once, and token 860 comes next. It is a relatively likely token. It might not have been the only possibility, there could have been many others, and indeed in the real training data 860 does follow 91.
- Append it, ask for the third token, sample, get 287, which again matches the training data.
- Do it again for a fourth, and again.
- Sample a fifth time and get 13659, the token
article, where the training data had 3962,post. So "viewing a single article" instead of "viewing a single post."
That last divergence is the whole point. These systems are stochastic. We are sampling, we are flipping coins, and sometimes we luck out and reproduce a small chunk of the text in the training set, but sometimes we get a token that was not verbatim part of any document in the training data. So we get remixes of the data we saw in training. Once a different token makes it in and you sample the next one from there, you very quickly start to generate token streams that are very different from the streams in the training documents. Statistically they will have similar properties, but they are not identical to your training data. They are kind of inspired by it.
Why article specifically? You might imagine that article is a relatively likely token in the context of "bar viewing single," and that the word followed that context window somewhere in the training documents to some extent, and we just happened to sample it here.
So inference is just predicting from these distributions one at a time, continuing to feed back tokens and get the next one, always flipping these coins, and depending on how lucky or unlucky you get you might land on very different kinds of patterns.
The practical division of labor follows from this. Downloading the internet and tokenizing it is a preprocessing step you do a single time. Then you start training networks, and in practical cases you would try to train many different networks of different settings, arrangements and sizes, so you do a lot of training. Then once you have a network you are happy with, you take the model and you do inference.
Which is what you are doing when you are talking to ChatGPT. That model was trained by OpenAI probably many months ago, they have a specific set of weights that work well, and when you are talking to the model all of that is just inference. There is no more training. Those parameters are held fixed and you are giving it some tokens and it is completing token sequences, and that is what you see generated.
GPT-2: what training actually looks and feels like (0:31:09)
For a concrete example of training and inference he picks the one he is particularly fond of: OpenAI's GPT-2. GPT stands for generatively pretrained Transformer, and this is the second iteration of the series. When you talk to ChatGPT today the model underlying all of the magic of that interaction is GPT-4, the fourth iteration.
GPT-2 was published in 2019 in this paper, and the reason he likes it is that it is the first time that a recognizably modern stack came together. All the pieces of GPT-2 are recognizable today by modern standards. Everything has just gotten bigger.
The details he pulls out:
- GPT-2 was a Transformer neural network, just like the ones you would work with today.
- 1.6 billion parameters. Modern Transformers would have a lot closer to a trillion, or several hundred billion.
- Maximum context length of 1,024 tokens. When sampling windows from the dataset we never take more than 1,024, so when predicting the next token you never have more than 1,024 tokens of context to make that prediction with. Tiny by modern standards: today context lengths are a lot closer to a couple hundred thousand, or maybe even a million, so you have far more tokens of history and can make a much better prediction.
- Trained on approximately 100 billion tokens. Also fairly small. FineWeb has 15 trillion, so 100 billion is quite small.
He then mentions that he reproduced GPT-2 for fun as part of llm.c, and points at his writeup of the run on GitHub. The economics are the striking part:
- The cost of training GPT-2 in 2019 was estimated at approximately $40,000.
- His run took about one day and about $600. (The writeup itself is titled with the exact figures: one 8xH100 node, 24 hours, $672.)
- And he was not even trying too hard. He thinks you could really bring this down to about $100 today.
Two reasons the costs collapsed. Number one, the datasets have gotten a lot better, and the way we filter them, extract them and prepare them has gotten a lot more refined, so the data is of just a lot higher quality. But really the biggest difference is that our computers have gotten much faster in terms of the hardware, and the software for running these models and squeezing out all the speed the hardware can give has also gotten much better, as everyone has focused on these models and tried to run them very very quickly.
Then he does something most explainers skip: he slides a live training run onto the screen so you can see what this actually looks and feels like as a researcher.
- Every single line is one update to the model, where we change its parameters by a little bit so it is better at predicting the next token in a sequence.
- Every single line is improving the prediction on 1 million tokens. We took a million tokens out of the dataset and tried to improve the prediction of each of them simultaneously, and made one update for all of that.
- The number to watch closely is the loss. A single number telling you how well your network is performing right now, constructed so that low loss is good. You watch it decrease as you make more updates, which corresponds to making better predictions on the next token. As a neural network researcher this is the number you are watching while you twiddle your thumbs and drink coffee and make sure it looks good.
- 1 million tokens per update, about 7 seconds per update, 32,000 total steps of optimization. 32,000 steps at a million tokens each is about 33 billion tokens that will be processed.
- At the moment he is filming, the run is at step 420 of 32,000, so a bit more than 1 percent done, because he has only been running it for 10 or 15 minutes.
And every 20 steps he has configured the optimization to do inference, so you can watch the model's output get better in real time. At 1 percent through training it is not yet very coherent, but it already has a little bit of local coherence. He reads the sample out loud, gibberish with the shape of English:
since she is mine it's a part of the information should discuss my father great companions Gordon showed me sitting over at
Then he scrolls up to step one, 20 updates in, where the output is completely random, because the model has had 20 updates to its parameters and is still essentially a random network. Set those side by side and the 1 percent sample looks like real progress. If you waited the entire 32,000 steps the model would be generating fairly coherent English with the token stream running correctly. That run needs about a day or two more. At this stage you just make sure the loss is decreasing, everything is looking good, and you wait.
The compute, the gold rush, and what the data centers are actually doing
He is not running this on his laptop. That would be way too expensive, and the network is just too large. All of it runs on a computer out there in the cloud, and he shows you which one.
- An 8x H100 node. Eight H100s in a single node, a single computer, which he is renting.
- He rents from Lambda, though many other companies provide the service. On demand pricing for an 8x Nvidia H100 GPU machine comes to $3 per GPU per hour.
- He shows what one H100 physically looks like. You slot it into your computer, and GPUs are a perfect fit for training neural networks because the computation is very expensive but displays a lot of parallelism, so you can have many independent workers all working at the same time on the matrix multiplication that is under the hood.
- And they stack. One GPU, eight GPUs in a node, multiple nodes into an entire data center or an entire system. The bigger data centers are of course much much more expensive.
This is where he names the economics out loud. All the big tech companies really desire these GPUs so they can train these language models, because the models are so powerful, and that is fundamentally what has driven the stock price of Nvidia to $3.4 trillion. This is the gold rush: getting the GPUs, getting enough of them so they can all collaborate to perform this optimization.
And what are they all collaborating on? Predicting the next token on a dataset like FineWeb. That is the computational workload that is extremely expensive. The more GPUs you have, the more tokens you can try to predict and improve on, the faster you process the dataset, the faster you iterate, the bigger a network you can train. He points at a news article from about a month before filming about Elon Musk getting 100,000 GPUs in a single data center, all of them extremely expensive, all of them taking a ton of power, and all of them just trying to predict the next token in the sequence and improve the network by doing so, probably getting a lot more coherent text than his run a lot faster.
Llama 3.1 base: playing with an internet document simulator (0:42:52)
He does not have a couple of ten or hundred million dollars to spend on a really big model, but luckily we can turn to the big tech companies who train these models routinely and release some of them when they are done. They spend a huge amount of compute, and then they release the network at the end of the optimization, which is very useful.
But not many of them release base models, and the distinction matters enormously here.
What is a base model? The model that comes out at the end of pretraining. It is a token simulator. An internet text token simulator. And it is not by itself useful yet, because what we want is an assistant. We want to ask questions and have it respond with answers, and these models will not do that. They just create remixes of the internet. They dream internet pages. So base models are not very often released, because they are only step one of a few other steps still needed to get to an assistant.
A few releases have been made, though. GPT-2 released its 1.6 billion (he corrects himself on camera: 1.5 billion) parameter model back in 2019, and that is a base model.
What does it actually mean to release a model? Two things:
- The code. Usually Python, describing in detail the sequence of operations the model makes. The sequence of steps taken in the Transformer, which is the forward pass of the neural network. This is just computer code, usually just a couple hundred lines of it, not that crazy, fairly understandable and usually fairly standard. He shows the GPT-2 repository on GitHub for this.
- The parameters. What is not standard, and where the actual value is. For GPT-2 that is roughly 1.5 billion numbers. One single list of 1.5 billion numbers, the precise and good setting of all the knobs such that the tokens come out well.
For something more modern he turns to Llama 3, released and trained by Meta, documented in a detailed paper. Where GPT-2 was 1.6 billion parameters on 100 billion tokens, Llama 3 is 405 billion parameters trained on 15 trillion tokens, in very much the same way, just much much bigger. The biggest base model they released is Llama 3.1 405B base. He notes in passing, as foreshadowing for the next section of the video, that they also released the instruct model, where instruct means it is an assistant you can ask questions and get answers from.
His favorite place to interact with base models is Hyperbolic, which serves Llama 3.1 405B base. You register, you make sure you have selected the base model and not the instruct one, and he sets max tokens down to 128 so as not to waste compute. Fundamentally what happens here is identical to the inference loop he just walked through: it continues the token sequence of whatever prefix you give it.
Demo one: this is not an assistant. He types what is 2 plus 2. It does not say "oh, it's four, what else can I help you with," because the question gets tokenized and those tokens just act as a prefix, and the model gets the probability for the next token. It is a glorified autocomplete. A very very expensive autocomplete, continuing according to the statistics of what it saw in its training documents, which are basically web pages. He hits enter and it actually does answer the question, then wanders off into philosophical territory. He copies and pastes the same prompt again and it goes off somewhere else entirely. Two things proven in one demo: it is not an assistant yet, and it is a stochastic system. The same prefix gives a different answer every time, because we get a probability distribution and sample from it and always get different samples.
Demo two: the knowledge is in there. Even though the model is not useful by itself for a lot of applications yet, it is still very useful, because in the task of predicting the next token it has learned a lot about the world and stored all that knowledge in the parameters. The internet web pages are compressed into the weights. Think of 405 billion parameters as a kind of compression of the internet, a zip file, except it is not lossless compression, it is lossy compression. We are left with a gestalt of the internet, and we can generate from it.
To elicit that knowledge he primes it: Here's my top 10 list of the top landmarks to see in Paris:. The model starts a list and gives him landmarks, with a lot of information. And here he plants the seed of the hallucination section:
You might not be able to fully trust some of the information. This is all just a recollection of internet documents, so the things that occur very frequently in the internet data are probably more likely to be remembered correctly compared to things that happen very infrequently. He says plainly that he does not have the expertise to verify the list is roughly correct. The knowledge is not precise and exact. It is vague, probabilistic and statistical. The information is not stored explicitly in any of the parameters. It is all just recollection.
Demo three: regurgitation. He goes to the Wikipedia page for zebra, copies one sentence, pastes it in and hits enter. The model produces an exact regurgitation of the Wikipedia entry, reciting it purely from memory, memory that lives in its parameters. "There are three living species, etc etc." He checks a few sentences down the completion to see whether it has strayed yet: still on track. Still on track. It will eventually deviate across those 512 tokens because it will not be able to remember exactly, but it has huge chunks of it memorized.
Why does this happen? These models can be extremely good at memorization, and usually this is not what you want in the final model. This is regurgitation, and it is usually undesirable to cite things directly that you have trained on. The mechanism is in the data mix: for documents deemed to be of very high quality as a source, Wikipedia being the canonical example, you preferentially sample from those sources when you train, so the model has probably done a few epochs on this data, meaning it has seen the page maybe 10 times or so. It is like you reading a text a hundred times and then being able to recite it, except these models can be a lot more efficient per presentation than humans. Ten viewings and it has remembered the article exactly in its parameters.
Demo four: the knowledge cutoff, and hallucination in its purest form. He goes to the Llama 3 paper, navigates to the pretraining data section, and reads out that the dataset has a knowledge cutoff at the end of 2023. So the model has not seen documents after that point, and certainly has not seen anything about the 2024 election and how it turned out.
So he primes it with tokens from the future: The Republican Party ticket ... Trump ... president of the United States from 2017 .... The model has to guess at the running mate and who the ticket ran against.
- Sample one: the running mate is Mike Pence, and the ticket runs against Hillary Clinton and Tim Kaine.
- Sample two, identical prompt, resampled: the running mate is Ron DeSantis, and they run against Joe Biden and Kamala Harris.
Two different parallel universes from the same prompt. "All of what we're seeing here is what's called hallucination. The model is just taking its best guess in a probabilistic manner."
Demo five: in context learning. Even though this is a base model and not an assistant, it can still be used in practical applications if you are clever with your prompt design. He builds a few shot prompt: ten lines, each one an English word, a colon, and the Korean translation. Then an eleventh line, teacher:, and a completion of just five tokens.
These models have in context learning abilities. As it reads this context it is learning in place that there is some kind of algorithmic pattern going on in the data, and it knows to continue that pattern. It takes on the role of a translator, and the completion gives the correct Korean word for teacher. Nothing was trained for this. It emerged because continuing patterns is what next token prediction rewards. So you can build apps by being clever with your prompting even when all you have is a base model.
Demo six: an assistant made out of nothing but a prompt. There is a clever way to instantiate a whole language model assistant by prompting alone, and the trick is to structure the prompt to look like a web page that is a conversation between a helpful AI assistant and a human, and then let the model continue the conversation.
To write the prompt he turned to ChatGPT itself, which he calls kind of meta: he told it he wanted to create an LLM assistant but all he had was the base model, so could it please write his prompt. What it came up with is actually quite good. It opens "Here's a conversation between an AI assistant and a human. The AI assistant is knowledgeable, helpful, capable of answering a wide variety of questions," and then, crucially, it does not stop at a description. It works much better if you make it a few shot prompt, so there are several turns of human and assistant conversation laid out before the real query.
He pastes the whole thing into the base model, adds Human: and why is the sky blue, then Assistant: and runs it. Out comes "The sky appears blue due to the phenomenon called Rayleigh scattering, etc etc." The base model is just continuing the sequence, but because the sequence looks like a conversation it takes on the role. It is a little subtle: when it finishes the assistant turn it just hallucinates the next question from the human and keeps going on and on. But the task is accomplished. For contrast he refreshes and puts why is the sky blue in on its own, with no scaffolding, and of course it does not work. You just get more questions. Because that is what a page of questions looks like on the internet.
Pretraining to post-training (0:59:23)
He zooms out and summarizes the first stage. We wish to train LLM assistants like ChatGPT. We have now covered the pretraining stage, and what it comes down to is: take internet documents, break them up into tokens, these atoms of little text chunks, and predict token sequences using neural networks. The output of the entire stage is the base model, which is the setting of the parameters of the network, and that base model is an internet document simulator on the token level. It generates token sequences with the same kind of statistics as internet documents, and we saw it can be used in some applications.
But we need to do better. We want an assistant. We want to ask questions and have the model give us answers. So we hand the base model, the internet document simulator, off to post training.
The cost structure flips completely here, and this is worth holding onto. Pretraining in practice can take roughly three months of training on many thousands of computers. The post training stage will typically be much shorter, like three hours. All of the massive data centers, all of the heavy compute and millions of dollars, are the pretraining stage. Post training is slightly cheaper but still extremely important, and it is where the LLM becomes an assistant.
Stage two: supervised finetuning
Post training data: conversations (1:01:06)
What we want now is not to sample internet documents but to give answers to questions. So we start thinking about conversations, which can be multi turn, in the simplest case between a human and an assistant. He puts three examples on screen:
- Human says "what is 2 plus 2," the assistant should respond with something like "2 plus 2 is 4."
- Human follows up with "what if it was a star instead of a plus," and the assistant responds accordingly.
- A third example where the assistant has some personality, where it is kind of nice.
- And a third category: when a human asks for something we do not wish to help with, we produce what is called a refusal. "We cannot help with that."
So we want to think through how an assistant should interact with a human, and we want to program the assistant and its behavior in these conversations. But because this is neural networks we are not going to be programming these explicitly in code. Everything is done through neural network training on datasets. So we are going to be implicitly programming the assistant by creating datasets of conversations.
An actual dataset could have hundreds of thousands of conversations that are multi turn and very long and cover a diverse breadth of topics. He shows three. The assistant is being programmed by example.
And then the question that drives the whole section: where does this data come from? It comes from human labelers. We give human labelers some conversational context and ask them to give the ideal assistant response in that situation. A human writes out the ideal response for an assistant in any situation, and then we get the model to train on this and imitate those kinds of responses.
The mechanics are almost anticlimactic. Take the base model from pretraining. Throw out the dataset of internet documents. Substitute the dataset of conversations. Continue training the model on the new data. The model very rapidly adjusts and learns the statistics of how this assistant responds to human queries, and later during inference we can prime the assistant and get a response, and it will be imitating what the human labelers would do in that situation. We are using the exact same algorithm, the exact same everything, except we are swapping out the dataset.
How a conversation becomes a token sequence
Everything in these models has to be turned into tokens, so how do we turn conversations into token sequences? We need to design an encoding, and he reaches for a good analogy: this is kind of similar to the TCP/IP packet on the internet. There are precise rules and protocols for how you represent information and how everything is structured together, laid out in a way that is written out on paper and that everyone can agree on. The same thing is now happening in LLMs. We need data structures and rules around how conversations get encoded and decoded to and from tokens.
So he goes back to TikTokenizer and pastes in the two turn conversation. It looks ugly but it is actually relatively simple, and the whole user and assistant exchange ends up being 49 tokens. A one dimensional sequence of 49 tokens, and these are the tokens.
All the different LLMs have a slightly different format or protocol, and it is a little bit of a wild west right now, but GPT-4o does it as follows:
- A special token
<|im_start|>, which is short for imaginary monologue start. (He admits on camera that he does not actually know why it is called that.) - Then you specify whose turn it is, for example
user, which is token 428. - Then the internal monologue separator,
<|im_sep|>. - Then the exact question, as the tokens of the question.
- Then you close it with
<|im_end|>, the end of the imaginary monologue.
The important thing to mention, and the detail that is easy to miss: <|im_start|> is not text. It is a special token that gets added, a new token, and it has never been trained on so far. It is a new token that we create in the post training stage. These special tokens get introduced and interspersed with text so that the model learns that this is the start of a turn, and whose turn it is, and this is what the user says, and then the user ends, and now it is a new start of a turn and it is the assistant's, and these are the tokens of what the assistant says.
The specific details are not what matters. What matters is that conversations, which we think of as a structured object, end up being turned via some encoding into one dimensional sequences of tokens, and because it is a one dimensional sequence of tokens we can apply all the stuff that we applied before. We are just predicting the next token in a sequence, just like before.
At inference, when you are on ChatGPT and you have a dialogue, what happens on the servers is that your new turn gets appended, then they put in <|im_start|>assistant<|im_sep|> and end it right there. They construct that context and start sampling from the model. That is the stage where the model is asked what is a good first token, a good second token, a good third token, and the LLM takes over and creates a response. It will not be identical to anything in the dataset, but it will have the flavor of the kind of conversation that was in the dataset.
InstructGPT, and the labeling instructions nobody reads
The first paper in this direction, and the first time OpenAI talked publicly about how you take language models and finetune them on conversations, is the 2022 InstructGPT paper. He takes you through it.
Stop one is section 3.4, where they describe the human contractors they hired, in this case from Upwork or through Scale AI, to construct these conversations. There are human labelers involved whose job it is professionally to create these conversations, and they are asked to come up with the prompts and then also to complete the ideal assistant responses.
The real prompts the labelers came up with, from the paper:
- "List five ideas for how to regain enthusiasm for my career"
- "What are the top 10 science fiction books I should read next"
- "Translate this sentence to Spanish"
Stop two is the excerpt of labeling instructions, and this is the part of the video that reframes everything downstream. The company developing the language model writes up labeling instructions for how the humans should create ideal responses. The excerpt on screen, at a high level, is asking people to be helpful, truthful and harmless, and not to answer questions the company does not want the system to handle later in ChatGPT.
He pauses on the scale of this document, because the excerpt is misleadingly short. "Usually they are not this short. Usually there are hundreds of pages, and people have to study them professionally." Then they write out the ideal assistant responses following those instructions. This is a very human heavy process as described in the paper.
The InstructGPT dataset was never actually released by OpenAI. But there are open source reproductions trying to follow the same setup and collect their own data, and the one he is familiar with is OpenAssistant, one of what he thinks are many examples. These were people on the internet asked to create conversations similar to what OpenAI did with human labelers. He reads an actual entry: someone wrote the prompt "can you write a short introduction to the relevance of the term monopsony in economics, please use examples," and then the same person or potentially a different person wrote out the ideal assistant response. Then the conversation continues with "now explain it to a dog," and a simpler explanation follows.
That becomes the label, and we train on it. And here is the thing that makes the whole approach work, because of course we cannot possibly cover all the prompts people will ask in the future. If you have a dataset of a few of these examples, the model during training starts to take on this persona of the helpful truthful harmless assistant, and it is all programmed by example. These are all examples of behavior, and if you have enough of them, like 100,000, and you train on it, the model starts to understand the statistical pattern and takes on this personality.
At test time, if you ask the exact same question that was in the training set, the answer may be recited as exactly what was in there. But more likely the model will do something of a similar vibe, having understood that this is the kind of answer that you want.
How the data is actually made now: UltraChat and SFT mixtures
He is careful to say the state of the art has advanced in the two or three years since InstructGPT. It is not very common for humans to be doing all the heavy lifting just by themselves anymore, because we now have language models, and those language models are helping us create these datasets. It is very rare that people will literally just write out the response from scratch. It is a lot more likely they will use an existing LLM to come up with an answer and then edit it. LLMs have started to permeate this post training stack and are used pervasively to help create these massive datasets of conversations.
UltraChat is his example of a more modern conversation dataset. It is to a very large extent synthetic, though he believes there is some human involvement and is careful to say he could be wrong about the exact mix. Usually there will be a little bit of human and a huge amount of synthetic help. The numbers have moved: these datasets now have millions of conversations, mostly synthetic, probably edited to some extent by humans, spanning a huge diversity of areas. These are fairly extensive artifacts by now, and there are SFT mixtures, mixtures of lots of different types and sources, partially synthetic and partially human.
But roughly speaking, nothing structural has changed. We still have SFT datasets, they are made up of conversations, and we are training on them just like we did before.
"You're talking to an average labeler" (1:17:02)
This is the reframing the video is most quoted for, and he builds it carefully.
He wants to dispel a little bit of the magic of talking to an AI. When you go to ChatGPT and give it a question and hit enter, what comes back is statistically aligned with what is happening in the training set, and those training sets really just have a seed in humans following labeling instructions.
So what are you actually talking to?
"It's not coming from some magical AI. Roughly speaking, it's coming from something that is statistically imitating human labelers, which comes from labeling instructions written by these companies."
The way to hold it in your head: imagine the answer given to you from ChatGPT is a simulation of a human labeler. It is like asking what would a human labeler say in this kind of a conversation. And this human labeler is not just a random person from the internet, because these companies actually hire experts. When you are asking questions about code, the labelers involved in creating those conversation datasets will usually be educated expert people, and you are asking a question of a simulation of those people.
"You're not talking to a magical AI, you're talking to an average labeler. This average labeler is probably fairly highly skilled, but you're talking to kind of like an instantaneous simulation of that kind of a person that would be hired in the construction of these datasets."
Then one more specific example to land it. He asks ChatGPT to recommend the top five landmarks to see in Paris and hits enter. What is coming out? It is not some kind of magical AI that has gone out and researched all the landmarks and then ranked them using its infinite intelligence. It is a statistical simulation of a labeler hired by OpenAI.
Two cases follow, and the second is the interesting one:
- If this specific question is in the post training dataset somewhere at OpenAI, then you are very likely to see an answer very very similar to what that human labeler put down for those five landmarks. And how did that labeler come up with it? They went on the internet and did their own little research for 20 minutes and came up with a list.
- If the query is not in the post training dataset, then what you are getting is a little bit more emergent. The model statistically understands that the landmarks in this training set are usually the prominent ones, the ones people usually want to see, the ones very often talked about on the internet. And remember the model already has a ton of knowledge from pretraining, so it has probably seen a ton of conversations about Paris, about landmarks, about the kinds of things people like to see. It is the pretraining knowledge combined with the post training dataset that results in this kind of imitation.
| Stage | Data it trains on | Who produces the label | What it costs | What it buys |
|---|---|---|---|---|
| Pretraining | 44 TB of filtered web text, 15 trillion tokens (FineWeb scale) | Nobody. The label is simply the token that actually came next | Roughly 3 months on many thousands of computers, millions of dollars | Knowledge, and a base model that simulates internet documents |
| Supervised finetuning | Conversations. Hundreds of thousands to millions of them, now largely synthetic and human edited | Hired human labelers writing the ideal response, following an instruction document that runs to hundreds of pages | Roughly 3 hours. Same algorithm, same everything, different dataset | An assistant, which is an imitation of a labeler |
| RL, verifiable domains | Practice problems where the final answer is known but the solution path is not | Nobody writes the path. The model generates thousands of attempts and keeps what reached the answer | Can run for tens or hundreds of thousands of steps | Thinking. Emergent chains of thought, and strategies no human taught it |
| RLHF, unverifiable domains | Creative prompts. Jokes, poems, summaries, with human orderings of five rollouts each | A separate reward model, trained to imitate human orderings, standing in for a human | A few hundred to a thousand updates, then you crop it and ship | A small improvement. Gameable, so not RL in the magical sense |
LLM psychology
He names this section himself. What are the emergent cognitive effects of the training pipeline that we have for these models? Everything in this part of the video follows from the mechanics he has just built up, which is what makes it convincing rather than anecdotal.
Where hallucinations come from (1:20:32)
You are probably familiar with model hallucinations. It is when LLMs make stuff up, totally fabricate information. It is a big problem with LLM assistants, it existed to a large extent with early models from many years ago, and the problem has gotten a bit better because of mitigations he is about to go into. But first, where do they come from?
He puts three perfectly reasonable training conversations on screen:
- "Who is Tom Cruise?" Well, Tom Cruise is a famous American actor and producer, etc.
- "Who is John Barrasso?" This turns out to be a US senator.
- "Who is Genghis Khan?" Well, Genghis Khan was, and so on.
Now notice what the human writing those answers did. The human either knows who this person is, or they researched them on the internet, and then they wrote a response with the confident tone of an answer. Three for three.
So at test time he asks about Orson Kovats, a name he made up at random and does not think belongs to a real person.
Here is the trap. The assistant will not just tell you it does not know, even if the language model itself might know inside its features, inside its activations, inside its brain sort of. Some part of the network may well represent that this is not someone it is familiar with. But saying "I don't know who this is" is not going to happen, because the model statistically imitates its training set, and in the training set questions of the form "who is X" are confidently answered with the correct answer. So it takes on the style of the answer, does its best, gives you statistically the most likely guess, and basically makes stuff up. These models do not have access to the internet, they are not doing research, they are statistical token tumblers trying to sample the next token in the sequence.
To demonstrate, he opens the Hugging Face inference playground and on purpose picks on Falcon 7B Instruct, an older model from a few years back that suffers from hallucinations more visibly. "Who is Orson Kovats?"
- Run one: "Orson Kovats is an American author and science fiction writer." Totally false.
- Resample: "Orson Kovats is a fictional character from a 1950s TV show." Total nonsense.
- Resample again: he is a former minor league baseball player.
The model does not know, so it gives lots of different answers. It starts with the tokens "who is Orson Kovats, assistant," gets probabilities, and samples from them, and the stuff it comes up with is statistically consistent with the style of the answer in its training set. You and I experience that as made up factual knowledge. The model is just imitating the format of the answer, and it is not going to go off and look it up.
For contrast he asks the same question of the state of the art ChatGPT, which is smarter in a specific way: he briefly sees "searching the web" flash on screen, which is tool use and which the video covers next. So he asks again with "don't use any tools", and gets back that there is no well known historical or public figure named Orson Kovats. This model knows that it does not know, and it tells you. So hallucinations have somehow improved even though they are clearly an issue in older models.
Mitigation one: teach the model what its own ignorance feels like
Clearly we need examples in the dataset where the correct answer for the assistant is that the model does not know about some particular fact. But we only need those answers to be produced in the cases where the model actually does not know. So the question becomes: how do we know what the model knows or does not know?
We can empirically probe the model to figure that out. And he walks through exactly how Meta did it for Llama 3, in the section of their paper they call factuality. The procedure interrogates the model to find the boundary of its knowledge, and then adds examples to the training set where, for the things the model does not know, the correct answer is that it does not know. He flags up front that this sounds like a very easy thing to do in principle, and roughly fixes the issue.
Why it works is the interesting part, and it follows from the network internals he showed earlier. The model might actually have a pretty good model of its self knowledge inside the network. You might imagine there is a neuron somewhere that lights up when the model is uncertain. The problem is that the activation of that neuron is not currently wired up to the model actually saying in words that it does not know. So even though the internals know, because some neurons represent it, the model will not surface that. It will instead take its best guess so that it sounds confident, just like it sees in its training set. So we need to interrogate the model and allow it to say "I don't know" in the cases where it does not know.
The procedure, step by step, with his live demo:
- Take a random document from the training set and take a paragraph. He picks Dominik Hašek, the Wikipedia featured article of the day, at random.
- Use an LLM to construct questions about that paragraph. He does it in ChatGPT: "here's a paragraph from this document, generate three specific factual questions based on this paragraph and give me the questions and the answers." LLMs are already good enough to reframe this information, and crucially, because the information is in the context window it does not have to rely on its memory, so it works pretty well with fairly high accuracy. He gets questions like "which team did he play for" and "how many cups did he win," with answers.
- Interrogate the model with those questions. He uses Mistral 7B as his stand in for the model being probed. First question: which team? The model says the Buffalo Sabres, which is right. So the model knows.
- Compare the model's answer with the correct answer programmatically. Models are good enough to do this automatically, so there are no humans involved here. You take the answer from the model and use another LLM as a judge to check whether it is correct.
- Repeat a few times, because one sample is noise. He asks three times and gets Buffalo Sabres, Buffalo Sabres, Buffalo Sabres. The model knows, so everything is great.
- Now the second question: how many Stanley Cups did he win? The correct answer is two. The model claims four, which does not match. Resample: it makes something up again. Resample: it says he did not win during his career. Three for three wrong, so the model does not know, and it is making stuff up.
- Create a new conversation in the training set. When the question is "how many Stanley Cups did he win," the answer is "I'm sorry, I don't know" or "I don't remember." And that is the correct answer for this question, because we interrogated the model and we saw that that is the case.
Do this for many different types of questions across many different types of documents and you are giving the model an opportunity, in its training set, to refuse based on its knowledge. With just a few examples the model gets the chance to learn the association between knowledge based refusal and that internal neuron of uncertainty somewhere in its network that we presume exists. Empirically this turns out to be probably the case. It can learn that when the uncertainty neuron is high, it actually does not know and it is allowed to say so. If you have those examples in your training set, this is a large mitigation for hallucination, and that is roughly why ChatGPT is able to do what it did on Orson Kovats.
Mitigation two: tools, and the search tokens
We can do much better than saying we do not know. We can give the LLM an opportunity to be factual and actually answer the question.
What do you and I do when asked a factual question we do not know? We go off and search, use the internet, figure out the answer. We can do the exact same thing with these models, and the reason the analogy holds is the distinction that becomes the most useful mental model in the whole video:
"Knowledge in the parameters of the neural network is a vague recollection. The knowledge in the tokens that make up the context window is the working memory."
Think of the knowledge inside the billions of parameters as something you read a month ago. If you keep reading something you will remember it, and the model remembers that, but if it is something rare you probably do not have a really good recollection. What you and I do is look it up, and when you look something up you are refreshing your working memory with the information so you can retrieve it and talk about it.
So we need an equivalent mechanism for the model, and we build it out of new tokens. He introduces two and a protocol for using them:
- Instead of answering from memory, the model can emit the special token
<SEARCH_START>. - Then the query, which is what will go to Bing in the case of OpenAI, or Google search, or something like that.
- Then
<SEARCH_END>.
And then the mechanism outside the model takes over. The program that is sampling from the model, running the inference, sees the special end token and instead of sampling the next token it pauses generating. It goes off, opens a session with Bing, pastes the search query in, gets all the text that is retrieved, maybe re-represents it with some other special tokens, and copy pastes that text straight into the context window.
Now that text from the web search is inside the context window. It feeds into the neural network. It is not a vague recollection anymore. It is data the model has in its context window, directly available. So as it samples the new tokens afterwards it can reference very easily the data that was pasted in.
How do you teach the model to use the tool correctly? The same way you teach it everything else: training sets. You need a bunch of conversations that show the model by example how to use web search, what the settings are where you use it, what that looks like, how you start a search. With a few thousand examples of that in your training set the model does a pretty good job, and it knows how to structure its queries. And because of the pretraining dataset and its understanding of the world, it kind of understands what a web search is and has a pretty good native understanding of what makes a good search query. So it all just works, and you only need a little bit of a few examples to show it how to use the new tool.
This is exactly what he saw happen earlier with Orson Kovats: the ChatGPT language model decided this is some kind of rare individual, and instead of giving him an answer from its memory it sampled a special token that does a web search. The little flash of "using the web tool," then two seconds of waiting, then the response, and crucially the response creates references and cites sources. The URLs and the text of those web pages were all stuffed invisibly between the search tokens, and the model is reading that text and citing it.
He runs the control experiment as well. Asked "how many Stanley Cups did Dominik Hašek win," ChatGPT decided that it knows the answer and had the confidence to say he won twice, relying on its memory, presumably because it has enough confidence in its weights and activations that this is retrievable. Then he asks the same query with web search on, and it goes off, finds a bunch of sources, everything gets pasted in, and it cites the Wikipedia article that is the source of the information for us as well.
So: tools are the second mitigation for hallucinations and factuality, and the model determines when to search.
The practical consequence: put it in the context window
He stresses the psychology point one more time, because it changes how you should prompt. The stuff we remember is our parameters. The stuff we just experienced a few seconds or minutes ago is our context window, and that context window is being built up as you have a conscious experience around you.
The practical implication, demonstrated live:
- He asks ChatGPT "can you summarize chapter one of Jane Austen's Pride and Prejudice?" This is a perfectly fine prompt and the answer is relatively reasonable, but only because ChatGPT has a pretty good recollection of a famous work. It has probably seen a ton of stuff about it, there are forums about the book, it has probably read versions of the book, and it kind of remembers.
- A much better prompt is "can you summarize for me chapter one of Jane Austen's Pride and Prejudice, and I am attaching it below for your reference," then a delimiter, then the chapter pasted in from some website.
Why the second is better: when it is in the context window the model has direct access to it. It does not have to recall it. So the summary can be expected to be significantly higher quality. And he closes the analogy honestly: you and I would work the same way. You would produce a much better summary if you had reread the chapter before you had to summarize it.
Knowledge of self (1:41:46)
The next psychological quirk is one he sees on the internet constantly. People ask LLMs "what model are you" and "who built you," and the question is a little bit nonsensical.
Why it is nonsensical follows from the fundamentals. This thing is not a person. It does not have a persistent existence in any way. It boots up, processes tokens, and shuts off, and it does that for every single person. It builds up a context window of conversation and then everything gets deleted. This entity is restarted from scratch every single conversation. It has no persistent self, it has no sense of self. It is a token tumbler, and it follows the statistical regularities of its training set.
So by default, if you ask, you get pretty random answers. He picks on Falcon again, which first evades the question with "talented engineers and developers," then says "I was built by OpenAI based on the GPT-3 model." Totally making stuff up.
And here he makes a correction that a lot of people get wrong. Many would take that answer as evidence that this model was somehow trained on OpenAI data. He does not actually think that is necessarily true. The reason: if you do not explicitly program the model to answer these kinds of questions, what you get is its statistical best guess at the answer. This model had an SFT data mixture of conversations, and during finetuning it understood that it was taking on the personality of a helpful assistant, but it was never told exactly what label to apply to itself. Meanwhile the pretraining stage took the documents from the entire internet, and ChatGPT and OpenAI are very prominent in those documents. So what is actually likely happening is that this is its hallucinated label for what it is. Its self identity is that it is ChatGPT by OpenAI, and it is only saying that because there is a ton of data on the internet of answers like this that actually came from ChatGPT.
There are two ways to override this as a developer, and he shows both.
Way one: hardcode it in the SFT data. He pulls up OLMo from the Allen Institute for AI, which he likes specifically because it is fully open source, paper and everything, which is nice. He opens its SFT mixture, which is the finetuning conversation data. There is a total of 1 million conversations in the mixture, and one component is labeled as hardcoded. Click into it and there are 240 conversations. They are exactly what you would expect: "Tell me about yourself," and the assistant says "I'm OLMo, an open language model developed by the Allen Institute for Artificial Intelligence, I'm here to help." "What is your name." Cooked up, hardcoded questions and the correct answers to give. Take 240 conversations like that, put them in your training set, finetune on it, and the model will be expected to parrot this stuff later. Do not give it this, and it is probably ChatGPT by OpenAI.
Way two: the system message. In these conversations, as well as turns between human and assistant, there is sometimes a special message called the system message at the very beginning of the conversation. So it is not just human and assistant, there is a system too. In the system message you can hardcode and remind the model that it is a model developed by OpenAI, that its name is ChatGPT 4o, what date it was trained, what its knowledge cutoff is. It documents the model a little bit, and then that gets inserted into your conversations. When you go to ChatGPT you see a blank page, but actually the system message is hidden in there and those tokens are in the context window.
Both routes are the same kind of thing: invisible tokens in the context window that remind the model of its identity. And he is blunt about what that means. "It's all just kind of cooked up and bolted on in some way. It's not actually like really deeply there in any real sense as it would be for a human."
Models need tokens to think (1:46:56)
This is the section with the most practical payoff, and he sets it up as a puzzle rather than a lecture.
The setup: suppose we are building out a conversation to enter into our training set of conversations, teaching the model how to solve simple math problems. The prompt is:
Emily buys 3 apples and 2 oranges. Each orange costs $2. The total cost of all the fruit is $13. What is the cost of each apple?
Simple question. He puts two answers on screen, left and right. Both are correct. Both say the answer is 3. But one of them is a significantly better answer for the assistant than the other. If he were a data labeler creating one of these, one would be a really terrible answer and the other would be okay. He invites you to pause the video and work out which, and he warns: if you use the wrong one, your model will actually be really bad at math. This is the sort of thing you would be careful with in your labeling documentation when training people to create ideal responses.
The key is to remember that when models are training and also inferencing, they work in a one dimensional sequence of tokens from left to right. He puts up the picture he has in his own mind: the token sequence evolving left to right, and to produce each next token we feed all the tokens so far into the network and get probabilities out. This is the exact same picture as the web demo of the Transformer from earlier.
And now the crucial constraint. There is a finite number of layers of computation. The demo network has only one, two, three layers of attention and MLP. A typical modern state of the art network would have more like 100 layers. But that is it: only about 100 layers of computation to go from the previous token sequence to the probabilities for the next token.
"There's a finite amount of computation that happens here for every single token, and you should think of this as a very small amount of computation."
That amount is roughly fixed for every single token in the sequence. Not perfectly: the more tokens you feed in, the more expensive the forward pass, but not by much. So the good mental model is a fixed amount of compute in that box for every single token, and it cannot possibly be too big, because there are not that many layers going top to bottom.
Therefore: you cannot imagine the model doing arbitrary computation in a single forward pass to get a single token. We have to distribute our reasoning and our computation across many tokens.
So why is the left answer worse? Imagine going left to right, emitting tokens one at a time. The model has to say "The answer is", then the dollar sign, and right there we are expecting it to cram all of the computation of this problem into that single token. It has to emit the correct answer, 3. And then once 3 has been emitted, everything that follows is post hoc justification, because the answer is already created and already in the context window. It is not actually being calculated in those tokens.
"If you are answering the question directly and immediately, you are training the model to try to basically guess the answer in a single token, and that is just not going to work because of the finite amount of computation that happens per token."
The right answer is significantly better because we are distributing the computation across the answer. We get the model to slowly come to the answer from left to right, producing intermediate results: the total cost of the oranges is 4, so 13 minus 4 is 9, so 9 divided by 3 is 3. Each of those calculations is by itself not that expensive. We are guessing a little bit at the difficulty the model is capable of in any single token, and there can never be too much work in any one token, because then the model will not be able to do it later at test time. By the time it is near the end it has all the previous results in its working memory and it is much easier for it to determine that the answer is 3.
He notes in passing that in your prompts you usually do not have to think about this explicitly, because the people at OpenAI have labelers who worry about it and make sure the answers are spread out. So when he asks this question in ChatGPT it goes very slowly: okay, let's define our variables, set up the equation. "These are not for you. These are for the model. If the model is not creating these intermediate results for itself, it's not going to be able to reach three."
Then he runs the experiment to prove the constraint is real, by being a bit mean to the model.
- Attempt one. Same prompt, plus "answer the question in a single token, just immediately give me the answer, nothing else." It worked. It actually produced two tokens, because the dollar sign is its own token, but it got the correct answer in a single forward pass of the network. The numbers here are very simple.
- Attempt two: make it harder. "Emily buys 23 apples and 177 oranges," same structure, same instruction. The model answered 5, which is not correct. It failed to do all of the calculation in a single forward pass. It could not go from the input tokens to the result in one go through the network.
- Attempt three. "Now don't worry about the token limit and just solve the problem as usual." Now it produces all the intermediate results, simplifies, each intermediate calculation is much easier, all of the tokens are correct, and it arrives at the solution, which is 7. It just could not squeeze that work into a single forward pass.
And then his own practical caveat, which is the honest one. If he were actually solving this in his day to day life he would not trust that all the intermediate calculations are correct either. So what he would actually do is say "use code." Code is one of the tools ChatGPT can use, and instead of the model doing mental arithmetic, which he does not fully trust, especially if the numbers get really big, it writes a program.
"We're using neural networks to do mental arithmetic, kind of like you doing mental arithmetic in your brain. It might just screw up some of the intermediate results. It's actually kind of amazing that it can even do this kind of mental arithmetic. I don't think I could do this in my head."
The mechanism is the same as web search. There is a special tool, and the model will not actually generate the result tokens itself. It writes the program, that program gets sent to a different part of the computer that just runs it, the result comes back, and the model gets access to the result. He can inspect that the code is correct, the Python interpreter does the arithmetic, and he personally trusts that a lot more, because it came out of a Python program, which has a lot more correctness guarantees than the mental arithmetic of a language model.
His summary of the section, in his own words: models need tokens to think. Distribute your computation across many tokens. Ask models to create intermediate results, and whenever you can, lean on tools and tool use instead of letting the model do it all in its memory.
Counting, and why "use code" works
He has one more example of the same constraint, in counting. Models are not very good at counting, for the exact same reason: you are asking for way too much in a single individual token.
The demo is "how many dots are below," followed by a block of dots. ChatGPT says "there are" and then just tries to solve the problem in a single token. In a single token it has to count the dots in its context window, in a single forward pass of the network, where there is very little computation available.
And when he looks at what the model actually sees, the problem gets worse. In TikTokenizer, a group of about 20 dots is a single token. The next group is another token. Then they break up differently for reasons that have to do with the details of the tokenizer. So the model basically sees a handful of token IDs, and from those token IDs it is expected to count.
ChatGPT's answer is 161. The correct answer is 177.
Then "use code," and he flags that it is actually kind of subtle and kind of interesting why this should work at all. It is not that he gave the model a calculator. It is that he broke the problem into problems that are easier for the model. He knows the model cannot do mental counting. But he knows the model is actually pretty good at copy pasting. So when he says use code, the model creates a Python string containing the dots, and the task of copy pasting the input is very simple, because the model sees it as just those four or so token IDs, and unpacking them into dots is easy. Then it calls Python's .count() and comes back with 177, correct.
"The Python interpreter is doing the counting. It's not the model's mental arithmetic doing the counting."
Same lesson. Models need tokens to think. Do not rely on their mental arithmetic. If you need counting, always ask them to lean on the tool.
Tokenization revisited: models struggle with spelling (2:01:11)
He promised to loop back around to tokenization, and here is the payoff. Models are not very good at all kinds of spelling related tasks, and the reason is that they do not see the characters. They see tokens. Their entire world is about tokens, which are these little text chunks. They do not see characters like our eyes do, so very simple character level tasks often fail.
The demo: give it the string ubiquitous and ask it to print only every third character starting with the first one. So start with u, then count 1, 2, 3 and q should be next, and so on. The model gets it wrong.
His hypothesis has two parts. The mental arithmetic is failing a little bit. But the more important issue is that when you put ubiquitous into TikTokenizer, it is three tokens. You and I see "ubiquitous" and can easily access the individual letters, because we see them, and with the word in the working memory of our visual field we can really easily index into every third letter. The model does not have access to the individual letters. It sees three tokens. And remember that these models are trained from scratch on the internet, so the model has to discover how many of all these different letters are packed into all these different tokens.
Then an honest aside about why we put up with this at all. The reason we even use tokens is mostly for efficiency. He notes that a lot of people are interested in deleting tokens entirely, that we should really have character level or byte level models. It is just that that would create very long sequences, and people do not know how to deal with that right now. So while we have the token world, any kind of spelling task is not actually expected to work super well.
The fix is the same fix as always. He says use code, and expects it to work, because the task of copy pasting ubiquitous into the Python interpreter is much easier and then we are leaning on Python to manipulate the characters of the string. It indexes into every third character and returns the right letters (u, q, t, s), which he confirms looks correct.
And then the famous one. "How many Rs are there in strawberry?" This went viral many times, and by the time of filming the models now get it correct and say there are three. But for a very long time all the state of the art models would insist that there are only two. This caused a lot of ruckus, because why are the models so brilliant that they can solve math olympiad questions but cannot count Rs in strawberry?
His answer is the one he has been building toward for two hours. Number one, the models do not see characters, they see tokens. Number two, they are not very good at counting. So here we are combining the difficulty of seeing the characters with the difficulty of counting, and that is why the models struggled with this. He adds a candid caveat about the specific case: by now he thinks OpenAI may have hardcoded the answer, or he is not sure what they did, but that specific query now works.
Jagged intelligence (2:04:53)
There are a bunch of other little sharp edges and he does not want to go through all of them or give a comprehensive analysis of every way the models fall short. He just wants to make the point that there are some jagged edges here and there, and that while a few of them make sense, some of them will not make as much sense, and you are left scratching your head even if you understand in depth how these models work.
His example is the one everybody has seen. The models are not very good at very simple questions like "what is bigger, 9.11 or 9.9?" This is shocking to a lot of people, because these models can solve complex math problems and answer PhD grade physics, chemistry and biology questions much better than he can, but sometimes they fall short on something this simple.
He runs it. The model says 9.11 is bigger than 9.9 and justifies it in some way, then flips its decision later in the same response. He is careful about reproducibility: he does not believe this is very reproducible. Sometimes it flips around its answer, sometimes it gets it right, sometimes it gets it wrong. He tries again and this time it does not even correct itself at the end.
So how is it that the model can do great at olympiad grade problems and then fail on this? He calls this one a bit of a head scratcher, and he is scrupulous about his sourcing here: a bunch of people studied this in depth, he has not actually read the paper, but what he was told by that team is that when you scrutinize the activations inside the network and look at which features and neurons turn on and off, a bunch of neurons light up that are usually associated with Bible verses.
So his reading is that the model is kind of reminded that these almost look like Bible verse markers, and in a Bible verse setting 9.11 would come after 9.9, so the model somehow finds it cognitively very distracting. Even here, where it is actually trying to justify the answer and come to it with math, it still ends up with the wrong answer. "It basically just doesn't fully make sense, and it's not fully understood."
His conclusion, which is the practical spine of the last third of the video:
"Treat this as what it is, which is a stochastic system that is really magical but that you can't also fully trust. You want to use it as a tool, not as something that you let rip on a problem and copy paste the results."
Stage three: reinforcement learning
From supervised finetuning to reinforcement learning (2:07:28)
He takes stock before the last stage. We have covered two major stages. In pretraining we train on internet documents and get a base model, an internet document simulator, which takes many months to train on thousands of computers and is a lossy compression of the internet. Extremely interesting, but not directly useful, because we do not want to sample internet documents. We want to ask questions of an AI and have it respond. So we construct an assistant through post training, specifically supervised finetuning, which is algorithmically identical to pretraining. Nothing changes. The only thing that changes is the dataset. Instead of internet documents we curate millions of conversations on all kinds of diverse topics between a human and an assistant, and fundamentally those conversations are created by humans: humans write the prompts, humans write the ideal responses, based on labeling documentation. In the modern stack this is not done fully manually, there is a lot of help from the models themselves, but fundamentally it is all still coming from human curation at the end.
Then we shifted gears into the cognitive implications: that the assistant will hallucinate if you do not take mitigations, that the models are quite impressive and can do a lot in their head but can lean on tools to become better, web search to hallucinate less and bring in recent information, or a code interpreter so the LLM can write code and actually run it and see the results.
Now the last and major stage. Reinforcement learning is still thought of as under the umbrella of post training, but it is the third major stage, and it is a different way of training language models.
He adds an organizational detail that is easy to miss and quietly explains a lot about how these labs work. Inside companies like OpenAI these are all separate teams. There is a team doing data for pretraining and a team doing the training for pretraining. There is a team doing all the conversation generation, and a different team doing the supervised finetuning. And there is a team for reinforcement learning. It is a handoff of these models: you get your base model, then you finetune it to be an assistant, then you go into reinforcement learning.
The motivation, and the analogy that carries the whole stage: going to school.
Just like you went to school to become really good at something, we want to take large language models through school. And when you work with textbooks you will see three major classes of information in them. He pulls a totally random book off the internet, some kind of organic chemistry text, to make the point concretely.
- Exposition. Most of the text, the meat of it, background knowledge and context. As you read through the words of the exposition, that is roughly equivalent to training on that data. It is where we build a knowledge base and get a sense of the topic. This is pretraining.
- Problems with their worked solutions. A human expert, in this case the author of the book, has given us not just a problem but also worked through the solution. The solution is equivalent to having the ideal response for an assistant. The expert is showing us how to solve the problem in its full form, and as we read the solution we are training on the expert data, and later we can try to imitate the expert. This is supervised finetuning.
- Practice problems. Usually many at the end of each chapter. We know practice problems are critical for learning, because of what they get you to do: they get you to practice yourself and discover ways of solving these problems yourself. What you are given is a problem description and the final answer, usually in the answer key, but not the solution. You know the answer you are trying to get to, you have the problem statement, and you are trying out many different things and seeing what gets you to the final solution best. In the process you lean on the background information from pretraining and maybe a little bit of imitation of human experts. This is reinforcement learning.
Why we cannot write the solutions for it
Before showing the mechanics he spends seven minutes on why this stage is necessary at all, and this is the most underrated argument in the video.
He goes back to the Emily fruit problem, with the token view up in TikTokenizer because he wants to remind you again that we are always working with one dimensional token sequences, and this is the native view of the LLM. This is what it actually sees. It sees token IDs.
He puts four candidate solutions on screen. All four reach the answer 3. Some set up a system of equations. Some just talk through it in English. Some skip right through to the solution. ChatGPT, given the question, defines a system of variables and does its little thing.
And then the observation that justifies the whole stage:
"If I am the human data labeler that is creating a conversation to be entered into the training set, I don't actually really know which of these conversations to add to the dataset."
He separates the two purposes that have been tangled together. The first purpose of a solution is to reach the right answer. The second, secondary purpose is presentation for the human, because we assume the person wants to see the intermediate steps and wants it presented nicely. Set presentation aside and focus only on reaching the answer. Which of these four is the optimal solution for the LLM to reach the right answer?
He does not know. As a human labeler he would not know. And he walks through why, using the compute budget per token he established earlier:
- One of the candidates is very nice because it is very few tokens, so it takes a short amount of time to get to the answer. But right at the step where it does the division in a single token, it is asking for a lot of computation to happen on that one individual token. So maybe this is a bad example to give the LLM, because it incentivizes skipping through the calculations very quickly, and it is going to make mistakes in the mental arithmetic.
- Maybe spreading it out more would work better. Maybe setting it up as an equation would be better. Maybe talking through it would be better. "We fundamentally don't know."
And the reason we do not know is not a gap in anyone's knowledge. It is structural:
"What is easy for you or I as human labelers, what's easy for us or hard for us, is different than what's easy or hard for the LLM. Its cognition is different."
Some of the token sequences that are trivial for him might be too much of a leap for the LLM. Conversely, many of the tokens he is writing out might be just trivial to the LLM, and we are wasting tokens on things that are trivial.
He generalizes this past the math example, and this is the sharpest version of the point. It is a very pervasive issue, because our knowledge is not the LLM's knowledge. The LLM has a ton of knowledge, a PhD in math and physics and chemistry and whatnot. In many ways it actually knows more than he does, and he is potentially not utilizing that knowledge in its problem solving. Conversely, he might be injecting a bunch of knowledge in his solutions that the LLM does not know in its parameters, and those are sudden leaps that are very confusing to the model. Our cognitions are different.
"We are not in a good position to create these token sequences for the LLM. They're useful by imitation to initialize the system, but we really want the LLM to discover the token sequences that work for it. It needs to find for itself what token sequence reliably gets to the answer given the prompt, and it needs to discover that in the process of reinforcement learning and of trial and error."
Reinforcement learning, mechanically (2:14:42)
He goes back to the Hugging Face inference playground and picks Gemma 2, the 2 billion parameter model, noting that two billion is very very small, a tiny model, but okay for the demo.
The mechanics are quite simple. We need to try many different kinds of solutions and see which ones work well. So: take the prompt, run the model, the model generates a solution, inspect the solution. We know the correct answer is $3.
- Attempt one. The model gets it correct, says $3.
- Delete, rerun. Attempt two. The model solves it in a slightly different way, and gets it correct. Every single attempt will be a different generation, because these models are stochastic: at every single token there is a probability distribution and we sample from it, so we go down slightly different paths.
- Attempt three. Again a slightly different solution, again correct.
And then the scale that makes it work. In practice you might sample thousands of independent solutions, or even a million solutions, for just a single prompt. Some will be correct and some will not. What we want to do is encourage the solutions that lead to correct answers.
He switches to the cartoon diagram, and the numbers here are his. One prompt, many different solutions tried in parallel. Fifteen solutions. Only four of them got the right answer, shown in green; the rest are red. He flags honestly that this specific prompt is not the best example, because it is trivial and even a two billion parameter model always gets it right, so he asks you to exercise some imagination and suppose the green ones are good and the red ones are bad.
Now, whatever token sequences happened in the red solutions, something went wrong somewhere, and that was not a good path. Whatever happened in the green ones, things went pretty well, and we want to do more things like that in prompts like this. The way we encourage that behavior in the future is to train on these sequences.
And here is the break from everything before it:
"These training sequences now are not coming from expert human annotators. There's no human who decided that this is the correct solution. This solution came from the model itself. The model is practicing here."
It tried out a few solutions, four of them worked, and now the model trains on them. This corresponds to a student looking at their own solutions and thinking: okay, this one worked really well, this is how I should be solving these kinds of problems.
There are many ways to tweak the methodology, but the simplest version of the core idea is to take the single best solution out of the four, which he highlights in yellow. That is the one that not only led to the right answer but maybe had other nice properties: maybe it was the shortest, maybe it looked nicest in some ways, and there are other criteria you could imagine. We decide that is the top solution, we train on it, and after the parameter update the model is slightly more likely to take that path in this kind of setting in the future.
Then the scale again, because one prompt is nothing. We run many different diverse prompts across lots of math problems and physics problems and wherever else, tens of thousands of prompts, with thousands of solutions per prompt, all happening at the same time. And as we iterate this process:
"The model is discovering for itself what kinds of token sequences lead it to correct answers. It's not coming from a human annotator. The model is playing in this playground and it knows what it's trying to get to, and it's discovering sequences that work for it. These are sequences that don't make any mental leaps, they seem to work reliably and statistically, and they fully utilize the knowledge of the model as it has it."
His one line summary: "It's basically a guess and check. We're going to guess many different types of solutions, we're going to check them, and we're going to do more of what worked in the future."
And the SFT model still matters in this picture, which is why the stages are sequential rather than alternatives. It initializes the model into the vicinity of the correct solutions. It gets the model to write out solutions, maybe gives it an understanding of setting up a system of equations, maybe it talks through a solution. So it gets you into the neighborhood. But reinforcement learning is where everything gets dialed in.
So the high level process for how we train large language models is, in short, very similar to how we train children, with one difference: children go through the chapters of books and do all these different types of training exercises within each chapter, whereas when we train AIs we do it stage by stage, depending on the type of that stage. First we read all the exposition in all the textbooks at the same time and build a knowledge base. Then we look at all the worked solutions from human experts across all the textbooks and get an SFT model that can imitate the experts, but does so kind of blindly, just doing its best guess, trying to mimic statistically the expert behavior. Then in the last stage we do all the practice problems across all the textbooks and we get the RL model.
Why RL is the frontier and not yet standard
He is careful to put this stage in context, because it is the one where the field is still moving.
The first two stages, pretraining and supervised finetuning, have been around for years. They are very standard and everyone does them, all the different LLM providers. It is this last stage, the RL training, that is a lot more early in its process of development and is not standard yet in the field.
And he is honest about why: he skipped over a ton of little details. The high level idea is very simple, trial and error learning, but there are a ton of details and little mathematical nuances to exactly how you pick the solutions that are the best, how much you train on them, what the prompt distribution is, how to set up the training run such that this actually works. A lot of little details and knobs attached to a core idea that is very very simple. Getting the details right is not trivial.
Which is why companies like OpenAI and the other LLM providers have experimented internally with reinforcement learning finetuning for LLMs for a while, but have not talked about it publicly. It was all kind of done inside the company.
And that is why the DeepSeek paper was such a big deal.
DeepSeek R1 (2:27:47)
The paper from DeepSeek, a company in China, talked very publicly about reinforcement learning finetuning for large language models, how incredibly important it is, and how it brings out a lot of reasoning capabilities in the models. It reinvigorated the public interest in using RL for LLMs and gave a lot of the details that are needed to reproduce their results and actually get the stage to work.
He takes you through what happens when you apply RL correctly.
Figure 2 of the paper, the quantitative result. This is the accuracy of solving mathematical problems, measured on AIME. He pulls up the competition web page so you can see the kinds of problems involved, and invites you to pause the video to read them. In the beginning the models are not doing very well, but as you update the model with many thousands of steps the accuracy continues to climb. The models are improving and solving these problems with higher accuracy as you do this trial and error on a large dataset.
But the qualitative result is more incredible than the quantitative one, and this is the part of the paper that mattered.
Another figure shows that later in the optimization the average length per response goes up. The model is using more tokens to get its higher accuracy results. It is learning to create very very long solutions. So why are the solutions long? He scrolls to the qualitative examples and reads what the model learned to do. It is an emergent property of the optimization: it just discovers that this is good for problem solving.
What it starts doing looks like this:
"Wait wait wait, that's an aha moment I can flag here. Let's reevaluate this step by step to identify the correct sum."
So what is the model doing? It is reevaluating steps. It has learned that it works better for accuracy to try out lots of ideas, try something from different perspectives, retrace, reframe, backtrack. It is doing a lot of the things that you and I are doing in the process of problem solving for mathematical questions. But it is rediscovering what happens in your head, not what you put down on the solution.
And that is the crucial distinction, because of the argument he built in the previous section:
"There is no human who can hardcode this stuff in the ideal assistant response. This is only something that can be discovered in the process of reinforcement learning, because you wouldn't know what to put here."
So the model learns what we call chains of thought in your head, and it is an emergent property of the optimization. That is what is bloating up the response length, and that is also what is increasing the accuracy.
"What's incredible here is basically the model is discovering ways to think. It's learning what I like to call cognitive strategies of how you manipulate a problem and how you approach it from different perspectives, how you pull in some analogies, how you try out many different things over time, check a result from different perspectives. Extremely incredible to see this emerge in the optimization without having to hardcode it anywhere. The only thing we've given it are the correct answers."
Then he runs it live on the Emily problem. DeepSeek R1 is available on chat.deepseek.com, and you have to make sure the DeepThink button is turned on to get the R1 model. He pastes the fruit problem in and reads the output, and the difference from the GPT-4o response is obvious on the page. Where the SFT style answer mimics an expert solution, R1 does this:
"Okay, let me try to figure this out. So Emily buys three apples and two oranges, each orange cost $2, total is 13, I need to find out blah blah blah." ... "Wait a second, let me check my math again to be sure." ... "Yep, all that checks out. I think that's the answer, I don't see any mistakes. Let me see if there's another way to approach the problem, maybe setting up an equation. Let the cost of one apple be..." ... "Yep, same answer. So definitely each apple is $3. All right, confident that that's correct."
Then, once the thinking process is done, it writes up the nice solution for the human and boxes in the correct answer at the bottom. So the two purposes he separated earlier, correctness and presentation, are now visibly separated in the output itself. "As you're reading this you can't escape thinking that this model is thinking."
"This is what's coming from the reinforcement learning process. This is what's bloating up the length of the token sequences. They're doing thinking and they're trying different ways. This is what's giving you higher accuracy in problem solving, and this is where we are seeing these aha moments."
And then the practical note on where to run it. Some people are a little bit nervous about putting very sensitive data into chat.deepseek.com, because this is a Chinese company. But DeepSeek R1 is an open weights model, available for anyone to download and use. You will not be able to run the full model in full precision on a MacBook or a local device, because it is fairly large, but many companies are hosting the full largest model, and the one he likes is Together AI. Sign up, go to playgrounds, select DeepSeek R1 in the chat; there are many other state of the art models in the dropdown. It is similar to the Hugging Face inference playground, but Together usually hosts all the state of the art models. The default settings will often be okay. Because the model was released by DeepSeek, what you get here should be basically equivalent in principle, identical in terms of the power of the model, quantitatively and qualitatively, just with different sampling randomness, and this one is coming from an American company.
Which models are thinking models, and which are not
Back in ChatGPT, he reads the dropdown. The models that say they use advanced reasoning, like o1, o3-mini and o3-mini-high, are referring to the fact that they were trained by reinforcement learning with techniques very similar to those of DeepSeek R1, per public statements of OpenAI employees. Those are thinking models trained with RL.
And the models you get in the free tier, GPT-4o and GPT-4o-mini, you should think of as mostly SFT models. They do not do the thinking you see in the RL models. There is a little bit of reinforcement learning involved with them, which he covers in the RLHF section, but they are mostly SFT models and he thinks you should think about it that way. (Note that o1 and o3-mini were current when this was filmed in early 2025, and the lineup has moved on since; the distinction between a thinking model and an SFT model is the durable part.)
He picks o3-mini-high and runs it, and flags that those models may not be available unless you pay a ChatGPT subscription of either $20 per month or $200 per month for some of the top models. It says "reasoning" and starts to work.
And then an important caveat about what you are allowed to see. What appears in the OpenAI web interface is not exactly the chains of thought the model produced. Under the hood the model produces them, but OpenAI chooses not to show the exact chains of thought and shows little summaries of them instead. His read on why: they are worried about the distillation risk, that someone could come in and imitate those reasoning traces and recover a lot of the reasoning performance by just imitating the chains of thought. So they hide them and show only summaries. You are not getting exactly what you would get in DeepSeek with respect to the reasoning itself, though the written up solution is equivalent.
On performance, he gives a measured answer rather than a verdict. These models and the DeepSeek models are currently roughly on par. It is kind of hard to tell because of the evaluations, but if you are paying $200 per month to OpenAI, he believes some of those models still look better. DeepSeek R1 is still a very solid choice for a thinking model that is available to you, on that website or any other, because the model is open weights and you can just download it.
His own usage pattern, which is the most useful calibration in this part of the video: empirically about 80 to 90 percent of his use is just GPT-4o. If you have a prompt that requires advanced reasoning you should probably use or at least try the thinking models, but for a simpler knowledge based or factual question it is overkill. There's no need to think 30 seconds about some factual question. When he comes across a very difficult problem in math or code he reaches for a thinking model, but then he has to wait a bit longer, because they are thinking.
Finally he points at Google AI Studio, with a parenthetical that is pure Karpathy: it looks really busy and really ugly, because "Google's just unable to do this kind of stuff well." But choose the model Gemini 2.0 Flash Thinking Experimental 01-21 and that is also an early experimental thinking model from Google. He gives it the same problem, clicks run, it does something similar, and comes out with the right answer. So Google also offers a thinking model. Anthropic, at the time of filming, does not.
"This is kind of like the frontier development of these LLMs. I think RL is kind of like this new exciting stage, but getting the details right is difficult, and that's why all these models and thinking models are currently experimental as of 2025, very early 2025."
AlphaGo, and move 37 (2:42:07)
The connection he wants to make is that the discovery that reinforcement learning is an extremely powerful way of learning is not new to the field of AI, and one place we have already seen it demonstrated is the game of Go, where DeepMind famously developed AlphaGo. There is a movie about it you can watch.
He goes to the AlphaGo paper and scrolls to a plot he calls really interesting and kind of familiar, because we are rediscovering the same thing in the more open domain of arbitrary problem solving instead of the closed specific domain of Go. The plot shows Elo rating in the game of Go, with a line for Lee Sedol, an extremely strong human player, and it compares two things:
- A model trained by supervised learning, imitating human expert players. Get a huge amount of games played by expert players and try to imitate them, and you are going to get better, but then you top out and you never quite get better than the very top players. You are never going to reach Lee Sedol, because you are just imitating human players. "You can't fundamentally go beyond a human player if you're just imitating human players."
- A model trained by reinforcement learning, which is significantly more powerful. For Go that means the system is playing moves that empirically and statistically lead to winning the game. AlphaGo plays against itself and uses RL to create rollouts. The exact same diagram as the LLM case, except there is no prompt, because it is just a fixed game of Go. It tries out lots of plays, and the games that lead to a win, instead of to a specific answer, are reinforced and made stronger. The system learns the sequences of actions that empirically and statistically lead to winning. Reinforcement learning is not going to be constrained by human performance, and it overcomes even the top players like Lee Sedol.
He adds a nice honest aside on the plot itself: probably they could have run it longer and just chose to crop it at some point, because this costs money. And then the connecting line: "we're only starting to see hints of this diagram in larger language models for reasoning problems."
"We're not going to get too far by just imitating experts. We need to go beyond that, set up these little game environments, and let the system discover reasoning traces or ways of solving problems that are unique and that just basically work well."
And then move 37. On uniqueness: notice that when you are doing reinforcement learning, nothing prevents you from veering off the distribution of how humans are playing the game. He goes back to the AlphaGo search results and one of the suggested modifications is "move 37."
Move 37 refers to a specific point in time where AlphaGo played a move no human expert would play. The probability of this move being played by a human player was evaluated to be about 1 in 10,000. A very rare move. But in retrospect it was a brilliant move. AlphaGo, in the process of reinforcement learning, discovered a strategy of playing that was unknown to humans and is in retrospect brilliant.
He recommends a YouTube video of the reactions and analysis, and plays the audio of the commentators realizing what they are looking at. The moment is on YouTube:
"That's a very, that's a very surprising move. I thought, I thought it was, I thought it was a mistake."
People were freaking out because it is a move a human would not play, that AlphaGo played because in its training this move seemed to be a good idea. It just happens not to be the kind of thing humans would do. And that is again the power of reinforcement learning.
Then he extrapolates, carefully flagging it as speculation. In principle we can see the equivalent of that if we continue scaling this paradigm in language models, and what that looks like is unknown. What does it mean to solve problems in such a way that even humans would not be able to get it? How can you be better at reasoning or thinking than humans? How can you go beyond just a thinking human?
- Maybe it means discovering analogies that humans would not be able to create.
- Maybe it is a new thinking strategy. Which is kind of hard to think through.
- Maybe it takes a whole new language that is not even English. Maybe the model discovers its own language that is a lot better at thinking, because the model is unconstrained to even stick with English.
"In principle the behavior of the system is a lot less defined. It is open to do whatever works, and it is open to also slowly drift from the distribution of its training data, which is English."
But all of that can only be done if we have a very large and diverse set of problems in which these strategies can be refined and perfected. And that is where a lot of the frontier LLM research is going right now: trying to create the kinds of prompt distributions that are large and diverse. He calls them game environments in which the LLMs can practice their thinking. It is like writing practice problems. We have to create practice problems for all domains of knowledge, and if we have tons of them, the models will be able to reinforcement learn on them and create these kinds of diagrams, but in the domain of open thinking instead of a closed domain like Go.
Reinforcement learning from human feedback (2:48:26)
There is one more section within reinforcement learning, and it is learning in unverifiable domains.
Everything we have looked at so far is in a verifiable domain, meaning any candidate solution can be scored very easily against a concrete answer. The answer is 3, and we can easily score solutions against 3. Either we require the models to box in their answers and check equality with whatever is in the box, or we use an LLM judge, which looks at a solution and the answer and scores the solution for whether it is consistent. LLMs are empirically good enough at the current capability to do this fairly reliably. In either case we have a concrete answer, we are just checking solutions against it, and we can do this automatically with no humans in the loop.
The problem is unverifiable domains. Usually these are creative tasks: write a joke about pelicans, write a poem, summarize a paragraph. In these domains it becomes harder to score different solutions.
He demonstrates with jokes, and the demonstration is funny mostly because of how it fails. He asks ChatGPT for pelican jokes:
"So much stuff in their beaks because they don't believe in backpacks."
"Why don't pelicans ever pay for their drinks? Because they always bill it to someone else."
"Haha. Okay." His verdict: these models are obviously not very good at humor. And then an aside worth keeping: "Actually I think it's pretty fascinating, because I think humor is secretly very difficult, and the models have the capability, I think."
So we can generate lots of jokes, that is fine. The problem is how do we score them? In principle we could get a human to look at all of them, just like he did. But work out the arithmetic of reinforcement learning:
- You are going to be doing many thousands of updates.
- For each update you want to be looking at say a thousand prompts.
- For each prompt you want to be looking at hundreds or thousands of different generations.
Concretely, with his cartoon numbers: 1,000 updates, each on 1,000 prompts, with 1,000 rollouts per prompt scored. We could run RL with that setup just fine, if we had infinite human time. The problem:
"In the process of doing this I will need to ask a human to evaluate a joke a total of 1 billion times. And so that's a lot of people looking at really terrible jokes."
This is an unscalable strategy. We need an automatic one. And the solution was proposed in the paper that introduced reinforcement learning from human feedback, a paper from OpenAI at the time, and he notes in passing that many of those people are now co founders at Anthropic.
The core trick is indirection. We involve humans just a little bit, and the way we cheat is to train a whole separate neural network that we call a reward model, which imitates human scores. Ask humans to score some rollouts, then imitate those human scores with a neural network, and that network becomes a simulator of human preferences. Now that we have a simulator we can do RL against it. Instead of asking a real human, we ask a simulated human for their score of a joke. Once we have a simulator we can query it as many times as we want and the whole process is automatic.
The simulator is not going to be a perfect human, but if it is at least statistically similar to human judgment, you might expect that this will do something. And in practice indeed it does.
Here is how training the reward model works, with his own cartoon numbers:
- A prompt, "write a joke about pelicans," and five separate rollouts, five different jokes.
- Ask a human to order the jokes from best to worst. One is the funniest, then two, three, four, and five is the worst. We ask humans to order rather than give scores directly, because it is an easier task. It is easier for a human to give an ordering than to give precise scores. That ordering is the human's entire contribution to the training process.
- Separately, ask the reward model to score the same jokes. The reward model is a whole separate neural network, completely separate, probably also a Transformer, but it is not a language model in the sense that it generates diverse language. It is just a scoring model. It takes two inputs, the prompt and a candidate joke, and outputs a single number, a score, ranging for example from 0 to 1, where 0 is the worst and 1 is the best. At some stage in training it might give these five jokes scores like 0.1 and 0.8 and so on.
- Compare the reward model's scores with the human's ordering, set up a loss function, calculate a correspondence, and update the model.
The intuition, with his numbers:
- The joke the human ranked funniest got 0.8 from the model. The model kind of agreed, that is a relatively high score, but it should have been even higher, so after an update maybe it grows to 0.81.
- The joke the human ranked number two got only 0.1 from the model. A massive disagreement. That score needs to be much higher, so after an update it might grow a lot more, to maybe 0.15.
- The joke the human ranked worst got a fairly high number from the model, so after the update it should come down to maybe 0.35.
So we are doing what we did before. Slightly nudging the predictions using a neural network training process, trying to make the reward model scores be consistent with the human ordering. As we update it on human data it becomes a better and better simulator of the scores and orders that humans provide, and then it becomes the simulator of human preferences that we do RL against.
And critically, the human cost collapses. We are not asking humans a billion times. Maybe 1,000 prompts with five rollouts each, so about 5,000 jokes that humans have to look at in total, and all they do is give the ordering.
The upside, and the discriminator generator gap
He gives RLHF its due before taking it apart.
The first upside is that it lets us run reinforcement learning at all in arbitrary domains, including unverifiable ones: summarization, poem writing, joke writing, any other creative writing, in domains outside math and code.
The second is that empirically it works. When you apply RLHF correctly the models you get are just a little bit better. But he is scrupulous about the explanation: "I have a top answer for why that might be, but I don't actually know that it is super well established on why this is." So he offers his best guess.
His best guess is the discriminator generator gap. In many cases it is significantly easier to discriminate than to generate, for humans. In supervised finetuning we are asking humans to generate the ideal assistant response, and in many cases the ideal response is very simple to write. But in summarization or poem writing or joke writing, how is a human labeler supposed to give the ideal response? It requires creative human writing.
RLHF sidesteps this, because we get to ask people a significantly easier question. They are not asked to write poems directly. They are given five poems from the model and asked to order them. That is a much easier task for a human labeler.
"This allows a lot more higher accuracy data, because we're not asking people to do the generation task, which can be extremely difficult. We're just trying to get them to distinguish between creative writings and find the ones that are best, and that is the signal that humans are providing: just the ordering."
Then the system in RLHF discovers the kinds of responses that would be graded well by humans, and that step of indirection allows the models to become a bit better. So: it lets us run RL, it empirically results in better models, and it lets people contribute their supervision without having to do extremely difficult tasks.
The downside, and "the the the the the"
Unfortunately RLHF also comes with significant downsides.
The main one is structural. We are doing reinforcement learning not with respect to humans and actual human judgment, but with respect to a lossy simulation of humans. That lossy simulation could be misleading, because it is just a simulation, just a language model outputting scores, and it might not perfectly reflect the opinion of an actual human with an actual brain in all the possible different cases.
But there is something even more subtle and devious going on that dramatically holds RLHF back as a technique we could scale to significantly smart systems:
"Reinforcement learning is extremely good at discovering a way to game the model, to game the simulation."
The reward model is a Transformer, a massive neural net with billions of parameters, imitating humans in a simulated way. And these are massive complicated systems. Billions of parameters outputting a single score. It turns out that there are ways to game these models. You can find inputs that were not part of their training set and these inputs inexplicably get very high scores, but in a fake way.
What that looks like in practice, with his numbers. Run RLHF for 1,000 updates, which is a lot of updates, and you might expect your jokes to be getting better, that you are getting real bangers about pelicans. That is not exactly what happens.
"In the first few hundred steps the jokes about pelicans are probably improving a little bit, and then they actually dramatically fall off the cliff and you start to get extremely nonsensical results."
Specifically, the top joke about pelicans starts to be "the the the the the." Which makes no sense. Why should that be a top joke? But when you take "the the the the the" and plug it into your reward model, you would expect a score of zero, and actually the reward model loves it. It will tell you that this is a score of 1.0. This is a top joke.
It makes no sense, but it is because these models are just simulations of humans, massive neural nets, and you can find inputs that get into the part of the input space that gives you nonsensical results at the top. These are adversarial examples, specific little inputs that go between the nooks and crannies of the model and give nonsensical results.
And then the move that does not work, which is the important part. You might imagine doing the obvious thing: "the the the the the" is obviously not a score of one, it is obviously a low score, so let's add it to the dataset and give it an ordering that is extremely bad. And indeed the model will learn that it should have a very low score, and will give it zero.
"The problem is that there will always be basically an infinite number of nonsensical adversarial examples hiding in the model. If you iterate this process many times and you keep adding nonsensical stuff to your reward model and giving it very low scores, you'll never win the game. Reinforcement learning, if you run it long enough, will always find a way to game the model."
Fundamentally this is because our scoring function is a giant neural net, and RL is extremely good at finding the ways to trick it.
So the operational conclusion is blunt. You run RLHF for maybe a few hundred updates, the model is getting better, and then you have to crop it, and you are done. You cannot run too much against this reward model, because the optimization will start to game it. You crop it, you call it, and you ship it. You can improve the reward model, but you come across these situations eventually at some point anyway.
"RLHF is not RL"
This is the line he wants you to leave with, and he explains precisely what he means by it.
"RLHF is not RL. And what I mean by that is, I mean, RLHF is RL obviously, but it's not RL in the magical sense. This is not RL that you can run indefinitely."
The contrast is with the verifiable domains. There, you either got the correct answer or you didn't, and the scoring function is much much simpler: you are just looking at the boxed area and seeing if the result is correct. It is very difficult to game these functions. Gaming a reward model is possible.
So in verifiable domains you can run RL indefinitely. You could run for tens of thousands, hundreds of thousands of steps, and discover all kinds of really crazy strategies that we might not even ever think of. In the game of Go there is no way to game the winning or losing of a game. We have a perfect simulator: we know where all the stones are placed and we can calculate whether someone has won. There is no way to game that, so you can do RL indefinitely and eventually beat even Lee Sedol. With a gameable model you cannot repeat this process indefinitely.
"I kind of see RLHF as not real RL, because the reward function is gameable. So it's more like in the realm of a little finetuning. It's a little improvement, but it's not something that is fundamentally set up correctly where you can insert more compute, run for longer, and get much better and magical results. It lacks magic."
It can finetune your model and get a better performance, and indeed GPT-4o has gone through RLHF, because it works well. It is just not RL in the same sense. "RLHF is like a little finetune that slightly improves your model."
| Verifiable domains | Unverifiable domains | |
|---|---|---|
| Examples he gives | Math problems, code, the game of Go | Write a joke about pelicans, write a poem, summarize a paragraph |
| How a solution gets scored | Equality against a concrete answer, usually boxed in, or an LLM judge, which is reliable enough at current capability | A reward model: a separate neural network, probably a Transformer, trained to imitate human orderings, outputting one score from 0 to 1 |
| Humans in the loop | None. Fully automatic | About 1,000 prompts at five rollouts each, giving orderings only, roughly 5,000 reads |
| How long you can run it | Indefinitely. Tens or hundreds of thousands of steps | A few hundred updates, then you crop it and ship |
| What breaks first | Nothing in the scoring function. There is no way to game "did you win the game of Go" | Reward hacking. "The the the the the" scores 1.0 and the jokes fall off a cliff |
| Can you patch the failure | Not applicable | No. Add that adversarial example to the dataset and RL finds another, infinitely |
| What it can reach | Past human performance. Move 37, and chains of thought no labeler could write | A small improvement over SFT. Not RL in the magical sense |
Closing the technical content: the Swiss cheese model
That is most of the technical content. He took us through the three major stages and paradigms, pretraining, supervised finetuning and reinforcement learning, and showed that they loosely correspond to the process we already use for teaching children. Pretraining is the basic knowledge acquisition of reading exposition. Supervised finetuning is looking at lots and lots of worked examples and imitating experts. Reinforcement learning is the practice problems.
The only difference is that we now have to effectively write textbooks for LLMs and AIs across all the disciplines of human knowledge, and in all the cases where we actually want them to work, like code and math and basically every other discipline. So we are in the process of writing textbooks for them, refining all the algorithms he has presented at a high level, and then doing a really really good job at the execution of training these models at scale and efficiently.
He flags the thing he deliberately did not cover, and marks it as serious rather than incidental. These are extremely large and complicated distributed jobs that have to run over tens of thousands or even hundreds of thousands of GPUs, and the engineering that goes into this is really at the state of the art of what's possible with computers at that scale. Very serious work, underlying what are ultimately very simple algorithms.
Then the theory of mind, and the thing he wants you to take away. These models are really good, they are extremely useful as tools for your work, and you should not trust them fully. Even though we have mitigations for hallucinations, the models are not perfect and they will still hallucinate. It has gotten better over time and will continue to get better. But they can.
And in addition to that, the Swiss cheese model of LLM capabilities, which is the image he wants in your mind:
"The models are incredibly good across so many different disciplines, but then fail randomly, almost, in some unique cases. So for example, what is bigger, 9.11 or 9.9? The model doesn't know, but simultaneously it can turn around and solve olympiad questions. And so this is a hole in the Swiss cheese, and there are many of them, and you don't want to trip over them."
So: do not treat these models as infallible. Check their work. Use them as tools, use them for inspiration, use them for the first draft, but work with them as tools and be ultimately responsible for the product of your work.
Preview of things to come (3:09:39)
A few bullet points on what you can expect coming down the pipe.
Multimodality. The models will very rapidly become multimodal. Everything above concerned text, but very soon we will have LLMs that can operate natively and very easily over audio, so they can hear and speak, and over images, so they can see and paint. We are already seeing the beginnings of all of this, but this will all be done natively inside the language model, enabling natural conversations.
And here is the part that makes it not a fundamental change, which is the nice payoff of having built the whole picture on tokens. As a baseline you can tokenize audio and images and apply the exact same approaches as everything above. It is not a fundamental change, we just have to add some tokens.
- For audio, you can look at slices of the spectrogram of the audio signal, tokenize those, and add more tokens that represent audio into the context window, and train on them just like above.
- For images, you use patches, tokenize the patches separately, and then an image is just a sequence of tokens. There is a lot of early work in this direction and it actually kind of works.
So we create streams of tokens representing audio and images as well as text, interleave them, and handle them all simultaneously in a single model.
Agents, and long running tasks. Currently most of the work is that we hand individual tasks to the models on a silver platter. Please solve this task for me, and the model does the little task, but it is up to us to organize a coherent execution of tasks to perform jobs. The models are not yet at the capability required to do this in a coherent, error correcting way over long periods of time. They cannot fully string together tasks to perform longer running jobs. But they are getting there, and this is improving.
What is probably going to happen: we start to see agents which perform tasks over time, and you supervise them, you watch their work, and they come up once in a while to report progress. Tasks that do not take a few seconds of response but many tens of seconds, or even minutes or hours.
But these models are not infallible, so all of this will require supervision. And he offers an analogy that has aged well: in factories people talk about the human to robot ratio for automation. He thinks we are going to see something similar in the digital space, where we talk about human to agent ratios, and humans become a lot more supervisors of agent tasks in the digital domain.
Pervasive and invisible. Everything is going to become a lot more pervasive and invisible, integrated into the tools and everywhere.
Computer use. Right now these models are not able to take actions on your behalf, but he calls out ChatGPT's launch of Operator as an early example of handing off control to the model to perform keyboard and mouse actions on your behalf, and says that is also something he finds very interesting.
And the research that is still to do: test time training. This is the most interesting of his forward looking points because it names a real structural gap rather than a capability gap.
Everything above has two major stages. First the training stage, where we tune the parameters to perform the tasks well. Then, once we get the parameters, we fix them and deploy the model for inference, and from there the model is fixed. It does not change anymore. It does not learn from all the stuff that it's doing at test time. It is a fixed number of parameters, and the only thing changing is the tokens inside the context window.
So the only type of test time learning the model has access to is the in context learning of its dynamically adjustable context window. And he thinks this is still different from humans, who actually are able to learn depending on what they are doing, especially when you sleep, when your brain is updating your parameters or something like that. There is no equivalent of that currently in these models and tools. There are a lot of wonky ideas still to be explored.
He thinks this will be necessary, and the argument is concrete. The context window is a finite and precious resource. Once we start to tackle very long running multimodal tasks and we are putting in videos, these token windows will start to grow extremely large, not thousands or even hundreds of thousands but significantly beyond that. And the only trick available to us right now is to make the context windows longer. "I think that approach by itself will not scale to actual long running tasks that are multimodal over time, and so I think new ideas are needed."
Keeping track of LLMs (3:15:15)
Three resources he has consistently used to stay up to date.
One: LM Arena. An LLM leaderboard that ranks all the top models, with the ranking based on human comparisons. Humans prompt the models and judge which one gives a better answer, they do not know which model is which, they are just looking at which answer is better, and you calculate a ranking from that. Clicking any model takes you to where that model is hosted.
He reads the board as it stood at the time of filming, and the detail he stops on is in the last column:
- Google Gemini is currently on top, with OpenAI right behind.
- DeepSeek is in position number three, and the reason this is a big deal is the license column. DeepSeek is an MIT license model. It is open weights. Anyone can use the weights, anyone can download them, anyone can host their own version and use it however they like. It is not a proprietary model you do not have access to. "This is kind of unprecedented, that a model this strong was released with open weights. So pretty cool from the team."
- Then a few more models from Google and OpenAI, then the usual suspects: xAI, then Anthropic with Sonnet at number 14, then Meta with Llama, which like DeepSeek is open weights but sits lower down.
And then the caveat, which is the useful part:
"I will say that this leaderboard was really good for a long time. I do think that in the last few months it's become a little bit gamed, and I don't trust it as much as I used to."
His evidence is his own sense of the field rather than a measurement, and he says so: empirically he feels like a lot of people are using Sonnet from Anthropic and that it is a really good model, but it is all the way down at number 14, and conversely he thinks not as many people are using Gemini but it is ranking really really high. So use this as a first pass, but try out a few of the models for your tasks and see which one performs better.
Two: the AI News newsletter, produced by swyx and friends, which he thanks them for maintaining. "AI News is not very creatively named, but it is a very good newsletter." Why it works: it is extremely comprehensive. Go to the archives and you see it is produced almost every other day, and some of it is written and curated by humans but a lot of it is constructed automatically with LLMs. So it is very comprehensive, and you are probably not missing anything major if you go through it. Of course you are probably not going to go through it, because it is so long, but he thinks the summaries all the way up top are quite good and have some human oversight.
Three: X. A lot of AI happens on X, so follow people you like and trust and get your latest and greatest there as well.
Where to find LLMs, and how to run them (3:18:34)
Four cases, and he is specific about which tool for which.
Proprietary models: go to the provider's website. For OpenAI that is ChatGPT. For Gemini it is either the Gemini site or AI Studio, and he cannot resist noting that "I think they have two for some reason that I don't fully understand. No one does."
Open weights models: go to an inference provider. His favorite is Together AI. Go to the playground and you can pick lots of different models, all of them open models of different types, and talk to them there.
Base models: go to Hyperbolic. This is the one that is actually hard to find, and he says so. "It's not as common to find base models even on these inference providers. They are all targeting assistants and chat." Even on Together he could not see base models. So for base models he usually goes to Hyperbolic, because they serve Llama 3.1 405B base, and he loves that model. And a small editorial wish: "I wish more people hosted base models, because they are useful and interesting to work with in some cases."
Smaller models, locally on your own machine. DeepSeek's biggest model you will not be able to run locally on a MacBook, but there are smaller versions of the DeepSeek model that are distilled, and you can also run models at smaller precision, not the native precision of fp8 on DeepSeek or bf16 on Llama but much lower than that. He says not to worry if you do not fully understand those details: the point is that you can run smaller distilled versions at even lower precision, fit them on your computer, and actually run pretty okay models on your laptop.
His tool for this is LM Studio, and his review of it is unsparing and funny.
"It kind of actually looks really ugly, and I don't like that it shows you all these models that are basically not that useful, like everyone just wants to run DeepSeek. So I don't know why they give you these 500 different types of models. They're really complicated to search for, and you have to choose different distillations and different precisions and it's all really confusing."
But once you understand how it works, which he says is a whole separate video, you can load up a model and talk to it. He loads Llama 3.2 Instruct 1B, asks for pelican jokes, asks for another one, and gets another one. All of that happens locally on your computer. We are not actually going to anyone else. This is running on the GPU on the MacBook Pro, which is very nice, and you can then eject the model when you are done, which frees up the RAM.
So LM Studio is probably his favorite, even though he thinks it has a lot of UI and UX issues and is really geared towards professionals almost. If you watch some videos on YouTube you can figure out how to use the interface.
The grand summary (3:21:46)
He loops all the way back to where he started. We go to ChatGPT, we enter a query, we hit go. What exactly is happening? What are we seeing? What are we talking to? How does this work?
He walks the whole path one more time, now that every piece has a name.
Your query is chopped up into tokens. He goes back to TikTokenizer and shows the place in the format where the user query goes. Your query goes into the conversation protocol format, the way we maintain conversation objects, so it gets inserted there, and then the whole thing ends up being just a one dimensional token sequence under the hood. ChatGPT saw that token sequence, and when you hit go it basically continues appending tokens into this list. It continues the sequence. It acts like a token autocomplete. He pastes the response back into the tokenizer so you can see the tokens it continued with.
And then the real question. Why are these the tokens that the model responded with? Where are they coming from? What are we talking to, and how do we program this system?
- Stage one, pretraining, which fundamentally has to do with knowledge acquisition from the internet into the parameters of the neural network.
- Stage two, supervised finetuning, is where the personality really comes in. A company like OpenAI curates a large dataset of conversations, say 1 million conversations across very diverse topics between a human and an assistant. And even though there is a lot of synthetic data generation used throughout this entire process and a lot of LLM help, fundamentally this is a human data curation task with lots of humans involved. Those humans are data labelers hired by OpenAI, given labeling instructions that they learn, whose task is to create ideal assistant responses for any arbitrary prompts. They are teaching the neural network by example how to respond.
So what is the thing that came back? His answer, the clearest statement of the idea in the whole video:
"This is the neural network simulation of a data labeler at OpenAI. So it's as if I gave this query to a data labeler at OpenAI, and this data labeler first reads all of the labeling instructions from OpenAI and then spends two hours writing up the ideal assistant response to this query and giving it to me. Now we're not actually doing that, right, because we didn't wait two hours. So what we're getting here is a neural network simulation of that process."
And we have to keep in mind that these neural networks do not function like human brains do. They are different. What is easy or hard for them is different from what is easy or hard for humans, so we really are just getting a simulation. He shows the token stream and the network with its activations and neurons in between one last time: a fixed mathematical expression that mixes inputs from tokens with parameters of the model and gets you the next token in a sequence, with a finite amount of compute for every single token. Whatever the humans write, the language model is imitating on this token level, with only that specific computation per token.
And as a result of that, and of the cognitive differences, the models suffer in a variety of ways and you have to be very careful with their use. They hallucinate. And there is the Swiss cheese model: there are holes in the cheese, and sometimes the models will just arbitrarily do something dumb. Even though they are doing lots of magical stuff, sometimes they just cannot. Maybe you are not giving them enough tokens to think. Maybe they are going to make stuff up because their mental arithmetic breaks. Maybe they are suddenly unable to count letters. Maybe they are unable to tell you that 9.11 is smaller than 9.9, and it looks kind of dumb. So it is a Swiss cheese capability and we have to be careful with it. And we saw the reasons for each of them.
But things change a little bit when you reach for a thinking model, and he is precise about why. GPT-4o basically does not do reinforcement learning. It does do RLHF, but RLHF is not RL. There is no time for magic in there, it is just a little bit of finetuning. The thinking models do use RL. They go through this third stage of perfecting their thinking process and discovering new thinking strategies and solutions to problem solving that look a little bit like your internal monologue in your head, and they practice that on a large collection of practice problems that companies like OpenAI create and curate and make available to the LLMs.
So when you talk to a thinking model, what you are seeing is not anymore just the straightforward simulation of a human data labeler. OpenAI is not showing us the under the hood thinking and the chains of thought underlying it, only a summary, but we know such a thing exists, and what we are getting is actually not just an imitation of a human data labeler. It is a function of thinking that was emergent in a simulation. It comes from the reinforcement learning process.
Two honest open questions he leaves on the table rather than answering:
- Does it transfer? The question he puts to the field: whether the thinking strategies developed inside verifiable domains transfer and are generalizable to other domains that are unverifiable, such as creative writing. "The extent to which that transfer happens is unknown in the field, I would say." We are not sure whether we can do RL on everything verifiable and see the benefits of that on things that are unverifiable.
- How early is this? "This reinforcement learning here is still way too new, primordial and nascent. So we're just seeing the beginnings of the hints of greatness in the reasoning problems." We are seeing something that is in principle capable of something like the equivalent of move 37, but not in the game of Go, in open domain thinking and problem solving. In principle this paradigm is capable of doing something really cool and new and exciting, something even that no human has thought of before. In principle these models are capable of analogies no human has had. He finds it incredibly exciting that these models exist, but again it is very early, and these are primordial models for now, and they will mostly shine in domains that are verifiable like math and code.
And then the close, which is the same advice he gave at the top, now with three and a half hours of mechanism behind it:
"It is an extremely exciting time to be in the field. Personally I use these models all the time daily, tens or hundreds of times, because they dramatically accelerate my work. I think a lot of people see the same thing, and I think we're going to see a huge amount of wealth creation as a result of these models. Be aware of some of their shortcomings. Even with RL models they're going to suffer from some of these. Use it as a tool in a toolbox, don't trust it fully, because they will randomly do dumb things. They will randomly hallucinate, they will randomly skip over some mental arithmetic and not get it right, they randomly can't count or something like that. So use them as tools in the toolbox, check their work, and own the product of your work. But use them for inspiration, for first draft, ask them questions, but always check and verify, and you will be very successful in your work if you do so."
Then, with the warmth that makes this video what it is: "I hope this video was useful and interesting to you. I hope you had fun. And it's already very long, so I apologize for that."
Key takeaways
- A base model is an internet document simulator, not an assistant. Every assistant you have used is a base model plus a thin, curated imitation layer. The base model will answer "what is 2 plus 2" with philosophy, and will recite the zebra Wikipedia article from memory, because that is what the internet looks like.
- Pretraining buys knowledge, supervised finetuning buys manners, reinforcement learning buys problem solving. Pretraining is three months on thousands of computers. Supervised finetuning is three hours with the identical algorithm and a different dataset. The expensive stage and the stage that gives you the personality are not the same stage.
- When you talk to a chat model you are talking to a statistical simulation of a human labeler following a company's labeling instruction document, which runs to hundreds of pages and tells them to be helpful, truthful and harmless. That explains the tone, the refusals, the hedging and the formatting better than any story about machine intelligence.
- Hallucination is the default behavior of a system trained to imitate confident answers. Every "who is X" in the training set is answered confidently, so the model answers confidently, even about names that do not exist. The two fixes: interrogate the model to find the boundary of its knowledge and train "I don't know" in at exactly those points, and give it tools so it can answer from text it can see.
- The parameters are a vague recollection. The context window is working memory. Weights are something you read a month ago. Anything in the context window is directly accessible. If you care about accuracy, paste the material in rather than asking the model to remember it.
- Models need tokens to think. Compute per token is roughly fixed and small, about 100 layers, so reasoning must be distributed across tokens. An answer that arrives immediately had one forward pass behind it, and everything after it is post hoc justification. The intermediate steps ChatGPT shows you are not for you, they are for the model.
- For arithmetic, counting or spelling, make it write code. Not because the model gets a calculator, but because copy pasting into a Python string is a task the model is genuinely good at, and the interpreter has correctness guarantees that mental arithmetic does not.
- Tokenization explains the spelling and counting failures. The model never sees characters. "Ubiquitous" is three tokens, twenty dots can be one token, and strawberry's Rs were hidden inside chunks. We keep tokens for efficiency, because character level models would create sequences nobody knows how to handle yet.
- Jagged intelligence is real and not fully understood. 9.11 versus 9.9 goes wrong in a way that traces to Bible verse neurons, and even knowing the mechanism does not make the failure predictable. Swiss cheese: excellent almost everywhere, with holes you do not want to trip over.
- Reinforcement learning on verifiable problems is qualitatively different from RLHF. The first can run indefinitely and discover strategies no human taught it, because a boxed answer cannot be gamed. The second runs a few hundred updates and gets cropped, because a reward model can always be gamed, and "the the the the the" will eventually score 1.0.
- Chains of thought were not designed, they emerged. DeepSeek R1's response lengths grew on their own because backtracking and rechecking raise accuracy. No human could have written those traces into a labeling guide, which is exactly the argument for the stage existing at all.
- Use them as tools in a toolbox. Check their work. Own the product of your work. He says it at the start and again at the end, and in between he spends three hours earning it.
Chapters
- 0:00:00 introduction
- 0:01:00 pretraining data (internet)
- 0:07:47 tokenization
- 0:14:27 neural network I/O
- 0:20:11 neural network internals
- 0:26:01 inference
- 0:31:09 GPT-2: training and inference
- 0:42:52 Llama 3.1 base model inference
- 0:59:23 pretraining to post-training
- 1:01:06 post-training data (conversations)
- 1:20:32 hallucinations, tool use, knowledge/working memory
- 1:41:46 knowledge of self
- 1:46:56 models need tokens to think
- 2:01:11 tokenization revisited: models struggle with spelling
- 2:04:53 jagged intelligence
- 2:07:28 supervised finetuning to reinforcement learning
- 2:14:42 reinforcement learning
- 2:27:47 DeepSeek-R1
- 2:42:07 AlphaGo
- 2:48:26 reinforcement learning from human feedback (RLHF)
- 3:09:39 preview of things to come
- 3:15:15 keeping track of LLMs
- 3:18:34 where to find LLMs
- 3:21:46 grand summary
Notable quotes
Lightly cleaned of filler only. Wording, order and timestamps are his.
"It is obviously magical and amazing in some respects. It's really good at some things, not very good at other things, and there's also a lot of sharp edges to be aware of." Andrej Karpathy, opening the video, 0:00:00
"The FineWeb dataset, which is fairly representative of what you would see in a production grade application, actually ends up being only about 44 terabytes of disk space. You can get a USB stick for like a terabyte very easily." On how small the filtered internet is, 0:02:04
"This sequence length is actually going to be a very finite and precious resource in our neural network." The sentence the whole of tokenization follows from, 0:09:14
"There's no memory in this expression. It's a fixed mathematical expression from input to output with no memory. It's stateless." On why you should not think of these as neurons, 0:24:38
"It's a token simulator. It's an internet text token simulator. They just create sort of remixes of the internet. They dream internet pages." On what a base model actually is, 0:43:10
"You can think of these 405 billion parameters as a kind of compression of the internet. You can think of it as kind of like a zip file, but it's not a lossless compression, it's a lossy compression." 0:49:56
"You're not talking to a magical AI, you're talking to an average labeler. This average labeler is probably fairly highly skilled, but you're talking to kind of like an instantaneous simulation of that kind of a person that would be hired in the construction of these datasets." The reframing the video is best known for, 1:18:03
"They don't have access to the internet, they're not doing research. These are statistical token tumblers, as I call them. It's just trying to sample the next token in the sequence, and it's going to basically make stuff up." On where hallucinations come from, 1:22:37
"Knowledge in the parameters of the neural network is a vague recollection. The knowledge in the tokens that make up the context window is the working memory." The single most useful mental model in the video, 1:39:34
"It has no persistent self, it has no sense of self. It's a token tumbler, and it follows the statistical regularities of its training set." On why asking a model what model it is makes no sense, 1:42:05
"These are not for you, these are for the model. If the model is not creating these intermediate results for itself, it's not going to be able to reach three." On the intermediate steps ChatGPT shows you, 1:53:48
"Models need tokens to think. Distribute your computation across many tokens. Ask models to create intermediate results, or whenever you can, lean on tools and tool use instead of allowing the models to do all of the stuff in their memory." 1:57:51
"What is easy for you or I as human labelers, what's easy for us or hard for us, is different than what's easy or hard for the LLM. Its cognition is different." The argument for why reinforcement learning has to exist, 2:17:54
"We really want the LLM to discover the token sequences that work for it. It needs to find for itself what token sequence reliably gets to the answer given the prompt." 2:19:28
"There is no human who can hardcode this stuff in the ideal assistant response. This is only something that can be discovered in the process of reinforcement learning, because you wouldn't know what to put here." On the emergent backtracking in DeepSeek R1, 2:31:48
"The model is discovering ways to think. It's learning what I like to call cognitive strategies of how you manipulate a problem and how you approach it from different perspectives. The only thing we've given it are the correct answers." 2:32:18
"You can't fundamentally go beyond a human player if you're just imitating human players." On the ceiling of supervised learning, from the AlphaGo plot, 2:43:31
"That's a very, that's a very surprising move. I thought, I thought it was a mistake." The commentary clip Karpathy plays on AlphaGo's move 37, 2:46:04
"In the process of doing this I will need to ask a human to evaluate a joke a total of 1 billion times. And so that's a lot of people looking at really terrible jokes." Why RLHF has to exist, 2:51:41
"Actually I think it's pretty fascinating, because I think humor is secretly very difficult, and the models have the capability, I think." After reading out two bad pelican jokes, 2:49:39
"When you take 'the the the the the' and you plug it into your reward model, you'd expect a score of zero, but actually the reward model loves this as a joke. It will tell you that this is a score of 1.0. This is a top joke. And this makes no sense." Reward hacking, demonstrated, 3:02:25
"RLHF is not RL. And what I mean by that is, I mean, RLHF is RL obviously, but it's not RL in the magical sense. This is not RL that you can run indefinitely." 3:05:00
"It's not RL in the sense that it lacks magic. It can finetune your model and get a better performance, but it's not something that is fundamentally set up correctly where you can insert more compute, run for longer, and get much better and magical results." 3:06:34
"The models are incredibly good across so many different disciplines, but then fail randomly, almost, in some unique cases. So this is a hole in the Swiss cheese, and there are many of them, and you don't want to trip over them." 3:08:38
"This is the neural network simulation of a data labeler at OpenAI. It's as if I gave this query to a data labeler at OpenAI, and this data labeler first reads all of the labeling instructions and then spends two hours writing up the ideal assistant response to this query and giving it to me." The grand summary, 3:24:27
"Use them as tools in the toolbox, check their work, and own the product of your work. But use them for inspiration, for first draft, ask them questions, but always check and verify, and you will be very successful in your work if you do so." The last substantive thing he says, 3:30:37
Resources mentioned
Data, tokenization and tooling
- FineWeb, the dataset he walks the pipeline through, and the Hugging Face writeup on how it was built, on Hugging Face
- Common Crawl, scouring the internet since 2007, 2.7 billion pages indexed as of 2024
- TikTokenizer, the site he uses for every token level demo in the video, with the GPT-4
cl100k_basetokenizer. The underlying library is OpenAI's tiktoken - Byte pair encoding and UTF-8, the two pieces of the tokenization story
- Brendan Bycroft's LLM visualization, the ~85,000 parameter Transformer he opens to show the forward pass
- Hugging Face inference playground, where he runs the Falcon, Mistral and Gemma demos
Models and papers
- The GPT-2 paper (2019) and the GPT-2 repository, his example of the first recognizably modern stack
- llm.c and his GPT-2 reproduction writeup, one 8xH100 node, 24 hours, $672, against the roughly $40,000 the original run cost in 2019. See also nanoGPT
- The Llama 3 herd of models paper, source of the 405B and 15 trillion token numbers, the end of 2023 knowledge cutoff, and the factuality procedure for mitigating hallucinations
- Llama 3.1 405B base, the base model he plays with, and Llama 3.2 1B Instruct, the one he runs locally
- InstructGPT (OpenAI, 2022), where section 3.4 names the contractors and the labeling instructions get excerpted
- OpenAssistant, the open reproduction of InstructGPT style conversation data
- UltraChat, his example of a modern, largely synthetic SFT mixture
- Falcon 7B Instruct, the older model he picks on for hallucinations and for the "I was built by OpenAI" answer
- Mistral 7B, the stand in model he interrogates about Dominik Hašek
- Gemma 2 2B, the tiny model he runs the reinforcement learning rollouts on
- OLMo from the Allen Institute for AI, fully open source, whose SFT mixture contains the 240 hardcoded identity conversations
- DeepSeek R1, the paper that made reinforcement learning finetuning public, with figure 2 on AIME accuracy and the emergent response length growth. Run it at chat.deepseek.com with DeepThink on, or on Together AI
- Deep reinforcement learning from human preferences, the paper that introduced RLHF, from OpenAI at the time, several of whose authors went on to co found Anthropic
- Mastering the game of Go with deep neural networks and tree search, the AlphaGo paper with the Elo plot, plus AlphaGo the movie and the move 37 moment on video
Hardware, hosting and running them yourself
- Nvidia H100, eight to a node, and the $3.4 trillion market cap he attributes to this exact workload
- Lambda, where he rents the 8x H100 node at $3 per GPU per hour
- Together AI for open weights models, Hyperbolic for base models including Llama 3.1 405B base
- LM Studio for running distilled, lower precision models locally on your own machine
- ChatGPT and OpenAI, Gemini and Google AI Studio with Gemini 2.0 Flash Thinking Experimental, DeepSeek
- ChatGPT Operator, his example of a model taking keyboard and mouse actions on your behalf
Keeping up
- LM Arena, the human preference leaderboard, with his caveat that it has become a little bit gamed
- AI News by swyx and friends, the comprehensive every other day newsletter
- X, where a lot of AI happens
People, places and oddities named along the way
- Labeling workforces: Upwork and Scale AI, per the InstructGPT paper
- Tom Cruise, John Barrasso and Genghis Khan, the three "who is" conversations, against the invented Orson Kovats
- Dominik Hašek, the Wikipedia featured article he picks at random to generate factuality probes from
- The zebra Wikipedia article, regurgitated verbatim by the base model, and chapter one of Jane Austen's Pride and Prejudice, pasted into the context window to beat recollection
- Rayleigh scattering, the answer the prompt built assistant gives to "why is the sky blue"
- Lee Sedol and the 2016 match, and the Elo rating plot that separates imitation from reinforcement
- Adversarial examples and reward hacking, the failure mode behind "the the the the the"
Related pages on this site
- Let's build the GPT tokenizer, the two hour version of the eight minutes he spends on tokenization here, and the source of most of the sharp edges in the second half of this video
- Let's reproduce GPT-2, the four hour build behind the llm.c run he shows on screen
- Attention in transformers, which opens the box he deliberately leaves closed
- RLHF, by John Schulman, the direct expansion of this video's last hour from the person who led the work
- Mechanistic interpretability, by Chris Olah, on the kind of activation analysis behind the Bible verse explanation for 9.11 versus 9.9
- Software is changing again, Karpathy's own follow on talk about what to build with all of this
Where this sits in the LLM Learning track
This is the opening lecture for a reason: it is the only video in the track that touches every stage of the stack in one pass, so everything after it has something to attach to. Start here even if you think you know the material.
What comes next opens the boxes this one leaves closed, roughly in the order he closed them. Attention in transformers takes the component Karpathy treats as a diagram with numbers flowing through it and walks it one matrix at a time. LLMs in five formulas supplies the quantitative frame for the things he gives you intuitively here, including the finite compute per token that drives the "models need tokens to think" section. Then the two long builds: the GPT tokenizer is the two hour version of the eight minutes above, and it is where every spelling and counting failure in this video stops being a curiosity and becomes something you have typed; reproduce GPT-2 is the four hour version of the training run he slides onto the screen at 0:34:58.
Two later videos in the track are direct expansions of single sections here. John Schulman on RLHF is the whole of the last hour of this video, argued at depth by the person who led that work, including the clearest statement anywhere of why imitation training teaches a model to hallucinate. And Chris Olah on mechanistic interpretability is the discipline behind the one explanation Karpathy cites secondhand and flags as unread: the Bible verse neurons lighting up on 9.11. When you get to them, come back to 1:20:32 and 2:04:53 and the handoff is obvious.
Where it stands
The reconstruction above is Karpathy's lecture in his own frame. Three notes on reading it almost two years later.
What has held up. Nearly all of the mechanism. Tokenization and its consequences, the compute budget per token, the labeler imitation framing, the parameters versus context window distinction, the difference between verifiable and unverifiable reward, and reward hacking are all still the right way to think about these systems, and they have become more load bearing rather than less. His forward looking bullets have aged unusually well: native multimodality arrived, long running agents with a human supervising a fleet arrived, and the human to agent ratio turned out to be a real unit of measure rather than an analogy.
What has moved. The model lineup is the dated part, and visibly so. o1 and o3-mini were the current thinking models when this was filmed in early 2025, GPT-4o was the default, and his line that Anthropic does not offer a thinking model was true that month and stopped being true shortly after. The claim that reinforcement learning is "not standard yet in the field" is the single most dated sentence in the video, since RL post training became standard almost immediately after. His LM Arena snapshot is a photograph of one week. Treat every model name as a timestamp and every mechanism as current.
What he flags himself, and it is worth noticing how often. He says he has not read the paper behind the 9.11 explanation and is relaying what a team told him. He says he does not know why RLHF empirically helps and offers his best guess, labeled as a guess. He says he could be wrong about how much of UltraChat is human. He says he does not know why the special token is called an imaginary monologue. He corrects his own parameter count on camera. He admits he lacks the expertise to verify the Paris landmarks the base model produced. For a video this confident about mechanism, the rate of unprompted uncertainty flagging is the reason the confident parts read as an argument rather than a sales pitch.


