At a glance
Ilya Sutskever walked onto a NeurIPS stage in Vancouver in December 2024 to collect a Test of Time award for Sequence to Sequence Learning with Neural Networks, the 2014 paper he wrote with Oriol Vinyals and Quoc Le. He did not give a victory lap. He put the slides from his own 2014 NeurIPS talk back on the screen, one at a time, and graded them: "a lot of the things in this work were correct, but some not so much."
Then, nine minutes in, he said the sentence the talk is remembered for. "Pre-training as we know it will unquestionably end." The reasoning is arithmetic, not prophecy. Compute grows through better hardware, better algorithms and larger clusters. Data does not grow, "because we have but one internet." Data, he said, is the fossil fuel of AI: laid down once, burned once, and we have achieved peak data.
The rest is him being careful about what he does and does not know. He hands the "what comes next" question to other people's speculations (agents, synthetic data, inference time compute) rather than claiming it himself. He offers one piece of evidence that a different scaling regime is even findable, and it is not from computer science: a chart of brain mass against body mass in which hominids sit on a visibly different slope from every other mammal. He closes on superintelligence, lists four properties he expects it to have, refuses to say how or when, and makes the point that gets the least attention and deserves more: a system that genuinely reasons is a system you cannot predict.
Then a Q&A that produced some of the most quoted moments of the conference, including the answer where he says he does not feel like the right person to comment on cryptocurrency, and the answer where he dismantles the phrase "out of distribution" by pointing out that our standard for generalization has moved so far in ten years that the word barely survives the trip.
Twenty four and a half minutes. Almost no slides. Almost nothing wasted.
The deep explanation
The photograph, and the frame he sets (0:01)
He opens with thanks: to the organizers for choosing the paper, and to "my incredible co-authors and collaborators, Oriol Vinyals and Quoc Le, who stood right before you a moment ago." The award was shared that year, which is why the stage was already warm.
Then the first image: a screenshot from a talk he gave ten years earlier, at NeurIPS 2014 in Montreal. "It was a much more innocent time," he says, and shows the two photographs side by side. "Here we are shown in the photos. This is the before. Here's the after, by the way. And now we've got more experienced, hopefully wiser."
That sets the frame for the whole first half, and it is an unusual frame for an award talk. He is not going to summarize the paper. He is going to put the original slides back up and audit them.
Because a lot of the things in this work were correct, but some not so much. And we can review them and we can see what happened and how it gently flowed to where we are today.
"Gently flowed" is doing real work in that sentence. The claim is continuity: nothing in the last decade was a break with 2014, it was the execution of 2014.
The whole paper in three bullet points (1:37)
Before the slides, the summary, which he gives as three lines and then stops:
But the summary of what we did is the following three bullet points. It's an autoregressive model trained on text. It's a large neural network and it's a large data set. And that's it.
That is the entire architecture of the modern frontier model, stated in 2014 and restated here as the thing that did not need to change. Everything that follows is him checking each bullet against the record.
Slide one: the deep learning hypothesis (1:45)
The first original slide is titled "the deep learning hypothesis." His verdict on it: "Not too bad."
What it said: if you have a large neural network with ten layers, it can do anything a human being can do in a fraction of a second.
He then unpacks why the hypothesis is phrased around a fraction of a second, because that framing looks arbitrary until you follow the chain. It is a three step argument resting on connectionism:
Well, if you believe the deep learning dogma, so to say, that artificial neurons and biological neurons are similar, or at least not too different, and you believe that real neurons are slow, then anything that we can do quickly, by we I mean human beings, I even mean just one human in the entire world. If there is one human in the entire world that can do some task in a fraction of a second, then a 10-layer neural network can do it, too, right? It follows. You just take their connections and you embed them inside your neural net, the artificial one.
The construction step is the striking part. It is not an empirical claim that training will succeed. It is a claim that a solution exists in the weight space: take the human's connections, embed them in the artificial net, and by construction the network can represent the function. Trainability is a separate problem.
He is explicit that the depth was a limitation rather than a theory: "We focused on 10-layer neural networks because this was the neural networks we knew how to train back in the day. If you could go beyond in your layers somehow, then you could do more. But back then we could only do 10 layers, which is why we emphasized whatever human beings can do in a fraction of a second."
Slide two: "our main idea", and the autoregressive bet (3:09)
The second original slide is headed "our main idea." He invites the room to spot what is on it: "you may be able to recognize two things, or at least one thing. You might be able to recognize that something auto-regressive is going on here."
Then he says what the slide actually claims, which is stronger than it looks:
This slide says that if you have an auto-regressive model, and it predicts the next token well enough, then it will in fact grab and capture and grasp the correct distribution over whatever over sequences that come next.
The three verbs piled up ("grab and capture and grasp") are him reaching for how total the claim is. Next token prediction done well enough does not approximate the distribution over sequences, it is the distribution over sequences.
His historical claim about it is carefully bounded. Not the first:
And this was a relatively new thing. It wasn't literally the first ever auto-regressive neural network, but I would argue it was the first auto-regressive neural network where we really believed that if you train it really well, then you will get whatever you want.
The belief was the novelty, not the mechanism. And the target they aimed the belief at was machine translation: "In our case, back then was the humble, today humble, then incredibly audacious task of translation." The reversal of adjectives in that sentence is the ten years in miniature. Translation went from audacious to humble while the method stayed the same.
The paper's own numbers, which the talk does not restate. Worth having beside the slide audit, from the published paper rather than the talk: on the WMT 2014 English to French task the LSTM scored BLEU 34.8 on the full test set, against 33.3 for the phrase based statistical system it was compared to, and rescoring that system's 1,000 best hypotheses with the LSTM pushed the score to 36.5, close to the then best result on the task. The paper also reports the trick that made it train: reversing the word order of the source sentence but not the target, which shortens the distance between corresponding words and made the optimization easier.
Slide three: ancient history, which is to say the LSTM (4:22)
Now I'm going to show you some ancient history that many of you might have never seen before. It's called the LSTM.
The LSTM, invented by Sepp Hochreiter and Jürgen Schmidhuber in 1997, gets the single best line of the talk as its introduction:
To those unfamiliar, an LSTM is the things that poor deep learning researchers did before transformers.
And then the technical reading of the diagram, which is genuinely useful and not just a joke:
And it's basically a ResNet but rotated in 90 degrees. So that's an LSTM. And it came before it, it's like kind of like a slightly more complicated ResNet. You can see there is your integrator, which is now called the residual stream, but you've got some multiplication going on. It's a little bit more complicated. But that's what we did. It was a ResNet rotated in 90 degrees.
Note the direction of the comparison. The LSTM came before ResNet, by eighteen years, and he is explaining the older thing in terms of the newer one because that is the vocabulary the room has now. The cell state is the residual stream rotated from the depth axis onto the time axis; the extra complication is the multiplicative gating that a residual block does not have.
Slide four: pipelining, one layer per GPU, and the only thing he says was wrong (4:44)
Another cool feature from that old talk that I want to highlight is that we used parallelization, but not just any parallelization. We used pipelining as witnessed by this one layer per GPU.
Pipeline parallelism meant assigning each LSTM layer to its own GPU and streaming minibatches through the stack. It was an unusual thing to build in 2014 and the slide was clearly shown with pride at the time. The 2024 verdict is three sentences long and is the only place in the talk where he says a decision was simply wrong:
Was it wise to pipeline? As we now know, pipelining is not wise. But we were not as wise back then. So we used that and we got a 3.5x speed up using eight GPUs.
Hold on to that number, because it is the one piece of hard engineering data in the talk and it is quietly damning. Eight GPUs returned 3.5 times the speed. That is 44 percent efficiency, and more than half the machine was idle at any moment, which is exactly the pipeline bubble problem. The honesty here matters: this is an award talk where the speaker volunteers that one quarter of his own highlighted slides was a mistake.
Slide five: the conclusion slide, where the scaling hypothesis got written down (5:35)
And the conclusion slide in some sense, the conclusion slide from the talk from back then is the most important slide because it spelled out what could arguably be the beginning of the scaling hypothesis.
What it said: if you have a very big data set and you train a very very big neural network, then success is guaranteed.
His assessment, with the hedge intact:
And one can argue, if one is charitable, that this indeed has been what's been happening.
"If one is charitable" is him declining to claim the decade as a vindication. But the structure of his retrospective makes the claim anyway. The field spent ten years executing that sentence.
| The 2014 slide | What it claimed | His 2024 verdict |
|---|---|---|
| The deep learning hypothesis | A ten layer neural network can do anything a human can do in a fraction of a second, because artificial neurons resemble slow biological ones. | HELD "Not too bad." The depth limit lifted, the reasoning did not need to. |
| Our main idea | An autoregressive model that predicts the next token well enough captures the true distribution over sequences. | HELD Not the first autoregressive net, but the first where "we really believed that if you train it really well, then you will get whatever you want." |
| The LSTM | A multilayered LSTM encoder and decoder, with a cell state carrying information across time. | REPLACED "The things that poor deep learning researchers did before transformers." A ResNet rotated 90 degrees, with gating. |
| Parallelization | Pipelining, one layer per GPU, for a 3.5x speedup on eight GPUs. | WRONG "As we now know, pipelining is not wise. But we were not as wise back then." |
| The conclusion slide | A very big data set plus a very very big neural network means success is guaranteed. | HELD, AND RAN THE DECADE "The most important slide ... arguably the beginning of the scaling hypothesis." |
| Connectionism (not a slide) | Intelligence emerges from many simple neuron like units connected and trained, not from designed symbolic structure. | THE ONE THAT TRULY HELD "The idea that truly stood the test of time ... the core idea of deep learning itself." |
The idea that truly stood the test of time (5:47)
Having audited the slides, he names the thing underneath them, and it is not an architecture:
I want to mention one other idea. And this is, I claim, the idea that truly stood the test of time. It's the core idea of deep learning itself. It's the idea of connectionism. It's the idea that if you allow yourself to believe that an artificial neuron is kind of sort of a like a biological neuron. Right? If you believe that one is kind of sort of like the other, then it gives you the confidence to believe that very large neural networks, they don't need to be literally human brain scale, they might be a little bit smaller, but you could configure them to do pretty much all the things that we do, human beings.
Note what connectionism supplies in his telling. Not a proof, not a method. Confidence. It is the permission structure that let a small number of people keep scaling something while the surrounding field thought the premise was silly.
And then, immediately, the caveat he almost forgot to include, delivered live:
There's still a difference. Oh, I forgot to put that in. There's still a difference because the human brain also figures out how to reconfigure itself. Whereas we are using the best learning algorithms that we have, which require as many data points as there are parameters. Human beings are still better in this regard.
That is a precise and load bearing admission, and it is easy to skate past because of how casually it arrives. "As many data points as there are parameters" is the sample efficiency of current training stated as a rule of thumb, and it is the exact constraint that the second half of the talk is about. The human brain reconfigures itself; gradient descent needs roughly a datapoint per weight. When he later lists "understand things from limited data" as a property of future systems, this is the deficiency he is pointing back at.
The age of pretraining, and the people he credits by name (7:12)
But all this led, so I claim, arguably, to the age of pre-training. And the age of pre-training is what we might say the GPT-2 model, the GPT-3 model, the scaling laws.
Three hedges in one sentence ("so I claim, arguably"), and then a specific and unhedged credit:
And I want to specifically call out my former collaborators, Alec Radford, also Jared Kaplan, Dario Amodei, for really making this work.
Those three names map exactly onto the artifacts he just listed. Alec Radford led GPT-2 and the GPT line. Jared Kaplan is first author on Scaling Laws for Neural Language Models. Dario Amodei is on that paper and ran the research organization that shipped GPT-3. All three had left OpenAI by the time of this talk, two of them to found Anthropic, which is worth noticing in a talk delivered six months after Sutskever himself left.
But that led to the age of pre-training. And this is what's been the driver of all of progress, all the progress that we see today. Extra-large neural networks. Extraordinarily large neural networks trained on huge data sets.
He says "all of progress" twice. There is no qualification on that one. Everything you have seen, he says, came from this.
"Pre-training as we know it will unquestionably end" (7:52)
The hinge of the talk arrives without a transition. He has just finished saying that pretraining drove all progress, and the next sentence is:
But pre-training as we know it will unquestionably end. Pre-training will end.
He says it twice, the second time stripped down. Then the argument, which is two lines of arithmetic:
Why will it end? Because while compute is growing through better hardware, better algorithms, and larger clusters. Right, all those things keep increasing your compute. All these things keep increasing your compute. The data is not growing because we have but one internet. We have but one internet.
(The caption track renders "compute is growing" as "computer is growing" at this point. He is talking about compute.)
The structure is a multiplication where one factor has a ceiling. Pretraining takes two inputs. Compute has three independent growth channels, each of which he names: better hardware, better algorithms, larger clusters. Data has none, because the supply is a single artifact that already exists. Repeat "we have but one internet" and the argument is finished; nothing else is needed.
Then the line that left the building and went everywhere:
You could even say, you can even go as far as to say that data is the fossil fuel of AI. It was like created somehow, and now we use it, and we've achieved peak data, and there'll be no more. We have to deal with the data that we have.
Three distinct claims are packed into that metaphor and they are worth separating, because the metaphor got quoted far more often than it got read. The data was laid down once by a process nobody is running again. It is consumed rather than reused. And the total is now in sight, which is what "peak data" means, borrowed straight from peak oil.
Watch the hedging around it. He does not assert the fossil fuel line, he works up to it in stages: "you could even say, you can even go as far as to say." He is flagging it as a figure of speech he is choosing to deploy, not as a finding. And he immediately limits the consequence:
Now, it's still still let us go quite far, but this is there's only one internet.
That sentence is the one the headlines dropped. The existing data still has room in it. The claim is about the horizon of a recipe, not a wall arriving on Tuesday.
What comes next, carefully credited to other people (8:57)
The transition into the forward looking half is a refusal to own it:
So, here I'll take a bit of liberty to speculate about what comes next. Actually, I don't need to speculate because many people are speculating, too, and I'll mention their speculations.
This matters for reading the rest of the talk honestly. The three candidates below are presented as the field's speculations that he is relaying, not as his roadmap. He gives each of them one or two sentences and declines to rank them.
Agents. Flagged as a word people are using rather than a plan anyone has:
You may have heard the phrase agents. It's common, and I'm sure that eventually something will happen, but people feel like something agents is the future.
"I'm sure that eventually something will happen" is about as thin as endorsement gets.
Synthetic data. Named as an open problem, not a solution:
More concretely, but also a little bit vaguely, synthetic data. But what does synthetic data mean? Figuring this out is a big challenge, and I'm sure that different people have all kinds of interesting progress there.
The construction is pointed. "More concretely, but also a little bit vaguely" is him saying the phrase sounds like a technique and is actually a question.
Inference time compute. The freshest of the three at the time, and the one he attaches to a shipped artifact:
And then inference time compute, or maybe what's been most recently most vividly seen in o1, the o1 model.
OpenAI o1 had been released three months before this talk, in September 2024. It is the only specific model anywhere in the forward looking section, and his summary of all three candidates is a shrug and a blessing:
These are all examples of things of people trying to figure out what to do after pre-training. And those are all very good things to do.
The biology slide, which is the real argument of the talk (9:52)
The three candidates above answer "what should people try next." They do not answer the harder objection sitting underneath the whole claim: if the one recipe we know how to scale runs out, why believe another recipe is out there to be found at all. The answer he gives is not from computer science, and the way he tells the story is most of its charm.
I want to mention one other example from biology, which I think is really cool. And the example is this. So, about many, many years ago at this conference also, I saw a talk where someone presented this graph. But the graph showed the relationship between the size of the body of a mammal and the size of their brain. In this case, it's in mass.
Then the thing he remembered about that old talk, which is the setup for the punchline:
And that talk, I remember vividly, they were saying, "Look, in biology everything is so messy, but here you have one rare example where there's a very tight relationship between the size of the body of the animal and their brain."
A tight relationship in a field with no tight relationships. That is why the graph stuck. And then the research method, stated with no embellishment at all:
And totally randomly, I became curious at this graph. And one of the early So, I went to Google to do research to look for this graph. And one of the images in Google images was this.
He went to Google Images. That is the provenance of the evidence in the most quoted AI talk of 2024, and he says so plainly rather than dressing it up as a literature review. There is then a live moment where the clicker misbehaves: "And the interesting thing in this image is you see like I don't know, is the mouse working? Oh, yeah, the mouse is working, right."
What the image shows, in his order: all the different mammals, then non-human primates, which are "basically the same thing," and then a third group.
But then you've got the hominids. And to my knowledge, hominids are like close relatives to the humans in evolution. Like the Neanderthals. There's a bunch of them. Ho- like it's called Homo habilis, maybe. There's a whole bunch. And they're all here.
Hominids including Neanderthals and, hedged with a "maybe," Homo habilis. He is not pretending to expertise he does not have, which is exactly why the next sentence lands:
And what's interesting is that they have a different slope on their brain-to-body scaling exponent. So, that's pretty cool. What that means is that there is a precedent, there is an example of biology figuring out some kind of different scaling. Something clearly is different. So, I think that is cool.
"A different slope on their brain-to-body scaling exponent" is a statement about allometry, the study of how biological quantities scale against body size. Everything else sits on one exponent. Hominids sit on another. Then, because he knows the room will ask, he reads the axes out loud:
And by the way, I want to highlight this x-axis is log scale. You see this is 100, this is 1,000, 10,000, 100,000, and likewise in grams. 1 g, 10 g, 100 g, 1,000 g.
The payoff is three sentences, and the middle one is the most important thing he says all day:
So, it is possible for things to be different. The things that we are doing, the things that we've been scaling so far is actually the first thing that we figured out how to scale. And without doubt, the field, everyone who's working here, will figure out what to do.
"The first thing that we figured out how to scale." Not the only one, not the best one. The first. The biology chart is there to make that sentence credible: evolution, with no researchers and no roadmap, stumbled onto a second exponent. A field of thousands of people, explicitly looking, should manage the same. That is why the chart earns the three minutes it gets, and it is also the strongest thing in the talk, because it turns a pessimistic claim about data into an optimistic claim about method without a word of hand waving.
A pause to say that what already happened is hard to convey (12:31)
Before the last section he stops and does something people in his position rarely bother with, which is to mark how strange the present is:
But, I want to talk here, I want to take a few minutes and speculate about the longer term, the longer term. Where are we all headed? Right? We're making all this progress. It's astounding progress. It's really I mean, those of you who've been in the field 10 years ago and you remember just how incapable everything has been, like yes, you can say even if you kind of say of course deep learning still to see it is just unbelievable. It's completely I can't convey that feeling to you.
The sentence falls apart as he says it, which is the point. Then the clean version:
You know, if you've joined the field in the last 2 years, then of course, you speak to computers and they talk back to you and they disagree and that's what computers are. But, it hasn't always been the case.
"They disagree" is the detail he picks. Not that they answer, that they push back. For anyone who joined after ChatGPT, that is simply what a computer is. For him it is the thing that cannot be conveyed.
Superintelligence, and the four properties (13:25)
But, I want to talk to you a little bit about superintelligence, just a bit. Because that is obviously where this field is headed. This is obviously what's being built here.
Two "obviously"s in two sentences, in a room full of researchers, most of whom would not describe their own job that way. He states it as a settled fact about the field rather than as his own position, which is a rhetorical choice worth noticing. His stated goal is modest and specific:
And the thing about superintelligence is that it will be different qualitatively from what we have. And my goal in the next minute to try to give you some concrete intuition of how it will be different. So that you yourself could reason about it.
He starts from the gap in today's systems rather than from the future ones:
So right now we have our incredible language models and their unbelievable chatbots and they can even do things, but they're also kind of strangely unreliable and they get confused when a while also having dramatically superhuman performance on evals. So it's really unclear how to reconcile this.
That is the honest statement of the 2024 situation, and the word "unclear" is him declining to resolve it. Superhuman on benchmarks, strangely unreliable in use, and no theory that explains both at once. Then the four properties, each with its hedge attached.
It will be agentic, really. With an immediate walk back of his own phrasing, live:
Those systems are actually going to be agentic in a real ways. Whereas right now the systems are not agents in any meaningful sense. Just very That might be too strong. They're very very slightly agentic. Just the beginning.
It will reason, and that is why it will be unpredictable. This is the most substantive idea in the second half and it gets less attention than the fossil fuel line:
It will actually reason. And by the way, I want to mention something about reasoning is that a system that reasons, the more it reasons, the more unpredictable it becomes. The more it reasons, the more unpredictable it becomes.
He says it twice, exactly, which is his tell for the sentences he wants carried out of the room. The argument for it runs back through the whole talk, all the way to slide one:
All the deep learning that we've been used to is very predictable because if you've been working on replicating human intuition essentially. It's like the gut feel. If you come back to the 0.1 second reaction time what kind of processing we do in our brains well it's our intuition. So we've endowed our AIs with some of that intuition.
That is the deep learning hypothesis closing its own loop. The thing a ten layer network could capture was the fraction of a second, which is intuition, which is by nature predictable, because it is the output of a shallow and stereotyped computation. Reasoning is not that. Reasoning is the part that takes longer than a tenth of a second, and its entire value is arriving somewhere you could not have guessed.
The evidence is one sentence, and it is unanswerable:
But reasoning, and you're seeing some early signs of that, reasoning is unpredictable and one reason to see that is because the chess AIs, the really good ones, are unpredictable to the best human chess players. So we will have to be dealing with AI systems that are incredibly unpredictable.
Chess engines are superhuman and narrow, and the best humans in the world still cannot anticipate their moves. There is no alignment trick that recovers predictability without giving up the reasoning, because they are the same property viewed from two sides.
It will understand things from limited data, and it will not get confused. Two sentences, both pointing straight back at the caveat he almost forgot in the connectionism section: "They will understand things from limited data. They will not get confused. All the things which are really big limitations."
And then the disclaimer that the headlines ignored:
I'm not saying how, by the way, and I'm not saying when. I'm saying that it will.
It will be self aware, which he treats as unremarkable. No mysticism offered, and none permitted:
And when all those things will happen together with self-awareness, because why not? Self-awareness is useful. It is part your ourselves are parts of our own world models.
Self awareness as an engineering consequence: a model of the world that contains the modeler is a better model of the world, so a system optimizing its world model will build one. That is the entire argument, delivered in under ten seconds, in a throwaway tone.
| Property | What he says about systems today | What he says about what is coming | The hedge he attaches, verbatim |
|---|---|---|---|
| Agency | "not agents in any meaningful sense", then corrected live to "very very slightly agentic. Just the beginning." | "actually going to be agentic in a real ways" | "That might be too strong." He interrupts himself to soften his own claim about today. |
| Reasoning | Deep learning so far replicates intuition, the gut feel, the 0.1 second reaction. Predictable by construction. | "It will actually reason." And therefore: "the more it reasons, the more unpredictable it becomes," said twice. | "you're seeing some early signs of that". The evidence offered is chess engines, which are already unpredictable to the best human players. |
| Sample efficiency | Current learning algorithms "require as many data points as there are parameters. Human beings are still better in this regard." | "They will understand things from limited data." | "I'm not saying how, by the way, and I'm not saying when. I'm saying that it will." |
| Reliability | "strangely unreliable and they get confused" while "also having dramatically superhuman performance on evals" | "They will not get confused." | "It's really unclear how to reconcile this." He states the paradox and does not resolve it. |
| Self awareness | Not discussed as present. | "because why not? Self-awareness is useful. It is part ... our selves are parts of our own world models." | Offered as a consequence rather than a milestone, in one sentence, with no emphasis at all. |
He finishes the sketch and then refuses to carry it any further:
When all those things come together, we will have systems of radically different qualities and properties that exist today. And of course, they will have incredible and amazing capabilities. But the kind of issues that come up with systems like this, and I'll just leave it as an exercise to to imagine.
Leaving it as an exercise to imagine is a deliberate non answer from the person who left OpenAI in May 2024 and founded Safe Superintelligence a month later, a company whose entire stated reason for existing is the issues he just declined to enumerate. He does not mention SSI once in the talk.
The closing line, delivered with a shrug audible in the recording:
And I would say that it's definitely also impossible to predict the future. Really, all kinds of stuff is possible. But on this uplifting note, I will conclude. Thank you so much.
The Q&A (16:31)
Four questions in eight minutes, and they produced several of the lines that got clipped and shared more than anything in the prepared talk. Worth having in full, because in each case the question shapes the answer.
Question one, at 16:44: are there other biological structures worth copying?
The asker wants to know whether, in 2024, there are parts of human cognition beyond the neuron that are worth exploring the way the neuron was.
His answer does two things. First, it hands the question back to anyone who has a real insight:
So, the way I'd answer this question is that if you are or someone is a person who has a specific insight about, "Hey, we are all being extremely silly because clearly the brain does something and we are not. And that's something that can be done," they should pursue it.
Then the audit of how much biology actually got used, which is the quotable part and a direct callback to his own connectionism section:
Like there's been a lot of desire to make biologically inspired AI. And you could argue on some level that biologically inspired AI is incredibly successful, which is all of the learning is biologically inspired AI. But on the other hand, the biological inspiration was very, very, very modest. It's like, let's use neurons. This is the full extent of the biological inspiration. Let's use neurons.
The whole of biologically inspired AI, in his accounting, is the decision to call the unit a neuron. Everything else was engineering. And the closing hedge is characteristically open: "And more detailed biological inspiration has been very hard to come by. But I wouldn't rule it out. I think if someone has a special insight, they might be able to see something and that would be useful."
Question two, at 17:45: will reasoning let a model auto correct its own hallucinations?
The longest question of the session, and the asker builds it carefully. The observation is about method: the way the field detects hallucination today is statistical, "some amount of standard deviations or whatever away from the mean," precisely because the models cannot be trusted to reason about their own output. The question is whether reasoning changes that:
Do you think that a model given reasoning will be able to correct itself, sort of auto correct itself? And that will be a core feature of future models so that there won't be as many hallucinations because the model will recognize when ... the model will be able to reason and understand when a hallucination is occurring. Does the question make sense?
The answer is the most direct one he gives all session:
Yes, and the answer is also yes. I think what you described is extremely highly plausible. Yeah, I mean you should check. I mean for yeah, it I wouldn't I wouldn't rule out that it might already be happening with some of the, you know, early reasoning models of today. I don't know. But longer term, why not?
"It might already be happening" in December 2024, three months after o1 shipped, is the most concretely forward looking thing in the whole talk, and he buries it in a hedge.
The exchange then turns on a word. The asker reaches for an analogy: "part of like Microsoft Word like autocorrect. It's a core feature." Sutskever will not take it:
Yeah, I just I mean I think calling it autocorrect is really doing a disservice. I think you are When you say autocorrect, you evoke like it's far grander than autocorrect, but other than but, you know, this point aside, the answer is yes.
That is him defending the thing he just agreed to from the framing that would make it sound small. A model that notices its own error is not spell check with extra steps.
Question three, at 19:48: do they need rights, and what incentive structure gets us there?
The asker opens by naming what the talk left out: "I loved the ending mysteriously leaving out do they replace us or are they, you know, superior? Do they need rights? You know, it's a new species of Homo sapien spawned intelligence." Then the actual question, which is about mechanism design rather than ethics: "How do you create the right incentive mechanisms for humanity to actually create it in a way that gives it the freedoms that we have as Homo sapiens?"
He endorses the question and declines the answer, in that order:
You know, I feel like this In some In some sense, those are those are the kind of questions that people should be reflecting on more. But to your question about what incentive structure should we create, I I don't feel that I know. I don't feel confident answering questions like this because uh it's like you're talking about creating some kind of a top-down structure government thing, I don't know.
The asker offers a specific mechanism: "It could be a cryptocurrency, too. There's Bittensor, you know, there's things." (Bittensor is a decentralized network that pays token rewards for contributed machine learning work. The caption track renders it as "Bit Tensor.") The response became one of the most circulated clips of the conference, mostly for its delivery:
I don't feel like I am the right person to comment on cryptocurrency.
And then, having declined, he answers anyway, and this is the part that got quoted far less often than the crypto line:
But you know, there is a chance, by the way, what you're describing will happen. That indeed we will have, you know, in some sense it's it's it's not a bad end result if you have AIs and all they want is to coexist with us and also just to have rights, maybe that will be fine. It's But I don't know. I mean, I think things are so incredibly unpredictable. I I hesitate to comment, but I encourage the speculation.
"Not a bad end result" is a load bearing phrase from someone whose company is named after the problem. And "I hesitate to comment, but I encourage the speculation" is the posture of the entire second half of the talk, stated outright.
Question four, at 21:49: do LLMs generalize multi hop reasoning out of distribution?
Asked by Shalev Lifshitz of the University of Toronto, who notes he works with Sheila McIlraith. (The caption track spells the name "Shalev Lifschitz" and renders "Sheila" without a surname.)
Sutskever rejects the shape of the question before answering it:
So, okay. The question assumes that the answer is yes or no, but the question should not be answered with a yes or no. Because what does it mean out of distribution generalization? What does it mean? What does it mean in distribution? And what does it mean out of distribution?
Then, because it is a test of time talk, he takes the opportunity to date the vocabulary:
Because it's a test of time talk, I'll say that long, long ago, before people were using deep learning, they were using things like string matching and n-grams for machine translation. People were using statistical phrase tables. Can you imagine? They had tens of thousands of lines of code of complexity which was I mean it's it was truly unfathomable.
(Statistical phrase tables and n-grams are the machinery the 2014 paper displaced, which makes the callback exact. The caption drops "lines of" from "tens of thousands of lines of code.")
And the argument lands:
And back then generalization meant is it literally not in this is the same word phrasing as in the data set. Now we may say, "Well, sure my model achieves this high score on I don't know math competitions, but maybe some discussion in some forum on the internet was about the same ideas and therefore it's memorized." Well, okay, you could say maybe it's in distribution, maybe it's memorization. But I also think that our standards for what counts as generalization have increased really quite substantially, dramatically, unimaginably if you keep track.
This is the best thing in the Q&A. The word generalization has not held still. In 2014 it meant "the exact string is not in the training set." In 2024 it means "a forum thread somewhere on the internet did not discuss a similar idea." Those are not the same test, they are not close to the same test, and a decade of progress is hiding inside the drift. Nobody announced the change, so nobody credits it.
His actual answer, once the terms are cleaned up, is a careful three part split:
And so I think the answer is to some degree probably not as well as human beings. I think it is true that human beings generalize much better. But at the same time they definitely generalize out of distribution to some degree. I hope it's a useful tautological answer.
Models generalize out of distribution. They do it worse than people. And the measuring stick keeps moving, which is why the yes or no version of the question has no answer.
The session ends there, cut short by the clock:
And unfortunately, we're out of time for this session. I have a feeling we could go on for the next 6 hours. But thank you so much Ilya for the talk.
The decade, using only the landmarks he names
Every entry below is something Sutskever points at in these twenty four minutes. Nothing has been added from the wider history, which is why there is no AlexNet and no Chinchilla here: he does not mention them, and the shape of his argument is clearer without them.
- 1997The LSTM. Sepp Hochreiter and Jürgen Schmidhuber publish Long Short-Term Memory. He will later describe it as "a ResNet but rotated in 90 degrees," explaining the older architecture with the newer one's vocabulary.
- Dec 2014NeurIPS in Montreal, "a much more innocent time." He presents Sequence to Sequence Learning with Neural Networks with Oriol Vinyals and Quoc Le. Three bullets: an autoregressive model trained on text, a large neural network, a large data set. Pipelining at one layer per GPU, 3.5x on eight GPUs. The conclusion slide writes down the scaling hypothesis.
- 2017Transformers arrive. He never names the paper, only the era boundary: the LSTM is "the things that poor deep learning researchers did before transformers." Architecture changed, the three bullets did not.
- 2019 to 2020The age of pretraining. GPT-2, GPT-3, and the scaling laws. He names Alec Radford, Jared Kaplan and Dario Amodei "for really making this work," and calls this "the driver of all of progress."
- Sep 2024o1. The only specific model in the forward looking half, cited as where inference time compute is "most recently most vividly seen."
- 13 Dec 2024This talk. The Test of Time award in Vancouver. Compute still grows on three channels, data grows on none, so pretraining as we know it will unquestionably end. The hominid chart is offered as proof that a second scaling regime can be found at all.
Key takeaways
- The whole 2014 paper was three bullet points, and all three survived. An autoregressive model trained on text, a large neural network, a large data set. The transformer replaced the LSTM without disturbing any of them.
- The deep learning hypothesis is a construction argument, not an empirical one. If artificial neurons resemble slow biological ones, then anything one human anywhere does in a fraction of a second is shallow, so a ten layer net can represent it: take that person's connections and embed them. Whether training finds those weights is a separate question, and the next slide is the bet that it does.
- Ten layers was a limitation, not a theory. "This was the neural networks we knew how to train back in the day."
- One decision gets graded as simply wrong, and it is the engineering one. Pipelining, one layer per GPU, 3.5x on eight GPUs. "As we now know, pipelining is not wise. But we were not as wise back then."
- The idea he says truly stood the test of time is connectionism, not any architecture. Its contribution was confidence: permission to believe a large enough network could do what people do.
- The caveat he almost forgot is the one the second half depends on. The brain reconfigures itself; current learning algorithms "require as many data points as there are parameters. Human beings are still better in this regard."
- Pretraining ends because of an asymmetry, not a plateau. Compute grows through better hardware, better algorithms and larger clusters. Data does not grow, because there is one internet. One factor of a product has a ceiling, so the product does.
- "Data is the fossil fuel of AI." Laid down once, consumed, peak already reached. He builds up to the phrase rather than asserting it, and immediately adds that the existing data "still let us go quite far."
- He does not claim the successor recipe. Agents, synthetic data and inference time compute are explicitly relayed as other people's speculations. Synthetic data in particular is framed as a question, not a technique.
- Hominid brain to body scaling is an existence proof, not an analogy about brains. Biology found a second exponent with nobody designing it. "The things that we've been scaling so far is actually the first thing that we figured out how to scale."
- Reasoning and predictability are the same property seen from two sides. Deep learning so far replicated intuition, which is predictable because it is shallow. Real reasoning is not, and the proof is that chess engines are unpredictable to the best human players.
- On superintelligence he refuses the two questions everyone wants answered. "I'm not saying how, by the way, and I'm not saying when. I'm saying that it will."
- Self awareness is offered as an engineering consequence. "Because why not? Self-awareness is useful. It is part ... our selves are parts of our own world models."
- The strongest idea in the Q&A is about the vocabulary, not the models. The standard for generalization has moved from "the exact phrasing is not in the data set" to "no forum thread discussed a similar idea," and the drift happened without anyone announcing it.
Chapters
The seven entries in bold are the video's own chapter markers, reproduced verbatim. Seven markers across twenty four minutes leaves most of the talk unlabelled, so everything else is a sub beat added here from the transcript's own clock. Two of his markers land slightly after the content they name, so the list below runs in strict time order rather than grouping sub beats under the marker they belong to.
- 0:00 Introduction and context
- 0:01 Thanks to the organizers, and to Oriol Vinyals and Quoc Le
- 0:35 The screenshot from the 2014 NeurIPS talk in Montreal, "a much more innocent time"
- 1:05 The frame: a ten year retrospective, and "some not so much"
- 1:23 Retrospective of 2014 work
- 1:37 The whole paper in three bullet points, and "that's it"
- 1:45 Slide one, the deep learning hypothesis: ten layers, a fraction of a second
- 2:37 Why a fraction of a second, and how the construction argument works
- 3:00 Why ten layers: it was all anyone could train
- 3:09 Slide two, "our main idea", and what the autoregressive claim actually says
- 3:41 Not the first autoregressive net, but the first that was believed
- 4:12 Translation, "today humble, then incredibly audacious"
- 4:22 Slide three, ancient history: the LSTM, "a ResNet but rotated in 90 degrees"
- 4:44 Slide four, parallelization by pipelining, one layer per GPU
- 5:17 The verdict: "pipelining is not wise", and the 3.5x on eight GPUs
- 5:32 The scaling hypothesis
- 5:35 Slide five, the conclusion slide, where the scaling hypothesis got written down
- 5:47 Connectionism, the idea he says truly stood the test of time
- 6:50 The caveat he forgot: the brain reconfigures itself, and sample efficiency
- 7:12 The age of pretraining: GPT-2, GPT-3, the scaling laws
- 7:30 Naming Alec Radford, Jared Kaplan and Dario Amodei
- 7:52 "Pre-training as we know it will unquestionably end"
- 8:03 The future of pre-training
- 8:05 Why: compute grows on three channels, data grows on none, "we have but one internet"
- 8:24 "Data is the fossil fuel of AI", and peak data
- 8:57 Speculation, credited to other people rather than claimed
- 9:08 Agents
- 9:20 Synthetic data, framed as a question rather than a technique
- 9:38 Inference time compute, and o1
- 9:52 Biology as inspiration
- 9:57 The talk he saw years ago at this same conference
- 10:28 "In biology everything is so messy, but here you have one rare example"
- 10:45 Going to Google Images to find the graph again
- 10:59 Mammals, non-human primates, and then the hominids
- 11:29 The different slope on the brain-to-body scaling exponent
- 12:01 Reading the axes: log scale, 100 to 100,000 grams, 1 to 1,000 grams
- 12:18 "The first thing that we figured out how to scale"
- 12:31 Superintelligence outlook
- 12:45 A pause: how hard it is to convey what already happened
- 13:03 "You speak to computers and they talk back to you and they disagree"
- 13:25 Superintelligence, "obviously where this field is headed"
- 13:50 Today's gap: superhuman on evals, strangely unreliable in use
- 14:12 Property one, agentic in a real way, immediately walked back
- 14:30 Property two, reasoning, and why reasoning means unpredictability
- 15:10 Chess engines, unpredictable to the best human players
- 15:25 Properties three and four: limited data, and not getting confused
- 15:33 "I'm not saying how, and I'm not saying when. I'm saying that it will"
- 15:37 Self awareness, "because why not?"
- 16:07 Leaving the consequences as an exercise, and the close
- 16:31 Q&A session
- 16:44 Q1: other biological structures worth exploring, and "let's use neurons"
- 18:10 Q2: will reasoning let a model auto correct its own hallucinations
- 19:18 "Yes, and the answer is also yes"
- 19:45 The Microsoft Word autocorrect analogy, and why he rejects it
- 20:00 Q3: do they need rights, and what incentive structure creates them
- 21:30 "I don't feel like I am the right person to comment on cryptocurrency"
- 22:05 Q4: Shalev Lifshitz on multi hop reasoning out of distribution
- 22:25 Taking apart the words "in distribution" and "out of distribution"
- 22:49 String matching, n-grams and statistical phrase tables, "can you imagine?"
- 23:30 How the standard for generalization moved without anyone announcing it
- 23:50 The answer: worse than humans, but genuinely out of distribution to some degree
- 24:20 Out of time, "I have a feeling we could go on for the next 6 hours"
Notable quotes
Because a lot of the things in this work were correct, but some not so much. And we can review them and we can see what happened and how it gently flowed to where we are today. Ilya Sutskever, setting the frame for the slide audit, 1:05
It's an autoregressive model trained on text. It's a large neural network and it's a large data set. And that's it. Ilya Sutskever, the entire 2014 paper in three bullet points, 1:37
If there is one human in the entire world that can do some task in a fraction of a second, then a 10-layer neural network can do it, too, right? It follows. You just take their connections and you embed them inside your neural net, the artificial one. Ilya Sutskever, the deep learning hypothesis as a construction argument, 2:37
To those unfamiliar, an LSTM is the things that poor deep learning researchers did before transformers. Ilya Sutskever, introducing the architecture the award winning paper was built on, 4:22
Was it wise to pipeline? As we now know, pipelining is not wise. But we were not as wise back then. So we used that and we got a 3.5x speed up using eight GPUs. Ilya Sutskever, the only decision in the paper he grades as wrong, 5:17
There's still a difference because the human brain also figures out how to reconfigure itself. Whereas we are using the best learning algorithms that we have, which require as many data points as there are parameters. Human beings are still better in this regard. Ilya Sutskever, the caveat he says out loud he forgot to put on the slide, 6:50
But pre-training as we know it will unquestionably end. Pre-training will end. Ilya Sutskever, the line the talk is remembered for, 7:52
The data is not growing because we have but one internet. We have but one internet. Ilya Sutskever, the whole argument in two sentences, said twice, 8:05
You could even say, you can even go as far as to say that data is the fossil fuel of AI. It was like created somehow, and now we use it, and we've achieved peak data, and there'll be no more. Ilya Sutskever, with the hedges that usually get cut off the front, 8:24
What that means is that there is a precedent, there is an example of biology figuring out some kind of different scaling. Ilya Sutskever, on the hominid slope, 11:29
The things that we are doing, the things that we've been scaling so far is actually the first thing that we figured out how to scale. Ilya Sutskever, the payoff of the biology slide, 12:18
If you've joined the field in the last 2 years, then of course, you speak to computers and they talk back to you and they disagree and that's what computers are. But, it hasn't always been the case. Ilya Sutskever, on what cannot be conveyed to anyone who arrived after ChatGPT, 13:03
A system that reasons, the more it reasons, the more unpredictable it becomes. The more it reasons, the more unpredictable it becomes. Ilya Sutskever, said twice in a row, which is his tell, 14:30
Reasoning is unpredictable and one reason to see that is because the chess AIs, the really good ones, are unpredictable to the best human chess players. Ilya Sutskever, the evidence, 15:10
I'm not saying how, by the way, and I'm not saying when. I'm saying that it will. Ilya Sutskever, on every property of superintelligence he just listed, 15:33
And when all those things will happen together with self-awareness, because why not? Self-awareness is useful. It is part your ourselves are parts of our own world models. Ilya Sutskever, disposing of the hardest question in the field in one sentence, 15:37
But on the other hand, the biological inspiration was very, very, very modest. It's like, let's use neurons. This is the full extent of the biological inspiration. Let's use neurons. Ilya Sutskever, in the Q&A, auditing how much biology deep learning actually took, 17:50
I don't feel like I am the right person to comment on cryptocurrency. Ilya Sutskever, after a questioner proposed Bittensor as the incentive mechanism for AI rights, 21:35
I mean, I think things are so incredibly unpredictable. I I hesitate to comment, but I encourage the speculation. Ilya Sutskever, the posture of the entire second half, stated outright, 22:00
But I also think that our standards for what counts as generalization have increased really quite substantially, dramatically, unimaginably if you keep track. Ilya Sutskever, on why "out of distribution" no longer means what it meant in 2014, 23:30
I hope it's a useful tautological answer. Ilya Sutskever, closing the last question of the session, 23:50
How the predictions have held up
Everything above is the talk. This section is the only assessment on the page, written in October 2026, and it is a reader's scorecard rather than anything Sutskever said.
The claim people argued about is not the claim he made. "Pre-training as we know it will unquestionably end" was read almost everywhere as "scaling has hit a wall, now." He said the opposite twice in the same breath: the data "still let us go quite far," and the end is a horizon for one recipe, not a date. Frontier labs did not stop pretraining, and the models kept improving, so by the headline reading the prediction looks wrong and by his actual reading it has not yet come due. That gap is worth noticing on a talk this heavily quoted.
The ordering of his three candidates turned out backwards, in the best way for him. He listed agents first, synthetic data second, and inference time compute last and most tentatively, attached to a model three months old. Inference time compute is the one that became a permanent axis of the field. Reasoning models and reinforcement learning after pretraining went from a curiosity to a central line of frontier spending in the two years since, though the actual compute split is an estimate rather than a disclosure: no major lab publishes how its budget divides between pretraining and post training, so every number in circulation on this is inferred.
Agents went from the word he was least willing to endorse to the dominant product category. "I'm sure that eventually something will happen" has aged into an understatement. Whether that validates the prediction or just the vocabulary is a fair argument.
Synthetic data is the one he framed most accurately. He refused to call it a solution and called the definition itself the hard part. Two years on it is used heavily in post training and is still argued about rather than settled, which is exactly the status he gave it.
The fossil fuel metaphor took the most criticism, and some of it lands. Critics have pointed out that training data is not a fixed geological deposit: people keep producing text, video and interaction logs, and data quality and filtering have mattered more than raw volume. His own framing makes room for this (he says peak data, not zero data) but the metaphor travelled further than the framing did, which is a hazard of coining a good line.
The idea with the longest shelf life is the one that got the least coverage. Reasoning implies unpredictability, and the two cannot be separated because they are the same property. That observation was made in passing in December 2024, before the reasoning model era properly began, and it has only become more awkward as the field has pushed test time reasoning harder while also asking for more predictable behaviour from the same systems. The chess engine example remains the cleanest statement of the tension anyone has offered.
On his own position. He never mentions Safe Superintelligence in the talk, which he co-founded in June 2024, a month after leaving OpenAI. It raised at a reported 30 billion dollar valuation in 2025, he became its chief executive when co-founder Daniel Gross left, Nvidia announced a partnership and investment in July 2026, and it has still shipped nothing. A lab organized around not shipping is a strange thing until you read it against this talk, where the speaker argues that the recipe everyone is shipping on has a horizon and that the next one has not been found yet.
Where this sits in the LLM Learning track
This opens Part 4, the final stage of the track, the part about where all of it goes. Everything earlier describes a recipe that works. This is the person who wrote the recipe's opening chapter saying out loud that its main ingredient is finite.
Read it against the Chinchilla and scaling section of Sasha Rush's Five Formulas, which quantifies the same data constraint from the other direction and gives you the arithmetic behind "one internet." Read the reasoning and unpredictability argument against John Schulman on reinforcement learning from human feedback, because the lever Schulman describes for making a model behave is the same lever that gets harder as the model reasons more. And the pair that closes the track is deliberate: this talk sets the open research question, and Karpathy's Software Is Changing Again is what gets built while the question stays open. Read them back to back.
Resources mentioned
- Sequence to Sequence Learning with Neural Networks, the 2014 paper receiving the award, also in the NeurIPS 2014 proceedings
- Announcing the NeurIPS 2024 Test of Time Paper Awards, which explains why two awards were given that year, and NeurIPS 2024 itself
- Oriol Vinyals and Quoc Le, his co-authors, who presented immediately before him
- Long Short-Term Memory, the 1997 Hochreiter and Schmidhuber paper, and the LSTM overview
- Deep Residual Learning for Image Recognition, the ResNet paper he uses to explain the LSTM backwards, and the residual network overview
- Attention Is All You Need, the transformer paper he refers to only as the era boundary
- Connectionism, the tradition he says truly stood the test of time, and the artificial neuron it rests on
- Autoregressive models, the first of his three bullet points
- Pipeline parallelism, the technique he says was not wise
- Alec Radford, Jared Kaplan and Dario Amodei, the three people he credits for the age of pretraining
- GPT-2, GPT-3 and Scaling Laws for Neural Language Models, the artifacts of that age
- OpenAI o1, the only specific model named in the forward looking half
- Allometry and brain size, the subject of the biology slide, with hominids, Neanderthals and Homo habilis as the group that leaves the line
- Statistical machine translation and n-grams, the phrase table machinery the 2014 paper displaced, measured on WMT 2014 with BLEU
- Bittensor, proposed from the floor as an incentive mechanism, and declined
- Shalev Lifshitz of the University of Toronto, who asked the last question, working with Sheila McIlraith
- Safe Superintelligence, the company he founded in June 2024 and never mentions, and OpenAI, which he left a month before that
Caption corrections. The automatic caption track mangles four things, corrected silently in the quotes above and flagged here instead: "computer is growing" should be "compute is growing" (8:05); "tens of thousands of code of complexity" should be "tens of thousands of lines of code" (22:49); "Bit Tensor" is Bittensor (21:30); and the last questioner, captioned "Shalev Lifschitz," is Shalev Lifshitz. The captions also render "n-grams" as "grams" and o1 as both "O1" and "0.1".


