youtube.nixfred.com nixfred.com

Ilya Sutskever: Sequence to sequence learning with neural networks: what a decade

Accepting a NeurIPS Test of Time award for the 2014 sequence to sequence paper, Ilya Sutskever puts his own slides from ten years earlier back on screen and grades them one by one: the deep learning hypothesis held, the autoregressive bet held, connectionism is the idea he says truly stood the test of time, and pipelining was simply unwise. Then the line the talk is remembered for: pretraining as we know it will unquestionably end, because compute grows on three channels while data grows on none. Data is the fossil fuel of AI and there is but one internet. He relays agents, synthetic data and inference time compute as other people's speculations rather than his roadmap, and offers hominid brain to body scaling as an existence proof that a second scaling regime can be found at all. The forward half lists four properties of superintelligence with every hedge attached, including the argument that gets the least attention: a system that genuinely reasons is a system you cannot predict. The Q&A, included in full here, produced several of the most quoted moments of the conference.

Published Dec 14, 2024 24:37 video 57 min read Added Jul 30, 2026 Open on YouTube →

At a glance

Ilya Sutskever walked onto a NeurIPS stage in Vancouver in December 2024 to collect a Test of Time award for Sequence to Sequence Learning with Neural Networks, the 2014 paper he wrote with Oriol Vinyals and Quoc Le. He did not give a victory lap. He put the slides from his own 2014 NeurIPS talk back on the screen, one at a time, and graded them: "a lot of the things in this work were correct, but some not so much."

Then, nine minutes in, he said the sentence the talk is remembered for. "Pre-training as we know it will unquestionably end." The reasoning is arithmetic, not prophecy. Compute grows through better hardware, better algorithms and larger clusters. Data does not grow, "because we have but one internet." Data, he said, is the fossil fuel of AI: laid down once, burned once, and we have achieved peak data.

The rest is him being careful about what he does and does not know. He hands the "what comes next" question to other people's speculations (agents, synthetic data, inference time compute) rather than claiming it himself. He offers one piece of evidence that a different scaling regime is even findable, and it is not from computer science: a chart of brain mass against body mass in which hominids sit on a visibly different slope from every other mammal. He closes on superintelligence, lists four properties he expects it to have, refuses to say how or when, and makes the point that gets the least attention and deserves more: a system that genuinely reasons is a system you cannot predict.

Then a Q&A that produced some of the most quoted moments of the conference, including the answer where he says he does not feel like the right person to comment on cryptocurrency, and the answer where he dismantles the phrase "out of distribution" by pointing out that our standard for generalization has moved so far in ten years that the word barely survives the trip.

Twenty four and a half minutes. Almost no slides. Almost nothing wasted.

The deep explanation

The photograph, and the frame he sets (0:01)

He opens with thanks: to the organizers for choosing the paper, and to "my incredible co-authors and collaborators, Oriol Vinyals and Quoc Le, who stood right before you a moment ago." The award was shared that year, which is why the stage was already warm.

Then the first image: a screenshot from a talk he gave ten years earlier, at NeurIPS 2014 in Montreal. "It was a much more innocent time," he says, and shows the two photographs side by side. "Here we are shown in the photos. This is the before. Here's the after, by the way. And now we've got more experienced, hopefully wiser."

That sets the frame for the whole first half, and it is an unusual frame for an award talk. He is not going to summarize the paper. He is going to put the original slides back up and audit them.

Because a lot of the things in this work were correct, but some not so much. And we can review them and we can see what happened and how it gently flowed to where we are today.

"Gently flowed" is doing real work in that sentence. The claim is continuity: nothing in the last decade was a break with 2014, it was the execution of 2014.

The whole paper in three bullet points (1:37)

Before the slides, the summary, which he gives as three lines and then stops:

But the summary of what we did is the following three bullet points. It's an autoregressive model trained on text. It's a large neural network and it's a large data set. And that's it.

That is the entire architecture of the modern frontier model, stated in 2014 and restated here as the thing that did not need to change. Everything that follows is him checking each bullet against the record.

Slide one: the deep learning hypothesis (1:45)

The first original slide is titled "the deep learning hypothesis." His verdict on it: "Not too bad."

What it said: if you have a large neural network with ten layers, it can do anything a human being can do in a fraction of a second.

He then unpacks why the hypothesis is phrased around a fraction of a second, because that framing looks arbitrary until you follow the chain. It is a three step argument resting on connectionism:

Well, if you believe the deep learning dogma, so to say, that artificial neurons and biological neurons are similar, or at least not too different, and you believe that real neurons are slow, then anything that we can do quickly, by we I mean human beings, I even mean just one human in the entire world. If there is one human in the entire world that can do some task in a fraction of a second, then a 10-layer neural network can do it, too, right? It follows. You just take their connections and you embed them inside your neural net, the artificial one.

The construction step is the striking part. It is not an empirical claim that training will succeed. It is a claim that a solution exists in the weight space: take the human's connections, embed them in the artificial net, and by construction the network can represent the function. Trainability is a separate problem.

PREMISE 1 Artificial neurons are similar to biological ones, or at least not too different PREMISE 2 Real neurons are slow so little depth fits inside a tenth of a second THEREFORE Anything ONE human anywhere can do in a fraction of a second ... ... IS REPRESENTABLE Take their connections, embed them in a 10 layer neural net Why ten layers, and not more: "We focused on 10-layer neural networks because this was the neural networks we knew how to train back in the day. If you could go beyond in your layers somehow, then you could do more." What the chain does NOT claim: that training will find those weights. It claims only that a solution exists in the weight space, by construction. Optimization is a separate problem, and the next slide is the bet that it is solvable.
Figure 1. The deep learning hypothesis as he states it at 1:37, laid out as the chain of inferences it actually is. The oddly specific "fraction of a second" is not a hedge, it is the consequence of premise two: if biological neurons are slow, a tenth of a second buys you only a shallow stack, and a shallow artificial stack should be able to hold the same function. Ten layers was not a principled number, it was the deepest thing anyone could reliably train in 2014.

He is explicit that the depth was a limitation rather than a theory: "We focused on 10-layer neural networks because this was the neural networks we knew how to train back in the day. If you could go beyond in your layers somehow, then you could do more. But back then we could only do 10 layers, which is why we emphasized whatever human beings can do in a fraction of a second."

Slide two: "our main idea", and the autoregressive bet (3:09)

The second original slide is headed "our main idea." He invites the room to spot what is on it: "you may be able to recognize two things, or at least one thing. You might be able to recognize that something auto-regressive is going on here."

Then he says what the slide actually claims, which is stronger than it looks:

This slide says that if you have an auto-regressive model, and it predicts the next token well enough, then it will in fact grab and capture and grasp the correct distribution over whatever over sequences that come next.

The three verbs piled up ("grab and capture and grasp") are him reaching for how total the claim is. Next token prediction done well enough does not approximate the distribution over sequences, it is the distribution over sequences.

His historical claim about it is carefully bounded. Not the first:

And this was a relatively new thing. It wasn't literally the first ever auto-regressive neural network, but I would argue it was the first auto-regressive neural network where we really believed that if you train it really well, then you will get whatever you want.

The belief was the novelty, not the mechanism. And the target they aimed the belief at was machine translation: "In our case, back then was the humble, today humble, then incredibly audacious task of translation." The reversal of adjectives in that sentence is the ten years in miniature. Translation went from audacious to humble while the method stayed the same.

The paper's own numbers, which the talk does not restate. Worth having beside the slide audit, from the published paper rather than the talk: on the WMT 2014 English to French task the LSTM scored BLEU 34.8 on the full test set, against 33.3 for the phrase based statistical system it was compared to, and rescoring that system's 1,000 best hypotheses with the LSTM pushed the score to 36.5, close to the then best result on the task. The paper also reports the trick that made it train: reversing the word order of the source sentence but not the target, which shortens the distance between corresponding words and made the optimization easier.

Slide three: ancient history, which is to say the LSTM (4:22)

Now I'm going to show you some ancient history that many of you might have never seen before. It's called the LSTM.

The LSTM, invented by Sepp Hochreiter and Jürgen Schmidhuber in 1997, gets the single best line of the talk as its introduction:

To those unfamiliar, an LSTM is the things that poor deep learning researchers did before transformers.

And then the technical reading of the diagram, which is genuinely useful and not just a joke:

And it's basically a ResNet but rotated in 90 degrees. So that's an LSTM. And it came before it, it's like kind of like a slightly more complicated ResNet. You can see there is your integrator, which is now called the residual stream, but you've got some multiplication going on. It's a little bit more complicated. But that's what we did. It was a ResNet rotated in 90 degrees.

Note the direction of the comparison. The LSTM came before ResNet, by eighteen years, and he is explaining the older thing in terms of the newer one because that is the vocabulary the room has now. The cell state is the residual stream rotated from the depth axis onto the time axis; the extra complication is the multiplicative gating that a residual block does not have.

Slide four: pipelining, one layer per GPU, and the only thing he says was wrong (4:44)

Another cool feature from that old talk that I want to highlight is that we used parallelization, but not just any parallelization. We used pipelining as witnessed by this one layer per GPU.

Pipeline parallelism meant assigning each LSTM layer to its own GPU and streaming minibatches through the stack. It was an unusual thing to build in 2014 and the slide was clearly shown with pride at the time. The 2024 verdict is three sentences long and is the only place in the talk where he says a decision was simply wrong:

Was it wise to pipeline? As we now know, pipelining is not wise. But we were not as wise back then. So we used that and we got a 3.5x speed up using eight GPUs.

Hold on to that number, because it is the one piece of hard engineering data in the talk and it is quietly damning. Eight GPUs returned 3.5 times the speed. That is 44 percent efficiency, and more than half the machine was idle at any moment, which is exactly the pipeline bubble problem. The honesty here matters: this is an award talk where the speaker volunteers that one quarter of his own highlighted slides was a mistake.

Slide five: the conclusion slide, where the scaling hypothesis got written down (5:35)

And the conclusion slide in some sense, the conclusion slide from the talk from back then is the most important slide because it spelled out what could arguably be the beginning of the scaling hypothesis.

What it said: if you have a very big data set and you train a very very big neural network, then success is guaranteed.

His assessment, with the hedge intact:

And one can argue, if one is charitable, that this indeed has been what's been happening.

"If one is charitable" is him declining to claim the decade as a vindication. But the structure of his retrospective makes the claim anyway. The field spent ten years executing that sentence.

The 2014 slideWhat it claimedHis 2024 verdict
The deep learning hypothesisA ten layer neural network can do anything a human can do in a fraction of a second, because artificial neurons resemble slow biological ones.HELD "Not too bad." The depth limit lifted, the reasoning did not need to.
Our main ideaAn autoregressive model that predicts the next token well enough captures the true distribution over sequences.HELD Not the first autoregressive net, but the first where "we really believed that if you train it really well, then you will get whatever you want."
The LSTMA multilayered LSTM encoder and decoder, with a cell state carrying information across time.REPLACED "The things that poor deep learning researchers did before transformers." A ResNet rotated 90 degrees, with gating.
ParallelizationPipelining, one layer per GPU, for a 3.5x speedup on eight GPUs.WRONG "As we now know, pipelining is not wise. But we were not as wise back then."
The conclusion slideA very big data set plus a very very big neural network means success is guaranteed.HELD, AND RAN THE DECADE "The most important slide ... arguably the beginning of the scaling hypothesis."
Connectionism (not a slide)Intelligence emerges from many simple neuron like units connected and trained, not from designed symbolic structure.THE ONE THAT TRULY HELD "The idea that truly stood the test of time ... the core idea of deep learning itself."
Figure 2. The slide audit, which is the whole first half of the talk. He reshows five of his own 2014 slides and grades each one. Four of the five hold; the one he calls simply unwise is the engineering decision, not the hypothesis. The row that is not a slide is the one he says truly stood the test of time.

The idea that truly stood the test of time (5:47)

Having audited the slides, he names the thing underneath them, and it is not an architecture:

I want to mention one other idea. And this is, I claim, the idea that truly stood the test of time. It's the core idea of deep learning itself. It's the idea of connectionism. It's the idea that if you allow yourself to believe that an artificial neuron is kind of sort of a like a biological neuron. Right? If you believe that one is kind of sort of like the other, then it gives you the confidence to believe that very large neural networks, they don't need to be literally human brain scale, they might be a little bit smaller, but you could configure them to do pretty much all the things that we do, human beings.

Note what connectionism supplies in his telling. Not a proof, not a method. Confidence. It is the permission structure that let a small number of people keep scaling something while the surrounding field thought the premise was silly.

And then, immediately, the caveat he almost forgot to include, delivered live:

There's still a difference. Oh, I forgot to put that in. There's still a difference because the human brain also figures out how to reconfigure itself. Whereas we are using the best learning algorithms that we have, which require as many data points as there are parameters. Human beings are still better in this regard.

That is a precise and load bearing admission, and it is easy to skate past because of how casually it arrives. "As many data points as there are parameters" is the sample efficiency of current training stated as a rule of thumb, and it is the exact constraint that the second half of the talk is about. The human brain reconfigures itself; gradient descent needs roughly a datapoint per weight. When he later lists "understand things from limited data" as a property of future systems, this is the deficiency he is pointing back at.

The age of pretraining, and the people he credits by name (7:12)

But all this led, so I claim, arguably, to the age of pre-training. And the age of pre-training is what we might say the GPT-2 model, the GPT-3 model, the scaling laws.

Three hedges in one sentence ("so I claim, arguably"), and then a specific and unhedged credit:

And I want to specifically call out my former collaborators, Alec Radford, also Jared Kaplan, Dario Amodei, for really making this work.

Those three names map exactly onto the artifacts he just listed. Alec Radford led GPT-2 and the GPT line. Jared Kaplan is first author on Scaling Laws for Neural Language Models. Dario Amodei is on that paper and ran the research organization that shipped GPT-3. All three had left OpenAI by the time of this talk, two of them to found Anthropic, which is worth noticing in a talk delivered six months after Sutskever himself left.

But that led to the age of pre-training. And this is what's been the driver of all of progress, all the progress that we see today. Extra-large neural networks. Extraordinarily large neural networks trained on huge data sets.

He says "all of progress" twice. There is no qualification on that one. Everything you have seen, he says, came from this.

"Pre-training as we know it will unquestionably end" (7:52)

The hinge of the talk arrives without a transition. He has just finished saying that pretraining drove all progress, and the next sentence is:

But pre-training as we know it will unquestionably end. Pre-training will end.

He says it twice, the second time stripped down. Then the argument, which is two lines of arithmetic:

Why will it end? Because while compute is growing through better hardware, better algorithms, and larger clusters. Right, all those things keep increasing your compute. All these things keep increasing your compute. The data is not growing because we have but one internet. We have but one internet.

(The caption track renders "compute is growing" as "computer is growing" at this point. He is talking about compute.)

The structure is a multiplication where one factor has a ceiling. Pretraining takes two inputs. Compute has three independent growth channels, each of which he names: better hardware, better algorithms, larger clusters. Data has none, because the supply is a single artifact that already exists. Repeat "we have but one internet" and the argument is finished; nothing else is needed.

PRE-TRAINING TAKES TWO INPUTS. ONLY ONE OF THEM SCALES. COMPUTE: THREE GROWTH CHANNELS, ALL NAMED IN THE TALK Better hardware Better algorithms Larger clusters COMPUTE keeps growing no ceiling in sight DATA: ONE SOURCE, ALREADY FINISHED The public internet "we have but one internet" PEAK DATA no second deposit, nothing is refilling it PRE-TRAINING = COMPUTE x DATA the recipe that drove, in his words, all of progress THE CLAIM When one factor of a product has a ceiling, the product has a ceiling. That is the entire argument. Not a claim that models stop improving, and not a claim that compute stops mattering. A claim about one recipe.
Figure 3. The argument at 7:52, drawn. Compute gets three independent growth channels, all three of which he names out loud. Data gets one source and a closed box. The asymmetry is the whole claim, and it is why he calls the ending unquestionable rather than likely.

Then the line that left the building and went everywhere:

You could even say, you can even go as far as to say that data is the fossil fuel of AI. It was like created somehow, and now we use it, and we've achieved peak data, and there'll be no more. We have to deal with the data that we have.

Three distinct claims are packed into that metaphor and they are worth separating, because the metaphor got quoted far more often than it got read. The data was laid down once by a process nobody is running again. It is consumed rather than reused. And the total is now in sight, which is what "peak data" means, borrowed straight from peak oil.

Watch the hedging around it. He does not assert the fossil fuel line, he works up to it in stages: "you could even say, you can even go as far as to say." He is flagging it as a figure of speech he is choosing to deploy, not as a finding. And he immediately limits the consequence:

Now, it's still still let us go quite far, but this is there's only one internet.

That sentence is the one the headlines dropped. The existing data still has room in it. The claim is about the horizon of a recipe, not a wall arriving on Tuesday.

What comes next, carefully credited to other people (8:57)

The transition into the forward looking half is a refusal to own it:

So, here I'll take a bit of liberty to speculate about what comes next. Actually, I don't need to speculate because many people are speculating, too, and I'll mention their speculations.

This matters for reading the rest of the talk honestly. The three candidates below are presented as the field's speculations that he is relaying, not as his roadmap. He gives each of them one or two sentences and declines to rank them.

Agents. Flagged as a word people are using rather than a plan anyone has:

You may have heard the phrase agents. It's common, and I'm sure that eventually something will happen, but people feel like something agents is the future.

"I'm sure that eventually something will happen" is about as thin as endorsement gets.

Synthetic data. Named as an open problem, not a solution:

More concretely, but also a little bit vaguely, synthetic data. But what does synthetic data mean? Figuring this out is a big challenge, and I'm sure that different people have all kinds of interesting progress there.

The construction is pointed. "More concretely, but also a little bit vaguely" is him saying the phrase sounds like a technique and is actually a question.

Inference time compute. The freshest of the three at the time, and the one he attaches to a shipped artifact:

And then inference time compute, or maybe what's been most recently most vividly seen in o1, the o1 model.

OpenAI o1 had been released three months before this talk, in September 2024. It is the only specific model anywhere in the forward looking section, and his summary of all three candidates is a shrug and a blessing:

These are all examples of things of people trying to figure out what to do after pre-training. And those are all very good things to do.

The biology slide, which is the real argument of the talk (9:52)

The three candidates above answer "what should people try next." They do not answer the harder objection sitting underneath the whole claim: if the one recipe we know how to scale runs out, why believe another recipe is out there to be found at all. The answer he gives is not from computer science, and the way he tells the story is most of its charm.

I want to mention one other example from biology, which I think is really cool. And the example is this. So, about many, many years ago at this conference also, I saw a talk where someone presented this graph. But the graph showed the relationship between the size of the body of a mammal and the size of their brain. In this case, it's in mass.

Then the thing he remembered about that old talk, which is the setup for the punchline:

And that talk, I remember vividly, they were saying, "Look, in biology everything is so messy, but here you have one rare example where there's a very tight relationship between the size of the body of the animal and their brain."

A tight relationship in a field with no tight relationships. That is why the graph stuck. And then the research method, stated with no embellishment at all:

And totally randomly, I became curious at this graph. And one of the early So, I went to Google to do research to look for this graph. And one of the images in Google images was this.

He went to Google Images. That is the provenance of the evidence in the most quoted AI talk of 2024, and he says so plainly rather than dressing it up as a literature review. There is then a live moment where the clicker misbehaves: "And the interesting thing in this image is you see like I don't know, is the mouse working? Oh, yeah, the mouse is working, right."

What the image shows, in his order: all the different mammals, then non-human primates, which are "basically the same thing," and then a third group.

But then you've got the hominids. And to my knowledge, hominids are like close relatives to the humans in evolution. Like the Neanderthals. There's a bunch of them. Ho- like it's called Homo habilis, maybe. There's a whole bunch. And they're all here.

Hominids including Neanderthals and, hedged with a "maybe," Homo habilis. He is not pretending to expertise he does not have, which is exactly why the next sentence lands:

And what's interesting is that they have a different slope on their brain-to-body scaling exponent. So, that's pretty cool. What that means is that there is a precedent, there is an example of biology figuring out some kind of different scaling. Something clearly is different. So, I think that is cool.

"A different slope on their brain-to-body scaling exponent" is a statement about allometry, the study of how biological quantities scale against body size. Everything else sits on one exponent. Hominids sit on another. Then, because he knows the room will ask, he reads the axes out loud:

And by the way, I want to highlight this x-axis is log scale. You see this is 100, this is 1,000, 10,000, 100,000, and likewise in grams. 1 g, 10 g, 100 g, 1,000 g.

Brain mass against body mass, rebuilt from the axes he reads out loud at 11:29 1 g 10 g 100 g 1,000 g 100 1,000 10,000 100,000 body mass, grams, LOG SCALE brain mass, grams, LOG SCALE a different slope on their brain-to-body scaling exponent mammals: "one rare example where there's a very tight relationship" non-human primates: "it's basically the same thing" hominids: Neanderthals, Homo habilis, "there's a whole bunch" The axis tick values are his, read out loud. The slopes are schematic: he states no exponent, only that the hominid group has one that every other mammal on the chart does not.
Figure 4. The slide the talk actually turns on, rebuilt from his own reading of the axes. Mammals occupy one slope, non-human primates a parallel one he waves off as basically the same thing, and hominids a visibly different and steeper one in the upper right. The payoff is not about brains. It is an existence proof that a scaling relationship can change exponent, found in a system nobody designed, which is his answer to the objection that the end of pretraining is the end of scaling.

The payoff is three sentences, and the middle one is the most important thing he says all day:

So, it is possible for things to be different. The things that we are doing, the things that we've been scaling so far is actually the first thing that we figured out how to scale. And without doubt, the field, everyone who's working here, will figure out what to do.

"The first thing that we figured out how to scale." Not the only one, not the best one. The first. The biology chart is there to make that sentence credible: evolution, with no researchers and no roadmap, stumbled onto a second exponent. A field of thousands of people, explicitly looking, should manage the same. That is why the chart earns the three minutes it gets, and it is also the strongest thing in the talk, because it turns a pessimistic claim about data into an optimistic claim about method without a word of hand waving.

A pause to say that what already happened is hard to convey (12:31)

Before the last section he stops and does something people in his position rarely bother with, which is to mark how strange the present is:

But, I want to talk here, I want to take a few minutes and speculate about the longer term, the longer term. Where are we all headed? Right? We're making all this progress. It's astounding progress. It's really I mean, those of you who've been in the field 10 years ago and you remember just how incapable everything has been, like yes, you can say even if you kind of say of course deep learning still to see it is just unbelievable. It's completely I can't convey that feeling to you.

The sentence falls apart as he says it, which is the point. Then the clean version:

You know, if you've joined the field in the last 2 years, then of course, you speak to computers and they talk back to you and they disagree and that's what computers are. But, it hasn't always been the case.

"They disagree" is the detail he picks. Not that they answer, that they push back. For anyone who joined after ChatGPT, that is simply what a computer is. For him it is the thing that cannot be conveyed.

Superintelligence, and the four properties (13:25)

But, I want to talk to you a little bit about superintelligence, just a bit. Because that is obviously where this field is headed. This is obviously what's being built here.

Two "obviously"s in two sentences, in a room full of researchers, most of whom would not describe their own job that way. He states it as a settled fact about the field rather than as his own position, which is a rhetorical choice worth noticing. His stated goal is modest and specific:

And the thing about superintelligence is that it will be different qualitatively from what we have. And my goal in the next minute to try to give you some concrete intuition of how it will be different. So that you yourself could reason about it.

He starts from the gap in today's systems rather than from the future ones:

So right now we have our incredible language models and their unbelievable chatbots and they can even do things, but they're also kind of strangely unreliable and they get confused when a while also having dramatically superhuman performance on evals. So it's really unclear how to reconcile this.

That is the honest statement of the 2024 situation, and the word "unclear" is him declining to resolve it. Superhuman on benchmarks, strangely unreliable in use, and no theory that explains both at once. Then the four properties, each with its hedge attached.

It will be agentic, really. With an immediate walk back of his own phrasing, live:

Those systems are actually going to be agentic in a real ways. Whereas right now the systems are not agents in any meaningful sense. Just very That might be too strong. They're very very slightly agentic. Just the beginning.

It will reason, and that is why it will be unpredictable. This is the most substantive idea in the second half and it gets less attention than the fossil fuel line:

It will actually reason. And by the way, I want to mention something about reasoning is that a system that reasons, the more it reasons, the more unpredictable it becomes. The more it reasons, the more unpredictable it becomes.

He says it twice, exactly, which is his tell for the sentences he wants carried out of the room. The argument for it runs back through the whole talk, all the way to slide one:

All the deep learning that we've been used to is very predictable because if you've been working on replicating human intuition essentially. It's like the gut feel. If you come back to the 0.1 second reaction time what kind of processing we do in our brains well it's our intuition. So we've endowed our AIs with some of that intuition.

That is the deep learning hypothesis closing its own loop. The thing a ten layer network could capture was the fraction of a second, which is intuition, which is by nature predictable, because it is the output of a shallow and stereotyped computation. Reasoning is not that. Reasoning is the part that takes longer than a tenth of a second, and its entire value is arriving somewhere you could not have guessed.

The evidence is one sentence, and it is unanswerable:

But reasoning, and you're seeing some early signs of that, reasoning is unpredictable and one reason to see that is because the chess AIs, the really good ones, are unpredictable to the best human chess players. So we will have to be dealing with AI systems that are incredibly unpredictable.

Chess engines are superhuman and narrow, and the best humans in the world still cannot anticipate their moves. There is no alignment trick that recovers predictability without giving up the reasoning, because they are the same property viewed from two sides.

It will understand things from limited data, and it will not get confused. Two sentences, both pointing straight back at the caveat he almost forgot in the connectionism section: "They will understand things from limited data. They will not get confused. All the things which are really big limitations."

And then the disclaimer that the headlines ignored:

I'm not saying how, by the way, and I'm not saying when. I'm saying that it will.

It will be self aware, which he treats as unremarkable. No mysticism offered, and none permitted:

And when all those things will happen together with self-awareness, because why not? Self-awareness is useful. It is part your ourselves are parts of our own world models.

Self awareness as an engineering consequence: a model of the world that contains the modeler is a better model of the world, so a system optimizing its world model will build one. That is the entire argument, delivered in under ten seconds, in a throwaway tone.

PropertyWhat he says about systems todayWhat he says about what is comingThe hedge he attaches, verbatim
Agency"not agents in any meaningful sense", then corrected live to "very very slightly agentic. Just the beginning.""actually going to be agentic in a real ways""That might be too strong." He interrupts himself to soften his own claim about today.
ReasoningDeep learning so far replicates intuition, the gut feel, the 0.1 second reaction. Predictable by construction."It will actually reason." And therefore: "the more it reasons, the more unpredictable it becomes," said twice."you're seeing some early signs of that". The evidence offered is chess engines, which are already unpredictable to the best human players.
Sample efficiencyCurrent learning algorithms "require as many data points as there are parameters. Human beings are still better in this regard.""They will understand things from limited data.""I'm not saying how, by the way, and I'm not saying when. I'm saying that it will."
Reliability"strangely unreliable and they get confused" while "also having dramatically superhuman performance on evals""They will not get confused.""It's really unclear how to reconcile this." He states the paradox and does not resolve it.
Self awarenessNot discussed as present."because why not? Self-awareness is useful. It is part ... our selves are parts of our own world models."Offered as a consequence rather than a milestone, in one sentence, with no emphasis at all.
Figure 5. The forward looking half of the talk, with the hedges restored. Every property in the middle column got quoted; most of the right hand column did not. The reasoning row is the one worth rereading, because the two sides of it are the same property: a system whose conclusions you can anticipate is not doing much reasoning.

He finishes the sketch and then refuses to carry it any further:

When all those things come together, we will have systems of radically different qualities and properties that exist today. And of course, they will have incredible and amazing capabilities. But the kind of issues that come up with systems like this, and I'll just leave it as an exercise to to imagine.

Leaving it as an exercise to imagine is a deliberate non answer from the person who left OpenAI in May 2024 and founded Safe Superintelligence a month later, a company whose entire stated reason for existing is the issues he just declined to enumerate. He does not mention SSI once in the talk.

The closing line, delivered with a shrug audible in the recording:

And I would say that it's definitely also impossible to predict the future. Really, all kinds of stuff is possible. But on this uplifting note, I will conclude. Thank you so much.

The Q&A (16:31)

Four questions in eight minutes, and they produced several of the lines that got clipped and shared more than anything in the prepared talk. Worth having in full, because in each case the question shapes the answer.

Question one, at 16:44: are there other biological structures worth copying?

The asker wants to know whether, in 2024, there are parts of human cognition beyond the neuron that are worth exploring the way the neuron was.

His answer does two things. First, it hands the question back to anyone who has a real insight:

So, the way I'd answer this question is that if you are or someone is a person who has a specific insight about, "Hey, we are all being extremely silly because clearly the brain does something and we are not. And that's something that can be done," they should pursue it.

Then the audit of how much biology actually got used, which is the quotable part and a direct callback to his own connectionism section:

Like there's been a lot of desire to make biologically inspired AI. And you could argue on some level that biologically inspired AI is incredibly successful, which is all of the learning is biologically inspired AI. But on the other hand, the biological inspiration was very, very, very modest. It's like, let's use neurons. This is the full extent of the biological inspiration. Let's use neurons.

The whole of biologically inspired AI, in his accounting, is the decision to call the unit a neuron. Everything else was engineering. And the closing hedge is characteristically open: "And more detailed biological inspiration has been very hard to come by. But I wouldn't rule it out. I think if someone has a special insight, they might be able to see something and that would be useful."

Question two, at 17:45: will reasoning let a model auto correct its own hallucinations?

The longest question of the session, and the asker builds it carefully. The observation is about method: the way the field detects hallucination today is statistical, "some amount of standard deviations or whatever away from the mean," precisely because the models cannot be trusted to reason about their own output. The question is whether reasoning changes that:

Do you think that a model given reasoning will be able to correct itself, sort of auto correct itself? And that will be a core feature of future models so that there won't be as many hallucinations because the model will recognize when ... the model will be able to reason and understand when a hallucination is occurring. Does the question make sense?

The answer is the most direct one he gives all session:

Yes, and the answer is also yes. I think what you described is extremely highly plausible. Yeah, I mean you should check. I mean for yeah, it I wouldn't I wouldn't rule out that it might already be happening with some of the, you know, early reasoning models of today. I don't know. But longer term, why not?

"It might already be happening" in December 2024, three months after o1 shipped, is the most concretely forward looking thing in the whole talk, and he buries it in a hedge.

The exchange then turns on a word. The asker reaches for an analogy: "part of like Microsoft Word like autocorrect. It's a core feature." Sutskever will not take it:

Yeah, I just I mean I think calling it autocorrect is really doing a disservice. I think you are When you say autocorrect, you evoke like it's far grander than autocorrect, but other than but, you know, this point aside, the answer is yes.

That is him defending the thing he just agreed to from the framing that would make it sound small. A model that notices its own error is not spell check with extra steps.

Question three, at 19:48: do they need rights, and what incentive structure gets us there?

The asker opens by naming what the talk left out: "I loved the ending mysteriously leaving out do they replace us or are they, you know, superior? Do they need rights? You know, it's a new species of Homo sapien spawned intelligence." Then the actual question, which is about mechanism design rather than ethics: "How do you create the right incentive mechanisms for humanity to actually create it in a way that gives it the freedoms that we have as Homo sapiens?"

He endorses the question and declines the answer, in that order:

You know, I feel like this In some In some sense, those are those are the kind of questions that people should be reflecting on more. But to your question about what incentive structure should we create, I I don't feel that I know. I don't feel confident answering questions like this because uh it's like you're talking about creating some kind of a top-down structure government thing, I don't know.

The asker offers a specific mechanism: "It could be a cryptocurrency, too. There's Bittensor, you know, there's things." (Bittensor is a decentralized network that pays token rewards for contributed machine learning work. The caption track renders it as "Bit Tensor.") The response became one of the most circulated clips of the conference, mostly for its delivery:

I don't feel like I am the right person to comment on cryptocurrency.

And then, having declined, he answers anyway, and this is the part that got quoted far less often than the crypto line:

But you know, there is a chance, by the way, what you're describing will happen. That indeed we will have, you know, in some sense it's it's it's not a bad end result if you have AIs and all they want is to coexist with us and also just to have rights, maybe that will be fine. It's But I don't know. I mean, I think things are so incredibly unpredictable. I I hesitate to comment, but I encourage the speculation.

"Not a bad end result" is a load bearing phrase from someone whose company is named after the problem. And "I hesitate to comment, but I encourage the speculation" is the posture of the entire second half of the talk, stated outright.

Question four, at 21:49: do LLMs generalize multi hop reasoning out of distribution?

Asked by Shalev Lifshitz of the University of Toronto, who notes he works with Sheila McIlraith. (The caption track spells the name "Shalev Lifschitz" and renders "Sheila" without a surname.)

Sutskever rejects the shape of the question before answering it:

So, okay. The question assumes that the answer is yes or no, but the question should not be answered with a yes or no. Because what does it mean out of distribution generalization? What does it mean? What does it mean in distribution? And what does it mean out of distribution?

Then, because it is a test of time talk, he takes the opportunity to date the vocabulary:

Because it's a test of time talk, I'll say that long, long ago, before people were using deep learning, they were using things like string matching and n-grams for machine translation. People were using statistical phrase tables. Can you imagine? They had tens of thousands of lines of code of complexity which was I mean it's it was truly unfathomable.

(Statistical phrase tables and n-grams are the machinery the 2014 paper displaced, which makes the callback exact. The caption drops "lines of" from "tens of thousands of lines of code.")

And the argument lands:

And back then generalization meant is it literally not in this is the same word phrasing as in the data set. Now we may say, "Well, sure my model achieves this high score on I don't know math competitions, but maybe some discussion in some forum on the internet was about the same ideas and therefore it's memorized." Well, okay, you could say maybe it's in distribution, maybe it's memorization. But I also think that our standards for what counts as generalization have increased really quite substantially, dramatically, unimaginably if you keep track.

This is the best thing in the Q&A. The word generalization has not held still. In 2014 it meant "the exact string is not in the training set." In 2024 it means "a forum thread somewhere on the internet did not discuss a similar idea." Those are not the same test, they are not close to the same test, and a decade of progress is hiding inside the drift. Nobody announced the change, so nobody credits it.

His actual answer, once the terms are cleaned up, is a careful three part split:

And so I think the answer is to some degree probably not as well as human beings. I think it is true that human beings generalize much better. But at the same time they definitely generalize out of distribution to some degree. I hope it's a useful tautological answer.

Models generalize out of distribution. They do it worse than people. And the measuring stick keeps moving, which is why the yes or no version of the question has no answer.

The session ends there, cut short by the clock:

And unfortunately, we're out of time for this session. I have a feeling we could go on for the next 6 hours. But thank you so much Ilya for the talk.

The decade, using only the landmarks he names

Every entry below is something Sutskever points at in these twenty four minutes. Nothing has been added from the wider history, which is why there is no AlexNet and no Chinchilla here: he does not mention them, and the shape of his argument is clearer without them.

Key takeaways

Chapters

The seven entries in bold are the video's own chapter markers, reproduced verbatim. Seven markers across twenty four minutes leaves most of the talk unlabelled, so everything else is a sub beat added here from the transcript's own clock. Two of his markers land slightly after the content they name, so the list below runs in strict time order rather than grouping sub beats under the marker they belong to.

Notable quotes

Because a lot of the things in this work were correct, but some not so much. And we can review them and we can see what happened and how it gently flowed to where we are today. Ilya Sutskever, setting the frame for the slide audit, 1:05

It's an autoregressive model trained on text. It's a large neural network and it's a large data set. And that's it. Ilya Sutskever, the entire 2014 paper in three bullet points, 1:37

If there is one human in the entire world that can do some task in a fraction of a second, then a 10-layer neural network can do it, too, right? It follows. You just take their connections and you embed them inside your neural net, the artificial one. Ilya Sutskever, the deep learning hypothesis as a construction argument, 2:37

To those unfamiliar, an LSTM is the things that poor deep learning researchers did before transformers. Ilya Sutskever, introducing the architecture the award winning paper was built on, 4:22

Was it wise to pipeline? As we now know, pipelining is not wise. But we were not as wise back then. So we used that and we got a 3.5x speed up using eight GPUs. Ilya Sutskever, the only decision in the paper he grades as wrong, 5:17

There's still a difference because the human brain also figures out how to reconfigure itself. Whereas we are using the best learning algorithms that we have, which require as many data points as there are parameters. Human beings are still better in this regard. Ilya Sutskever, the caveat he says out loud he forgot to put on the slide, 6:50

But pre-training as we know it will unquestionably end. Pre-training will end. Ilya Sutskever, the line the talk is remembered for, 7:52

The data is not growing because we have but one internet. We have but one internet. Ilya Sutskever, the whole argument in two sentences, said twice, 8:05

You could even say, you can even go as far as to say that data is the fossil fuel of AI. It was like created somehow, and now we use it, and we've achieved peak data, and there'll be no more. Ilya Sutskever, with the hedges that usually get cut off the front, 8:24

What that means is that there is a precedent, there is an example of biology figuring out some kind of different scaling. Ilya Sutskever, on the hominid slope, 11:29

The things that we are doing, the things that we've been scaling so far is actually the first thing that we figured out how to scale. Ilya Sutskever, the payoff of the biology slide, 12:18

If you've joined the field in the last 2 years, then of course, you speak to computers and they talk back to you and they disagree and that's what computers are. But, it hasn't always been the case. Ilya Sutskever, on what cannot be conveyed to anyone who arrived after ChatGPT, 13:03

A system that reasons, the more it reasons, the more unpredictable it becomes. The more it reasons, the more unpredictable it becomes. Ilya Sutskever, said twice in a row, which is his tell, 14:30

Reasoning is unpredictable and one reason to see that is because the chess AIs, the really good ones, are unpredictable to the best human chess players. Ilya Sutskever, the evidence, 15:10

I'm not saying how, by the way, and I'm not saying when. I'm saying that it will. Ilya Sutskever, on every property of superintelligence he just listed, 15:33

And when all those things will happen together with self-awareness, because why not? Self-awareness is useful. It is part your ourselves are parts of our own world models. Ilya Sutskever, disposing of the hardest question in the field in one sentence, 15:37

But on the other hand, the biological inspiration was very, very, very modest. It's like, let's use neurons. This is the full extent of the biological inspiration. Let's use neurons. Ilya Sutskever, in the Q&A, auditing how much biology deep learning actually took, 17:50

I don't feel like I am the right person to comment on cryptocurrency. Ilya Sutskever, after a questioner proposed Bittensor as the incentive mechanism for AI rights, 21:35

I mean, I think things are so incredibly unpredictable. I I hesitate to comment, but I encourage the speculation. Ilya Sutskever, the posture of the entire second half, stated outright, 22:00

But I also think that our standards for what counts as generalization have increased really quite substantially, dramatically, unimaginably if you keep track. Ilya Sutskever, on why "out of distribution" no longer means what it meant in 2014, 23:30

I hope it's a useful tautological answer. Ilya Sutskever, closing the last question of the session, 23:50

How the predictions have held up

Everything above is the talk. This section is the only assessment on the page, written in October 2026, and it is a reader's scorecard rather than anything Sutskever said.

The claim people argued about is not the claim he made. "Pre-training as we know it will unquestionably end" was read almost everywhere as "scaling has hit a wall, now." He said the opposite twice in the same breath: the data "still let us go quite far," and the end is a horizon for one recipe, not a date. Frontier labs did not stop pretraining, and the models kept improving, so by the headline reading the prediction looks wrong and by his actual reading it has not yet come due. That gap is worth noticing on a talk this heavily quoted.

The ordering of his three candidates turned out backwards, in the best way for him. He listed agents first, synthetic data second, and inference time compute last and most tentatively, attached to a model three months old. Inference time compute is the one that became a permanent axis of the field. Reasoning models and reinforcement learning after pretraining went from a curiosity to a central line of frontier spending in the two years since, though the actual compute split is an estimate rather than a disclosure: no major lab publishes how its budget divides between pretraining and post training, so every number in circulation on this is inferred.

Agents went from the word he was least willing to endorse to the dominant product category. "I'm sure that eventually something will happen" has aged into an understatement. Whether that validates the prediction or just the vocabulary is a fair argument.

Synthetic data is the one he framed most accurately. He refused to call it a solution and called the definition itself the hard part. Two years on it is used heavily in post training and is still argued about rather than settled, which is exactly the status he gave it.

The fossil fuel metaphor took the most criticism, and some of it lands. Critics have pointed out that training data is not a fixed geological deposit: people keep producing text, video and interaction logs, and data quality and filtering have mattered more than raw volume. His own framing makes room for this (he says peak data, not zero data) but the metaphor travelled further than the framing did, which is a hazard of coining a good line.

The idea with the longest shelf life is the one that got the least coverage. Reasoning implies unpredictability, and the two cannot be separated because they are the same property. That observation was made in passing in December 2024, before the reasoning model era properly began, and it has only become more awkward as the field has pushed test time reasoning harder while also asking for more predictable behaviour from the same systems. The chess engine example remains the cleanest statement of the tension anyone has offered.

On his own position. He never mentions Safe Superintelligence in the talk, which he co-founded in June 2024, a month after leaving OpenAI. It raised at a reported 30 billion dollar valuation in 2025, he became its chief executive when co-founder Daniel Gross left, Nvidia announced a partnership and investment in July 2026, and it has still shipped nothing. A lab organized around not shipping is a strange thing until you read it against this talk, where the speaker argues that the recipe everyone is shipping on has a horizon and that the next one has not been found yet.

Where this sits in the LLM Learning track

This opens Part 4, the final stage of the track, the part about where all of it goes. Everything earlier describes a recipe that works. This is the person who wrote the recipe's opening chapter saying out loud that its main ingredient is finite.

Read it against the Chinchilla and scaling section of Sasha Rush's Five Formulas, which quantifies the same data constraint from the other direction and gives you the arithmetic behind "one internet." Read the reasoning and unpredictability argument against John Schulman on reinforcement learning from human feedback, because the lever Schulman describes for making a model behave is the same lever that gets harder as the model reasons more. And the pair that closes the track is deliberate: this talk sets the open research question, and Karpathy's Software Is Changing Again is what gets built while the question stays open. Read them back to back.

Resources mentioned

Caption corrections. The automatic caption track mangles four things, corrected silently in the quotes above and flagged here instead: "computer is growing" should be "compute is growing" (8:05); "tens of thousands of code of complexity" should be "tens of thousands of lines of code" (22:49); "Bit Tensor" is Bittensor (21:30); and the last questioner, captioned "Shalev Lifschitz," is Shalev Lifshitz. The captions also render "n-grams" as "grams" and o1 as both "O1" and "0.1".

Full transcript
[00:00:01] I want to thank the organizers for choosing our paper for this award. It was very nice. And I also want to thank my incredible co-authors and collaborators, Oriol Vinyals and Quoc Le, who stood right before you a moment ago. And what you have here is an image, a screenshot, from [00:00:35] a similar talk 10 years ago at NeurIPS in 2014 in Montreal. And it was a much more innocent time. Here we are shown in the photos. This is the before. Here's the after, by the way. And now we've got more experienced, hopefully wiser. But here I'd like to talk a little bit about the work itself and maybe a [00:01:05] 10-year retrospective on it. Because a lot of the things in this work were correct, but some not so much. And we can review them and we can see what happened and how it gently flowed to where we are today. So, let's begin by talking about what we did. And the way we'll do it is by showing slides from the same talk 10 years ago. [00:01:37] But the summary of what we did is the following three bullet points. It's an autoregressive model trained on text. It's a large neural network and it's a large data set. And that's it. Now let's dive in into the details a little bit more. So, this was a slide 10 years ago. Not too bad. The deep learning hypothesis. And what we said here is that if you have a large neural network [00:02:07] with 10 layers, then it can do anything that a human being can do in a fraction of a second. Like why did we have this emphasis emphasis on things that human beings can do in a fraction of a second? Why this thing specifically? Well, if you believe the deep learning dogma, so to say, that artificial neurons and biological neurons are similar, or at least not too different, and you believe that real neurons are slow, then anything that we can do [00:02:37] quickly, by we I mean human beings, I even mean just one human in the entire world. If there is one human in the entire world that can do some task in a fraction of a second, then a 10-layer neural network can do it, too, right? It follows. You just take their connections and you embed them inside your neural net, the artificial one. So, this was the motivation. Anything that a human being can do in a fraction of a second, a big 10-layer 10-layer neural network can do, too. We focused on 10-layer neural networks because this was the neural networks we knew how to train back in the day. [00:03:09] If you could go beyond in your layers somehow, then you could do more. But back then we could only do 10 layers, which is why we emphasized whatever human beings can do in a fraction of a second. A different slide from the talk, a slide which says our main idea. And you may be able to recognize two things, or at least one thing. You might be able to recognize that something auto-regressive is going on here. What is it saying really? What does the slide really say? This slide says that if you have an [00:03:41] auto-regressive model, and it predicts the next token well enough, then it will in fact grab and capture and grasp the correct distribution over whatever over sequences that come next. And this was a relatively new thing. It wasn't literally the first ever auto-regressive neural network, but I would argue it was the first auto-regressive neural network where we really believed that if you train it really well, then you will get whatever you want. In our case, back then was the [00:04:12] humble, today humble, then incredibly audacious task of translation. Now I'm going to show you some ancient history that many of you might have never seen before. It's called the LSTM. To those unfamiliar, an LSTM is the things that poor deep learning researchers did before transformers. And it's basically a ResNet but rotated in 90°. So that's an LSTM. And it came before it it's like [00:04:44] it's like kind of like a slightly more complicated ResNet. You can see there is your integrator, which is now called the residual stream, but you've got some multiplication going on. It's a little bit more complicated. But that's what we did. It was a ResNet rotated in 90°. Another cool feature from that old talk that I want to highlight is that we used parallelization, but not just any parallelization. We used pipelining as witnessed by this one layer per GPU. [00:05:17] Was it wise to pipeline? As we now know, pipelining is not wise. But we were not as wise back then. So we used that and we got a 3.5x speed up using eight GPUs. And the conclusion slide in some sense, the conclusion slide from the talk from back then is the most important slide because it spelled out what could arguably be the beginning of the scaling hypothesis, right? That if you have a [00:05:47] very big data set and you train a very very big neural network, then success is guaranteed. And one can argue, if one is charitable, that this indeed has been what's been happening. I want to mention one other idea. And this is, I claim, the idea that truly stood the test of time. It's the core idea of deep learning itself. It's the idea of connectionism. It's the idea that if you allow yourself to believe [00:06:18] that an artificial neuron is kind of sort of a like a biological neuron. Right? If you believe that one is kind of sort of like the other, then it gives you the confidence to believe that very large neural networks, they don't need to be literally human brain scale, they might be a little bit smaller, but you could configure them to do pretty much all the things that we do, human beings. [00:06:50] There's still a difference. Oh, I forgot to put that in. There's still a difference because the human brain also figures out how to reconfigure itself. Whereas we are using the best learning algorithms that we have, which require as many data points as there are parameters. Human beings are still better in this regard. But all this led, so I claim, arguably, to the age of pre-training. And the age of pre-training is what we [00:07:21] might say the GPT-2 model, the GPT-3 model, the scaling laws. And I want to specifically call out my uh former collaborators, Alec Radford, also Jared Kaplan, Dario Amodei, for really making this work. But that led to the age of pre-training. And this is what's been the driver of all of progress, all the progress that we see today. Extra-large neural networks. Extraordinarily large neural networks [00:07:52] trained on huge data sets. But pre-training as we know it will unquestionably end. Pre-training will end. Why will it end? Because while computer is growing through better hardware, better algorithms, and larger clusters. Right, all those things keep increasing your compute. All these things keep increasing your compute. The data is not growing because we have but one internet. [00:08:24] We have but one internet. You could even say, you can even go as far as to say that data is the fossil fuel of AI. It was like created somehow, and now we use it, and we've achieved peak data, and there'll be no more. We have to deal with the data that we have. Now, it's still still let us go quite far, but this is there's only one internet. [00:08:57] So, here I'll take um a bit of liberty to speculate about what comes next. Actually, I don't need to speculate because many people are speculating, too, and I'll mention their speculations. You may have heard the phrase agents. It's common, and I'm sure that eventually something will happen, but people feel like something agents is the future. More concretely, but also a little bit vaguely, synthetic data. But what does synthetic data mean? [00:09:27] Figuring this out is a big challenge, and I'm sure that different people have all kinds of interesting progress there. And then inference time compute, or maybe what's been most recently most vividly seen in O1, the O1 model. These are all examples of things of people trying to figure out what to do after pre-training. And those are all very good things to do. I want to mention one other example from biology, [00:09:57] which I think is really cool. And the example is this. So, about many, many years ago at this conference also, I saw a talk where someone presented this graph. But the graph showed the relationship between the size of the body of the size of the body of a mammal and the size of their brain. In this case, it's in mass. And the that talk, I remember vividly, [00:10:28] they were saying, "Look, it's in biology everything is so messy, but here you have one rare example where there's a very tight relationship between the size of the body of the animal and their brain." And totally randomly, I became curious at this graph. And one of the early one of the early So, I went to Google to do research to to look for this graph. And one of the images in Google images was this. And the interesting thing in this image is you see like I don't know, is the mouse [00:10:59] working? Oh, yeah, the mouse is working, right. So, you've got these mammals, right? All the different mammals. Then you've got non-human primates. It's basically the same thing. But then you've got the hominids. And to my knowledge, hominids are like close relatives to the humans in evolution. Like the Neanderthals. There's a bunch of them. Ho- like it's called Homo habilis, maybe. There's [00:11:29] There's a whole bunch. And they're all here. And what's interesting is that they have a different slope on their brain-to-body scaling exponent. So, that's pretty cool. What that means is that there is a precedent, there is an example of biology figuring out some kind of different scaling. Something clearly is different. So, I think that is cool. And by the way, I want to highlight this x-axis is [00:12:01] log scale. You see this is 100, this is 1,000, 10,000, 100,000, and likewise in grams. 1 g, 10 g, 100 g, 1,000 g. So, it is possible for things to be different. The things that we are doing, the things that we've been scaling so far is actually the first thing that we figured out how to scale. And without doubt, the field, everyone who's working here, will figure out what to do. [00:12:31] But, I want to talk here, I want to take a few minutes and speculate about the longer term, the longer term. Where are we all headed? Right? We're making all this progress. It's an It's astounding progress. It's really I mean, those of you who've been in the field 10 years ago and you remember just how incapable everything has been, like yes, you can say even if you kind of say of course deep learning still to see it is just unbelievable. [00:13:03] It's completely I can't convey that feeling to you. You know, if you've joined the field in the last 2 years, then of course, you speak to computers and they talk back to you and they disagree and that's what computers are. But, it hasn't always been the case. But, I want to talk to you a little bit about superintelligence, just a bit. Because that is obviously where this field is headed. This is obviously what's being built here. And the thing about superintelligence is [00:13:33] that it will be different qualitatively from what we have. And my goal in the next minute to try to give you some concrete intuition of how it will be different. So that you yourself could reason about it. So right now we have our incredible language models and their unbelievable chatbots and they can even do things, but they're also kind of strangely unreliable and they get confused when a while also having dramatically superhuman performance on [00:14:04] evals. So it's really unclear how to reconcile this. But eventually sooner or later the following will be achieved. Those systems are actually going to be agentic in a real ways. Whereas right now the systems are not agents in any meaningful sense. Just very That might be too strong. They're very very slightly agentic. Just the beginning. It will actually reason. And by the way, I want to mention something about reasoning is that a system that reasons, the more it [00:14:36] reasons, the more unpredictable it becomes. The more it reasons, the more unpredictable it becomes. All the deep learning that we've been used to is very predictable because if you've been working on replicating human intuition essentially. It's like the gut feel. If you come back to the 0.1 second reaction time what kind of processing we do in our brains well it's our intuition. So we've endowed our AIs with some of that intuition. But reasoning, and you're seeing some early [00:15:07] signs of that, reasoning is unpredictable and one reason to see that is because the chess AIs, the really good ones, are unpredictable to the best human chess players. So we will have to be dealing with AI systems that are incredibly unpredictable. They will understand things from limited data. They will not get confused. All the things which are really big limitations. I'm not saying how, by the way, and I'm not saying when. I'm saying that it will. [00:15:37] And when all those things will happen together with self-awareness, because why not? Self-awareness is useful. It is part your ourselves are parts of our own world models. When all those things come together, we will have systems of radically different qualities and properties that exist today. And of course, they will have incredible and amazing capabilities. But the kind of issues that come up with systems like this, and I'll just leave it as an exercise to to imagine. [00:16:07] It's very different from what we are used to. And I would say that it's definitely also impossible to predict the future. Really, all kinds of stuff is possible. But on this uplifting note, I will conclude. Thank you so much. Um [00:16:44] Thank you. Um now in 2024, are there other biological structures that are part of human cognition that you think are worth exploring in a similar way or that you're interested anyway? So, the way I'd answer this question is that if you are or someone is a person who has a specific insight about, "Hey, [00:17:14] we are all being extremely silly because clearly the brain does something and we are not. And that's something that can be done, they should pursue it. I personally don't Well, depends on the level of obstruction you're looking at. Maybe I'll answer it this way. Like there's been a lot of desire to make biologically inspired AI. And you could argue on some level that biologically inspired AI is incredibly successful, which is all of the learning is biologically inspired AI. [00:17:45] But on the other hand, the biological inspiration was very, very, very modest. It's like, let's use neurons. This is the full extent of the biological inspiration. Let's use neurons. And more detailed biological inspiration has been very hard to come by. But I wouldn't rule it out. I think if someone has a special insight, they might be able to to see something and that would be useful. I have a question for you about sort of auto correct. Um so here [00:18:17] is Here's the question. You mentioned reasoning as being one of the core aspects of maybe the modeling in the future and maybe a differentiator. Um what we saw in some of the poster sessions is that hallucinations in today's models the way we're analyzing I mean maybe you correct me, you're the expert on this, but the way we're analyzing whether a model is hallucinating today without because we know of the dangers of models not being able to reason that we're [00:18:47] using a statistical analysis, let's say some amount of standard deviations or whatever away from the mean. In the future, wouldn't it Do you think that a model given reasoning will be able to correct itself, sort of auto correct itself? And that will be a core feature of future models so that there won't be as many hallucinations because the model will recognize when I maybe that's too esoteric of a question, but the model will be able to reason and understand when a hallucination is occurring. Does the question make sense? [00:19:18] Yes, and the answer is also yes. I think what you described is extremely highly plausible. Yeah, I mean you should check. I mean for yeah, it I wouldn't I wouldn't rule out that it might already be happening with some of the, you know, early reasoning models of today. I don't know. But longer term, why not? Yeah, I mean part part of like Microsoft Word like autocorrect. It's a you know, it's a it's a core feature. Yeah, I just I mean I think calling it autocorrect is really doing a [00:19:48] disservice. I think you are When you say autocorrect, you evoke like it's far grander than autocorrect, but other than but, you know, this point aside, the answer is yes. Thank you. Hi Ilya. I loved the ending mysteriously leaving out do they replace us or are they, you know, superior? Do they need rights? You know, it's a new species of Homo sapien [00:20:18] spawned intelligence. So maybe they need I mean I think the RL guy thinks they think you know, we need rights for these things. I have a unrelated question to that. How do you how do you create the right incentive mechanisms for humanity to actually create it in a way that gives it the freedoms that we have as Homo sapiens? You know, I feel like this In some In [00:20:48] some In some sense, those are those are the kind of questions that people should be reflecting on more. But to your question about what incentive structure should we create, I I don't feel that I know. I don't feel confident answering questions like this because uh it's like you're talking about creating some kind of a [00:21:19] top-down structure government thing, I don't know. It could be a cryptocurrency, too. Yeah. I mean There's Bit Tensor, you know, there's things. I don't feel like I am the right person to comment on cryptocurrency. But you know, there is a chance, by the way, what the what you're describing will happen. That indeed we will have, you know, in some sense it's it's it's not a bad [00:21:49] end result if you have AIs and all they want is to coexist with us and also just to have rights, maybe that will be fine. It's But I don't know. I mean, I think things are so incredibly unpredictable. I I hesitate to comment, but I encourage the speculation. Thank you. Uh and uh yeah, thank you for the talk. It's really awesome. Hi Anna. Thank you for the great talk. Uh my name is Shalev Lifschitz from University of Toronto. [00:22:19] Uh working with Sheila. Thanks for all the work you've done. I wanted to ask do you think LLMs generalize multi-hop reasoning out of distribution? So, okay. The question assumes that the answer is yes or no, but the question should not be answered with a yes or no. Because what does it mean out of distribution generalization? What does it mean? What does it mean in distribution? And what does it mean out [00:22:49] of distribution? Because it's a test of time talk, I'll say that long, long ago, before people were using deep learning, they were using things like string matching and grams for machine translation. People were using statistical phrase tables. Can you imagine? They had tens of thousands of code of complexity which was I mean it's it was truly unfathomable. And [00:23:19] back then generalization meant is it literally not in this is the same word phrasing as in the data set. Now we may say, "Well, sure my model achieves this high score on I don't know math competitions, but maybe the math maybe some discussion in some forum on the internet was about the same ideas and therefore it's memorized." Well, okay, you could say maybe it's in distribution, maybe it's memorization. But I also think that our standards for what counts as generalization have increased really [00:23:50] quite substantially, dramatically, unimaginably if you keep track. And so I think the answer is to some degree probably not as well as human beings. I think it is true that human beings generalize much better. But at the same time they definitely generalize out of distribution to some degree. I hope it's a useful tautological answer. Thank you. And [00:24:20] unfortunately, we're out of time for this session. I have a feeling we could go on for the next 6 hours. But thank you so much Ilya for the talk.