At a glance
Andrej Karpathy walks on stage at Y Combinator's AI Startup School, looks at a room full of bachelors, masters and PhD students about to enter the industry, and tells them the thing they most need to hear: software has not changed much on a fundamental level for 70 years, and then it changed twice in the last few years. Software 1.0 is the code you write. Software 2.0 is the weights of a neural network, which you do not write but rather produce by curating datasets and running an optimizer. Software 3.0 is the prompt, a program written in English and executed by a large language model. He wants everyone in that room fluent in all three, because real systems move between them.
The first half of the talk is about what this new computer actually is. He tries three analogies in sequence, each one honestly, each one with its limits named: LLMs as utilities, LLMs as fabs, and LLMs as operating systems. The third one is the one he thinks is right, and he pushes it until it yields a date. The LLM is the CPU, the context window is memory, tool use is the peripherals, and we are all thin clients talking to expensive centralized compute over a network under a time sharing scheme. That is the 1960s. The personal computing revolution for this technology has not happened, and nobody in the room knows what it looks like. Along the way he flags the one thing about this technology that is unprecedented: the direction of diffusion flipped, and consumers got it before governments and corporations did.
The second half is what to build. Partial autonomy products with an autonomy slider, built so the generation and verification loop between the model and the human spins as fast as possible. Make verification visual, because a GUI runs on the vision hardware in your head and reading text does not. Keep generation small, because a ten thousand line diff makes you the bottleneck again. Keep the AI on the leash. Build Iron Man suits rather than Iron Man robots. And start writing your documentation, your docs sites and your repositories for a brand new kind of reader: an agent, which is a computer that behaves like a person, and which cannot click a button.
He is extremely quotable throughout, and he is also funny about his own failures, including the week he lost to authentication and payments and the free credits that turned his own vibe coded app into "a major cost center in my life."
The talk, rebuilt
"Software is changing again," and he means again (0:00:00)
He is introduced as the former director of AI at Tesla, looks at the crowd, and says "Wow, a lot of people here." Then the setup: many of you are students, you are about to enter the industry, and this is an extremely unique and very interesting time to do it. Why? Because software is changing. Again.
The "again" is a joke on himself. "I say again because I actually gave this talk already. Um, but the problem is that software keeps changing. So I actually have a lot of material to create new talks." The serious version of the same sentence is the thesis of the whole keynote: software has not changed on such a fundamental level for 70 years, and then it has changed about twice, quite rapidly, in the last few years. Which means there is a huge amount of work to do, and a huge amount of software to write and rewrite.
He starts where you would start if you wanted to show someone all the software there is. The slide is Map of GitHub, a tool that lays out public repositories as a geography you can zoom into. These are the instructions to the computer for carrying out tasks in the digital space. Zoom in and you see the repositories, and all the code that has ever been written.
Software 1.0, Software 2.0, and now 3.0 (0:01:25)
A few years ago, looking at that map, he noticed that a new type of software was around, and he gave it a name. Software 1.0 is the code you write for the computer. Software 2.0 is neural networks, and in particular the weights of a neural network. You do not write that code directly. You tune the datasets and then you run an optimizer to create the parameters.
He is candid that the framing landed differently then than it does now. At the time neural nets were seen as just a different kind of classifier, like a decision tree. The framing fit the moment less well than it fits today.
What makes it real now is that Software 2.0 grew its own infrastructure, in exact parallel to the first era. GitHub is where Software 1.0 lives. Hugging Face is, in his words, "basically equivalent of GitHub in software 2.0." And there is a visualization for it too, the Model Atlas (paper, code), which is the Map of GitHub of model space. He points at the giant circle sitting at the center of it: those are the parameters of FLUX, the image generator. Anytime somebody fine tunes on top of a FLUX model, "you basically create a git commit in this space," and you get a different image generator out the other end. Same social graph, different artifact.
So: Software 1.0 is computer code that programs a computer. Software 2.0 is the weights that program neural networks. His example of the second one on screen is AlexNet, the image recognizer.
Then the pivot that gives the talk its title. Every neural network we had been familiar with until recently was a fixed function computer. Image in, categories out. What changed, and he thinks it is a quite fundamental change, is that neural networks became programmable, with large language models. "It's a new kind of a computer and so in my mind it's worth giving it a new designation of software 3.0. And basically your prompts are now programs that program the LLM."
And the programming language is English. He says it twice, because it still strikes him as remarkable: these programs are written in our native language.
His summary slide is sentiment classification done three ways. You can write some amount of Python to do sentiment classification. You can train a neural net to do it. Or you can prompt a large language model to do it, and the example on screen is a few shot prompt. Same task, three eras, and you program the computer "in a slightly different way" in each. He also notes what you can already see on GitHub: a lot of code there is not just code anymore, there is a bunch of English interspersed with it, which is a growing category of a new kind of code.
| Software 1.0 | Software 2.0 | Software 3.0 | |
|---|---|---|---|
| What the program is | Source code: instructions to the computer for carrying out tasks in the digital space | The weights of a neural network | The prompt |
| Who or what writes it | A person, line by line | An optimizer. The person curates the dataset | Anyone who can write English. "Suddenly everyone is a programmer" |
| How you change its behavior | Edit the logic | Change the data and run the optimizer again | Rewrite the instruction |
| Where it lives | GitHub, mapped by Map of GitHub | Hugging Face, mapped by the Model Atlas | Increasingly in the repository, as English interspersed with code |
| His example on screen | Python that classifies sentiment | AlexNet, and a trained sentiment classifier | A few shot prompt that classifies sentiment |
| Degrees of freedom | Total, and you own every bug | A fixed function computer: image in, categories out | Programmable in natural language |
| Barrier to entry | Five to ten years of study | Machine learning expertise plus data | You already speak the language |
Programming computers in English, and the C++ that got deleted (0:04:40)
"Not only is it a new programming paradigm, it's also remarkable to me that it's in our native language of English." When this blew his mind a few years ago he tweeted it, the tweet caught a lot of attention, and as of this talk it is still his pinned tweet: remarkably, we are now programming computers in English.
Then he reaches back to Tesla for the proof that a paradigm really can eat a stack. At Tesla they were working on Autopilot, trying to get the car to drive. The slide he showed at the time put the inputs to the car at the bottom, running up through a software stack to produce steering and acceleration. His observation then: there was a ton of C++ code in the Autopilot, which was the Software 1.0 code, and there were some neural nets in there doing image recognition.
What happened over time, as they made Autopilot better, is the part worth remembering. The neural network grew in capability and size, and the C++ code was deleted. Capability and functionality that had originally been written in 1.0 migrated to 2.0. His concrete example: all the stitching up of information across images from the different cameras and across time, which used to be explicit code, got done by a neural network instead, and they were able to delete a lot of code. "The software 2.0 stack quite literally ate through the software stack of the autopilot."
He thought that was remarkable then. He thinks we are watching the same thing again now, one layer up, with a new kind of software eating through the stack.
Which produces his first piece of advice to the room, and it is not "learn prompting." It is the opposite of specialization:
We have three completely different programming paradigms, and I think if you're entering the industry it's a very good idea to be fluent in all of them, because they all have slight pros and cons, and you may want to program some functionality in 1.0 or 2.0 or 3.0. Are you going to train a neural net? Are you going to just prompt an LLM? Should this be a piece of code that's explicit?
Those are decisions somebody has to make, and he thinks you should be able to transition fluidly between the three.
Analogy one: LLMs are utilities (0:06:10)
The first half of the talk proper is a single question. What is this new computer, and what does its ecosystem look like? He answers it by trying analogies and naming where each one breaks.
He starts from a line he says struck him years ago, from Andrew Ng, who he notes is speaking right after him at the same event: AI is the new electricity. He thinks it captures something real, because LLMs certainly feel like they have properties of utilities right now.
Work the analogy and it holds up in detail:
- The labs, and he names OpenAI, Gemini and Anthropic, spend capex to train the models. That is the equivalent of building out the grid.
- Then there is opex to serve that intelligence over APIs to everybody.
- Delivery is metered access. We pay per million tokens.
- And the demands we make of that API are utility demands, not software demands: low latency, high uptime, consistent quality.
The analogy even extends to the hardware of switching. In electricity you have a transfer switch, so you can move your source between grid, solar, battery and generator. For LLMs we have OpenRouter, and you can easily switch between the different models that exist.
Then he names the place the electricity analogy is actually better than the real thing. Because LLMs are software, they do not compete for physical space. "So it's okay to have basically like six electricity providers and you can switch between them, right? Because they don't compete in such a direct way." You cannot run six power grids into one house. You can route one application across six model providers without anybody digging up the street.
And then the line from this section that got quoted everywhere, which he delivers off the back of something that had happened days before the talk, when several of the major models went down at once and people found themselves unable to work:
When the state of the art LLMs go down, it's actually kind of like an intelligence brownout in the world. It's kind of like when the voltage is unreliable in the grid, and the planet just gets dumber.
He adds the part that makes it more than a joke: the effect scales with how much we rely on these models, which is already dramatic and which he expects to keep growing.
Analogy two: LLMs are fabs
The utility framing misses the capital intensity, so he tries a second one. LLMs have properties of semiconductor fabs, and the reason is the size of the capex. "It's not just like building some power station or something like that. You're investing a huge amount of money." On top of that the tech tree for the technology is growing quite rapidly, which puts us in a world of deep tech trees and research and development secrets centralizing inside the labs.
He names the limit of this analogy himself, immediately, and it is the sharpest thing in the section: "the analogy muddies a little bit also because, as I mentioned, this is software, and software is a bit less defensible because it is so malleable." A fab is a building full of machines nobody else can buy. A model is a file.
But the mapping is fun where it works, and he runs it:
- A 4 nanometer process node is something like a cluster with a certain max FLOPS.
- If you are using NVIDIA GPUs and doing only the software, not the hardware, that is the fabless model.
- If you are also building your own hardware and training on TPUs because you are Google, that is the Intel model. You own your fab.
Analogy three, the one he believes: LLMs are operating systems
"Actually I think the analogy that makes the most sense perhaps is that in my mind LLMs have very strong kind of analogies to operating systems."
The argument against the utility framing is that this is "not just electricity or water. It's not something that comes out of the tap as a commodity." These are increasingly complex software ecosystems. Not fungible units.
And once you look at the industry structure, the shape is familiar. You have a few closed source providers, which is Windows and macOS, and you have an open source alternative, which is Linux. For LLMs there are a few competing closed source providers, and the Llama ecosystem is, in his careful phrasing, "currently like maybe a close approximation to something that may grow into something like Linux." He flags that it is still very early, because these are just simple LLMs today, and the complexity is coming: it is not just about the model itself, it is about all the tool use and the multimodalities and how all of that works together.
Then he pushes the analogy down into the architecture, which is the part of this talk that ended up on the most slide decks:
- The LLM is the CPU. It is a new kind of computer, and the model sits where the processor sits.
- The context window is the memory. It is what the processor can directly address.
- The LLM orchestrates memory and compute for problem solving, using its available capabilities, which is exactly the job description of an operating system.
And the app distribution story matches too. If you want to download an app, say you go to VS Code and download it, you can run it on Windows, Linux or Mac. In the same way you can take an LLM app like Cursor and run it on GPT or Claude or Gemini. "Right? It's just a drop down." Portable applications over swappable kernels, selected from a menu.
The 1960s of LLMs, and why there is no personal computer yet (0:11:04)
The analogy keeps paying out, and the next thing it gives him is a date.
LLM compute is still very expensive for this new kind of computer. Expensive compute forces centralization, so the models live in the cloud. We are all thin clients interacting with them over the network. None of us has full utilization of these computers, so it makes sense to run them under time sharing, and the phrase he uses for what you and I are in that scheme is worth sitting with: "we're all just, you know, a dimension of the batch when they're running the computer in the cloud."
That is not a metaphor for the 1960s. That is a description of the 1960s. "This is very much what computers used to look like during this time. The operating systems were in the cloud. Everything was streamed around and there was batching."
So the personal computing revolution for this technology has not happened yet, because it is not economical. It does not make sense yet. But he notes that some people are trying, and he points at the one piece of consumer hardware that turns out to be a surprisingly good fit: Mac minis. The reason is specific and technical. If you are doing batch one inference, the workload is entirely memory bandwidth bound rather than compute bound, which is exactly the regime a machine with fast unified memory and modest compute is good at. "So this actually works." He calls these early indications, maybe, of personal computing, and then refuses to predict what it turns into: "it's not clear what this looks like. Maybe some of you get to invent what this is or how it works."
Then one more analogy, which is the one that explains why so much of the second half of the talk is about interfaces. Whenever he talks to ChatGPT or any LLM directly in text, he feels like he is talking to an operating system through the terminal. "It's just text. It's direct access to the operating system."
And a GUI has not been invented. Not in a general way. He is careful about this, because the obvious objection is that plenty of LLM apps have graphical interfaces. His distinction: "should ChatGPT have a GUI, like different than just the text bubbles? Certainly some of the apps that we're going to go into in a bit have GUIs, but there's no GUI across all the tasks, if that makes sense." A terminal with nicer bubbles is still a terminal. The general graphical shell for this computer is an open problem, and he leaves it open on purpose, in front of a room of people who might build it.
The one thing that is genuinely unprecedented: diffusion ran backwards (0:12:49)
Here he stops cataloguing similarities and names the single property of LLMs that does not match early computing at all, something he had written about separately because it struck him as so different: LLMs flipped the direction of technology diffusion.
The usual pattern is unmistakable once you list it. Electricity, cryptography, computing, flight, the internet, GPS. Transformative technologies, and in every case the first users were governments and corporations, because the thing was new and expensive, and only later did it diffuse to consumers.
LLMs ran the other way. With early computers it was all about ballistics and military use. With LLMs, in his words, "it's all about how do you boil an egg or something like that." And then, with evident delight at the absurdity of it:
It's really fascinating to me that we have a new magical computer and it's like helping me boil an egg. It's not helping the government do something really crazy like some military ballistics or some special technology.
He adds the structural half of the observation: corporations and governments are lagging behind the adoption of all of us. It is just backwards. And he thinks that fact informs where the first apps are and how we want to use the technology, which is a strategic point disguised as a joke about eggs.
His summary of the whole first half, close to verbatim, is a stack of four claims:
- LLMs are complicated operating systems, and "LLM labs" is accurate language for what the providers are.
- They are circa 1960s in computing terms, and we are redoing computing all over again.
- They are currently available via time sharing and distributed like a utility.
- What is new and unprecedented is that they are not in the hands of a few governments and corporations. They are in the hands of all of us, because we all already have a computer and it is all just software.
And then the line that lands the section, about ChatGPT arriving on everyone's devices at once: "it was beamed down to our computers, like billions of people, like instantly and overnight, and this is insane." He means that literally, and he turns it immediately into the invitation that the rest of the talk answers. "Now it is our time to enter the industry and program these computers. This is crazy."
The psychology of LLMs: people spirits (0:14:39)
Before you program something, he argues, you should know what it is like. So he spends a section on the psychology of these models, and the framing he gives them is the one people still quote:
The way I like to think about LLMs is that they're kind of like people spirits. They are stochastic simulations of people.
The mechanics underneath the metaphor are stated plainly. The simulator in this case is an autoregressive transformer. It is a neural net, it operates at the level of tokens, and it goes "chunk chunk chunk chunk," with an almost equal amount of compute spent on every single chunk. There are weights, and we fit those weights to all of the text we have on the internet. Because it was fit to humans, what comes out the other side has an emergent psychology that is humanlike. Not designed. Emergent.
Then he does the thing the rest of the talk depends on: he describes that psychology honestly, as a profile with both ends.
The superpower first. LLMs have encyclopedic knowledge and memory, and they can remember a lot more things than any single individual human can, because they read so many things. His reference here is Rain Man, which he interrupts himself to recommend: "I actually really recommend people watch. It's an amazing movie. I love this movie." Dustin Hoffman plays an autistic savant with almost perfect memory, who can read a phone book and remember all of the names and phone numbers. LLMs are very similar. "They can remember SHA hashes and lots of different kinds of things very, very easily." Superpowers, in some respects.
Then the cognitive deficits, and he lists them as deficits, not as temporary bugs:
- Hallucination. They make up stuff, and they do not have a very good internal model of self knowledge, or at least not a sufficient one. He notes this has gotten better, but not perfect.
- Jagged intelligence. This is the term that stuck. "They're going to be superhuman in some problem solving domains, and then they're going to make mistakes that basically no human will make." His two examples are the famous ones: a model insisting that 9.11 is greater than 9.9, or that there are two Rs in "strawberry." The point is not the specific failures, it is the shape of the surface. "There are rough edges that you can trip on."
- Anterograde amnesia. This one gets the longest treatment because it is the one that changes how you build. Think about a coworker who joins your organization. Over time that coworker learns the organization, gains a huge amount of context on it, goes home, sleeps, consolidates knowledge, and develops expertise. "LLMs don't natively do this, and this is not something that has really been solved in the R and D of LLMs." Which means the context window is not long term memory, it is working memory, "and you have to sort of program the working memory quite directly, because they don't just kind of like get smarter by default." He thinks a lot of people get tripped up by exactly this analogy.
- Gullibility, which is a security surface. LLMs are quite gullible and susceptible to prompt injection risks. They might leak your data. He gestures at the fact that there are many other security considerations beyond that one.
For the memory problem he has a second pair of movie references, and they are better than they first sound: Memento and 50 First Dates. "In both of these movies, the protagonists, their weights are fixed and their context windows get wiped every single morning, and it's really problematic to go to work or have relationships when this happens." The joke is funny and the mapping is exact, which is why it works. The weights are the person. The context window is the day.
The job he hands the room at the end of the section is to hold both halves at once:
You have to simultaneously think through this superhuman thing that has a bunch of cognitive deficits and issues, and yet they are extremely useful. So how do we program them, and how do we work around their deficits and enjoy their superhuman powers?
| The profile | What it gives you | What it costs you | What he says to do about it |
|---|---|---|---|
| Memory of the world | Encyclopedic knowledge, more than any individual human. Remembers SHA hashes. His reference: Rain Man | Hallucination, and insufficient self knowledge | Treat it as a savant, not an oracle. Verify |
| Capability surface | Superhuman in some problem solving domains | Jagged intelligence: 9.11 > 9.9, two Rs in "strawberry" | Expect rough edges you can trip on. There is no clean boundary to learn |
| Memory of you | A context window you can program directly | Anterograde amnesia. No consolidation, no expertise over time. His reference: Memento and 50 First Dates | Program the working memory yourself. Do not expect a coworker who learns the org |
| Instruction following | Programmable in English by anyone | Gullible. Prompt injection, data leakage | Treat it as a security surface, not a trust relationship |
Partial autonomy apps, and the anatomy of a good one (0:18:22)
He switches to opportunities, and he is explicit that what follows is not a comprehensive list, just the things he found interesting enough for this talk. The first one is what he calls partial autonomy apps.
He makes the case by asking why anybody would do the obvious dumb thing. Take coding. You can go to ChatGPT directly and start copy pasting code around, and copy pasting bug reports around, and copy pasting everything around. "Why would you do that? Why would you go directly to the operating system?" It makes a lot more sense to have an app dedicated to this, and many people in the room use Cursor. So does he.
Cursor is his worked example of an early LLM app, and he pulls four properties out of it that he thinks generalize across all LLM apps:
- The traditional interface survives. You notice first that there is still an interface that lets a human go in and do all the work manually, just as before. The LLM integration is in addition to that, and what it buys you is the ability to go in bigger chunks.
- The app does a ton of the context management. You are not the one assembling what the model sees.
- It orchestrates multiple calls to multiple models. In Cursor's case, under the hood, there are embedding models for all your files, the actual chat models, and models that apply diffs to the code. All of that is orchestrated for you.
- An application specific GUI, and he thinks its importance is underappreciated. The reason is not aesthetics, it is bandwidth, and he states it concretely. You do not want to talk to the operating system directly in text, because "text is very hard to read, interpret, understand," and because some of these actions should not be taken in text at all. "It's much better to just see a diff as like red and green change, and you can see what's being added, what's subtracted. It's much easier to just do command Y to accept or command N to reject. I shouldn't have to type it in text, right?" The function of the GUI is auditing: it "allows a human to audit the work of these fallible systems and to go faster."
- The autonomy slider. This is the one he names as a design primitive, and Cursor's version of it is four rungs you can point at: tab completion, where you are mostly in charge; command K to change a selected chunk of code; command L to change the entire file; and command I to "just let it rip, do whatever you want in the entire repo," which is the full autonomy agentic version. "You are in charge of the autonomy slider, and depending on the complexity of the task at hand you can tune the amount of autonomy that you're willing to give up for that task."
Then he shows the same anatomy in a product that has nothing to do with code, which is how you know it is a pattern and not a coding tool convention. Perplexity packages up a lot of the information, orchestrates multiple LLMs, has a GUI that lets you audit its work (it cites sources, and you can imagine inspecting them), and has an autonomy slider of its own with three rungs: quick search, research, or deep research, where you "come back 10 minutes later." Varying levels of autonomy you give up to the tool.
Then he turns the examples into homework for the room, and these four questions are the most directly actionable thing in the talk:
I feel like a lot of software will become partially autonomous. For many of you who maintain products and services, how are you going to make your products and services partially autonomous? Can an LLM see everything that a human can see? Can an LLM act in all the ways that a human could act? And can humans supervise and stay in the loop of this activity?
The reason the third question matters is the one he keeps returning to: "these are fallible systems that aren't yet perfect."
He also asks the awkward version of the question, the one that shows the problem is real and not solved: "What does a diff look like in Photoshop or something like that?" Red and green lines work for text. For a layered image edit, nobody knows what the auditable representation is. And he points at the sheer volume of retrofitting implied: "a lot of the traditional software right now, it has all these switches and all this kind of stuff that's all designed for humans. All of this has to change and become accessible to LLMs."
The generation and verification loop, and keeping the AI on the leash (0:23:40)
This is the design argument of the talk, and he says up front that he is not sure it gets as much attention as it should.
The structure of the work has changed shape. "We're now kind of cooperating with AIs, and usually they are doing the generation and we as humans are doing the verification." Which makes the whole system a loop with two participants, and the throughput of a loop is set by its slower half. "It is in our interest to make this loop go as fast as possible, so we're getting a lot of work done."
There are two ways to do that, and he numbers them.
Number one: speed up verification a lot. GUIs are extremely important to this, and the reason he gives is physiological rather than aesthetic. A GUI "utilizes your computer vision GPU in all of our head." Then the line:
Reading text is effortful and it's not fun, but looking at stuff is fun, and it's just a kind of like a highway to your brain.
So GUIs are very useful for auditing systems, and visual representations in general are the lever.
Number two: keep the AI on the leash. He thinks a lot of people are getting way over excited with AI agents, and the counterargument is arithmetic, not taste:
It's not useful to me to get a diff of 10,000 lines of code to my repo. Like, I'm still the bottleneck, right? Even though that 10,000 lines come out instantly, I have to make sure that this thing is not introducing bugs, and that it's doing the correct thing, and that there's no security issues.
An instant diff you cannot verify has not saved you time. It has moved the time somewhere less pleasant. He admits the slide for this part is not very good and apologizes for it, which is a nice moment, and then says the honest version of what he is doing: like many people in the room, he is still developing ways of using these agents in his own coding workflow.
His own practice, stated plainly, is three rules:
- He is always scared of getting way too big diffs.
- He always goes in small incremental chunks, making sure everything is good at each step.
- He wants to spin the loop very, very fast, working on a small chunk of a single concrete thing.
He also distinguishes the two modes he works in, and this is the distinction most of the discourse around this talk flattened. "If I'm just vibe coding, everything is nice and great. But if I'm actually trying to get work done, it's not so great to have an overreactive agent doing all this kind of stuff." Vibe coding and getting work done are different activities with different tolerances, and he applies the leash to the second one.
Then he points at a blog post he had read recently and thought was quite good, which develops best practices for working with LLMs, several of them about keeping the AI on the leash. The one he draws out is a causal chain worth memorizing, because it explains why the apparently slower approach is faster:
- Your prompt is vague.
- So the AI does not do exactly what you wanted.
- So verification fails.
- So you ask for something else, and now you are spinning.
"So it makes a lot more sense to spend a bit more time to be more concrete in your prompts, which increases the probability of successful verification, and you can move forward." Concreteness is not politeness toward the model. It is loop throughput.
He closes the section with his current side interest, which is education, and it is the clearest worked example of the leash in the whole talk. He does not think it works to go to ChatGPT and say "hey, teach me physics." "I don't think this works, because the AI is like, gets lost in the woods."
So for him it is two separate apps. There is an app for a teacher, which creates courses. And there is an app that takes courses and serves them to students. The design win is the thing in between: "we now have this intermediate artifact of a course that is auditable, and we can make sure it's good, we can make sure it's consistent, and the AI is kept on the leash with respect to a certain syllabus, a certain progression of projects." Put a reviewable artifact between the model and the Customer, and the model stops wandering.
Five years of Autopilot, one perfect 2013 demo, and the decade of agents (0:26:00)
"I'm no stranger to partial autonomy," he says, and the credential is five years at Tesla working on exactly this. Autopilot is a partial autonomy product and it shares a lot of the features he has been describing. Right there in the instrument panel is the GUI of the Autopilot, showing the driver what the neural network sees. And there was an autonomy slider: over the course of his tenure, they did "more and more autonomous tasks for the user."
Then he tells the story that is the single best argument in the talk against agent timelines, and he tells it against his own instincts at the time.
The first time he ever drove in a self driving vehicle was 2013. A friend who worked at Waymo offered to give him a drive around Palo Alto. He took a picture of it using Google Glass, which he notes "many of you are so young that you might not even know what that is," and which was all the rage at the time. They got in the car and went for about a 30 minute drive around Palo Alto highways and streets.
"And this drive was perfect. There was zero interventions. And this was 2013, which is now 12 years ago."
The honest part is what he concluded in the moment: "when I had this perfect drive, this perfect demo, I felt like, wow, self driving is imminent, because this just worked. This is incredible."
And then the 12 years:
Here we are 12 years later and we are still working on autonomy. We are still working on driving agents, and even now we haven't actually really solved the problem. Like, you may see Waymos going around and they look driverless, but there's still a lot of teleoperation and a lot of human in the loop of a lot of this driving. So we still haven't even declared success.
He is clear that he thinks it will succeed at this point. The claim is not that it fails, it is that it took a long time, and that a flawless demo told him almost nothing about how long. Which is the setup for the correction he actually came to deliver:
When I see things like "oh, 2025 is the year of agents," I get very concerned, and I kind of feel like, you know, this is the decade of agents. And this is going to be quite some time. We need humans in the loop. We need to do this carefully. This is software. Let's be serious here.
"This is software, let's be serious here" is the whole talk compressed into seven words. Software is tricky in the same way driving is tricky, and he has the receipts for both.
The Iron Man suit, not the Iron Man robot (0:27:52)
The last analogy, and the one he says he always thinks through, is the Iron Man suit. "I always love Iron Man. I think it's so correct in a bunch of ways with respect to technology and how it will play out."
What he loves about the suit specifically is that it is both things at once. It is an augmentation, and Tony Stark can drive it. And it is an agent: in some of the movies the suit is quite autonomous, flies around on its own, and goes and finds Tony. That duality is the autonomy slider. We can build augmentations or we can build agents, and we want to do a bit of both.
But at this stage, working with fallible LLMs, he gives the room a direction rather than a balance:
It's less Iron Man robots and more Iron Man suits that you want to build. It's less like building flashy demos of autonomous agents and more building partial autonomy products.
And these products, he says, have custom GUIs and UI/UX, and the reason they do is the loop: it is "done so that the generation verification loop of the human is very, very fast." But without losing sight of the fact that it is in principle possible to automate the work. So: there should be an autonomy slider in your product, and you should be thinking about how you can slide it and make your product more autonomous over time.
Vibe coding: everyone is now a programmer (0:29:06)
He switches gears to a dimension he thinks is genuinely unique. It is not just that there is a new programming paradigm that allows for autonomy in software. It is that it is programmed in English, which is a natural interface, "and suddenly everyone is a programmer, because everyone speaks natural language like English."
He calls this extremely bullish, very interesting, and completely unprecedented, and he quantifies the change by what it replaced: "it used to be the case that you need to spend five to 10 years studying something to be able to do something in software. This is not the case anymore."
Then, with perfect deadpan: "I don't know if by any chance anyone has heard of vibe coding."
The story he tells about the tweet is the most human stretch of the talk, and it is also a lesson about virality that has nothing to do with AI. He has been on Twitter for about 15 years at this point, "and I still have no clue which tweet will become viral and which tweet fizzles and no one cares." He thought this one would fizzle. "It was just like a shower of thoughts." Instead it became a total meme. "But I guess it struck a chord and it gave a name to something that everyone was feeling but couldn't quite say in words." And now there is a Wikipedia page and everything, which gets applause from the room, and which he receives with "yeah, this is like a major contribution now or something like that."
Then two things he loves and one thing he lost money on.
The kids. Tom Wolf of Hugging Face shared a video that Karpathy says he really loves: kids vibe coding. "I find that this is such a wholesome video. Like, how can you look at this video and feel bad about the future? The future is great." His read on it is a real prediction and not just sentiment: "I think this will end up being like a gateway drug to software development." And he puts his own position on the record: "I'm not a doomer about the future of the generation."
The iOS app. He tried vibe coding himself, because it is fun, and because it is the right tool "when you want to build something super duper custom that doesn't appear to exist and you just want to wing it because it's a Saturday." So he built an iOS app, and the detail that matters is that he cannot program in Swift. "I was really shocked that I was able to build like a super basic app." He declines to explain it, calls it really dumb, and gives the number that is the point: it was "just like a day of work," and it was running on his phone later that day. "I didn't have to read through Swift for like five days to get started."
MenuGen. The second one is live, and he tells the room they can try it. The problem it solves is his own: he shows up at a restaurant, reads through the menu, and has no idea what any of the things are. He needs pictures. That product did not exist, so he vibe coded it. You go to the site, you take a picture of a menu, and it generates the images for the dishes. Everyone gets five dollars in credits for free when they sign up.
Which leads to the best line in this section: "therefore, this is a major cost center in my life. So this is a negative revenue app for me right now. I've lost a huge amount of money on MenuGen."
And then the lesson, which is the real payload of the whole section and which he clearly wants the room to take seriously, because it is the reason the last part of the talk exists:
The fascinating thing about MenuGen for me is that the code, the vibe coding part, the code was actually the easy part. And most of it actually was when I tried to make it real, so that you can actually have authentication and payments and the domain name and Vercel deployment. This was really hard. And all of this was not code. All of this DevOps stuff was me in the browser clicking stuff, and this was extremely slow and took another week.
The demo worked on his laptop in a few hours. Making it real took a week. And the reason was not difficulty, it was annoyance.
His example on screen is adding Google login, and the Clerk integration instructions. He apologizes that the text is small on the slide, but the shape of it is the argument: a huge amount of step by step instructions telling him how to integrate this. "And this is crazy. Like it's telling me go to this URL, click on this dropdown, choose this, go to this, and click on that. And it's like telling me what to do. Like, a computer is telling me the actions I should be taking. Like, you do it. Why am I doing this? What the hell? I had to follow all these instructions. This was crazy."
Which is the hinge of the talk: "So I think the last part of my talk therefore focuses on, can we just build for agents? I don't want to do this work. Can agents do this?"
Building for agents: a new consumer of digital information (0:33:39)
The framing he opens with is a category claim, and it is the most immediately useful idea in the talk for anybody who maintains a product.
"There's a new category of consumer and manipulator of digital information. It used to be just humans through GUIs, or computers through APIs. And now we have a completely new thing." Agents. "They're computers, but they are humanlike kind of, right? They're people spirits. There's people spirits on the internet, and they need to interact with our software infrastructure."
Can we build for them? It is a new thing, so nobody has.
| Who is reading your product | How they arrive | What you already built for them | What they need that does not exist yet |
|---|---|---|---|
| Humans | A graphical interface | Everything. Buttons, dropdowns, screenshots, bold text, lists, pictures | Nothing. This is the case you have solved |
| Computers | An API | Documented endpoints, schemas, SDKs | Nothing. This is also solved |
| Agents | They read your site and your docs like a person, then act like a program | Nothing. Your docs say "click this button," and they cannot click | A file at a known location saying what this domain is. Markdown instead of HTML. Executable commands instead of click instructions. Repositories flattened into something ingestible |
His checklist for serving that third reader has five items, and each one comes with a company already doing it.
1. Tell the agent what your domain is, in a file. The precedent is robots.txt, which sits on your domain and instructs, or as he corrects himself, advises web crawlers on how to behave on your site. "In the same way you can have maybe llms.txt, a file which is just a simple markdown that's telling LLMs what this domain is about, and this is very readable to an LLM." The alternative is the status quo, and he is blunt about it: "if it had to instead get the HTML of your web page and try to parse it, this is very error prone and difficult and will screw it up and it's not going to work. So we can just directly speak to the LLM. It's worth it."
2. Serve your documentation as markdown. A huge amount of documentation is currently written for people, "so you will see things like lists and bold and pictures, and this is not directly accessible by an LLM." He names the early movers: Vercel and Stripe are already offering their documentation in markdown, and he says there are a few more he has seen. "Markdown is super easy for LLMs to understand. This is great."
3. Change the content of the docs, not just the format. This is the part people skip, and he is emphatic that the format is the easy half. "It's not just about taking your docs and making them appear in markdown. That's the easy part. We actually have to change the docs, because anytime your docs say 'click,' this is bad. An LLM will not be able to natively take this action right now." His example of someone doing the hard half: Vercel is replacing every occurrence of "click" with an equivalent curl command that your agent can run on your behalf.
4. Speak the protocol. "And then of course there's Model Context Protocol from Anthropic. And this is also another way, it's a protocol of speaking directly to agents as this new consumer and manipulator of digital information." He says he is very bullish on these ideas.
5. Make your repositories ingestible, and love the one URL trick. The other thing he really likes is the small tools that help ingest data in LLM friendly formats. His worked example is his own repository: when he goes to a GitHub repo like nanoGPT, he cannot feed that to an LLM and ask questions about it, because GitHub is a human interface. But change the URL from github.com to gitingest.com and "this will actually concatenate all the files into a single giant text, and it will create a directory structure," ready to be copy pasted into your favorite model. The more dramatic version is DeepWiki from Devin, which does not just dump the raw content: Devin does an analysis of the repository and builds whole documentation pages for it, which is even more helpful to paste into a model. "So I love all the little tools where you just change the URL and it makes something accessible to an LLM. This is all well and great, and I think there should be a lot more of it."
And in the middle of all that, the single most persuasive anecdote in the section, because it is a case where this already worked for him. He brings up 3Blue1Brown, who makes beautiful animation videos on YouTube, which gets applause. "Yeah, I love this library." The library is Manim (original repo), and Karpathy wanted to make his own animations. There is extensive documentation on how to use it, and he did not want to read it.
So I copy pasted the whole thing to an LLM and I described what I wanted, and it just worked out of the box. Like the LLM just vibe coded me an animation exactly what I wanted, and I was like, wow, this is amazing. So if we can make docs legible to LLMs, it's going to unlock a huge amount of use.
That is the proof. The docs were already good enough, in the sense that they were complete; what mattered was that they were pasteable.
He finishes the section with the honest counterargument, which he raises himself and then answers. It is absolutely possible that in the future LLMs will be able to go around and click things, and he immediately corrects himself: "this is not even future, this is today, they'll be able to go around and click stuff."
So why bother meeting them halfway? Two reasons. First, cost: clicking through interfaces is "still fairly expensive, I would say, to use, and a lot more difficult." Second, coverage: there will be a long tail of software that simply never adapts, because those are not "live player" repositories or pieces of digital infrastructure with anybody maintaining them, and for those we will need the click-capable agents and the URL rewriting tools. "But I think for everyone else, I think it's very worth kind of meeting in some middle point. So I'm bullish on both, if that makes sense."
Both halves of that answer, not one. The agents will learn to use human interfaces, and you should still publish markdown.
Summary: it is the 1960s, and it is time to build (0:38:14)
The closing is short and he delivers it as a stack, which is also a decent index of the talk:
- What an amazing time to get into the industry. We need to rewrite a ton of code, and a ton of code will be written by professionals and by coders.
- These LLMs are kind of like utilities, kind of like fabs, but especially like operating systems. And it is so early. "It's like the 1960s of operating systems," and a lot of the analogies cross over.
- They are fallible people spirits that we have to learn to work with, and in order to do that properly we need to adjust our infrastructure toward it.
- Build partial autonomy products, using the ways of working he described, and spin the generation and verification loop very quickly.
- A lot of code has to be written for the agents more directly.
And then the last thing he says, which is the Iron Man suit one more time, used as a forecast:
Going back to the Iron Man suit analogy, I think what we'll see over the next decade roughly is we're going to take the slider from left to right. And it's going to be very interesting to see what that looks like. And I can't wait to build it with all of you. Thank you.
Key takeaways
- Three paradigms, and the skill is fluency in all three. Software 1.0 is code, 2.0 is weights produced by an optimizer over a dataset you curated, 3.0 is a prompt written in English. His advice to people entering the industry is not to specialize in the newest one but to be able to decide, per piece of functionality, which paradigm it belongs in and to move between them fluidly.
- Software 2.0 already ate a production stack once. At Tesla, as Autopilot improved, the neural network grew and the C++ was deleted, including the multi camera and multi frame stitching that used to be explicit code. He expects 3.0 to do the same thing one layer up.
- Three analogies, and he names the limits of each. Utility captures the capex, opex, metered pricing and the demand for low latency and high uptime, plus the new failure mode of an intelligence brownout. Fab captures the capital intensity and the centralizing R and D secrets, but muddies because software is malleable and therefore less defensible. Operating system is the one he believes, because this is a complex software ecosystem rather than a commodity out of a tap.
- The OS mapping is literal and it dates the moment. Model as CPU, context window as directly addressed memory, tool use and multimodality as peripherals, provider as a swappable kernel under a portable app. Expensive centralized compute, thin clients, time sharing, and you are "a dimension of the batch." That is the 1960s, and the personal computing revolution has not happened, though batch one inference being memory bound makes a Mac mini a surprisingly good fit.
- No general GUI has been invented for this computer. Talking to a model in text is talking to an operating system through a terminal. Individual apps have interfaces, but there is no GUI across all the tasks, and he leaves that explicitly as an open invention.
- Diffusion ran backwards for the first time. Electricity, cryptography, computing, flight, the internet and GPS all reached governments and corporations first. LLMs reached billions of consumers first, and institutions are the laggards. He thinks that fact should inform where you look for the first applications.
- Design around the psychology, which has two ends. Encyclopedic savant memory on one side; hallucination, jagged intelligence, anterograde amnesia and prompt injection gullibility on the other. The context window is working memory you must program directly, not a colleague who gets smarter by going home and sleeping.
- Partial autonomy is the product shape, and the autonomy slider is the control. The anatomy, drawn from Cursor and Perplexity: keep the manual interface, manage the context for the user, orchestrate multiple models, give it an application specific GUI for auditing, and expose a slider the user tunes per task.
- The loop runs at the speed of its slower half, so reduce what the model does per cycle. Make verification visual, because a GUI runs on the vision hardware in your head. Keep generation small, because a ten thousand line instant diff puts you right back in the bottleneck. Be concrete in prompts, because vague prompts cause failed verification, which causes spinning.
- A perfect demo tells you nothing about the timeline. A flawless 30 minute zero intervention self driving ride in 2013 was followed by 12 more years of work that is still not finished. Hence "this is the decade of agents," not the year, and "this is software, let's be serious here."
- Build Iron Man suits rather than Iron Man robots, but build them knowing the suit is also an agent, and plan to move the slider rightward over the next decade.
- The demo is hours and the product is a week, and the week is not code. MenuGen worked on his laptop in a few hours; authentication, payments, the domain and the deployment took another week of clicking in a browser. That gap is the argument for the final section.
- There is a third reader of your software now, and it cannot click. Humans arrive through GUIs, computers through APIs, agents read like people and act like programs. Serve them: a markdown file at a known location saying what your domain is, docs in markdown, every "click this" replaced with a runnable command, a protocol like MCP, and repositories flattened into something pasteable.
Chapters
- 0:00:00 Intro
- 0:01:25 Software evolution: From 1.0 to 3.0
- 0:04:40 Programming in English: Rise of Software 3.0
- 0:06:10 LLMs as utilities, fabs, and operating systems
- 0:11:04 The new LLM OS and historical computing analogies
- 0:14:39 Psychology of LLMs: People spirits and cognitive quirks
- 0:18:22 Designing LLM apps with partial autonomy
- 0:23:40 The importance of human-AI collaboration loops
- 0:26:00 Lessons from Tesla Autopilot & autonomy sliders
- 0:27:52 The Iron Man analogy: Augmentation vs. agents
- 0:29:06 Vibe Coding: Everyone is now a programmer
- 0:33:39 Building for agents: Future-ready digital infrastructure
- 0:38:14 Summary: We're in the 1960s of LLMs, time to build
Notable quotes
Software has not changed much on such a fundamental level for 70 years, and then it's changed, I think, about twice quite rapidly in the last few years. And so there's just a huge amount of work to do, a huge amount of software to write and rewrite. Andrej Karpathy, the thesis of the whole talk, 1:02
It's a new kind of a computer, and so in my mind it's worth giving it a new designation of software 3.0. And basically your prompts are now programs that program the LLM. Andrej Karpathy, on why programmable neural networks deserve a new name, 3:05
Not only is it a new programming paradigm, it's also remarkable to me that it's in our native language of English. Andrej Karpathy, on the pinned tweet, 4:07
The software 2.0 stack quite literally ate through the software stack of the autopilot. Andrej Karpathy, on the C++ that got deleted at Tesla, 5:08
We have three completely different programming paradigms, and I think if you're entering the industry it's a very good idea to be fluent in all of them. Andrej Karpathy, his actual advice to the room, 5:39
When the state of the art LLMs go down, it's actually kind of like an intelligence brownout in the world. It's kind of like when the voltage is unreliable in the grid, and the planet just gets dumber. Andrej Karpathy, on the outage that happened days before the talk, 7:42
This is software, and software is a bit less defensible because it is so malleable. Andrej Karpathy, naming the limit of his own fab analogy, 8:12
This is not just electricity or water. It's not something that comes out of the tap as a commodity. These are now increasingly complex software ecosystems. Andrej Karpathy, on why the utility analogy is not enough, 9:15
The LLM is a new kind of a computer. It's kind of like the CPU equivalent. The context windows are kind of like the memory, and then the LLM is orchestrating memory and compute for problem solving. Andrej Karpathy, the operating system mapping, 10:15
You can take an LLM app like Cursor and you can run it on GPT or Claude or Gemini. It's just a drop down. Andrej Karpathy, on portable apps over swappable kernels, 10:46
We're all just, you know, a dimension of the batch when they're running the computer in the cloud. Andrej Karpathy, on what time sharing makes of us, 11:18
Whenever I talk to ChatGPT or some LLM directly in text, I feel like I'm talking to an operating system through the terminal. It's just text. It's direct access to the operating system. Andrej Karpathy, on the missing graphical interface, 11:48
It's really fascinating to me that we have a new magical computer and it's like helping me boil an egg. It's not helping the government do something really crazy like some military ballistics. Andrej Karpathy, on diffusion running backwards, 13:20
ChatGPT was beamed down to our computers, like billions of people, like instantly and overnight, and this is insane. Andrej Karpathy, on the one genuinely unprecedented thing, 14:20
The way I like to think about LLMs is that they're kind of like people spirits. They are stochastic simulations of people, and the simulator in this case happens to be an autoregressive transformer. Andrej Karpathy, the framing that outlived the talk, 14:50
They display jagged intelligence. So they're going to be superhuman in some problem solving domains, and then they're going to make mistakes that basically no human will make. Andrej Karpathy, on 9.11 being greater than 9.9, 16:21
In both of these movies, the protagonists, their weights are fixed and their context windows get wiped every single morning, and it's really problematic to go to work or have relationships when this happens. Andrej Karpathy, on Memento, 50 First Dates, and anterograde amnesia, 17:22
A GUI allows a human to audit the work of these fallible systems and to go faster. Andrej Karpathy, on why the interface is not cosmetic, 19:53
You are in charge of the autonomy slider, and depending on the complexity of the task at hand you can tune the amount of autonomy that you're willing to give up for that task. Andrej Karpathy, on Cursor's four rungs, 20:23
Reading text is effortful and it's not fun, but looking at stuff is fun, and it's just a kind of like a highway to your brain. Andrej Karpathy, on using the vision hardware in your head, 22:24
It's not useful to me to get a diff of 10,000 lines of code to my repo. I'm still the bottleneck, right? Even though that 10,000 lines come out instantly, I have to make sure that this thing is not introducing bugs. Andrej Karpathy, on keeping the AI on the leash, 22:56
If I'm just vibe coding, everything is nice and great. But if I'm actually trying to get work done, it's not so great to have an overreactive agent doing all this kind of stuff. Andrej Karpathy, on the two different modes, 23:28
I'm always scared to get way too big diffs. I always go in small incremental chunks. I want to make sure that everything is good. I want to spin this loop very, very fast. Andrej Karpathy, describing his own workflow, 23:58
I don't think it just works to go to ChatGPT and be like, "Hey, teach me physics." I don't think this works, because the AI gets lost in the woods. Andrej Karpathy, on why his education project is two apps and an auditable syllabus, 25:00
And this drive was perfect. There was zero interventions. And this was 2013, which is now 12 years ago. Andrej Karpathy, on the Waymo ride that convinced him self driving was imminent, 26:31
When I see things like "oh, 2025 is the year of agents," I get very concerned, and I kind of feel like this is the decade of agents. We need humans in the loop. We need to do this carefully. This is software. Let's be serious here. Andrej Karpathy, the correction he came to deliver, 27:34
It's less Iron Man robots and more Iron Man suits that you want to build. It's less like building flashy demos of autonomous agents and more building partial autonomy products. Andrej Karpathy, on what to build this year, 28:04
It used to be the case that you need to spend five to 10 years studying something to be able to do something in software. This is not the case anymore. Andrej Karpathy, on English as the interface, 29:06
I've been on Twitter for like 15 years at this point, and I still have no clue which tweet will become viral and which tweet fizzles and no one cares. I thought that this tweet was going to be the latter. Andrej Karpathy, on the tweet that coined vibe coding, 29:37
How can you look at this video and feel bad about the future? The future is great. I think this will end up being like a gateway drug to software development. Andrej Karpathy, on Tom Wolf's video of kids vibe coding, 30:42
Everyone gets $5 in credits for free when they sign up, and therefore this is a major cost center in my life. So this is a negative revenue app for me right now. I've lost a huge amount of money on MenuGen. Andrej Karpathy, on his own vibe coded product, 31:44
The code was actually the easy part, and most of it actually was when I tried to make it real, so that you can actually have authentication and payments and the domain name and Vercel deployment. This was really hard, and all of this was not code. Andrej Karpathy, on the week that followed the few hours, 32:16
It's telling me go to this URL, click on this dropdown, choose this, go to this, and click on that. A computer is telling me the actions I should be taking. Like, you do it. Why am I doing this? What the hell? Andrej Karpathy, on integrating Google login, and the hinge of the talk, 33:17
There's people spirits on the internet, and they need to interact with our software infrastructure. Can we build for them? Andrej Karpathy, naming the new consumer of digital information, 33:48
It's not just about taking your docs and making them appear in markdown. That's the easy part. We actually have to change the docs, because anytime your docs say "click," this is bad. Andrej Karpathy, on the half of the job everyone skips, 35:55
I copy pasted the whole thing to an LLM and I described what I wanted, and it just worked out of the box. The LLM just vibe coded me an animation exactly what I wanted. Andrej Karpathy, on pasting the entire Manim documentation into a model, 35:23
I love all the little tools where you just change the URL and it makes something accessible to an LLM. Andrej Karpathy, on gitingest and DeepWiki, 36:57
Going back to the Iron Man suit analogy, I think what we'll see over the next decade roughly is we're going to take the slider from left to right. And I can't wait to build it with all of you. Andrej Karpathy, the closing line, 39:00
Resources mentioned
The talk and the speaker
- Andrej Karpathy, his site, and his account on X where both of the tweets below live. Background: his Wikipedia entry, and he is introduced on stage as former director of AI at Tesla
- Y Combinator, its AI Startup School, the YC library page for this talk, and the YC YouTube channel
- Software 2.0, the 2017 essay where he coined the term he builds on here
- The pinned tweet: we are now programming computers in English
- The vibe coding tweet, the one he expected to fizzle, and the Wikipedia page it produced
- Andrew Ng, source of "AI is the new electricity," and the speaker scheduled immediately after him at the same event
The maps of the three eras
- Map of GitHub, the zoomable atlas of public repositories he opens with, and GitHub itself
- Hugging Face, his "equivalent of GitHub in software 2.0"
- Model Atlas, the visualization of model space, with its paper and code
- FLUX from Black Forest Labs, the image generator whose parameters sit at the center of that atlas
- AlexNet, his example of a Software 2.0 artifact, and Python, his example of a Software 1.0 one
The model providers and the ecosystem
- OpenAI, Google DeepMind's Gemini and Anthropic, the labs he names as spending the capex
- OpenRouter, his equivalent of an electrical transfer switch
- Llama, the candidate for the Linux role, against the closed source Windows and macOS of this analogy
- NVIDIA GPUs (the fabless model), Google TPUs (the Intel model, where you own your fab)
- Mac mini, his one piece of evidence for personal LLM computing, because batch one inference is memory bound
- ChatGPT and Claude, named as the drop down alongside Gemini
- Time sharing, the 1960s scheduling model he says we are living inside again
The psychology references
- Autoregressive transformer, the simulator underneath the "people spirits"
- Rain Man with Dustin Hoffman, for encyclopedic savant memory, and he recommends watching it
- Memento and 50 First Dates, for fixed weights and a context window wiped every morning
- Hallucination and prompt injection, the two deficits with security consequences
The partial autonomy products
- Cursor, his primary worked example, with the tab / command K / command L / command I autonomy slider
- Perplexity, the same anatomy outside coding, with quick search / research / deep research
- VS Code, the app portability analogy
- Tesla Autopilot, the partial autonomy product he worked on for five years, GUI in the instrument panel included
- Waymo, whose car gave him a zero intervention 30 minute ride around Palo Alto in 2013, photographed on Google Glass
- Photoshop, his example of software where nobody knows what a diff looks like yet
- Iron Man and Tony Stark, the suit that is both augmentation and agent
The vibe coding projects
- MenuGen, live, the app that turns a photo of a restaurant menu into pictures of the dishes, and his own write up of building it
- Swift, the language he cannot program and shipped an iOS app in anyway, in a day
- Clerk, whose Google login instructions are the slide that provoked the final section
- Vercel, the deployment that took part of the week, and later the company he praises for rewriting "click" as curl
- Tom Wolf of Hugging Face, who shared the video of kids vibe coding
Building for agents
- robots.txt, the precedent, and llms.txt, the proposal he endorses. Live examples of the pattern: vercel.com/llms.txt and docs.stripe.com/llms.txt
- Vercel's docs and Stripe's docs, the two early movers he names for markdown documentation
- curl, the executable replacement for "click this button"
- Model Context Protocol from Anthropic, the protocol for speaking directly to agents
- gitingest, which flattens a repository into one pasteable text file when you swap it in for github.com, demonstrated on his own nanoGPT
- DeepWiki from Devin (Cognition), which has the agent analyze a repository and write documentation pages for it
- 3Blue1Brown (channel) and Manim (original repo), the documentation he pasted wholesale into a model and got a working animation out of
One resource on screen stays unidentified: at 0:24:28 he shows a blog post of LLM best practices that he had "read recently and thought was quite good," the one with the vague prompt to failed verification to spinning argument. He does not name it on stage and the slide text is not legible, so it is not linked here rather than guessed at.
Where this sits in the LLM Learning track
This is the closer of the LLM Learning track, and the only one of the twelve written for somebody deciding what to build rather than learning how the thing works. Everything before it opens the box: the whole stack in one sitting, attention one matrix at a time, the tokenizer and the model built from an empty file, the alignment lever, the look inside. This talk closes the box, hands the thing back to you as a component, and asks the product question.
Read it as a pair with the video immediately before it, Ilya Sutskever on what a decade of sequence to sequence taught him, rather than simply after it. The two set up the field's two live questions and they are different questions. Sutskever's is where the next increment of capability comes from, given that compute keeps growing and data does not. Karpathy's is what you build with the capability already sitting on the table, given that it hallucinates, forgets everything every morning, and can be talked into leaking your data. Together they are the honest state of play: nobody knows how far the curve goes, and there is a decade of product work available regardless of the answer.
It also sits naturally against the "Ship something with it" stage just upstream. Jeremy Howard's ladder is the how, the evals talk is how you find out whether it worked, and Building Effective Agents in LangGraph is the same autonomy question Karpathy asks here, answered in code: reach for a workflow first and build an agent only when the path genuinely cannot be written down. That is the autonomy slider as an engineering decision rather than a product one, and the two pages argue the same thing from opposite ends.
And several abstractions here are only load bearing if you have already seen the machinery underneath. "The context window is the memory" is a throwaway metaphor until you have watched attention actually address it. "Software 2.0 is the weights" is a slogan until you have run the optimizer that produces them. The track is ordered so that by the time you arrive here, every analogy in this talk cashes out into something you have already built.
An honest footnote
The strongest part of this talk is the part Karpathy argues against his own interest. He is the person who coined vibe coding, and he spends the longest single stretch of the keynote explaining why he does not use it for real work, why a ten thousand line diff is a liability, and why a flawless demo in 2013 was followed by twelve years of unfinished work. That is a load bearing correction delivered to exactly the audience most likely to ignore it, and the discourse that followed the talk mostly kept the phrase and dropped the leash.
The weakest part is the operating system analogy, not because it is wrong but because it is seductive. The mapping is clean enough that it invites you to extrapolate the rest of computing history onto it: if we are in the 1960s, then a personal computing revolution and a graphical shell and a software industry are all simply scheduled. Karpathy does not actually claim that. He names the pieces he can map, says the GUI has not been invented and that it is not clear what personal computing looks like here, and invites the room to invent it. The honest reading of his analogy is that it describes the current constraint, expensive centralized compute allocated by time sharing, and not a timeline.
Two claims are worth tracking rather than accepting. The first is that English is the programming language, which is true at the level of the interface and much less true at the level of the artifact: the thing that actually makes an LLM product work is usually a scaffold of ordinary code around the prompt, and he says as much when he describes Cursor orchestrating embedding models and diff appliers under the hood. The second is "the decade of agents," which was a useful corrective in mid 2025 and has aged into a claim with real content, since the measurable thing is whether autonomy products are still shipping with humans in the verification loop. His own test is the right one to apply: not whether the demo works, but whether anybody has declared success.
A note on names. The automatic captions underneath this page mangle most of the proper nouns in the talk, and the spellings here are corrected against the real artifacts: Karpathy for "Carpathy," Anthropic for "Enthropic," Andrew Ng for "Anduring," ChatGPT for "Chach" and "Chaship," Claude for "cloud," AlexNet for "Alexet," Waymo for "Whimo," Manim for "Manon," 3Blue1Brown for "three blue one brown," Vercel for "Versell," gitingest for "get ingest," MenuGen and menugen.app for "menu genen" and "menu.app", GUI for "guey," stochastic for "stoastic," anterograde for "entrograde," Memento and 50 First Dates for "Momento" and "51st dates," llms.txt for "lm.txt txt," and vibe coding for the five different ways the caption track spells it. Quotes on this page are cleaned of those transcription errors and of pure filler, and are otherwise his words in his order.


