At a glance
Chris Olah came to the inaugural San Francisco Alignment Workshop with three things he thought were worth saying, and he said all three in forty minutes while his slides refused to advance. The first was a short course in what mechanistic interpretability actually is: a neural network is a compiled binary with no source code, and the job is to recover the source. The second was awkward enough that he flagged it as awkward, a colleague of many years who one day told him, matter of fact, that obviously all interpretability research is worthless, and an open invitation to the skeptics in the room to come argue with him at office hours. The third was superposition, which he calls the question on which the whole safety payoff of the field will rise or fall.
The spine of the talk is the vision work. Features, weights, circuits: three kinds of object, and the claim that weights in isolation say nothing while weights between features you already understand are full of readable structure. He walks a dozen of them, from a dog head detector wired into a dog head plus neck detector to a car detector that wants windows at the top and wheels at the bottom. Then he breaks his own picture. Most neurons are polysemantic, and the hypothesis he offers is that the model is simulating a larger, sparser network that got projected down and folded on top of itself, which is why five features can live in two dimensions and why no neuron means one thing.
The back half is the honest accounting. What a feature is, which he says twice he cannot define and will not define in terms of human concepts. How Transformers differ, through the residual stream and attention heads and the induction heads whose formation shows up as a visible bump in the loss curve of every multilayer model he has looked at. What he wants out of safety, stated as a quantifier over all situations the model could be in. Where he is worried (superposition) and where he is not (scalability). And at the end, with no safety argument attached at all, the reason he says actually moves him: gradient descent grows beautiful structure, and somebody should go look at it.
Three things worth saying, and a laptop that would not cooperate
The talk opens twice. Olah thanks the room, says it is a privilege to be there, starts explaining that he spent time before the event asking himself what would really be useful to say, and then stops: "wait, this is totally not what I am presenting." The slide on the screen is his, just not the one he is on. "There is a great delay between" the laptop and the projector, and that delay runs through the entire talk as a running joke he keeps losing.
What he settled on was three things.
One: a flavor of what mechanistic interpretability is. He is explicit that this is for the part of the audience that is less familiar, and that interpretability is not one field. "There's a lot of different kinds of interpretability work going on," and he is only going to show one kind.
Two: the extent to which this work is worthless. This is the one he calls "a little bit more awkward to talk about," and the caption track drops the operative word twice, both times leaving the sentence hanging where the judgment should be.
I was very struck a few years ago. I had a colleague I'd worked with for many years, and one day they just matter of fact told me that, well, obviously all interpretability research is [the caption drops the word here]. And that really stuck with me, because it seemed to me that probably, you know, it was unusual for somebody to say such a thing so bluntly and directly, but it might be the case that quite a few people believed that. Chris Olah, on the second thing he wanted to say, 1:00
His conclusion is that a talk is the wrong instrument for that argument, because if you already believe it, nothing in the next forty minutes will land. So he converts it into an invitation. He has an office hour style session later in the day, and he would like the skeptics specifically.
To the extent that you sort of are skeptical of this kind of work, I would not be insulted. I would be privileged and delighted if you were willing to come and sort of trust me to talk about your doubts, and see if we can make some progress towards the truth together. Chris Olah, 1:31
Three: superposition. He gives it top billing before he has defined it, and the reason he gives for putting it in front of this particular room is not just that it matters.
The third thing I wanted to talk about is a challenge we call superposition, which increasingly seems to me like the question on which the impact of mechanistic interpretability for safety is going to rise or fall. Chris Olah, 2:02
He also pitches it as a research opportunity with unusually good economics: a question "that is very amenable to research without access to lots of compute," and one with "very rich mathematical structure." That is a deliberate recruiting line in a workshop full of people deciding what to work on next, and it lands differently from the usual "we need more people" because it names the constraint most of the room actually has.
The thirty second pitch: a binary with no source code
He calls it "like a 30 second pitch," and it is the cleanest framing of the field anyone has: the comparison is not biology, it is software engineering.
In regular software you have planning documents, maybe some goals, maybe a design doc. From those you write source code. Then you compile the source code into a binary. Three artifacts, and the middle one is the one humans read.
In a neural network you have a training objective you are optimizing, which is the goal. And then you "sort of directly turn that into neural network parameters, into a trained neural network." Two artifacts. There was "never anything like source code in the middle."
Really the goal of mechanistic interpretability is to somehow take those neural network parameters and turn them into something like source code. Chris Olah, 3:34
Then he pushes on the analogy, which he says is "actually a pretty deep analogy," and the mapping is item by item.
| In a computer program | In a neural network | What that means here |
|---|---|---|
| Variables | A neuron, or a linear combination of neurons, representing a feature | The named things the computation is about. Sometimes a neuron is one. Often it is not, and that is the whole problem the second half of the talk is about |
| Program state | The activations of a layer | The values the variables hold on this particular input |
| Processor or VM it runs on | The architecture | Fixed, written by humans, already understood |
| The compiled binary | The parameters | What training hands you. Complete, executable, and unreadable |
| The source code | missing | "The thing that a computer program typically does have that a neural network does not is the source code, and so that's what we would like to get" |
Middle ground between trivial and impossible
Before any results, he spends a minute on a framing move, and it is aimed at a specific failure of imagination he keeps encountering. People treat interpretability either as a thing that should be trivial and easy, or as a thing that is going to be impossible. He thinks that dichotomy is strange.
This is also where the slides fully break down. Mid sentence: "it's like, oh come on, update slide, please, please change. Yes, but it's on the right thing on the laptop, the laptop is going and switching, it's just not showing the right slide." A beat later he gives up on them: "well, I guess I will try to talk without my slides." They come back a few seconds later on their own.
The position he lands on is that interpretability may be merely very hard. His two calibration points are both deliberately unflattering to the optimists.
Harder than the Linux kernel. Any interesting real neural network is "probably harder to understand or reverse engineer than, say, a very complicated program like the Linux kernel." He flags that he does not know much about reverse engineering normal software, and that reverse engineering the kernel cold, knowing nothing about it, "seems like it would be a very hard challenge." That is the scale of difficulty he is claiming for networks. Not impossible. That hard.
Cells are also hard. The second point is the shape of the argument rather than the size of it: "no one's like, oh, biology, it's not easy therefore it must be impossible for us to understand cells." Difficulty is not an impossibility proof. "There's room for something to be merely extremely difficult in between."
And from that, the research strategy for the whole program: do not try to understand an entire network.
I think the goal of trying to understand an entire neural network is very hard, but we might be able to understand just portions of neural networks and then gradually grow that little portion, so we might be able to go and slowly, sort of rigorously, grow out from understanding tiny little portions to understanding larger chunks. Chris Olah, on the only strategy he thinks is available, 6:08
ConvNets, and the three kinds of object
To make any of this concrete he picks a model family: convolutional vision networks. The reason is partly biographical. "I was very fortunate for a number of years to work here, with our hosts, leading the interpretability team here at the time," he says, which is the single most understated line in the talk. The team was the Clarity team at OpenAI, which he founded and led before co founding Anthropic, and the work he is about to show is the Circuits thread in Distill, the journal he also co founded. He mentions none of that. He says they "made a lot of progress on understanding convnets" and moves on.
The way he likes to think about it is in terms of three kinds of object.
- Features. "The easiest way to think about this would be a neuron," and the examples he gives are a car detector, a curve detector, an edge detector, a floppy ear detector. The working definition at this point in the talk is deliberately loose: "some kind of articulable property of the input." He will spend six minutes later in the talk taking that definition apart.
- Weights, which connect features together. In a convnet the weights are not a single number per pair but a grid, "because we have to talk about relative positions, you know, does it excite something at this offset or that offset."
- Circuits, "which are the combinations of weights and features together." Features are the variables; circuits are the lines of code.
Curve detectors, which survive being taken seriously
His existence proof for the first object is curve detectors. In InceptionV1 "there's all these curve detectors, in fact these exist sort of in every vision model we've looked at." They look for curves at a particular orientation, and the reason he reaches for them rather than something flashier is that they have been audited harder than anything else in the field.
It turns out these hold up to extremely rigorous investigation. In fact we wrote an entire two papers just on curve detectors, and you can do all kinds of things to test these really are curve detectors, up to the point of going and rewriting the entire circuit from scratch from understanding it and re implanting them. Chris Olah, on how much evidence is available for one family of neurons, 7:08
The two papers are Curve Detectors, which is the evidence that the units are what they look like, and Curve Circuits, which is the part he describes as rewriting from scratch: the circuit was reimplemented by hand from the understanding of it and then put back into the model in place of the original, and it works. That is the strongest form of claim available about a piece of a network, because a reimplementation that functions is not a story about the weights, it is a replacement for them.
Weights alone are boring; weights between features are not
Then he does the move that the whole vision program rests on. Look at a weight in isolation, input neuron to output neuron, and "yeah, the weights by themselves, you know, they don't say that much." Nothing is legible there. But contextualize the same weights with features you already understand on both ends and they turn into readable code.
His first example is a dog head facing to the right with a long snout, wired into a dog head plus neck detector. The weight grid is positive at the offset where a head should be if the neck is there, and not on the other side.
You can see that now, all of a sudden, the weights make a lot of sense. We're having the dog head detector go and excite the dog head plus neck detector if it's on the right side of the dog, where the head should be. Chris Olah, 8:09
The second is the car detector, and it is the one that does the most work in the talk because he comes back to it twice more in the Q and A as his example of an explanation that cannot be simplified: a car detector that "goes and looks for windows on the top and wheels on the bottom."
4c:447 only as a joke about unhelpful variable names; that it is specifically the car detector, assembled from window detector 4b:237 and wheel detector 4b:373, comes from the OpenAI Microscope announcement, which uses the same circuit as its motivating example.Feature visualizations are variable names, not evidence
Someone asks, while he is mid walkthrough, what the pictures actually are. His answer is a careful downgrade of his own tool. The images are feature visualizations, "just created by optimizing the input," starting from random noise and pushing it until the neuron fires hard. And then, unprompted, the caveat:
The claim here isn't that that is, like, decisive evidence about what a neuron is doing. It's more like a variable name that we can put there. Chris Olah, on feature visualization, 9:10
The joke that follows is the best argument for the practice. The alternative to a picture is the unit's actual name, and the actual name is 4c:447. He has spent so long with InceptionV1 that he knows what 4c:447 is off the top of his head, and he knows that is not a transferable property. A visualization is "very suggestive of what it does and a little bit of evidence." If you want confidence you do the curve detector treatment: look at everything, run the synthetic stimuli, read the weights.
A dozen circuits in ninety seconds
Then he just lists them, fast, because the volume is the argument. All of these are from the Circuits thread on vision models, and the catalogue of early units is written up in An Overview of Early Vision in InceptionV1.
- High low frequency detectors. Units that look for high frequency texture on one side of the receptive field and low frequency on the other. They are built "by going and piecing together a bunch of neurons that represent high frequency and a bunch that represent low frequency in the previous layer." And the detail he clearly loves: "as the high low frequency detectors rotate, so do the weights. So there's actually just this beautiful structure." That family of rotating weight structure is the subject of Naturally Occurring Equivariance in Neural Networks.
- Center surround detectors, assembled out of color contrast detectors.
- Black and white versus color detection, which he calls "a more interesting circuit," building black and white versus color detectors and then assembling those into center surround units as well.
- Curve detectors from earlier curve detectors, and this one is the cleanest piece of readable logic in the set: the earlier curve detectors excite the later ones "when they match the orientation, and inhibit when they have the opposite orientation. And again the weights rotate as the features rotate."
- A triangle detector assembled from edge detectors.
- A small circle detector assembled from very early curve detectors. He tees this one up like a magician: "we have, wait for it, a small circle detector."
- Lines, built from a whole bunch of color contrast detectors.
- Sophisticated boundary detectors, "being constructed from all kinds of different cues."
The summary he wants taken away is deliberately modest about coverage and immodest about what is there. The method for reading weights in context is written up in Visualizing Weights.
These are just a handful of examples. Neural network weights, once you start to contextualize them, are full of structure. Also lots of things we don't understand, but lots of structure. Chris Olah, closing the vision walkthrough, 11:12
The first real pushback: is a picture an explanation?
The question from the floor is the right one, and it is essentially whether the whole thing is circular. If you explain a car detector in terms of a window detector and a wheel detector, and you only know what those are because of another picture, what has been established?
His answer has two parts, and both matter for the rest of the talk.
First, the pictures are not the explanation. "These visualizations are just variable names. The actual thing that's analogous to, like, i equals one, or something like this, is the weights." The weights are "like the assembly instructions, or the code, of the computer program." The picture is the label on the variable; the weight grid is the statement.
Second, the regress is real and it terminates. "Then you could ask, okay well how do we trust those? Then we need to go back another step, and so on." Which is exactly how you read an unfamiliar program: you ground each definition in the ones below it.
It starts to look a lot like a computer program, where we have an understanding where we can base it on our understanding of previous variables, and if we want to go and understand those and really carefully understand those we have to go back further and further and further. And we can go and trace things back all the way to the input if we want. It's a lot of work. Chris Olah, 13:03
He takes one more question, sees the room filling with hands, and calls it: "I'm kind of tempted to ask to hold questions maybe to the end, because I do want to cover a lot of ground and I'm a little nervous that we won't get through it." The office hours offer comes back out.
Superposition: the picture he just painted is too optimistic
He says it himself, in the handoff: "the picture that I painted for you so far is a bit overly optimistic in an important way." The important way is that networks are also full of neurons that do not behave like the ones he just showed.
Neural networks are also full of what we call polysemantic neurons, neurons that respond to many unrelated features. Chris Olah, 14:04
The framing of the surprise is worth keeping, because it is the opposite of how this is usually told. The remarkable thing is not that some neurons are a mess. It is that any of them are clean.
In some ways it's miraculous that neural networks, when they have a privileged basis, when they have neurons, they have an activation function, that you get lots of neurons that just seem to correspond to meaningful features. But you also get many that don't. Chris Olah, on which half of the observation needs explaining, 14:04
The hypothesis
The hypothesis for why is superposition, and he states it as a want on the model's part: "the model wants to represent more features than it has neurons, possibly many more features than it has neurons." Once that is true the geometry is forced. "As soon as you have that, of course you can't align all the features with neurons, because there's only n neurons and you want to represent more things."
Then he takes it to its conclusion, which is the sentence that reorganizes everything before it.
If you take that really seriously, what it starts to suggest is that the model is actually sort of simulating a larger neural network. There is some larger, sparser neural network, and then it gets projected down and folded on top of itself to go and create the network that we actually observe. And then of course the neurons don't correspond to features, they correspond to linear combinations of features. Chris Olah, the central claim of the talk, 14:34
The polysemantic neuron is not a defect and not a mystery under this account. It is a projection artifact. You were reading an axis of the small network and hoping it was a variable of the large one.
What the toy models show
He is careful about the evidential status. "It's hard to really directly study this in real models, but it turns out we can show that this is exactly what happens in toy models, and then there's suggestive evidence that this happens in large models." The toy models are Toy Models of Superposition, published five months before this talk, and the experiment he describes is the sparsity sweep: features of varying importance, and a knob for how often they occur.
If you just start at zero percent sparsity, if the features are just completely dense, you only get as many features as you have dimensions. But as you make the features sparser, then because they're probably not going to go and co occur and interfere with each other, the model starts to be able to go and pack more things in, and eventually you end up with five features in a two dimensional space. Chris Olah, on the sparsity sweep, 15:35
Five in two is a specific number, and it is the whole hypothesis in miniature. Nothing about the network changed except how often the features show up. Sparsity is what buys the packing, because two features that are almost never both present can share a neighborhood of directions and almost never collide.
And then the result that makes superposition a problem rather than a curiosity about storage: it is not just representation.
It turns out you can actually do computation as well in superposition. You can in some sense have a computational graph that's higher dimensional, project it down, and actually do useful computation while holding everything in superposition. Chris Olah, 15:35
Which features you get to see, and which you do not
A consequence he draws out, and it is the most uncomfortable one in the talk, is that the clean neurons are a biased sample. If superposition is real, the features that got an axis to themselves got it for a reason.
What you start to think, when you observe these neurons that are sort of monosemantic, is that those are particularly important and common features, and then the things that we don't observe, the sort of less important, more dispersed things, are in superposition. Chris Olah, 16:06
So the part of the model that is legible is the part that is frequent and important. The rare machinery is exactly the machinery that is hiding, which is a bad property for a method whose purpose is finding things you did not think to look for. That is the step from "superposition is interesting" to "superposition is the problem."
Unless you understand the superposition structure, there will always be the potential for unknown behaviors to just sort of suddenly occur. This is something that I'm very worried about. And you won't be able to understand really what the weights are actually doing. Chris Olah, 16:06
He then draws the boundary of the claim honestly, which is characteristic. He is not saying nothing can be done without solving it: "I think there are other valuable things one can do within mechanistic interpretability without this, but I think that we will surrender a great deal if it turns out that we can't go and address this somehow." And he expects the problem is not his alone: "I think it probably holds for a lot of other approaches as well."
The consolation prize is that the problem is gorgeous
Having called superposition the thing that could sink the field, he spends thirty seconds on why he enjoys it anyway, and this is the recruiting pitch from the intro cashing out.
It turns out though that superposition is full of beautiful mathematical structure. It is deeply connected to compressed sensing. It turns out that in toy models at least, these features organize themselves into polyhedra, into regular polyhedra, which is kind of wild, or uniform polyhedra. It's very not obvious that they should do this. It turns out that the learning dynamics involve these weird, like, electron jumping like behaviors that are very strange. It's very mysterious. Chris Olah, 17:09
Three concrete hooks in one breath: a connection to compressed sensing, feature geometry that lands on uniform polyhedra with no one asking it to, and learning dynamics where features appear to jump between configurations like electrons between shells rather than sliding smoothly. None of it needs a cluster.
"I don't have a super compelling definition of what a feature is"
The question he takes next is the one that threatens the foundation, and he does not defend the foundation.
I actually think that defining what we mean by feature is extremely hard, and I don't have a super compelling definition of what a feature is. Chris Olah, 17:40
What he offers instead is a chain of partial answers, in increasing order of how much he believes them.
The intuition pump. "Intuitively what we mean is something like a curve detector, or a car detector, or things like this. And sometimes those end up corresponding to neurons, but sometimes we can show that they correspond to linear combinations of neurons."
The place where it is exact. In the toy models the ambiguity vanishes, because the features are an input to the experiment rather than an output: "we can construct toy problems where we do know what features are and where that's exactly defined, and then they will go and represent them in this way." That is why he keeps returning to that paper. It is the only setting where the thing being claimed about real networks has a ground truth.
The definition he rejects. One obvious move is to say a feature is a human understandable property of the input that the network represents. He will not take it: "I don't want humans to be sort of involved in the definition." The cost of that definition is that it makes the existence of a feature contingent on a person recognizing it, which forecloses the possibility he most wants to keep open.
The definition he actually holds. Stated as a belief he cannot yet formalize, which is a rare register for a technical talk.
The thing that I sort of in my heart believe a feature is, but I don't know how to properly define this, is something like: a feature is a fundamental unit of the neural network's computation. And those don't necessarily correspond to neurons. They sometimes do, and they often appear not to. Chris Olah, 18:40
How Transformers are different
He flags that he is going to be brief here, and why: "this actually becomes quite messy and ends up with a lot of detailed mathematics, which I think is actually very cool mathematics, but it's a little bit harder to go through and talk." The long version is A Mathematical Framework for Transformer Circuits. What he gives instead is two architectural differences, plus a third observation he mentions almost in passing and which is the one the rest of the talk hangs on.
The third one first: there is more superposition. "You just see a lot of superposition in Transformers. You see a lot more than you do in vision models." He does not explain why here. He explains why forty minutes later, in the final stretch of Q and A, and that answer is the most quantitative thing in the talk.
Difference one: the residual stream. It is not quite a ResNet, he says, and the reason it matters is what the addition implies. Because every layer reads from and writes directly into the same stream, projecting in and out of it rather than passing a value hand to hand, there are "sort of these implicit latent weights connecting neurons across layers, or attention heads across layers, or all these things." Two components in distant layers can be in communication through the stream without anything in the architecture naming that connection. "That creates a lot of interesting structure that doesn't exist in quite the same way in any convnets really, because even ResNets aren't exactly like this."
Difference two: attention heads. They are a genuinely new kind of object for this kind of analysis, not just more neurons. "Attention heads kind of create a new sort of fundamental unit of mechanistic interpretability, similar to features, weights or activations," and with them come attentional features, of which he says there are "lots of very rich and interesting" examples. Here the slides fail him again on cue: "the other is attention heads. Oh no, apparently I'm not gonna talk about that more in a second."
Induction heads, and the one he finds most interesting
The example he picks is the one that connects a single attention pattern to a macroscopic property of training. If you look at attention patterns you often see "these weird off diagonals. They're very striking."
It turns out this corresponds to a particular type of attention head called an induction head, which searches through for previous cases where something happens, and then looks forward one step, and then goes and copies that and goes and outputs that. Chris Olah, 21:12
This mechanism is how a model does a certain kind of in context learning: you saw this token before, here is what came next, do that again. And he thinks a generalized version of it is doing much more than that: "in fact it turns out that sort of a generalized version, it seems to be the driver of in context learning, or at least there's a non trivial amount of evidence for that hypothesis."
Then the part that makes induction heads special among all the circuits in the talk. They are the one piece of microscopic machinery whose arrival you can see from the outside.
In fact these are so important that they create a bump in the loss curve when you train Transformers. It's just a visible bump in every Transformer loss curve that I've seen, if you have more than two layers, because that's what's necessary for induction. Chris Olah, 21:42
Two layers is not an arbitrary threshold. An induction head has to compose with an earlier head to work, so the mechanism cannot exist in a one layer model, and the bump appears exactly when it becomes possible. The effect also reaches the scaling laws: "if you look at the scaling laws, the original scaling laws paper, you'll see that there's a point where there's sort of a divergence from the original trend, and that corresponds to the models not having induction heads and then forming them."
These inductions are so important they appear in the macroscopic picture. We have this microscopic picture that we're developing here, and it bridges all the way to the macroscopic. Chris Olah, on why he keeps telling this particular story, 21:42
Do they all form at once?
Yes, and the questioner clearly already knew that, because the follow up is about the mechanism. Olah confirms it: "a massive number of induction heads simultaneously occur," and there seems to be a deep reason.
There's a number of pieces that have to exist for an induction head circuit to form, and once you start to have those ingredients you then get an extremely fast feedback loop, and then all of them form simultaneously because the ingredients are in place. Chris Olah, 22:43
What that looks like empirically: "if you just track the existence of induction heads in a Transformer over time, they will just sort of all form in a very small window," and then maybe evolve a bit further, with the occasional late one, "but the vast majority of them form at once."
He adds a control that is the clearest evidence the bump and the circuit are the same event. "We have a slightly similar experiment where you can go and set up an architecture that very easily learns induction heads, and then induction heads form right at the start." In that setup the formation happens inside the first, extremely steep part of the loss curve, so there is nothing to see: "you shouldn't see an analogous bump in the loss curve." Move the circuit's arrival and the macroscopic signature moves with it.
What he actually wants out of safety
There are many ways interpretability might help with safety, and he deliberately picks one, the one he finds motivating. He starts from the shape of the guarantee rather than from any technique.
It seems to me like the thing that we most want out of safety is something like the ability to make statements of a kind, you know, something vaguely like: for all the situations the model could be in, the model will never deliberately do X. We want to be able to say something like that. Chris Olah, 23:45
Read that as the logic it is. It is a universal quantifier over situations, and no amount of testing gets you one, because testing is a finite collection of examples and the statement is about everything that is not in your collection. That is the entire argument for looking inside, and he never has to say it out loud.
His best guess at how such a statement could ever be earned is where features and circuits come back as the load bearing pieces.
My present best guess as to how you could achieve that is that you want to be able to say something like: okay, there don't exist features such that the model will deliberately, that will participate in the model deliberately doing X. And that's going to then be some kind of claim about the circuits that feature participates in. Chris Olah, 24:15
He immediately marks the distance. "This is a very ambitious and wild and kind of crazy thing to be aiming for. I don't mean to say that this is anywhere near being able to go and make this kind of statement, but this is kind of a spiritual North Star, I think, for the most ambitious way that mechanistic interpretability could help with safety."
Two challenges, and he is only worried about one
Superposition. Stated here as an access problem: "how can we actually access the features? How can we know that we've got everything? How can we rule out these places where the model might have something, a feature that activates very rarely, and it's just some direction in activation space where it behaves totally differently?" That last clause is the safety relevant failure in a single sentence. A rare direction with different behavior is exactly the thing the universal quantifier was supposed to exclude, and it is exactly the thing superposition hides.
Scalability. "If these models are so large, and the successes of mechanistic interpretability so far are so small, they're these tiny little circuits that we're understanding, how can we hope to get to these large models?"
And then the verdict, which is not symmetric.
I am very worried about superposition. I am more optimistic about scalability. Chris Olah, 25:15
The reasons for the optimism are three concrete directions, all of which have since become their own subfields:
- Automate it with AI. "Trying to use AI to automate things."
- Exploit modularity or graph structure in models. If the computation graph has large scale organization, you do not have to read it uniformly.
- Find motifs and symmetries that collapse the work. "There might also be sort of interesting motifs, we've seen this in vision models, that can massively simplify models once you understand them, if you start to understand certain symmetries." One understood symmetry retires a whole family of units at once, the way the rotating weights of the curve detectors did.
But there is a dependency between the two challenges, and it runs the wrong way for anybody who wants to work on the easier one.
Unfortunately I think this is a really challenging problem to work on right now, because we don't actually have access to the units that we want to understand. If we think the features are there and they're all in superposition, they're a giant mess, and it's very difficult to start working on how you can even actually really effectively scale. Chris Olah, on why scalability is hard to attack today, 25:46
You cannot design a method for reading a million circuits while you are unsure what a circuit is made of. Superposition is not just the risk, it is the blocker on the thing he is optimistic about. "My real fear is this challenge of superposition."
The other reason he does this, which has nothing to do with safety
He checks the clock, notes he is "right on schedule," and spends the last ninety seconds on something he labels as emotional rather than argumentative. It is the most quoted part of the talk for good reason.
The claim is about what kind of science this is. Machine learning gets compared to math and physics, where elegance means compression: short laws, clean derivations. He thinks the better comparison is biology, where elegance means structure you did not design and could not have predicted.
In the case of biology, evolution creates awe inspiring complexity in nature. You know, we have this system that goes and generates all this beautiful structure. And in a kind of similar way, gradient descent, it seems to me, creates beautiful, mind boggling structure. It grows, and it creates all these beautiful circuits that have beautiful symmetries in them, and in toy models it arranges its features into regular polyhedra. Chris Olah, 26:48
Then the part that is a methodological commitment disguised as enthusiasm: the mess is provisional. "There's messy stuff, but then you discover sometimes that that messy stuff actually was just a really beautiful thing that you didn't understand." High low frequency detectors were incomprehensible before they were trivial. That record is the reason he reads an unexplained unit as unexplained rather than as noise.
So a belief that I have, that is perhaps only semi rational: these models are just full of beautiful structure, if we're just willing to put in the effort to go and find it. And I think that's the thing that I find actually most emotionally motivating about this work. Chris Olah, closing the prepared talk, 27:48
Where he tells you to start
He closes with credit and three pointers. The credit first: "all this work was done in collaboration with many others, including others at different institutions," and then, as the only self identification in forty minutes, "I'm at Anthropic."
The three places to start, in his order:
- The original Circuits thread, "which made a really serious attempt to reverse engineer one vision model, InceptionV1." The entry point is Zoom In: An Introduction to Circuits.
- The Transformer Circuits thread, "which has been trying to do analogous work on Transformers."
- Neel Nanda's annotated reading list, which he singles out: "which actually might be the best jumping off point." The version current at the time of the talk is An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers; it has since been replaced by a second edition.
Q and A, in order
The last eleven minutes are questions, and they are not filler. Two of them produce material that is not anywhere in the prepared talk, including the only numbers he gives all day.
Modularity: the branches are already there, you just have to look
The question is about whether you could build models that are easier to interpret, by making them modular on purpose. His answer starts with an accident from 2012.
In the Krizhevsky et al. 2012 paper, the one that is usually called AlexNet, the model was split across two GPUs that "only talk to each other every other step," for hardware reasons. The striking phenomenon: look at the first layer and the two halves specialized. "There's all these Gabor black and white Gabor filters on one side, and color contrast detectors on the other." Look at the next layer and you see similar things.
It generalizes. "Lots of vision models have branches, and like InceptionV1 has branches, and it turns out that like this branch over here has all of the black and white versus color and color contrast units, this one's full of curve detectors and some sort of fur and eye related detectors and some boundary detectors, and this one's full of 3D geometry detectors." That is Branch Specialization in the Circuits thread.
So one answer to the question is: fine, build models with branches like that. But then he gives the better answer, and it reverses the premise of the question. You do not have to impose the structure, because it is already in there.
It turns out that the structure is implicitly there. You can just go and take the magnitude of the connections between features in one layer and features in another layer, and go and do a singular value decomposition, and you discover that actually all the same structure was implicitly there. Chris Olah, on finding modularity in a model nobody made modular, 30:27
Where it stops working is the tell. "In practice this becomes hard on the later layers of vision models, you don't see this as much, and in language models we don't see it." His explanation is the same one as always, and it is the sharpest statement of what superposition actually does to a model.
I think that's because of superposition. The best case for superposition is to go and fold two completely unrelated computational graphs on top of each other. Chris Olah, 30:58
Two unrelated graphs are the ideal thing to superpose, because they never need to be active at the same time, so folding them together costs almost nothing in interference. Which means the models with the most modular structure available to find are exactly the models where superposition has the most to gain by hiding it. Hence the optimism that goes with the complaint: "one of the things I'm sort of most optimistic about is that if you can solve superposition you'll discover that there's just lots of implicit modularity and other kinds of organizational structure to these graphs." The evidence he offers for that bet is the early layers of vision models, "kind of the best proxy I have for a model that has very little superposition," where the structure is visible.
What is a feature, take two: Lakatos, and the weights are the code
Another questioner comes back to the definition problem, and this time he gives the epistemology instead of the definition. First, the self aware version of the problem:
You might be like, oh that's kind of wild, right? Like Chris is spending his whole career on this thing and he doesn't even know what the definition that he's talking about is. But actually I think that's often very productive. Chris Olah, 32:54
The authority he cites is Imre Lakatos and Proofs and Refutations, which he describes as "this play about a bunch of math students going and being confused about definitions and progressively finding more and more useful definitions for things." That is a precise citation for what he is doing: the definition of a feature is expected to be an output of the research program, not an input to it.
Then the reason he will not accept the convenient definition, stated more directly than the first time:
I don't want to define features in terms of human understandable concepts, because I want to allow for the possibility that these models have concepts that we don't understand. Chris Olah, 33:25
What he is after is "in some sense pulling apart the fundamental units of computation, whatever they are," and he adds the honest empirical footnote: "so far in the cases where we've done this, they have been things that we've understood."
This leads to the best clarification in the talk, because it corrects a reading of his own software analogy that most people take away from it. "Turning them into code" does not mean producing Python.
The weights are the code. For the car detector that I showed you earlier, there's no simpler explanation than just showing you, you know, there's a window detector, there's a wheel detector, it wants to see the windows at the top, it wants to see the wheels at the bottom. That is the most understandable depiction of that that there is. Chris Olah, 33:55
And the condition on that, which is the thesis of the entire vision section in one line: "but you couldn't have understood if I didn't contextualize it." The work is not translation. The work is contextualization, "and then very painstakingly trying to go through it."
His example of what that process actually feels like is the high low frequency detectors. Nobody understood them at first. In retrospect they are embarrassingly simple: high frequency on one side, low frequency on the other, "and this occurs all the time at boundaries of objects, because you have something that's out of focus in a background, or the background is busy and [the object] is a solid color." A concept the model invented, that no human would have thought to name, that turns out to be an obvious thing to want once you see it.
The advantages over neuroscience
He takes a detour he clearly enjoys, because the comparison to neuroscience comes up constantly and he thinks it is backwards.
I think that we have an enormous advantage over neuroscience. In fact I wrote a whole article about all the things that make my life so much better than neuroscientists'. Chris Olah, 34:56
The article is Interpretability vs Neuroscience, posted in March 2021. Then he enumerates, from memory, for the people who have not thought about it.
| What he can do | What it rules out |
|---|---|
| Ground truth access to the neuron's computational structure | No inference from indirect measurement. The thing itself is on disk |
| The activations of every neuron, non destructively, for as many stimuli as he wants | No electrode budget, no choosing which cells to reach, no limit on the stimulus set, no killing the sample to read it |
| The ground truth connectome | No reconstruction project. He stresses it is better than a connectome: "not just which neurons are connected to each other, but we know which ones excite and inhibit each other" |
| Gradients through the structure, and optimization through it | No black box probing. This is also what makes feature visualization possible at all |
| The same model again and again | No individual variation. Neuroscience has to switch between organisms "which might be different" |
| Handing the exact model to a colleague | "We can go and ask questions of exactly the same model and talk about the same neurons." Findings compose because the object is shared |
| Far faster experiment cycles | No wet lab. This is the one he spends the least time on and that changes the most in practice |
Why language models carry so much more superposition
This is the question that pays the best, because it is the one place he puts numbers on the problem. He names two factors, and says language models go the wrong way on both.
Factor one: the feature importance curve. Order every feature the model could learn by importance, "which is like something like how much it reduces the loss," and look at the shape of the curve. "The flatter that curve is, the more superposition you're going to have. The steeper it falls off, the less superposition you're going to have." A steep curve means a handful of features carry nearly all the value, so a model with limited dimensions can simply keep the top few and drop the tail. A flat curve means the thousandth feature is nearly as valuable as the tenth, and dropping the tail is expensive, so the model packs instead.
Factor two: sparsity. And here he gives the contrast that makes it vivid, using the room he is standing in.
If you look at this thing behind me and you ask how many vertical edges are there in just a single image, you know, there are a lot of vertical edges. On the other hand, a lot of these language models know who I am. How often do I come up in text? One in 100 million tokens, one in a billion tokens? I don't know. I don't think I'm very common. Those actually feel like they're probably too small a number. Chris Olah, on the sparsity gap between vision and language, 37:32
An edge detector in early vision fires many times per image. A "Chris Olah" feature fires on the order of once per hundred million tokens, and he suspects even that overstates it. That is the gap that makes superposition pay in a language model, because the rarer the feature, the cheaper it is to pack.
He adds two more pieces of the same answer:
- Language models get dense fast. "Because you start with word embeddings and then you very quickly go to more sophisticated things, language models very rapidly go to denser features." The monosemantic neurons you do find in a language model are the common things: "common constructions of multiple words, or things like this,
would notfor instance." Which is exactly the prediction from earlier: what you can see one neuron at a time is what is frequent. - The concept space is simply bigger. "The space of concepts you want a language model to represent is just so much bigger than like an ImageNet model, and so you just have this much flatter, longer feature importance curve."
Are language neurons harder to read than vision neurons?
No, and he thinks this is a red herring. His evidence is tooling. At OpenAI they built Microscope, which shows every significant unit of eight vision models; there is "a similar internal thing for language models," which the captions render as lexoscope. Both show you the dataset examples that drive a unit.
And then the one genuine advantage language has, which he is openly grateful for:
One really cool thing you can do in language models that you can't do in image models is you can just edit the text and see how the neuron activations change in real time. That's like a thing that I'm very grateful for. Chris Olah, 39:04
His conclusion: a language neuron is not harder to understand or recognize. "Maybe it takes a little bit more time, because you can maybe more quickly visually go through a bunch of visual dataset examples than language ones, but I think not in a very significant way. So I think that's not the primary thing going on." The difficulty with language models is superposition, not legibility.
The last question, and a flat "I have no idea"
He is out of time, takes one more, and the question is partly inaudible on the recording: something about human concepts, and how superposition would change if the model's concepts were less human like.
He gives the structural part of the answer. Superposition "depends on the gap between the sparsity of the neurons and the sparsity of the fundamental underlying data and the fundamental underlying features," and if those features really are extremely rare, "one in a million tokens, one in 10 million tokens, then there's still an enormous, enormous sparsity gap." So the packing pressure does not go away. The part he cannot answer is the other factor, "how do you think the feature importance curve changes," and he does not pretend otherwise.
The basic answer is: I have no idea. And I can give you a bunch of conceptual models for trying to reason about this, but I still then have no idea. Chris Olah, the last thing he says before thanking the room, 40:36
Key takeaways
- A trained network is a binary with no source code. Planning docs, source, binary in software; training objective, parameters in a network. Variables map to features, state maps to activations, the VM maps to the architecture, the binary maps to the parameters, and the source code row is empty. Interpretability is the project of filling it in.
- Interpretability is plausibly merely very hard. His own calibration is that reverse engineering a real network is harder than reverse engineering the Linux kernel cold, and that this says nothing about possibility. Difficulty is not an impossibility proof, which is why the strategy is to understand small portions rigorously and grow them.
- Features, weights, circuits. Features are articulable properties of the input, weights connect them and in a convnet are grids indexed by relative position, and circuits are weights and features together. A weight in isolation says nothing. A weight between two understood features is a statement.
- Curve detectors are the audited case. They appear in every vision model they looked at, carry two full papers of evidence, and were reimplemented from scratch and re implanted into the model. A circuit you can rewrite and reinstall is a circuit you understand.
- Feature visualizations are variable names, not proof. They are inputs optimized from noise to drive a unit. The alternative label is
4c:447. The actual evidence is the weights. - Most neurons are polysemantic, and the surprise is that any are not. Monosemantic neurons exist because a network has a privileged basis at all, and the ones that are clean are the frequent, important features. The hidden ones are the rare ones.
- Superposition: the model simulates a larger, sparser network, projected down and folded on top of itself. With dense features you get as many features as dimensions. With sparse features you get five in two dimensions. Computation, not just storage, survives the folding.
- Superposition is the field's load bearing problem. Without it you cannot rule out a rare direction where the model behaves differently, and you cannot read the weights. It also blocks the scalability work he is otherwise optimistic about, because you cannot design methods for circuits whose units you do not have.
- He cannot define a feature, and refuses the easy definition. Not "a human understandable property of the input", because that forecloses concepts we do not have. His working belief is "a fundamental unit of the network's computation", which he cannot yet formalize. Lakatos is the precedent.
- Transformers differ in two structural ways plus one empirical one. The residual stream creates implicit latent weights between distant layers. Attention heads are a new unit with their own features. And there is simply far more superposition than in vision models.
- Induction heads bridge micro and macro. Search back for a previous occurrence, step forward one token, copy. They need at least two layers, they appear to drive in context learning, they all form at once in a narrow window, and their formation is visible as a bump in the loss curve and a divergence in the scaling laws.
- The safety target is a universal quantifier. For all situations the model could be in, it will never deliberately do X. No finite behavioral test reaches that statement. His route to it is the nonexistence of features that would participate in doing X, plus a claim about their circuits.
- Language models superpose more for two reasons that point the same way. The feature importance curve is flatter and longer, because the concept space is bigger, and the features are far sparser: many vertical edges per image, versus a feature for his own name at perhaps one in 100 million to one in a billion tokens.
- Modularity is already in there. Branch specialization in AlexNet and InceptionV1 was an accident of hardware and architecture, and an SVD on the magnitudes of between layer connections recovers the same structure in models nobody made modular. It disappears in language models, which he reads as superposition folding unrelated graphs together.
- The non safety reason is that gradient descent grows beautiful structure. Closer to the elegance of biology than of math or physics, and the track record says an unexplained unit is unexplained rather than noise.
Chapters
- 0:00 Welcome, and what would be useful to say to this room
- 0:30 The slides are already wrong, and stay wrong
- 1:00 Thing two: the colleague who said all interpretability research is worthless
- 1:31 An invitation to the skeptics, at office hours
- 2:02 Thing three: superposition, the question it all rises or falls on
- 2:33 Why it is a good problem if you have no compute
- 3:03 The thirty second pitch: software has source code, networks do not
- 3:34 The goal, stated: parameters into something like source code
- 4:04 Variables, state, VM, binary, and the missing fifth row
- 4:35 Middle ground between trivial and impossible, and a slide revolt
- 5:05 Merely very hard: harder than reverse engineering the Linux kernel
- 5:36 Nobody says cells are impossible
- 6:08 The strategy: understand portions, then grow them
- 6:38 ConvNets, the host organization, and three kinds of object
- 7:08 Curve detectors, in every vision model they looked at
- 7:39 Weights on their own are not interesting
- 8:09 A dog head wired into a dog head plus neck
- 8:40 Pose invariant dog heads, and the car detector
- 9:10 Feature visualizations are variable names, not decisive evidence
- 9:40 The joke about unit 4c:447
- 10:11 High low frequency detectors, and weights that rotate with the feature
- 10:41 Center surround, black and white versus color, curves from curves
- 11:12 Triangles, a small circle detector, lines
- 11:42 Boundary detectors, and the takeaway: full of structure
- 12:33 Question: is the picture the explanation? The weights are the code
- 13:03 Tracing trust back to the input
- 13:34 Holding the rest of the questions
- 14:04 Polysemantic neurons, and the superposition hypothesis
- 14:34 The model is simulating a larger, sparser network folded on itself
- 15:04 Hard in real models, exact in toy models
- 15:35 Zero sparsity, then five features in two dimensions
- 16:06 Computation in superposition, and which features stay visible
- 16:37 Why this is the challenge: unknown behaviors, unreadable weights
- 17:09 Compressed sensing, regular polyhedra, electron jumps
- 17:40 Question: what is a feature? "I don't have a super compelling definition"
- 18:10 Toy problems where a feature is exactly defined
- 18:40 Why he refuses a human centered definition
- 19:11 How Transformers are different
- 19:41 The residual stream
- 20:11 Implicit latent weights across layers
- 20:42 Attention heads as a new fundamental unit, and attentional features
- 21:12 Induction heads: search back, step forward, copy
- 21:42 The bump in every loss curve, and the kink in the scaling laws
- 22:13 Question: do they all form at once?
- 22:43 Ingredients, then an extremely fast feedback loop
- 23:13 The control: an architecture that forms them immediately, with no bump
- 23:45 What we most want out of safety, as a statement about all situations
- 24:15 Restating it in features and circuits: the spiritual North Star
- 24:45 Two challenges: superposition and scalability
- 25:15 Very worried about one, more optimistic about the other
- 25:46 Why scalability is hard to even work on yet
- 26:17 Right on schedule, and the other motivation
- 26:48 Evolution, gradient descent, and awe inspiring complexity
- 27:18 Polyhedra, mess that turns out to be structure, a semi rational belief
- 27:48 Thanks, collaborators, and "I'm at Anthropic"
- 28:18 Where to start: Circuits, Transformer Circuits, Neel Nanda's list
- 29:26 Q and A: modularity, and the two GPUs of Krizhevsky 2012
- 29:57 Branch specialization in InceptionV1
- 30:27 The structure is implicitly there: take an SVD of the connections
- 30:58 Why you do not see it in language models
- 31:28 Solve superposition and the modularity shows up
- 32:54 Q and A: what is a feature, take two, and Lakatos
- 33:25 Allowing concepts humans do not have
- 33:55 The weights are the code, and contextualization is the work
- 34:26 High low frequency detectors: baffling, then obvious
- 34:56 Q and A: the enormous advantage over neuroscience
- 35:31 Every activation, the connectome, and the weights themselves
- 36:01 Gradients, the same model twice, shared neurons, faster cycles
- 36:31 Q and A: why language models carry more superposition
- 37:01 The feature importance curve, flat versus steep
- 37:32 Vertical edges in one image versus one in a billion tokens
- 38:02 Word embeddings, denser features, and the common monosemantic ones
- 38:33 A bigger concept space, a flatter curve
- 39:04 Microscope, the language model equivalent, and editing the text live
- 39:34 The last question
- 40:06 The sparsity gap, and an honest "I have no idea"
This chapter map was built for this page from the transcript's own timings, because the video publishes no chapters of its own. Every timestamp above is clickable and seeks the embedded player.
Notable quotes
I had a colleague I'd worked with for many years, and one day they just matter of fact told me that, well, obviously all interpretability research is [the caption drops the word here]. And that really stuck with me. Chris Olah, on the second of his three topics, 1:00
To the extent that you sort of are skeptical of this kind of work, I would not be insulted. I would be privileged and delighted if you were willing to come and sort of trust me to talk about your doubts, and see if we can make some progress towards the truth together. Chris Olah, turning a talk he cannot give into an office hour, 1:31
The third thing I wanted to talk about is a challenge we call superposition, which increasingly seems to me like the question on which the impact of mechanistic interpretability for safety is going to rise or fall. Chris Olah, 2:02
Really the goal of mechanistic interpretability is to somehow take those neural network parameters and turn them into something like source code. Chris Olah, the thirty second pitch, 3:34
These are just a handful of examples. Neural network weights, once you start to contextualize them, are full of structure. Also lots of things we don't understand, but lots of structure. Chris Olah, after a dozen vision circuits in ninety seconds, 11:12
These visualizations are just variable names. The actual thing that's analogous to, like, i equals one, is the weights. The weights are like the assembly instructions, or the code, of the computer program. Chris Olah, answering whether a picture counts as an explanation, 12:33
In some ways it's miraculous that neural networks, when they have a privileged basis, when they have neurons, that you get lots of neurons that just seem to correspond to meaningful features. But you also get many that don't. Chris Olah, on which half of polysemanticity needs explaining, 14:04
There is some larger, sparser neural network, and then it gets projected down and folded on top of itself to go and create the network that we actually observe. And then of course the neurons don't correspond to features, they correspond to linear combinations of features. Chris Olah, stating the superposition hypothesis, 14:34
If you just start at zero percent sparsity, if the features are just completely dense, you only get as many features as you have dimensions. But as you make the features sparser, the model starts to be able to go and pack more things in, and eventually you end up with five features in a two dimensional space. Chris Olah, on the toy models sparsity sweep, 15:35
Unless you understand the superposition structure, there will always be the potential for unknown behaviors to just sort of suddenly occur. This is something that I'm very worried about. Chris Olah, 16:06
I actually think that defining what we mean by feature is extremely hard, and I don't have a super compelling definition of what a feature is. Chris Olah, on the central term of his field, 17:40
The thing that I sort of in my heart believe a feature is, but I don't know how to properly define this, is something like: a feature is a fundamental unit of the neural network's computation. Chris Olah, 18:40
It turns out this corresponds to a particular type of attention head called an induction head, which searches through for previous cases where something happens, and then looks forward one step, and then goes and copies that and goes and outputs that. Chris Olah, defining induction heads, 21:12
It's just a visible bump in every Transformer loss curve that I've seen, if you have more than two layers, because that's what's necessary for induction. Chris Olah, on seeing a circuit form from outside the model, 21:42
It seems to me like the thing that we most want out of safety is something like the ability to make statements of a kind, something vaguely like: for all the situations the model could be in, the model will never deliberately do X. Chris Olah, on the shape of the guarantee, 23:45
I am very worried about superposition. I am more optimistic about scalability. Chris Olah, 25:15
In the case of biology, evolution creates awe inspiring complexity in nature. And in a kind of similar way, gradient descent, it seems to me, creates beautiful, mind boggling structure. Chris Olah, 26:48
So a belief that I have, that is perhaps only semi rational: these models are just full of beautiful structure, if we're just willing to put in the effort to go and find it. And I think that's the thing that I find actually most emotionally motivating about this work. Chris Olah, closing the prepared talk, 27:48
You might be like, oh that's kind of wild, right? Like Chris is spending his whole career on this thing and he doesn't even know what the definition that he's talking about is. But actually I think that's often very productive. Chris Olah, in Q and A, 32:54
The weights are the code. There's no simpler explanation than just showing you there's a window detector, there's a wheel detector, it wants to see the windows at the top, it wants to see the wheels at the bottom. Chris Olah, on what "turning it into code" actually means, 33:55
I think that we have an enormous advantage over neuroscience. In fact I wrote a whole article about all the things that make my life so much better than neuroscientists'. Chris Olah, 34:56
A lot of these language models know who I am. How often do I come up in text? One in 100 million tokens, one in a billion tokens? I don't know. Chris Olah, on the sparsity of a feature for himself, 37:32
The basic answer is: I have no idea. And I can give you a bunch of conceptual models for trying to reason about this, but I still then have no idea. Chris Olah, the last answer of the session, 40:36
A note on the captions, since the names matter here
The transcript below is the automatic caption track, and it mangles the vocabulary of this field harder than usual, which matters more than normal because these are terms of art with precise referents. Corrected once here rather than silently throughout, in the order they first appear:
- "mechanistic and purple", "interpoly", "mechanistic interprety", "mechanistic interability", "interpreter work" are all mechanistic interpretability.
- "confidence", "contents", "Content" are convnets, convolutional networks.
- "nearly very hard" and "nearly extremely difficult" are almost certainly merely very hard and merely extremely difficult, which is the opposite of what the caption says and is the actual claim of that passage.
- "super possession", "Superstition", "supervision", "Civilization", "supervisions", "super session", "reciprocision" are all superposition. The "supervision" substitution is the dangerous one, because "I am very worried about supervision" reads as a sentence about oversight when it is a sentence about superposition.
- "poly semantic" is polysemantic; "monosomatic" and "monosynthetic" are monosemantic.
- "compressed something" is compressed sensing.
- "courgeski 2012" is Krizhevsky et al. 2012, the AlexNet paper.
- "lackatosh, proof and reputation" is Imre Lakatos, Proofs and Refutations.
- "Neil Nanda" is Neel Nanda.
- "ground Trace connectum" is ground truth connectome; "fern eye related detectors" is fur and eye related detectors; "3D geometry reactors" is 3D geometry detectors; "pose ozone variant" is pose invariant; "intentional features" is attentional features.
- "non-preal amount of evidence" is a non trivial amount of evidence.
- At 1:00 and again at 1:31 the caption simply drops the word that carries the judgment, both times at the end of "all interpretability research is ___". The quotes on this page mark the gap rather than guess at it.
Quotes on this page are verbatim from the caption track with filler words removed, and nothing else changed except where a bracket marks a gap.
What happened next, eight months later
Olah ends the talk with superposition unresolved and names it as the thing the field rises or falls on. He does not describe a solution, and there was not one to describe in February 2023: nothing in this talk mentions dictionary learning or sparse autoencoders, so if you came here looking for that part of the Anthropic story, it is not in this video.
It arrived in October 2023, as Towards Monosemanticity: Decomposing Language Models With Dictionary Learning, published eight months after this talk. The approach is exactly the undo operation that the superposition hypothesis implies: if a layer's activations are a sparse combination of many more feature directions than there are neurons, then train a much wider sparse autoencoder to reconstruct those activations under a sparsity penalty, and the directions it learns should be the features the network folded together. They were more interpretable than the raw neurons, and the features that came out were nameable, specific things.
That is the sequel rather than the talk, and it is worth knowing which is which. Watch this video for the problem statement, stated by the person who named the problem, with the full argument for why it is the one that matters.
Where this sits in the LLM Learning track
This closes the Make it behave stage of the LLM Learning track, and it is the other half of a pair. John Schulman on reinforcement learning from human feedback comes immediately before it, and shapes a model's behavior from the outside by rewarding what you want and penalizing what you do not. This talk asks the complementary question: can you open the model and check what is in there?
Put next to each other, the two talks are the only two strategies anyone has for making a system trustworthy, and the honest summary of the pairing is uncomfortable. The outside in approach is deployed in every frontier model today and is imprecise by construction, because it can only ever reward the cases you thought to test. The inside out approach is the only one that could ever support the statement Olah actually wants, the one quantified over all situations the model could be in, and as of this talk it cannot reach the units it needs.
From here the track turns practical, with Jeremy Howard's hacker's guide to language models opening the Ship something with it stage. Worth carrying forward from this page: when you later build evaluations for a system you are shipping, you are doing exactly the behavioral testing whose limits Olah spends this talk circling. That is not a reason to skip the evaluations. It is a reason to know what they cannot tell you.
Resources mentioned
The talk and the event
- Looking Inside Neural Networks with Mechanistic Interpretability, the official recording page, which includes its own transcript
- FAR.AI, which runs and publishes the series, and the San Francisco Alignment Workshop, the inaugural one, listed as February 27 to 28, 2023. FAR.AI dates this recording February 26, 2023, a day before the listed dates
- Alignment Workshop, the series home
The speaker and the places this work is published
- Chris Olah, and his Wikipedia entry for the career he does not mention in the talk
- Anthropic, where he says he is, and where the Transformer Circuits work is done
- OpenAI, the host organization he refers to as "here", where he founded and led the Clarity interpretability team
- Distill, the journal he co founded, where the Circuits thread was published
- Transformer Circuits Thread, the continuation of that work on Transformers
The Circuits thread, on vision models
- Zoom In: An Introduction to Circuits, the entry point
- Circuits thread index, the serious attempt to reverse engineer InceptionV1
- Curve Detectors and Curve Circuits, the two papers on curve detectors, the second being the one that rebuilt the circuit from scratch and re implanted it
- An Overview of Early Vision in InceptionV1, the catalogue including high low frequency detectors, center surround units and boundary detectors
- Naturally Occurring Equivariance in Neural Networks, the weights that rotate with the feature
- Visualizing Weights, the method behind every weight picture in the talk
- Branch Specialization, the modularity answer in Q and A
- Feature Visualization, what the pictures are and how they are made
- Going Deeper with Convolutions, the InceptionV1 paper
- ImageNet Classification with Deep Convolutional Neural Networks, Krizhevsky, Sutskever and Hinton 2012, the two GPU model whose split specialized by itself
- OpenAI Microscope, the tool for browsing every significant unit of eight vision models, and the source of the
4c:447car detector decomposition
The Transformer Circuits thread
- Mechanistic Interpretability, Variables, and the Importance of Interpretable Bases, the long form of the source code analogy he opens with
- A Mathematical Framework for Transformer Circuits, the residual stream and attention head mathematics he skips
- In context Learning and Induction Heads, induction heads, the loss curve bump and the scaling law divergence
- Toy Models of Superposition, the sparsity sweep, five features in two dimensions, computation in superposition and the polyhedra. Also on arXiv
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learning, published eight months after this talk, which is where the superposition problem gets its first real attack
Named by him, or needed to follow him
- Neel Nanda's annotated reading list, which he calls possibly the best jumping off point: the version current at the time, and the second edition that replaced it
- Interpretability vs Neuroscience, the article behind the list of advantages in Q and A
- Scaling Laws for Neural Language Models, Kaplan et al., the original scaling laws paper with the divergence he points at
- Proofs and Refutations by Imre Lakatos, his precedent for working with a term you cannot yet define
- Compressed sensing, uniform polyhedra, singular value decomposition and Gabor filters, the mathematics he name checks in passing
- The Linux kernel, his yardstick for a hard reverse engineering problem
- Mechanistic interpretability, the field this talk helped name


