youtube.nixfred.com nixfred.com

Chris Olah - Looking Inside Neural Networks with Mechanistic Interpretability

Chris Olah's case for treating a neural network as an object to study rather than a black box to probe from outside: a trained model is a compiled binary with no source code, features are its variables, and circuits are the weights between features you already understand. He reads a dozen circuits straight out of InceptionV1, then breaks his own picture with polysemantic neurons and the superposition hypothesis, that a model simulates a larger, sparser network projected down and folded on top of itself, which is why five features fit in two dimensions and why no neuron means just one thing. He calls superposition the question the whole safety payoff rises or falls on, states the guarantee he actually wants as a claim quantified over every situation the model could be in, and closes on the motivation he says moves him most: gradient descent grows beautiful structure, and somebody should go look at it.

Published Sep 1, 2023 40:59 video 68 min read Added Jul 30, 2026 Open on YouTube →

At a glance

Chris Olah came to the inaugural San Francisco Alignment Workshop with three things he thought were worth saying, and he said all three in forty minutes while his slides refused to advance. The first was a short course in what mechanistic interpretability actually is: a neural network is a compiled binary with no source code, and the job is to recover the source. The second was awkward enough that he flagged it as awkward, a colleague of many years who one day told him, matter of fact, that obviously all interpretability research is worthless, and an open invitation to the skeptics in the room to come argue with him at office hours. The third was superposition, which he calls the question on which the whole safety payoff of the field will rise or fall.

The spine of the talk is the vision work. Features, weights, circuits: three kinds of object, and the claim that weights in isolation say nothing while weights between features you already understand are full of readable structure. He walks a dozen of them, from a dog head detector wired into a dog head plus neck detector to a car detector that wants windows at the top and wheels at the bottom. Then he breaks his own picture. Most neurons are polysemantic, and the hypothesis he offers is that the model is simulating a larger, sparser network that got projected down and folded on top of itself, which is why five features can live in two dimensions and why no neuron means one thing.

The back half is the honest accounting. What a feature is, which he says twice he cannot define and will not define in terms of human concepts. How Transformers differ, through the residual stream and attention heads and the induction heads whose formation shows up as a visible bump in the loss curve of every multilayer model he has looked at. What he wants out of safety, stated as a quantifier over all situations the model could be in. Where he is worried (superposition) and where he is not (scalability). And at the end, with no safety argument attached at all, the reason he says actually moves him: gradient descent grows beautiful structure, and somebody should go look at it.

Three things worth saying, and a laptop that would not cooperate

The talk opens twice. Olah thanks the room, says it is a privilege to be there, starts explaining that he spent time before the event asking himself what would really be useful to say, and then stops: "wait, this is totally not what I am presenting." The slide on the screen is his, just not the one he is on. "There is a great delay between" the laptop and the projector, and that delay runs through the entire talk as a running joke he keeps losing.

What he settled on was three things.

One: a flavor of what mechanistic interpretability is. He is explicit that this is for the part of the audience that is less familiar, and that interpretability is not one field. "There's a lot of different kinds of interpretability work going on," and he is only going to show one kind.

Two: the extent to which this work is worthless. This is the one he calls "a little bit more awkward to talk about," and the caption track drops the operative word twice, both times leaving the sentence hanging where the judgment should be.

I was very struck a few years ago. I had a colleague I'd worked with for many years, and one day they just matter of fact told me that, well, obviously all interpretability research is [the caption drops the word here]. And that really stuck with me, because it seemed to me that probably, you know, it was unusual for somebody to say such a thing so bluntly and directly, but it might be the case that quite a few people believed that. Chris Olah, on the second thing he wanted to say, 1:00

His conclusion is that a talk is the wrong instrument for that argument, because if you already believe it, nothing in the next forty minutes will land. So he converts it into an invitation. He has an office hour style session later in the day, and he would like the skeptics specifically.

To the extent that you sort of are skeptical of this kind of work, I would not be insulted. I would be privileged and delighted if you were willing to come and sort of trust me to talk about your doubts, and see if we can make some progress towards the truth together. Chris Olah, 1:31

Three: superposition. He gives it top billing before he has defined it, and the reason he gives for putting it in front of this particular room is not just that it matters.

The third thing I wanted to talk about is a challenge we call superposition, which increasingly seems to me like the question on which the impact of mechanistic interpretability for safety is going to rise or fall. Chris Olah, 2:02

He also pitches it as a research opportunity with unusually good economics: a question "that is very amenable to research without access to lots of compute," and one with "very rich mathematical structure." That is a deliberate recruiting line in a workshop full of people deciding what to work on next, and it lands differently from the usual "we need more people" because it names the constraint most of the room actually has.

The thirty second pitch: a binary with no source code

He calls it "like a 30 second pitch," and it is the cleanest framing of the field anyone has: the comparison is not biology, it is software engineering.

In regular software you have planning documents, maybe some goals, maybe a design doc. From those you write source code. Then you compile the source code into a binary. Three artifacts, and the middle one is the one humans read.

In a neural network you have a training objective you are optimizing, which is the goal. And then you "sort of directly turn that into neural network parameters, into a trained neural network." Two artifacts. There was "never anything like source code in the middle."

Really the goal of mechanistic interpretability is to somehow take those neural network parameters and turn them into something like source code. Chris Olah, 3:34

Then he pushes on the analogy, which he says is "actually a pretty deep analogy," and the mapping is item by item.

In a computer programIn a neural networkWhat that means here
VariablesA neuron, or a linear combination of neurons, representing a featureThe named things the computation is about. Sometimes a neuron is one. Often it is not, and that is the whole problem the second half of the talk is about
Program stateThe activations of a layerThe values the variables hold on this particular input
Processor or VM it runs onThe architectureFixed, written by humans, already understood
The compiled binaryThe parametersWhat training hands you. Complete, executable, and unreadable
The source codemissing"The thing that a computer program typically does have that a neural network does not is the source code, and so that's what we would like to get"
Figure 1. The analogy he opens with, line by line in his order. Four of the five rows already exist inside a trained model; interpretability is the project of producing the fifth. The row that does double duty is the first one, because by the end of the talk the honest entry there is not "a neuron" but "a direction in activation space that we cannot yet reliably isolate." The same analogy is written up at length in Mechanistic Interpretability, Variables, and the Importance of Interpretable Bases.

Middle ground between trivial and impossible

Before any results, he spends a minute on a framing move, and it is aimed at a specific failure of imagination he keeps encountering. People treat interpretability either as a thing that should be trivial and easy, or as a thing that is going to be impossible. He thinks that dichotomy is strange.

This is also where the slides fully break down. Mid sentence: "it's like, oh come on, update slide, please, please change. Yes, but it's on the right thing on the laptop, the laptop is going and switching, it's just not showing the right slide." A beat later he gives up on them: "well, I guess I will try to talk without my slides." They come back a few seconds later on their own.

The position he lands on is that interpretability may be merely very hard. His two calibration points are both deliberately unflattering to the optimists.

Harder than the Linux kernel. Any interesting real neural network is "probably harder to understand or reverse engineer than, say, a very complicated program like the Linux kernel." He flags that he does not know much about reverse engineering normal software, and that reverse engineering the kernel cold, knowing nothing about it, "seems like it would be a very hard challenge." That is the scale of difficulty he is claiming for networks. Not impossible. That hard.

Cells are also hard. The second point is the shape of the argument rather than the size of it: "no one's like, oh, biology, it's not easy therefore it must be impossible for us to understand cells." Difficulty is not an impossibility proof. "There's room for something to be merely extremely difficult in between."

And from that, the research strategy for the whole program: do not try to understand an entire network.

I think the goal of trying to understand an entire neural network is very hard, but we might be able to understand just portions of neural networks and then gradually grow that little portion, so we might be able to go and slowly, sort of rigorously, grow out from understanding tiny little portions to understanding larger chunks. Chris Olah, on the only strategy he thinks is available, 6:08

ConvNets, and the three kinds of object

To make any of this concrete he picks a model family: convolutional vision networks. The reason is partly biographical. "I was very fortunate for a number of years to work here, with our hosts, leading the interpretability team here at the time," he says, which is the single most understated line in the talk. The team was the Clarity team at OpenAI, which he founded and led before co founding Anthropic, and the work he is about to show is the Circuits thread in Distill, the journal he also co founded. He mentions none of that. He says they "made a lot of progress on understanding convnets" and moves on.

The way he likes to think about it is in terms of three kinds of object.

Curve detectors, which survive being taken seriously

His existence proof for the first object is curve detectors. In InceptionV1 "there's all these curve detectors, in fact these exist sort of in every vision model we've looked at." They look for curves at a particular orientation, and the reason he reaches for them rather than something flashier is that they have been audited harder than anything else in the field.

It turns out these hold up to extremely rigorous investigation. In fact we wrote an entire two papers just on curve detectors, and you can do all kinds of things to test these really are curve detectors, up to the point of going and rewriting the entire circuit from scratch from understanding it and re implanting them. Chris Olah, on how much evidence is available for one family of neurons, 7:08

The two papers are Curve Detectors, which is the evidence that the units are what they look like, and Curve Circuits, which is the part he describes as rewriting from scratch: the circuit was reimplemented by hand from the understanding of it and then put back into the model in place of the original, and it works. That is the strongest form of claim available about a piece of a network, because a reimplementation that functions is not a story about the weights, it is a replacement for them.

Weights alone are boring; weights between features are not

Then he does the move that the whole vision program rests on. Look at a weight in isolation, input neuron to output neuron, and "yeah, the weights by themselves, you know, they don't say that much." Nothing is legible there. But contextualize the same weights with features you already understand on both ends and they turn into readable code.

His first example is a dog head facing to the right with a long snout, wired into a dog head plus neck detector. The weight grid is positive at the offset where a head should be if the neck is there, and not on the other side.

You can see that now, all of a sudden, the weights make a lot of sense. We're having the dog head detector go and excite the dog head plus neck detector if it's on the right side of the dog, where the head should be. Chris Olah, 8:09

The second is the car detector, and it is the one that does the most work in the talk because he comes back to it twice more in the Q and A as his example of an explanation that cannot be simplified: a car detector that "goes and looks for windows on the top and wheels on the bottom."

CIRCUIT 1 · a dog head, placed where a neck expects it dog head detector, facing right + 0 weights, by relative position dog head plus neck detector excites at the one offset where a head belongs, not on the other side CIRCUIT 2 · windows on top, wheels on the bottom window detector car body detector wheel detector +++ + +++ one weight grid per input feature car detector InceptionV1 4c:447 the spatial arrangement IS the definition
Figure 2. Two of the circuits he reads out loud, drawn as he describes them. The point is not the pictures of dogs and cars, it is that each arrow carries a grid of weights indexed by relative position, and once both ends are features you understand, the grid is a sentence. Olah mentions the unit number 4c:447 only as a joke about unhelpful variable names; that it is specifically the car detector, assembled from window detector 4b:237 and wheel detector 4b:373, comes from the OpenAI Microscope announcement, which uses the same circuit as its motivating example.

Feature visualizations are variable names, not evidence

Someone asks, while he is mid walkthrough, what the pictures actually are. His answer is a careful downgrade of his own tool. The images are feature visualizations, "just created by optimizing the input," starting from random noise and pushing it until the neuron fires hard. And then, unprompted, the caveat:

The claim here isn't that that is, like, decisive evidence about what a neuron is doing. It's more like a variable name that we can put there. Chris Olah, on feature visualization, 9:10

The joke that follows is the best argument for the practice. The alternative to a picture is the unit's actual name, and the actual name is 4c:447. He has spent so long with InceptionV1 that he knows what 4c:447 is off the top of his head, and he knows that is not a transferable property. A visualization is "very suggestive of what it does and a little bit of evidence." If you want confidence you do the curve detector treatment: look at everything, run the synthetic stimuli, read the weights.

A dozen circuits in ninety seconds

Then he just lists them, fast, because the volume is the argument. All of these are from the Circuits thread on vision models, and the catalogue of early units is written up in An Overview of Early Vision in InceptionV1.

The summary he wants taken away is deliberately modest about coverage and immodest about what is there. The method for reading weights in context is written up in Visualizing Weights.

These are just a handful of examples. Neural network weights, once you start to contextualize them, are full of structure. Also lots of things we don't understand, but lots of structure. Chris Olah, closing the vision walkthrough, 11:12

The first real pushback: is a picture an explanation?

The question from the floor is the right one, and it is essentially whether the whole thing is circular. If you explain a car detector in terms of a window detector and a wheel detector, and you only know what those are because of another picture, what has been established?

His answer has two parts, and both matter for the rest of the talk.

First, the pictures are not the explanation. "These visualizations are just variable names. The actual thing that's analogous to, like, i equals one, or something like this, is the weights." The weights are "like the assembly instructions, or the code, of the computer program." The picture is the label on the variable; the weight grid is the statement.

Second, the regress is real and it terminates. "Then you could ask, okay well how do we trust those? Then we need to go back another step, and so on." Which is exactly how you read an unfamiliar program: you ground each definition in the ones below it.

It starts to look a lot like a computer program, where we have an understanding where we can base it on our understanding of previous variables, and if we want to go and understand those and really carefully understand those we have to go back further and further and further. And we can go and trace things back all the way to the input if we want. It's a lot of work. Chris Olah, 13:03

He takes one more question, sees the room filling with hands, and calls it: "I'm kind of tempted to ask to hold questions maybe to the end, because I do want to cover a lot of ground and I'm a little nervous that we won't get through it." The office hours offer comes back out.

Superposition: the picture he just painted is too optimistic

He says it himself, in the handoff: "the picture that I painted for you so far is a bit overly optimistic in an important way." The important way is that networks are also full of neurons that do not behave like the ones he just showed.

Neural networks are also full of what we call polysemantic neurons, neurons that respond to many unrelated features. Chris Olah, 14:04

The framing of the surprise is worth keeping, because it is the opposite of how this is usually told. The remarkable thing is not that some neurons are a mess. It is that any of them are clean.

In some ways it's miraculous that neural networks, when they have a privileged basis, when they have neurons, they have an activation function, that you get lots of neurons that just seem to correspond to meaningful features. But you also get many that don't. Chris Olah, on which half of the observation needs explaining, 14:04

The hypothesis

The hypothesis for why is superposition, and he states it as a want on the model's part: "the model wants to represent more features than it has neurons, possibly many more features than it has neurons." Once that is true the geometry is forced. "As soon as you have that, of course you can't align all the features with neurons, because there's only n neurons and you want to represent more things."

Then he takes it to its conclusion, which is the sentence that reorganizes everything before it.

If you take that really seriously, what it starts to suggest is that the model is actually sort of simulating a larger neural network. There is some larger, sparser neural network, and then it gets projected down and folded on top of itself to go and create the network that we actually observe. And then of course the neurons don't correspond to features, they correspond to linear combinations of features. Chris Olah, the central claim of the talk, 14:34

The polysemantic neuron is not a defect and not a mystery under this account. It is a projection artifact. You were reading an axis of the small network and hoping it was a variable of the large one.

What the toy models show

He is careful about the evidential status. "It's hard to really directly study this in real models, but it turns out we can show that this is exactly what happens in toy models, and then there's suggestive evidence that this happens in large models." The toy models are Toy Models of Superposition, published five months before this talk, and the experiment he describes is the sparsity sweep: features of varying importance, and a knob for how often they occur.

If you just start at zero percent sparsity, if the features are just completely dense, you only get as many features as you have dimensions. But as you make the features sparser, then because they're probably not going to go and co occur and interfere with each other, the model starts to be able to go and pack more things in, and eventually you end up with five features in a two dimensional space. Chris Olah, on the sparsity sweep, 15:35

Five in two is a specific number, and it is the whole hypothesis in miniature. Nothing about the network changed except how often the features show up. Sparsity is what buys the packing, because two features that are almost never both present can share a neighborhood of directions and almost never collide.

And then the result that makes superposition a problem rather than a curiosity about storage: it is not just representation.

It turns out you can actually do computation as well in superposition. You can in some sense have a computational graph that's higher dimensional, project it down, and actually do useful computation while holding everything in superposition. Chris Olah, 15:35

0% sparsity: 2 features, 2 dimensions sparse features: 5 features, still 2 dimensions feature A feature B one feature per neuron, zero interference A B C D E each neuron axis now answers to several features: polysemanticity sparsity rises
Figure 3. His toy model result drawn at the size he quotes it: two features in two dimensions when features are dense, five in the same two dimensions once they are sparse. The left panel is the only regime in which reading one neuron at a time is a valid way to read a model. Everything the talk calls polysemanticity is the right panel seen through the vertical and horizontal axes.

Which features you get to see, and which you do not

A consequence he draws out, and it is the most uncomfortable one in the talk, is that the clean neurons are a biased sample. If superposition is real, the features that got an axis to themselves got it for a reason.

What you start to think, when you observe these neurons that are sort of monosemantic, is that those are particularly important and common features, and then the things that we don't observe, the sort of less important, more dispersed things, are in superposition. Chris Olah, 16:06

So the part of the model that is legible is the part that is frequent and important. The rare machinery is exactly the machinery that is hiding, which is a bad property for a method whose purpose is finding things you did not think to look for. That is the step from "superposition is interesting" to "superposition is the problem."

Unless you understand the superposition structure, there will always be the potential for unknown behaviors to just sort of suddenly occur. This is something that I'm very worried about. And you won't be able to understand really what the weights are actually doing. Chris Olah, 16:06

He then draws the boundary of the claim honestly, which is characteristic. He is not saying nothing can be done without solving it: "I think there are other valuable things one can do within mechanistic interpretability without this, but I think that we will surrender a great deal if it turns out that we can't go and address this somehow." And he expects the problem is not his alone: "I think it probably holds for a lot of other approaches as well."

The consolation prize is that the problem is gorgeous

Having called superposition the thing that could sink the field, he spends thirty seconds on why he enjoys it anyway, and this is the recruiting pitch from the intro cashing out.

It turns out though that superposition is full of beautiful mathematical structure. It is deeply connected to compressed sensing. It turns out that in toy models at least, these features organize themselves into polyhedra, into regular polyhedra, which is kind of wild, or uniform polyhedra. It's very not obvious that they should do this. It turns out that the learning dynamics involve these weird, like, electron jumping like behaviors that are very strange. It's very mysterious. Chris Olah, 17:09

Three concrete hooks in one breath: a connection to compressed sensing, feature geometry that lands on uniform polyhedra with no one asking it to, and learning dynamics where features appear to jump between configurations like electrons between shells rather than sliding smoothly. None of it needs a cluster.

"I don't have a super compelling definition of what a feature is"

The question he takes next is the one that threatens the foundation, and he does not defend the foundation.

I actually think that defining what we mean by feature is extremely hard, and I don't have a super compelling definition of what a feature is. Chris Olah, 17:40

What he offers instead is a chain of partial answers, in increasing order of how much he believes them.

The intuition pump. "Intuitively what we mean is something like a curve detector, or a car detector, or things like this. And sometimes those end up corresponding to neurons, but sometimes we can show that they correspond to linear combinations of neurons."

The place where it is exact. In the toy models the ambiguity vanishes, because the features are an input to the experiment rather than an output: "we can construct toy problems where we do know what features are and where that's exactly defined, and then they will go and represent them in this way." That is why he keeps returning to that paper. It is the only setting where the thing being claimed about real networks has a ground truth.

The definition he rejects. One obvious move is to say a feature is a human understandable property of the input that the network represents. He will not take it: "I don't want humans to be sort of involved in the definition." The cost of that definition is that it makes the existence of a feature contingent on a person recognizing it, which forecloses the possibility he most wants to keep open.

The definition he actually holds. Stated as a belief he cannot yet formalize, which is a rare register for a technical talk.

The thing that I sort of in my heart believe a feature is, but I don't know how to properly define this, is something like: a feature is a fundamental unit of the neural network's computation. And those don't necessarily correspond to neurons. They sometimes do, and they often appear not to. Chris Olah, 18:40

How Transformers are different

He flags that he is going to be brief here, and why: "this actually becomes quite messy and ends up with a lot of detailed mathematics, which I think is actually very cool mathematics, but it's a little bit harder to go through and talk." The long version is A Mathematical Framework for Transformer Circuits. What he gives instead is two architectural differences, plus a third observation he mentions almost in passing and which is the one the rest of the talk hangs on.

The third one first: there is more superposition. "You just see a lot of superposition in Transformers. You see a lot more than you do in vision models." He does not explain why here. He explains why forty minutes later, in the final stretch of Q and A, and that answer is the most quantitative thing in the talk.

Difference one: the residual stream. It is not quite a ResNet, he says, and the reason it matters is what the addition implies. Because every layer reads from and writes directly into the same stream, projecting in and out of it rather than passing a value hand to hand, there are "sort of these implicit latent weights connecting neurons across layers, or attention heads across layers, or all these things." Two components in distant layers can be in communication through the stream without anything in the architecture naming that connection. "That creates a lot of interesting structure that doesn't exist in quite the same way in any convnets really, because even ResNets aren't exactly like this."

Difference two: attention heads. They are a genuinely new kind of object for this kind of analysis, not just more neurons. "Attention heads kind of create a new sort of fundamental unit of mechanistic interpretability, similar to features, weights or activations," and with them come attentional features, of which he says there are "lots of very rich and interesting" examples. Here the slides fail him again on cue: "the other is attention heads. Oh no, apparently I'm not gonna talk about that more in a second."

Induction heads, and the one he finds most interesting

The example he picks is the one that connects a single attention pattern to a macroscopic property of training. If you look at attention patterns you often see "these weird off diagonals. They're very striking."

It turns out this corresponds to a particular type of attention head called an induction head, which searches through for previous cases where something happens, and then looks forward one step, and then goes and copies that and goes and outputs that. Chris Olah, 21:12

1. search back for a previous occurrence of this token ... A B ... ... ... A current token 2. step forward one 3. copy whatever followed it, and output that predict B
Figure 4. The induction head, as he describes it: search back, step forward one, copy. Step one is the off diagonal stripe that makes these heads visible in an attention pattern in the first place. The reason this one circuit gets a third of his Transformer section is that it is the only mechanism in the talk whose formation is visible from outside the model, as a bump in the loss curve and as a kink in the scaling laws. The full treatment is In context Learning and Induction Heads.

This mechanism is how a model does a certain kind of in context learning: you saw this token before, here is what came next, do that again. And he thinks a generalized version of it is doing much more than that: "in fact it turns out that sort of a generalized version, it seems to be the driver of in context learning, or at least there's a non trivial amount of evidence for that hypothesis."

Then the part that makes induction heads special among all the circuits in the talk. They are the one piece of microscopic machinery whose arrival you can see from the outside.

In fact these are so important that they create a bump in the loss curve when you train Transformers. It's just a visible bump in every Transformer loss curve that I've seen, if you have more than two layers, because that's what's necessary for induction. Chris Olah, 21:42

Two layers is not an arbitrary threshold. An induction head has to compose with an earlier head to work, so the mechanism cannot exist in a one layer model, and the bump appears exactly when it becomes possible. The effect also reaches the scaling laws: "if you look at the scaling laws, the original scaling laws paper, you'll see that there's a point where there's sort of a divergence from the original trend, and that corresponds to the models not having induction heads and then forming them."

These inductions are so important they appear in the macroscopic picture. We have this microscopic picture that we're developing here, and it bridges all the way to the macroscopic. Chris Olah, on why he keeps telling this particular story, 21:42

Do they all form at once?

Yes, and the questioner clearly already knew that, because the follow up is about the mechanism. Olah confirms it: "a massive number of induction heads simultaneously occur," and there seems to be a deep reason.

There's a number of pieces that have to exist for an induction head circuit to form, and once you start to have those ingredients you then get an extremely fast feedback loop, and then all of them form simultaneously because the ingredients are in place. Chris Olah, 22:43

What that looks like empirically: "if you just track the existence of induction heads in a Transformer over time, they will just sort of all form in a very small window," and then maybe evolve a bit further, with the occasional late one, "but the vast majority of them form at once."

He adds a control that is the clearest evidence the bump and the circuit are the same event. "We have a slightly similar experiment where you can go and set up an architecture that very easily learns induction heads, and then induction heads form right at the start." In that setup the formation happens inside the first, extremely steep part of the loss curve, so there is nothing to see: "you shouldn't see an analogous bump in the loss curve." Move the circuit's arrival and the macroscopic signature moves with it.

What he actually wants out of safety

There are many ways interpretability might help with safety, and he deliberately picks one, the one he finds motivating. He starts from the shape of the guarantee rather than from any technique.

It seems to me like the thing that we most want out of safety is something like the ability to make statements of a kind, you know, something vaguely like: for all the situations the model could be in, the model will never deliberately do X. We want to be able to say something like that. Chris Olah, 23:45

Read that as the logic it is. It is a universal quantifier over situations, and no amount of testing gets you one, because testing is a finite collection of examples and the statement is about everything that is not in your collection. That is the entire argument for looking inside, and he never has to say it out loud.

His best guess at how such a statement could ever be earned is where features and circuits come back as the load bearing pieces.

My present best guess as to how you could achieve that is that you want to be able to say something like: okay, there don't exist features such that the model will deliberately, that will participate in the model deliberately doing X. And that's going to then be some kind of claim about the circuits that feature participates in. Chris Olah, 24:15

He immediately marks the distance. "This is a very ambitious and wild and kind of crazy thing to be aiming for. I don't mean to say that this is anywhere near being able to go and make this kind of statement, but this is kind of a spiritual North Star, I think, for the most ambitious way that mechanistic interpretability could help with safety."

THE STATEMENT WE WANT For all the situations the model could be in, the model will never deliberately do X. There do not exist features such that the model will participate in deliberately doing X which is then a claim about the circuits those features participate in CHALLENGE 1: SUPERPOSITION can we access the features at all, and know that we have them all? his verdict: very worried CHALLENGE 2: SCALABILITY tiny understood circuits versus enormous frontier models his verdict: more optimistic
Figure 5. The safety case in his own order, with his own confidence attached. The first box is the only one of the three that a behavioral test could even pretend to address, and it is a universal quantifier over situations, which no finite test can reach. The asymmetry in the two verdicts is the whole reason superposition got top billing in the opening minutes: a scalability problem is an engineering schedule, while a supervision problem means you do not have the variables yet.

Two challenges, and he is only worried about one

Superposition. Stated here as an access problem: "how can we actually access the features? How can we know that we've got everything? How can we rule out these places where the model might have something, a feature that activates very rarely, and it's just some direction in activation space where it behaves totally differently?" That last clause is the safety relevant failure in a single sentence. A rare direction with different behavior is exactly the thing the universal quantifier was supposed to exclude, and it is exactly the thing superposition hides.

Scalability. "If these models are so large, and the successes of mechanistic interpretability so far are so small, they're these tiny little circuits that we're understanding, how can we hope to get to these large models?"

And then the verdict, which is not symmetric.

I am very worried about superposition. I am more optimistic about scalability. Chris Olah, 25:15

The reasons for the optimism are three concrete directions, all of which have since become their own subfields:

But there is a dependency between the two challenges, and it runs the wrong way for anybody who wants to work on the easier one.

Unfortunately I think this is a really challenging problem to work on right now, because we don't actually have access to the units that we want to understand. If we think the features are there and they're all in superposition, they're a giant mess, and it's very difficult to start working on how you can even actually really effectively scale. Chris Olah, on why scalability is hard to attack today, 25:46

You cannot design a method for reading a million circuits while you are unsure what a circuit is made of. Superposition is not just the risk, it is the blocker on the thing he is optimistic about. "My real fear is this challenge of superposition."

The other reason he does this, which has nothing to do with safety

He checks the clock, notes he is "right on schedule," and spends the last ninety seconds on something he labels as emotional rather than argumentative. It is the most quoted part of the talk for good reason.

The claim is about what kind of science this is. Machine learning gets compared to math and physics, where elegance means compression: short laws, clean derivations. He thinks the better comparison is biology, where elegance means structure you did not design and could not have predicted.

In the case of biology, evolution creates awe inspiring complexity in nature. You know, we have this system that goes and generates all this beautiful structure. And in a kind of similar way, gradient descent, it seems to me, creates beautiful, mind boggling structure. It grows, and it creates all these beautiful circuits that have beautiful symmetries in them, and in toy models it arranges its features into regular polyhedra. Chris Olah, 26:48

Then the part that is a methodological commitment disguised as enthusiasm: the mess is provisional. "There's messy stuff, but then you discover sometimes that that messy stuff actually was just a really beautiful thing that you didn't understand." High low frequency detectors were incomprehensible before they were trivial. That record is the reason he reads an unexplained unit as unexplained rather than as noise.

So a belief that I have, that is perhaps only semi rational: these models are just full of beautiful structure, if we're just willing to put in the effort to go and find it. And I think that's the thing that I find actually most emotionally motivating about this work. Chris Olah, closing the prepared talk, 27:48

Where he tells you to start

He closes with credit and three pointers. The credit first: "all this work was done in collaboration with many others, including others at different institutions," and then, as the only self identification in forty minutes, "I'm at Anthropic."

The three places to start, in his order:

  1. The original Circuits thread, "which made a really serious attempt to reverse engineer one vision model, InceptionV1." The entry point is Zoom In: An Introduction to Circuits.
  2. The Transformer Circuits thread, "which has been trying to do analogous work on Transformers."
  3. Neel Nanda's annotated reading list, which he singles out: "which actually might be the best jumping off point." The version current at the time of the talk is An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers; it has since been replaced by a second edition.

Q and A, in order

The last eleven minutes are questions, and they are not filler. Two of them produce material that is not anywhere in the prepared talk, including the only numbers he gives all day.

Modularity: the branches are already there, you just have to look

The question is about whether you could build models that are easier to interpret, by making them modular on purpose. His answer starts with an accident from 2012.

In the Krizhevsky et al. 2012 paper, the one that is usually called AlexNet, the model was split across two GPUs that "only talk to each other every other step," for hardware reasons. The striking phenomenon: look at the first layer and the two halves specialized. "There's all these Gabor black and white Gabor filters on one side, and color contrast detectors on the other." Look at the next layer and you see similar things.

It generalizes. "Lots of vision models have branches, and like InceptionV1 has branches, and it turns out that like this branch over here has all of the black and white versus color and color contrast units, this one's full of curve detectors and some sort of fur and eye related detectors and some boundary detectors, and this one's full of 3D geometry detectors." That is Branch Specialization in the Circuits thread.

So one answer to the question is: fine, build models with branches like that. But then he gives the better answer, and it reverses the premise of the question. You do not have to impose the structure, because it is already in there.

It turns out that the structure is implicitly there. You can just go and take the magnitude of the connections between features in one layer and features in another layer, and go and do a singular value decomposition, and you discover that actually all the same structure was implicitly there. Chris Olah, on finding modularity in a model nobody made modular, 30:27

Where it stops working is the tell. "In practice this becomes hard on the later layers of vision models, you don't see this as much, and in language models we don't see it." His explanation is the same one as always, and it is the sharpest statement of what superposition actually does to a model.

I think that's because of superposition. The best case for superposition is to go and fold two completely unrelated computational graphs on top of each other. Chris Olah, 30:58

Two unrelated graphs are the ideal thing to superpose, because they never need to be active at the same time, so folding them together costs almost nothing in interference. Which means the models with the most modular structure available to find are exactly the models where superposition has the most to gain by hiding it. Hence the optimism that goes with the complaint: "one of the things I'm sort of most optimistic about is that if you can solve superposition you'll discover that there's just lots of implicit modularity and other kinds of organizational structure to these graphs." The evidence he offers for that bet is the early layers of vision models, "kind of the best proxy I have for a model that has very little superposition," where the structure is visible.

What is a feature, take two: Lakatos, and the weights are the code

Another questioner comes back to the definition problem, and this time he gives the epistemology instead of the definition. First, the self aware version of the problem:

You might be like, oh that's kind of wild, right? Like Chris is spending his whole career on this thing and he doesn't even know what the definition that he's talking about is. But actually I think that's often very productive. Chris Olah, 32:54

The authority he cites is Imre Lakatos and Proofs and Refutations, which he describes as "this play about a bunch of math students going and being confused about definitions and progressively finding more and more useful definitions for things." That is a precise citation for what he is doing: the definition of a feature is expected to be an output of the research program, not an input to it.

Then the reason he will not accept the convenient definition, stated more directly than the first time:

I don't want to define features in terms of human understandable concepts, because I want to allow for the possibility that these models have concepts that we don't understand. Chris Olah, 33:25

What he is after is "in some sense pulling apart the fundamental units of computation, whatever they are," and he adds the honest empirical footnote: "so far in the cases where we've done this, they have been things that we've understood."

This leads to the best clarification in the talk, because it corrects a reading of his own software analogy that most people take away from it. "Turning them into code" does not mean producing Python.

The weights are the code. For the car detector that I showed you earlier, there's no simpler explanation than just showing you, you know, there's a window detector, there's a wheel detector, it wants to see the windows at the top, it wants to see the wheels at the bottom. That is the most understandable depiction of that that there is. Chris Olah, 33:55

And the condition on that, which is the thesis of the entire vision section in one line: "but you couldn't have understood if I didn't contextualize it." The work is not translation. The work is contextualization, "and then very painstakingly trying to go through it."

His example of what that process actually feels like is the high low frequency detectors. Nobody understood them at first. In retrospect they are embarrassingly simple: high frequency on one side, low frequency on the other, "and this occurs all the time at boundaries of objects, because you have something that's out of focus in a background, or the background is busy and [the object] is a solid color." A concept the model invented, that no human would have thought to name, that turns out to be an obvious thing to want once you see it.

The advantages over neuroscience

He takes a detour he clearly enjoys, because the comparison to neuroscience comes up constantly and he thinks it is backwards.

I think that we have an enormous advantage over neuroscience. In fact I wrote a whole article about all the things that make my life so much better than neuroscientists'. Chris Olah, 34:56

The article is Interpretability vs Neuroscience, posted in March 2021. Then he enumerates, from memory, for the people who have not thought about it.

What he can doWhat it rules out
Ground truth access to the neuron's computational structureNo inference from indirect measurement. The thing itself is on disk
The activations of every neuron, non destructively, for as many stimuli as he wantsNo electrode budget, no choosing which cells to reach, no limit on the stimulus set, no killing the sample to read it
The ground truth connectomeNo reconstruction project. He stresses it is better than a connectome: "not just which neurons are connected to each other, but we know which ones excite and inhibit each other"
Gradients through the structure, and optimization through itNo black box probing. This is also what makes feature visualization possible at all
The same model again and againNo individual variation. Neuroscience has to switch between organisms "which might be different"
Handing the exact model to a colleague"We can go and ask questions of exactly the same model and talk about the same neurons." Findings compose because the object is shared
Far faster experiment cyclesNo wet lab. This is the one he spends the least time on and that changes the most in practice
Figure 6. The list he rattles off at 35:31, with what each advantage removes. His own write up adds three more that he skips here: weight tying, which collapses the number of distinct units to study in early vision; feature visualization as a causal instrument, because an input built from scratch to drive a neuron contains only what drove it; and interventions, where ablating a unit or editing a circuit is finer grained than optogenetics. His summary of the whole comparison: "it's just this, like, truly enormous difference."

Why language models carry so much more superposition

This is the question that pays the best, because it is the one place he puts numbers on the problem. He names two factors, and says language models go the wrong way on both.

Factor one: the feature importance curve. Order every feature the model could learn by importance, "which is like something like how much it reduces the loss," and look at the shape of the curve. "The flatter that curve is, the more superposition you're going to have. The steeper it falls off, the less superposition you're going to have." A steep curve means a handful of features carry nearly all the value, so a model with limited dimensions can simply keep the top few and drop the tail. A flat curve means the thousandth feature is nearly as valuable as the tenth, and dropping the tail is expensive, so the model packs instead.

Factor two: sparsity. And here he gives the contrast that makes it vivid, using the room he is standing in.

If you look at this thing behind me and you ask how many vertical edges are there in just a single image, you know, there are a lot of vertical edges. On the other hand, a lot of these language models know who I am. How often do I come up in text? One in 100 million tokens, one in a billion tokens? I don't know. I don't think I'm very common. Those actually feel like they're probably too small a number. Chris Olah, on the sparsity gap between vision and language, 37:32

An edge detector in early vision fires many times per image. A "Chris Olah" feature fires on the order of once per hundred million tokens, and he suspects even that overstates it. That is the gap that makes superposition pay in a language model, because the rarer the feature, the cheaper it is to pack.

He adds two more pieces of the same answer:

dimensions the model actually has everything to the right is still worth representing, so a flat curve has to pack rather than drop early vision: steep, short language models: flat, long features, ordered by importance importance: how much it reduces the loss his other axis, sparsity, moves the same way: vertical edges in a single image, many per image, versus a feature for his own name, which he guesses at one in 100 million to one in a billion tokens.
Figure 7. The feature importance curve, drawn as the shape claim it is. No scale is marked because he gives none; the whole argument is in the slope. Steep and short means a model can keep the head of the distribution and discard the tail, which is the regime where one neuron can afford to mean one thing. Flat and long means the tail is still worth money, and since the dimension budget does not move, the model buys the tail on credit by sharing directions. Both of his factors, curve shape and sparsity, point the same way for language, which is why the clean neuron stories in this talk are all from vision.

Are language neurons harder to read than vision neurons?

No, and he thinks this is a red herring. His evidence is tooling. At OpenAI they built Microscope, which shows every significant unit of eight vision models; there is "a similar internal thing for language models," which the captions render as lexoscope. Both show you the dataset examples that drive a unit.

And then the one genuine advantage language has, which he is openly grateful for:

One really cool thing you can do in language models that you can't do in image models is you can just edit the text and see how the neuron activations change in real time. That's like a thing that I'm very grateful for. Chris Olah, 39:04

His conclusion: a language neuron is not harder to understand or recognize. "Maybe it takes a little bit more time, because you can maybe more quickly visually go through a bunch of visual dataset examples than language ones, but I think not in a very significant way. So I think that's not the primary thing going on." The difficulty with language models is superposition, not legibility.

The last question, and a flat "I have no idea"

He is out of time, takes one more, and the question is partly inaudible on the recording: something about human concepts, and how superposition would change if the model's concepts were less human like.

He gives the structural part of the answer. Superposition "depends on the gap between the sparsity of the neurons and the sparsity of the fundamental underlying data and the fundamental underlying features," and if those features really are extremely rare, "one in a million tokens, one in 10 million tokens, then there's still an enormous, enormous sparsity gap." So the packing pressure does not go away. The part he cannot answer is the other factor, "how do you think the feature importance curve changes," and he does not pretend otherwise.

The basic answer is: I have no idea. And I can give you a bunch of conceptual models for trying to reason about this, but I still then have no idea. Chris Olah, the last thing he says before thanking the room, 40:36

Key takeaways

Chapters

This chapter map was built for this page from the transcript's own timings, because the video publishes no chapters of its own. Every timestamp above is clickable and seeks the embedded player.

Notable quotes

I had a colleague I'd worked with for many years, and one day they just matter of fact told me that, well, obviously all interpretability research is [the caption drops the word here]. And that really stuck with me. Chris Olah, on the second of his three topics, 1:00

To the extent that you sort of are skeptical of this kind of work, I would not be insulted. I would be privileged and delighted if you were willing to come and sort of trust me to talk about your doubts, and see if we can make some progress towards the truth together. Chris Olah, turning a talk he cannot give into an office hour, 1:31

The third thing I wanted to talk about is a challenge we call superposition, which increasingly seems to me like the question on which the impact of mechanistic interpretability for safety is going to rise or fall. Chris Olah, 2:02

Really the goal of mechanistic interpretability is to somehow take those neural network parameters and turn them into something like source code. Chris Olah, the thirty second pitch, 3:34

These are just a handful of examples. Neural network weights, once you start to contextualize them, are full of structure. Also lots of things we don't understand, but lots of structure. Chris Olah, after a dozen vision circuits in ninety seconds, 11:12

These visualizations are just variable names. The actual thing that's analogous to, like, i equals one, is the weights. The weights are like the assembly instructions, or the code, of the computer program. Chris Olah, answering whether a picture counts as an explanation, 12:33

In some ways it's miraculous that neural networks, when they have a privileged basis, when they have neurons, that you get lots of neurons that just seem to correspond to meaningful features. But you also get many that don't. Chris Olah, on which half of polysemanticity needs explaining, 14:04

There is some larger, sparser neural network, and then it gets projected down and folded on top of itself to go and create the network that we actually observe. And then of course the neurons don't correspond to features, they correspond to linear combinations of features. Chris Olah, stating the superposition hypothesis, 14:34

If you just start at zero percent sparsity, if the features are just completely dense, you only get as many features as you have dimensions. But as you make the features sparser, the model starts to be able to go and pack more things in, and eventually you end up with five features in a two dimensional space. Chris Olah, on the toy models sparsity sweep, 15:35

Unless you understand the superposition structure, there will always be the potential for unknown behaviors to just sort of suddenly occur. This is something that I'm very worried about. Chris Olah, 16:06

I actually think that defining what we mean by feature is extremely hard, and I don't have a super compelling definition of what a feature is. Chris Olah, on the central term of his field, 17:40

The thing that I sort of in my heart believe a feature is, but I don't know how to properly define this, is something like: a feature is a fundamental unit of the neural network's computation. Chris Olah, 18:40

It turns out this corresponds to a particular type of attention head called an induction head, which searches through for previous cases where something happens, and then looks forward one step, and then goes and copies that and goes and outputs that. Chris Olah, defining induction heads, 21:12

It's just a visible bump in every Transformer loss curve that I've seen, if you have more than two layers, because that's what's necessary for induction. Chris Olah, on seeing a circuit form from outside the model, 21:42

It seems to me like the thing that we most want out of safety is something like the ability to make statements of a kind, something vaguely like: for all the situations the model could be in, the model will never deliberately do X. Chris Olah, on the shape of the guarantee, 23:45

I am very worried about superposition. I am more optimistic about scalability. Chris Olah, 25:15

In the case of biology, evolution creates awe inspiring complexity in nature. And in a kind of similar way, gradient descent, it seems to me, creates beautiful, mind boggling structure. Chris Olah, 26:48

So a belief that I have, that is perhaps only semi rational: these models are just full of beautiful structure, if we're just willing to put in the effort to go and find it. And I think that's the thing that I find actually most emotionally motivating about this work. Chris Olah, closing the prepared talk, 27:48

You might be like, oh that's kind of wild, right? Like Chris is spending his whole career on this thing and he doesn't even know what the definition that he's talking about is. But actually I think that's often very productive. Chris Olah, in Q and A, 32:54

The weights are the code. There's no simpler explanation than just showing you there's a window detector, there's a wheel detector, it wants to see the windows at the top, it wants to see the wheels at the bottom. Chris Olah, on what "turning it into code" actually means, 33:55

I think that we have an enormous advantage over neuroscience. In fact I wrote a whole article about all the things that make my life so much better than neuroscientists'. Chris Olah, 34:56

A lot of these language models know who I am. How often do I come up in text? One in 100 million tokens, one in a billion tokens? I don't know. Chris Olah, on the sparsity of a feature for himself, 37:32

The basic answer is: I have no idea. And I can give you a bunch of conceptual models for trying to reason about this, but I still then have no idea. Chris Olah, the last answer of the session, 40:36

A note on the captions, since the names matter here

The transcript below is the automatic caption track, and it mangles the vocabulary of this field harder than usual, which matters more than normal because these are terms of art with precise referents. Corrected once here rather than silently throughout, in the order they first appear:

Quotes on this page are verbatim from the caption track with filler words removed, and nothing else changed except where a bracket marks a gap.

What happened next, eight months later

Olah ends the talk with superposition unresolved and names it as the thing the field rises or falls on. He does not describe a solution, and there was not one to describe in February 2023: nothing in this talk mentions dictionary learning or sparse autoencoders, so if you came here looking for that part of the Anthropic story, it is not in this video.

It arrived in October 2023, as Towards Monosemanticity: Decomposing Language Models With Dictionary Learning, published eight months after this talk. The approach is exactly the undo operation that the superposition hypothesis implies: if a layer's activations are a sparse combination of many more feature directions than there are neurons, then train a much wider sparse autoencoder to reconstruct those activations under a sparsity penalty, and the directions it learns should be the features the network folded together. They were more interpretable than the raw neurons, and the features that came out were nameable, specific things.

That is the sequel rather than the talk, and it is worth knowing which is which. Watch this video for the problem statement, stated by the person who named the problem, with the full argument for why it is the one that matters.

Where this sits in the LLM Learning track

This closes the Make it behave stage of the LLM Learning track, and it is the other half of a pair. John Schulman on reinforcement learning from human feedback comes immediately before it, and shapes a model's behavior from the outside by rewarding what you want and penalizing what you do not. This talk asks the complementary question: can you open the model and check what is in there?

Put next to each other, the two talks are the only two strategies anyone has for making a system trustworthy, and the honest summary of the pairing is uncomfortable. The outside in approach is deployed in every frontier model today and is imprecise by construction, because it can only ever reward the cases you thought to test. The inside out approach is the only one that could ever support the statement Olah actually wants, the one quantified over all situations the model could be in, and as of this talk it cannot reach the units it needs.

From here the track turns practical, with Jeremy Howard's hacker's guide to language models opening the Ship something with it stage. Worth carrying forward from this page: when you later build evaluations for a system you are shipping, you are doing exactly the behavioral testing whose limits Olah spends this talk circling. That is not a reason to skip the evaluations. It is a reason to know what they cannot tell you.

Resources mentioned

The talk and the event

The speaker and the places this work is published

The Circuits thread, on vision models

The Transformer Circuits thread

Named by him, or needed to follow him

Full transcript
[00:00:00] welcome Chris thank you so um it's really a privilege uh to be here speaking to you today and before I came I spent some time sort of asking myself what what would it really be useful and valuable and interesting for me to say to you today because it why did it go back to [00:00:30] wait this is totally not what I am presenting um uh uh I mean this is my slide but it's not on the there okay there we go we are on the right oh no there is a great delay between well so uh yeah so I I was I was really wondering you know what what would it be valuable and useful for me to talk to you about and it felt to me like like one thing which was was very natural as well you know I could talk a little bit about what mechanistic and purple is about because you know for those of you there's a lot of different kinds of interpreter work going on and for those [00:01:00] of you who are less familiar I could try to give you a bit of a flavor and that might be useful um but uh it also seems to me and this is a little bit more Awkward to talk about that um it might be useful to talk about you know the extent to which this work is um because uh you know I'm I was I was very struck a few years ago I had a colleague um I'd worked with for many years um and one day they just matter of fact told me that you know well obviously all interpoly research is um and uh I that really stuck with me [00:01:31] because it seemed to me that probably um you know they were it was unusual for somebody to say such a thing so sort of so bluntly and directly but you know it might be the case that quite a few people believed that um and to the extent that you believe that I can't say anything you stole this talk um because you will not believe it um unfortunately I think this is going to be difficult for me to address um in a talk format but it turns out that I'm going to be doing some kind of sort of office hour-like thing uh later and so to the extent that you um you sort of are skeptical of this kind of work um I would not be insulted I would be privileged and delighted if you were [00:02:02] willing to come and sort of trust me to to talk about your doubts um and see if we can make some progress towards the truth together um the third thing I wanted to talk about is a challenge we call superposition um which increasingly seems to me like the question on which the impact of mechanistic interpretability for safety is going to rise or fall um and I want to highlight it because it seems like such a a central and important question to me um but I also want to highlight it to you because I believe it as an attention [00:02:33] a question that might be worth your time not only because of its importance um but because I believe that's a question that is very amenable to research without access to lots of compute um and which also has very rich mathematical structure so I will talk about some other things I'll talk about some interesting interesting fun results at some points um but these are really the three things that I felt like I could say that might be useful and important to you okay so we'll start with a little bit of introduction to what mechanistic type really is um [00:03:03] there's like a 30 second play um so in in regular software engineering you know we have we have planning documents so we might write some some goals or have some kind of design doc for our software and then we write some source code and then we compile that source code into um a compiled binary um but for neural networks um you know we we might go and have we have a training objective that we're optimizing that's our goal um and then we sort of directly turn that into neural network parameters into a trained neural network and there was [00:03:34] never anything like source code in the middle and so really the goal of mechanistic interpretability is to somehow take those neural network parameters and turn them into something like source code um and I think this is actually a pretty deep analogy so um you know we can think of this in lots of ways computer programs have variables I think in a neural network that's roughly analogous to something like a neuron or a linear combination of neurons that represent a feature um a computer program has a state I think that's kind of analogous to the activations of a neural network layer [00:04:04] um the neural network has a processor or VM that it runs on that's I think the neural network architecture um and then we have at the end this this binary and that's kind of analogous to the neural network parameters but the thing that a computer program typically does have that a neural network does not is the source code and so that's what we would like to get um I think an interesting uh Point here is that there's there's Middle Ground between trivial and impossible so I feel like when we're talking about interpretability we often I hear people [00:04:35] either sort of treated like a thing that should be trivial and easy um or a thing that's going to be impossible but it seems to me um that that that's sort of a strange dichotomy and I it in fact feels to me again this is just it's like oh come on update slide please please change yes but it's on the right thing on the laptop the laptop is going and switching it's just not showing the right slide on the uh yeah um [00:05:05] uh well I guess I will I will try to talk without my slides um it seems to me that it might be the case that interpretability is nearly very hard but not impossible and that you can have something between trivial and impossible so um I you know I often imagine that even you know sort of any any interesting real neural network is probably harder to understand or here we go or a reverse engineer um then um uh a um uh then then say a very complicated program like the Linux kernel you know uh like if I if I was [00:05:36] trying to reverse on here the Linux kernel without knowing anything about it I mean I don't know much about reverse engineering normal software um but that seems like would be a very hard Challenge and I I suspect that that's that's when I'm when I talk about you know trying to reverse engineer neural network I think we're talking about you know a very significant challenge of that kind um and that it might be might be possible similarly you know no one's like oh you know biology you know it it's not easy therefore it must be impossible for us to understand cells you know there's room for something to be nearly extremely difficult in between and that being the case [00:06:08] um you know we'd like something that could maybe make it a little bit easier um easier and so you know I think the goal of trying to understand an entire neural network is very hard but we might be able to understand just portions of neural networks and then gradually grow that little portion so we might be able to go and slowly sort of and rigorously grow out from understanding tiny little portions to understanding um larger chunks um okay so um um so something that we could talk about is we could we could try to make this a little bit more concrete by by talking about a particular kind of neural [00:06:38] network and I um I think that confidence are maybe a good place to start um I was very fortunate for a number of years to work here um with our hosts leading the interpretability team here at the time and um we uh we made a lot of progress on understanding contents and I think a way that um I like to think about this is in terms of three kinds of objects um so we have uh features um which are uh say the easiest way to think about this would be a neuron and it might be something like a car detector or a curve detector uh or an [00:07:08] edge detector or a floppy ear detector um some kind of articulable property of the input um weights which connect together features and then circuits which are the combinations of weights and features together um so as an example um there it turns out that an Inception V1 there's all these curve detectors in fact these exist sort of in every Vision model we've looked at um they look for uh sort of just just curves um and it turns out these hold up to extremely rigorous investigation in fact we wrote an entire two papers [00:07:39] um just on curve detectors um and you can do all kinds of things to test these really are curve detectors up to the point of going and rewriting the entire circuit from scratch from understanding it and re-implanting them um and so that's very interesting because that gives us a way to think about about one unit in five things now another thing we can ask about are weights and remember that in a Content the weights are going to be um you know a grid because we have to talk about relative positions you know does it excite something at this offset or that offset um but if we just look at the weights in [00:08:09] isolation they're they're not that interesting right we have the input neuron and the output neuron but yeah the weights by themselves you know they don't say that much but if we then go and connect them to features that we understand um suddenly that becomes very different so here we have um a dog head facing to the right with a long snout and on the other side it connects to a dog head plus neck detector and you can see that now all of a sudden the weights make a lot of sense we're attaching um we're having the the dog head detector go and excite the dog head boss neck detector if it's on the [00:08:40] right side of the dog where the head should be so that's very interesting he puts it on the other side um and it turns out that once you contextualize weights like this neural network weights are full of structure uh so we have dog heads being attached to necks um we have uh uh pose ozone variant dog head detectors being constructed um from dog heads facing in different orientations they converge the the snout is in the same place we have [00:09:10] um a car detector um that goes and looks for Windows on the top and wheels on the bottom um yes yeah so these are uh all uh feature visualizations they're just created by optimizing the input and so you go and you take random noise and you optimize them to go and cause the neuron to fire um the claim here isn't that that is like evident that is like decisive evidence about what a neuron is doing it's more like a variable name that we can put there because you know being [00:09:40] like and this is 4c447 um you know that's not a very useful like I mean I spent so long with this model that I I know at 4c47 that you went um uh and so uh having having a variable name um that's very suggestive of what it does and a little bit of evidence I think there are important ways in future in which feature visualization provides evidence for what a neuron is doing um is is helpful for representing these circuits of course that you really want to be confident you want to look at all kinds of things um and do the kind of detailed uh investigation that I described for the curve detector foreign [00:10:11] we have high low frequency detectors and here the these these units they go and they look for high frequency patterns on one side of the receptor field and low frequency patterns on the other and you can see actually that as they um they just do that by going and piecing together a bunch of neurons that represent high frequency and a bunch that represent low frequency in the previous layer um and as the high low frequency detectors rotate so do the weights so there's actually just this beautiful structure um uh here we have um uh a bunch of color uh sort of color contrast detectors being used to create [00:10:41] assembled into Center surround detectors um we have uh a more interesting circuit that does black and white versus color detection um going and constructing black and white versus color detectors and assembling them into um again these kind of Center surround units curve detectors um are built from earlier curve detectors and you can look at as they're assembled they look for the the earlier curve detectors excite the later curve detectors when they match the orientation and inhibit when they have the opposite orientation and again the weights rotate as the features [00:11:12] rotate um here we have a triangle detector being assembled from Edge detectors we have wait for it a small circle detector being assembled from very early curve detectors um a whole bunch of color contrast detectors being used to create lines and the thing that I want to take away you to take away from all this is these are just a handful of examples neural network weights once you start to contextualize them are full of structure also lots of things we don't understand [00:11:42] but lots of structure um sophisticated boundary detectors being constructed from all kinds of different cues um so all all kinds of things okay so that is the basic picture now unfortunately yes sure sure yeah yeah this is great [00:12:33] organization is something else designing yeah so my response would be that these these these socializations are just variable names the actual thing that's analogous to like I equals one or something like this is the weights um the weights are like the assembly instructions or the code of the computer program um and so here we're defining a car [00:13:03] detector in terms of a window detector and a wheel detector and the car body detector and then you could ask okay well how do we trust those then we need to go back another step and so on and then the stress to look a lot like a computer program where you know we we have an understanding where we can base it on our understanding of previous variables and if we want to go and understand those and really carefully understand those we have to go back further and further and further um and we we can go and Trace things back all the way to the input if we want it's a lot of work um but yeah I think that's the answer so these are these are just the functioning as variable names um they're useful variable names they do provide certain kinds of causal evidence [00:13:34] um but um I see that we're getting a lot of raised hands I'm kind of tempted to ask to hold questions maybe to the end because I do want to cover a lot of ground and I'm a little nervous that we won't get through it and again I'll also be available for sort of office hour type things to go and answer any questions you have that I don't address during this talk um so I want to now talk about superposition um if the slides will update um which they might not um and you see uh the picture that I painted for you so far is a bit overly [00:14:04] optimistic and an important way which is that neural networks are also full of what we call poly semantic neurons neurons that respond to many unrelated features so surprisingly like in some ways it's miraculous that neural networks um when they when they have a privileged basis when they have neurons they have an activation function that you get lots of neurons that just seem to correspond to meaningful features but you also get many that don't um and one hypothesis for why this is um is that the features are in super possession the model wants to represent more features than it has neurons [00:14:34] possibly many more features than it has on our own then as soon as you have that of course you can't align all the features with neurons because there's only there's only n neurons and you want to represent more things um and so this is a very a very frustrating situation and in fact the picture gets kind of crazier where if you take that really seriously when it starts to stressed is that the model is actually sort of simulating a larger neural network there is some larger sparse or neural network and then it gets projected down and folded on top of itself to go and create the network that we actually observe and then of course [00:15:04] the neurons don't correspond to Features they correspond to your linear combinations of features um and so it's hard to really directly study this in Real Models um but it turns out we can show that this is exactly what happens in toy models um and then there's suggestive evidence that this happens in large models um so um it turns out that the essential thing is the sparsity of the features where if you have a bunch of features of varying importance and then you can see a detailed uh descriptions experiment in the the toy models paper [00:15:35] um if you just start at at zero percent sparsity if the features are just completely dense um you only get as many features as you have dimensions but as you make the features sparser um then because they're probably not going to go and co-occur and interfere with each other the model starts to be able to go and pack more things in and eventually you end up with five features in a two-dimensional space um it turns out you can actually do computation as well in supervisions you can actually in some sense have a have a computational graph that's higher dimensional project it down and actually do useful computation while holding [00:16:06] everything in superposition so this makes it very hard when this is true for us to go and understand things and then the the what you start to think when you observe these these neurons that are sort of monosomatic is that those are particularly important and common features and then the things that we don't observe are the sort of less important dispersal things are in superposition um so I think this is a really fundamental challenge this kind of work um unless you understand the Superstition structure there will always be unknown the potential for unknown behaviors to just sort of suddenly occur [00:16:37] um this is something that I'm very worried about and you won't you won't be able to understand really what the weights um are actually doing um so this seems to me like sort of the challenge for mechanistic interprety right now um and I think that um a lot of our our success for for doing useful work is going to rise and fall on whether we can make progress on that problem and I think it probably holds for a lot of other approaches as well I don't mean to say that nothing can be done without this I think there are other valuable things one can do um within mechanistic interpretability without this but I think that we will surrender a great deal if it turns out [00:17:09] that we can't go and address this somehow um it also turns out though that Civilization is full of beautiful mathematical structure it is deeply connected to compressed something um it turns out that in toy models at least these features organize themselves into polyhedra into regular polyhedra um which is is kind of wild or uniform polyhedra um it's very not obvious that they should do this it turns out that the learning Dynamics involve these weird like electron jumping like behaviors that are very strange um it's very mysterious [00:17:40] um I see that there is a question um if it yeah so I actually think that defining what we mean by feature is extremely hard and I don't have a super compelling definition of what a feature is but uh intuitively what we mean is something like a curve detector or a carb detector [00:18:10] or I think car detector or things like this and sometimes those end up corresponding to neurons but sometimes we can show that they they correspond to linear combinations of neural so we can construct toy problems where we do know what features are and where that's exactly defined and then they will go and represent them in this way um so I think the toy model is probably the best the best explanation I can give of this um in a lot of ways because in that case we know exactly what a feature is and it does just go and play out this way I think what a feature is sort of in a in a completely generative General model [00:18:40] um is is harder to go in and say especially like like one definition you could give is that a feature is sort of a human understandable property of the input that the network represents um but I think that that kind of I don't want humans to be sort of involved in the definition um another like the thing that I sort of in my heart believe a feature is but I don't know how to properly Define this as something like a feature is like a fundamental unit of the neural networks computation and those don't necessarily correspond in neurons they sometimes do um and they they often appear not to apparent often it seems like they instead are represented by these linear combinations of neurons but I think I [00:19:11] think the toy models paper where we just thought of the problem such that there's an obvious thing that is a feature and then you get this Behavior as the it's sort of the best demonstration I can give of how those things can can decouple okay so I wanted to talk briefly about how Transformers are different and I'm only going to talk about this for a little bit because this actually becomes quite messy and and ends up with a lot of detailed mathematics which I think it's actually very cool about mathematics but it um it's a little bit harder to go through and talk um but I I'd highlight two things in particular that make Transformers very [00:19:41] different um well maybe there's a third one which is that you just that you see a lot of superposition in Transformers you see a lot more than you do in Vision models um but I think there's two very deep architectural defenses so one is that we have a residual stream um it's a little bit even different from a resnet we're in a trans well I'll talk about this more in a second and the other is attention heads oh no apparently I'm not gonna talk about that more in a second so um uh the original residual stream just because you directly add um and and and sort of project in and out of it [00:20:11] um for all of your layers there's sort of these implicit latent uh weights connecting uh you know neurons in across layers or attention heads across layers or all these things and so that creates a lot of interesting structure um that doesn't doesn't exist in quite the same way in any contents really um uh because even even resonance aren't exactly like this differently um there's also um attention heads which we can talk about more and um attention has kind of create a new uh sort of fundamental unit of of [00:20:42] mechanistic interprealty um similar to uh to feature weights or activations my thing is as attentional features um and I won't talk about this too much but there's lots of very rich and interesting intentional features and um this is a one that I find particularly interesting if you if you look at um because often see these these weird off diagonals they're very striking if you look at attention patterns um and it turns out this corresponds to a particular type of attention head called an induction head which searches [00:21:12] through for previous cases where something happens and then looks forward one step and then goes and copies that and goes and outputs that um so this allows the model to sort of do in context uh learning of a kind in fact it turns out that sort of a generalized version it seems to be the driver of in context learning um or at least there's there's a non-preal amount of evidence for that hypothesis um in fact these are so important that they correspond and they create a bump in the loss curve when you train Transformers it's just a visual bump in in every transform or loss curve that I've seen if you have more than two [00:21:42] layers because that's what's the necessary for induction it's deformed they also cause actually a deviation in the scaling loss um so if you look at the scaling laws the original scaling laws paper you'll see that there's a point where the there's sort of a Divergence from the original Trend and that corresponds to the models not having induction heads and then forming them um so they think so this is kind of interesting thing we've got a these inductions are so important they appear in the macroscopic picture like we have this microscopic picture that we're developing here and and it Bridges all the way to the macroscopic um okay yeah uh now so far I've just [00:22:13] been talking about in her really um sort of broadly but it's worth me briefly painting a bit of extra for how this might connect to safety uh yes Joshua oh yeah it does yeah a massive number of induction heads simultaneously occur um this if there seems to be a deep reason for this basically um it there's there's a number of pieces [00:22:43] that have to exist for an induction head circuit to form and once you start to have those ingredients you then get an extremely fast feed feedback loop um and then all of them form simultaneously because the ingredients are in place and they perhaps evolve a bit further from there but there's this sort of sudden place where all the like if you look at if you just track the existence of induction heads in a Transformer over time they will just sort of all form in a very small window um and maybe maybe further evolve or maybe you'll still like get a couple lip form afterwards but the vast majority of them form at once and so yeah [00:23:13] yeah we have a slightly similar experiment where you can go and set up an architecture that very easily learns induction heads and then induction heads form right at the start um so you can uh it's a I haven't quite visualized the Lost score but it's it's like right at the beginning so it'd be in the extremely steep regime and you won't be able to see this clearly but it's like it would happen like right here um and then yeah you don't see an analogous I guess we can we can check this but I yeah you shouldn't see an analogous bump in the Lost curve um okay so I want to briefly talk about [00:23:45] um how so there are lots of ways in which interpreter might contribute to safety um I just want to talk about one that seems to me uh like a particularly important one and um a one that I find very motivating um and so it seems to me like the thing that we most want out of safety is something like the ability to to make statements of a kind you know something vaguely like for all the situations the model could be in the model will never deliberately do X we want to be able to say something like that [00:24:15] um and my present best guess as to how you could achieve that is that you want to be able to say something like okay they don't exist features let's say such that the model will deliberately that will participate in the model deliberately doing X and that's going to then be some kind of Claim about the circuits that feature participates in so this is this is a very ambitious and wild and kind of crazy thing to be aiming for I don't I don't mean to say that this is anywhere anywhere near being able to go and make this kind of statement but this is kind of a spiritual North Star I think um for the most ambitious way that it that mechanistic interability could help [00:24:45] with safety um and I think there are two major challenges to this working the first is um supervision how can we actually access the features how can we know that we've got everything how can we rule out these these places where the model might have something a feature that activates very rarely and it's just some direction in angular space where it behaves totally differently how can we address that and then separately um scalability if these models are so large and the success of The mechanistic Interpreter so far are so small they're these these tiny little circuits that [00:25:15] we're understanding how can we hope to get to these large models um so I am very worried about supervision I am more optimistic about scalability um I think that there are a bunch of really good ideas for approaching um these include trying to use AI to automate things um I think trying to exploit modularity or graph structure and models um and I think uh there might also be sort of interesting motifs we've seen this in Vision models that can massively simplify models once you understand them [00:25:46] um if you start to understand certain symmetries um unfortunately I think this is a really challenging problem to work on right now um because we don't actually have access to the units that we want to understand and we're sort of trying you know we don't if we don't if we think the features are there and they're all in Superstition they're a giant mess and it's very difficult to start working on how you can I even actually really effectively scale um but I I do feel pretty optimistic about this and my my real fear is this challenge of super session okay um I guess I'm getting to the end of my [00:26:17] a lot of time so I'm right on schedule um I wanted to briefly just talk about one more brief thing which is so far I've you know I've talked about intern probably and I sort of talked about why you might care about it for safety but I feel like at an emotional level something I also find really motivating isn't just this goal of safety but this belief that I have that neural networks if we take them seriously as objects of Investigation are full of beautiful structure um and I sometimes think that actually maybe you know the Elegance of machine learning is in some ways more like the [00:26:48] Elegance of biology or or perhaps at least as much as like the Elegance of biology as math or physics um so in the case of biology Evolution creates all inspiring complexity in nature you know we have this this system that goes and generates all this beautiful structure and in a kind of similar way gradient descent it seems to me creates beautiful [00:27:18] mind-boggling structure it grows and it creates all these beautiful circuits that have beautiful symmetries in them um and it in toy models that arrange its features into regular polyhedra and there's all of this just like it's just sort of too good to be true in some ways it's full of all this amazing stuff when we and and then there's messy stuff but then you you discover sometimes that that messy stuff actually was just a really beautiful thing that you didn't understand and so a belief that I have that if only perhaps semi-rational so these models [00:27:48] are just full of beautiful structure if we're just willing to put in the effort to go and find it and I think that's the thing that I find actually most emotionally motivating about this work uh yeah so thank you all for your time um I want to emphasize that all this work uh was uh done in collaboration with many others including others um at different institutions um I'm I'm at anthropic [00:28:18] um and if you're interested in this work um you should look at the original circuits thread which made a really serious attempt to reverse engineer one visual Vision model Inception V1 the Transformer circuits thread which has been trying to do Center analogous work on Transformers um and Neil Nanda has this really nice annotated reading list which actually might be the best uh jumping off point so thank you great uh [00:28:48] yeah better [00:29:26] so my my answer would be my thinking on this mostly comes from this this paper that I was lucky to work on with a number of people at open AI um so you might be familiar that um you know in the original courgeski 2012 paper that's really striking phenomenon where you look at the um the first layer of the model and you see that there's all these Gabor black and white Gabor filters on one side and color contrast detectors and the other the model is set up such that you have these sort of two gpus that only talk to each other every other step and it turns out if you look at the next layer you see similar similar things and [00:29:57] it turns out this is actually a very general phenomenon lots of models Vision models have branches and um uh like Inception V1 has branches and like it turns out that like this Branch over here has all of the uh like black and white versus color and color contrast units this one's full of curve detectors and some sort of fern eye related detectors and some boundary detectors and this one's full of 3D geometry reactors and just like the probabilities are like definitely it's it's definitely going and organizing things like this so one thing to say is okay well we could create branches like [00:30:27] that but then it actually turns out that if you look at other models and you can just visualize which features say in the first layer oh I think this is my timer to remind me for the I think you can know that it made a sound when I ran out um uh it turns out that the um the structure is implicitly there you can just go and do uh take the magnitude of of the connections between features in one layer and features in another layer and go and and do a singular value decomposition and you discover that that actually all the same structure was implicitly there um and so I think that's kind of thing [00:30:58] uh is just there if we're if we if we look for it and we we can just sort of probably find pull it out um in practice this becomes hard on the later layers of vision models you don't don't see this as much and in language models we don't see it but I think that's because of superstition I think the thing that superposition sort of the best case reciprocision is to go and fold two completely unrelated computational graphs on top of each other and so my guess is in fact that um and in fact one of the things I'm sort of most optimistic about is that if you can solve superposition you'll discover that there's that there's just lots of [00:31:28] like implicit modularity um and other kinds of organizational structure to these graphs um that we can just see because I think the early layers of vision models are kind of the best proxy I have for a model that has very little superposition um and that seems to be the case there hi yes [00:32:13] actually the smart one to say exact ly we want to understand oh details the thing that would concerned me a little bit is understandable foreign [00:32:54] yeah that's a great question um so earlier you might have heard me sort of uh dissembling on um or or trying being confused about what is a feature um and I think this is kind of kind of the question here where um I feel like I there's a thing that I'm trying to point at and I don't know what it means and actually um you know you might be like oh that's kind of wild right like Chris Chris is like spending his whole career on this thing and he doesn't even know what the definition that he's talking about is but um actually I think that's often very productive um there's this book by lackatosh proof and reputation it's just this play about [00:33:25] um you know a bunch of math students going and and being confused about definitions and progressively finding more and more useful definitions for things um but uh yeah like I don't want to Define features in terms of human understandable Concepts because I want to allow for the possibility that these models have um have Concepts that we we don't understand so what I really am looking for in terms of having human well having having sort of reverse engineering these things is in some sense pulling apart the fundamental units of computation whatever they are um so far in the cases where we've done this they have been [00:33:55] things that we've understood and understanding the algorithms that drive that and when I when I when I say turning them into code I don't so much mean turning them into like python code or something like this the weights are the code like for the for the door the for the um the the the car detector that I showed you earlier um like there's no simpler explanation than just showing you you know there's a window detector there's a wheel detector you know it wants to see the windows at the top it wants to see the wheels at the bottom that is the most understandable depiction of that that there that there is and so uh the thing [00:34:26] but but you couldn't have understood if I didn't contextualize it right and so there's this difficult work of going and contextualizing what these things are going and undoing this problem with superposition to the extent that it is and then very painstakingly trying to go through it and when we find Concepts that we don't understand and initially like um in Vision models there's these high low frequency detectors we initially didn't understand them um in retrospect they're super simple all they are is they look for a something that's high frequency on one side low frequency on the other and this occurs all the time at boundaries of objects because you have [00:34:56] um sort of uh like something is out of focus in a background um or you know the the background is is busy and I'm a solid color or something like this okay that's one answer one other thing that I really sort of want to say here is I think that we have a enormous advantage over Neuroscience in fact I wrote a whole article about all the things that make my life so much better than neuroscientists and I do think [00:35:31] but I I but but but just for the for the for people who haven't thought about this I just want to enumerate some of the advantages we have um so we have access to the ground truth of whatever neurons computational structure is we can get access to the activations of every neuron um in a non-destructive way and collect them for as many stimuli as we want we have access to the ground Trace connectum we know how every neuron [00:36:01] connects together in fact not only we know the connectome we know the weights of the connectome it's not just which neurons are connected to each other but we know which ones excite and inhibit each other um we can go and take gradients through the structure and optimize through it um uh we can go and work with the same model again and again rather than having to go and switch between models which might be different and when I study a model I can give that model potentially to you or to one of my colleagues and we can go and ask questions of exactly the same model and talk about the same neurons um [00:36:31] there are the experience Cycles so much faster there are so many things um so it's just this like truly enormous difference um so I think the fundamental thing um is my present theory for him with us is there there's two things there's the feature importance curve which is if you [00:37:01] imagine that there's you know order all the features that the model could learn in terms of their importance and which is like something like how much it reduces the loss or something like this and look at that curve and so the flatter that curve is the more superposition you're going to have the steeper it falls off um the the Lesser position you're going to have and the second thing is the sparsity and I think that on both of these uh language models go the wrong direction um so on uh on sparsity um Vision models when you start with something especially with early Vision right you know it's dealing with things like an edge detector if you look at [00:37:32] this thing behind me and you ask how many vertical edges are there in just a single image you know um there are a lot of vertical edges on the other hand if you are if you're like you know um like a lot of these language models know who I am how often do I come up in uh in text 100 million tokens one in a billion tokens I don't know like it's I I don't think I'm very common those actually feel like they're probably too small a number or something like this um so there's just a massive massive discrepancy between the the frequency [00:38:02] um and therefore the sparsity of features an early vision and the sort of linguistic features and because you start with word embeddings and then you very quickly go to more sophisticated things language models very rapidly go to um a denser features um you do find monosynthetic neurons and language models and they all tend to be things that are very common um so you're like oh you know um uh like common common constructions of multiple words or or things like this um would not for instance or things like this um so and and then I think also just the [00:38:33] space of Concepts you want a language model to represent is just so much bigger than like an imagenet model or something like this and and so you just have this much flatter longer um uh feature importance curve I don't think so um I found that when you have like we have a tool so we have an open AI we built this this tool microscope we have a similar internal thing for language models lexoscope um it shows you all the data set examples one really cool thing you can do in language models that you [00:39:04] can't do in image models is you can just edit the text and see how the neuron activation change in real time that's that's like a thing that I'm very grateful for um and so actually I don't think like when you have Interpol division uh language neurons I don't think that they're harder to understand um or recognize um maybe maybe it takes a little bit more time because you can maybe more quickly visually go through like a bunch of visual data set examples and language ones but I think not in a very significant way so I think that's not the not the primary thing going on um I believe I may be out of time one more one more okay remember I will be available [00:39:34] um at this kind of office hours thing and also just like very happy to chat with people all all day um yes there are some research ER but on the other hand I assumed that a person Concepts can you speculate what [00:40:06] yeah I really don't know um so one thing that's interesting is that supervision sort of depends on the gap between the sparsity of the the neurons and the sparsity of the sort of fundamental underlying data and the fundamental underlying features and so if you think that those features are extremely sparse like they occur you know one in a million tokens one in 10 million tokens then just the sport there's still an enormous enormous sparsity Gap um I think the more interesting or fundamental thing is something like how do you think the feature importance curve changes um uh and [00:40:36] I don't know like and well yeah perhaps this is something to talk about more offline I think I think the fun the basic answer is I have no idea and I can like give you a bunch of conceptual models for trying to reason about this um but I still then have no idea thank you so much for all of your time thank you Chris amazing