40:59Chris Olah - Looking Inside Neural Networks with Mechanistic Interpretability
Chris Olah's case for treating a neural network as an object to study rather than a black box to probe from outside: a trained model is a compiled binary with no source code, features are its variables, and circuits are the weights between features you already understand. He reads a dozen circuits straight out of InceptionV1, then breaks his own picture with polysemantic neurons and the superposition hypothesis, that a model simulates a larger, sparser network projected down and folded on top of itself, which is why five features fit in two dimensions and why no neuron means just one thing. He calls superposition the question the whole safety payoff rises or falls on, states the guarantee he actually wants as a claim quantified over every situation the model could be in, and closes on the motivation he says moves him most: gradient descent grows beautiful structure, and somebody should go look at it.