58:01Large Language Models in Five Formulas
Sasha Rush's tutorial takes the opposite approach to hype: pick the five places where large language models can actually be measured, bounded, or forecast, and be honest that the rest is hard. Perplexity for generation, attention for memory, the GEMM for efficiency, Chinchilla for scaling, and RASP for reasoning. He builds each one from the ground up with the numbers on screen: the Wall Street Journal perplexity ladder from 10,000 down to 20.5, a 3x3 matrix multiply that costs 54 global memory reads done naively and 18 done in blocks, BERT Base and PaLM at opposite ends of the token to parameter trade off, and RASP programs that compile into real Transformer weights. Every section ends with him saying exactly where that formula stops being useful.