youtube.nixfred.com nixfred.com
Creator

Sasha Rush

Associate professor at Cornell Tech and researcher at Hugging Face. Turns the parts of large language models we can actually measure into clean formulas.

1video

← All videos

58:01
Sasha Rush

Large Language Models in Five Formulas

Sasha Rush's tutorial takes the opposite approach to hype: pick the five places where large language models can actually be measured, bounded, or forecast, and be honest that the rest is hard. Perplexity for generation, attention for memory, the GEMM for efficiency, Chinchilla for scaling, and RASP for reasoning. He builds each one from the ground up with the numbers on screen: the Wall Street Journal perplexity ladder from 10,000 down to 20.5, a 3x3 matrix multiply that costs 54 global memory reads done naively and 18 done in blocks, BERT Base and PaLM at opposite ends of the token to parameter trade off, and RASP programs that compile into real Transformer weights. Every section ends with him saying exactly where that formula stops being useful.

AIDeep LearningScienceJan 30, 2024