youtube.nixfred.com nixfred.com
Topic

Deep Learning

Everything tagged Deep Learning.

11videos

← All videos

15:27
AI Revolution

China Just Shocked Everyone With a 10 Trillion Parameter AI Model

Three frontier model stories inside 48 hours, walked in order. The Financial Times reported that ByteDance is training a model with as many as 10 trillion parameters, which would put it above the industry estimates for Anthropic's Mythos 5 at roughly 8 trillion and Fable 5 around 5 trillion, and roughly 3.5x Moonshot's Kimi K3 at 2.8 trillion. Reuters relayed the report and said directly it could not verify it. Meta shipped the beta of Muse Code, a terminal coding agent powered by Muse Spark 1.2, with persistent background agents, a replayable event log, and contributor pricing 12.5x cheaper on input than standard. OpenAI made GPT 5.6 Luna free with unlimited text for roughly a billion users while a much larger flagship, Astra, sits at release candidate under the checkpoint name Mu4. The page rebuilds every number and caveat, and separates what is announced from what is reported, estimated, and leaked.

AIBusinessDeep LearningAug 8, 2026
2:07:14
Cosmo Explains

Everything About Machine Learning Explained Slowly (For Sleep)

The entire history of machine learning told slowly and in order, from Ramon Llull's rotating discs and Ada Lovelace's objection to the multimodal systems hundreds of millions of people use today. Cosmo Explains walks the mechanisms as well as the people: how a McCulloch Pitts neuron computes, how Rosenblatt made one learn, why XOR broke the perceptron and froze the field, how backpropagation uses the chain rule to assign blame three layers deep, why overfitting made SVMs and maximum margins the respectable choice for two decades, and how data, GPUs and a few simple architectural tricks converged on AlexNet in 2012. From there it is ResNet's skip connection, the transformer throwing away recurrence, next word prediction buying grammar and reasoning for free, CLIP and diffusion, and finally Goodhart's law and the alignment problem. The register is built for sleep; the content is a full course.

AIDeep LearningScienceJul 25, 2026
3:31:23
Andrej Karpathy

Deep Dive into LLMs like ChatGPT

Andrej Karpathy's general audience deep dive into the full training stack behind ChatGPT, in one sitting: pretraining on a filtered crawl of the internet, tokenization, what the base model actually is (an internet document simulator), supervised fine tuning that turns it into an assistant, and reinforcement learning as the third and least mature stage. The second half is LLM psychology: why hallucinations happen and how they get patched, why models need tokens to think, why they cannot spell or compare 9.11 to 9.9, and what emergent chains of thought in DeepSeek R1 mean. It ends with where to find and run these models yourself.

AIDeep LearningFeb 5, 2025
24:37
Ilya Sutskever

Ilya Sutskever: Sequence to sequence learning with neural networks: what a decade

Accepting a NeurIPS Test of Time award for the 2014 sequence to sequence paper, Ilya Sutskever puts his own slides from ten years earlier back on screen and grades them one by one: the deep learning hypothesis held, the autoregressive bet held, connectionism is the idea he says truly stood the test of time, and pipelining was simply unwise. Then the line the talk is remembered for: pretraining as we know it will unquestionably end, because compute grows on three channels while data grows on none. Data is the fossil fuel of AI and there is but one internet. He relays agents, synthetic data and inference time compute as other people's speculations rather than his roadmap, and offers hominid brain to body scaling as an existence proof that a second scaling regime can be found at all. The forward half lists four properties of superintelligence with every hedge attached, including the argument that gets the least attention: a system that genuinely reasons is a system you cannot predict. The Q&A, included in full here, produced several of the most quoted moments of the conference.

AIDeep LearningScienceDec 14, 2024
4:01:26
Andrej Karpathy

Let's reproduce GPT-2 (124M)

Four hours, empty file to trained model. Karpathy writes the GPT-2 124M architecture to match the released weights exactly, then spends most of the video making it fast: TF32 and bfloat16, torch.compile, flash attention, choosing vocabulary sizes with nice powers of two, then the optimization recipe from the GPT-3 paper (AdamW settings, gradient clipping, cosine schedule with warmup, weight decay, gradient accumulation to reach a half million token batch) and finally distributed training across eight GPUs. It ends with a real run on FineWeb-Edu that beats the published GPT-2 124M numbers.

AIDeep LearningJun 9, 2024
26:09
3Blue1Brown

Attention in transformers, step-by-step | Deep Learning Chapter 6

Chapter 6 of the 3Blue1Brown deep learning series opens the attention block and walks through it one matrix at a time: queries, keys, the attention pattern and its softmax, masking so tokens cannot see the future, the value matrix and its low rank factorization, and finally multi headed attention running many of these in parallel. Grant Sanderson keeps a running parameter tally against GPT-3 the whole way, so the abstractions stay attached to real numbers. This is the page to read when you want the mechanism itself rather than a metaphor for it.

AIDeep LearningApr 7, 2024
2:13:34
Andrej Karpathy

Let's build the GPT Tokenizer

Karpathy builds a byte pair encoding tokenizer from scratch and argues that tokenization is the root of a startling amount of LLM weirdness: bad spelling, failed string reversal, worse performance in non English languages, arithmetic errors, the Python indentation problem in GPT-2, and the unspeakable SolidGoldMagikarp tokens. It starts at Unicode and UTF-8, works up through the BPE merge algorithm, the GPT-2 and GPT-4 splitting patterns, special tokens, and SentencePiece, then closes on how to choose a vocabulary size.

AIDeep LearningFeb 20, 2024
58:01
Sasha Rush

Large Language Models in Five Formulas

Sasha Rush's tutorial takes the opposite approach to hype: pick the five places where large language models can actually be measured, bounded, or forecast, and be honest that the rest is hard. Perplexity for generation, attention for memory, the GEMM for efficiency, Chinchilla for scaling, and RASP for reasoning. He builds each one from the ground up with the numbers on screen: the Wall Street Journal perplexity ladder from 10,000 down to 20.5, a 3x3 matrix multiply that costs 54 global memory reads done naively and 18 done in blocks, BERT Base and PaLM at opposite ends of the token to parameter trade off, and RASP programs that compile into real Transformer weights. Every section ends with him saying exactly where that formula stops being useful.

AIDeep LearningScienceJan 30, 2024
1:31:13
Jeremy Howard

A Hackers' Guide to Language Models

Jeremy Howard's code first tour of language models for people who want to build with them today. It starts with what a language model is and the three stage recipe he helped popularize with ULMFiT, moves to using the strongest hosted models well, then to the practical engineering: the API and function calling, running open models locally with quantization, retrieval augmented generation, and fine tuning a small model on a narrow task. The through line is that you learn this by running notebooks, not by reading about it.

AIDeep LearningDevOpsSep 24, 2023
40:59
FAR.AI

Chris Olah - Looking Inside Neural Networks with Mechanistic Interpretability

Chris Olah's case for treating a neural network as an object to study rather than a black box to probe from outside: a trained model is a compiled binary with no source code, features are its variables, and circuits are the weights between features you already understand. He reads a dozen circuits straight out of InceptionV1, then breaks his own picture with polysemantic neurons and the superposition hypothesis, that a model simulates a larger, sparser network projected down and folded on top of itself, which is why five features fit in two dimensions and why no neuron means just one thing. He calls superposition the question the whole safety payoff rises or falls on, states the guarantee he actually wants as a claim quantified over every situation the model could be in, and closes on the motivation he says moves him most: gradient descent grows beautiful structure, and somebody should go look at it.

AIDeep LearningScienceSep 1, 2023
1:03:32
Berkeley EECS

John Schulman - Reinforcement Learning from Human Feedback: Progress and Challenges

John Schulman, co-founder of OpenAI and lead of the ChatGPT RLHF work, gives the clearest available argument for why supervised fine tuning alone teaches a model to hallucinate and why reinforcement learning is the lever that can teach it to hedge instead. The core idea: imitation makes the model assert things it has no internal evidence for, while a reward that scores confident wrong answers below honest uncertainty makes calibration the optimal policy. He then walks through reward models, retrieval and citation as the route to verifiability, and the open problems he had not solved.

AIDeep LearningScienceApr 19, 2023