youtube.nixfred.com nixfred.com

Attention in transformers, step-by-step | Deep Learning Chapter 6

Chapter 6 of the 3Blue1Brown deep learning series opens the attention block and walks through it one matrix at a time: queries, keys, the attention pattern and its softmax, masking so tokens cannot see the future, the value matrix and its low rank factorization, and finally multi headed attention running many of these in parallel. Grant Sanderson keeps a running parameter tally against GPT-3 the whole way, so the abstractions stay attached to real numbers. This is the page to read when you want the mechanism itself rather than a metaphor for it.

Published Apr 7, 2024 26:09 video 42 min read Added Jul 30, 2026 Open on YouTube →

At a glance

Chapter 5 of the 3Blue1Brown Neural networks series followed a stream of data through a transformer and treated attention as a labelled box. Chapter 6 opens the box. Grant Sanderson takes the attention mechanism apart one matrix at a time: the query matrix that turns an embedding into a question, the key matrix that turns an embedding into an answer, the dot product grid those two produce, the softmax that normalizes it into an attention pattern, the mask that stops a token reading the future, the value matrix that actually moves meaning between positions, and finally multi headed attention, which is the same head run ninety six times in parallel with ninety six different learned specialties.

What makes the chapter land is the running parameter tally. Every matrix he introduces gets counted against the real dimensions of GPT-3, so the abstractions never float free of the machine: 12,288 for the embedding dimension, 128 for the key query space, about 1.5 million parameters per matrix, 6.3 million per head, 600 million per block, and just under 58 billion across the whole network. By the last minute you know not only what attention does but exactly what fraction of GPT-3 is spent doing it, which turns out to be about a third.

The example he follows the whole way is one sentence, "a fluffy blue creature roamed the verdant forest," and the narrow question of how the adjectives get baked into the noun they modify. He is explicit that he invented that example and that the real learned behavior is far harder to read. It is a scaffold for the arithmetic, not a claim about what any particular head does.

This page rebuilds the derivation in the video's order, keeps every number he states, and keeps the arithmetic he does on screen.

The deep explanation

Where chapter 6 picks up: an embedding is a lookup table with no idea of context

The recap he asks you to hold in mind is short. The model's goal is to take in a piece of text and predict what word comes next. The input text is broken into tokens, which are very often words or pieces of words. He simplifies throughout by pretending tokens are always just words, purely to make the examples easier to think about.

The first step in a transformer associates each token with a high dimensional vector, its embedding. The important idea carried over from chapter 5 is that directions in this high dimensional space of all possible embeddings can correspond with semantic meaning. The example there was gender: adding a certain step in the space takes you from the embedding of a masculine noun to the embedding of the corresponding feminine noun. That is one direction out of many you could imagine encoding numerous other aspects of a word's meaning.

And that gives the aim of the whole architecture, stated in one sentence: the transformer progressively adjusts these embeddings so that they do not merely encode an individual word, but bake in much, much richer contextual meaning.

He says up front that a lot of people find attention very confusing, and that it is fine if it takes time to sink in. He repeats that reassurance twice more at the two densest moments of the video, which tells you something about how he expects it to land.

Three moles, and the behavior we actually want

Before any matrix multiplication, he argues for the behavior. Consider three phrases:

You and I know the word mole means something different in each one, based on context. But after the first step of a transformer, the step that breaks up the text and associates each token with a vector, the vector associated with mole is identical in all three cases. The initial token embedding is effectively a lookup table with no reference to the context whatsoever. It is only in the next step that the surrounding embeddings get the chance to pass information into this one.

The picture to hold: there are multiple distinct directions in embedding space encoding the multiple distinct meanings of the word mole, and a well trained attention block calculates what you need to add to the generic embedding to move it toward one of those specific directions, as a function of the context.

His second example runs the same idea forward. Take the embedding of the word tower. Presumably some generic, non specific direction in the space, associated with lots of other large, tall nouns. If that word was immediately preceded by Eiffel, you would want the mechanism to update the vector so that it points somewhere that more specifically encodes the Eiffel tower, maybe correlated with vectors associated with Paris and France and things made of steel. If it was also preceded by the word miniature, the vector should be updated further still, so that it no longer correlates with large, tall things at all. Two words of context, two successive corrections to the same vector, and the second one partly undoes the first.

He then generalizes past word sense disambiguation. The attention block does not only refine the meaning of a word. It allows the model to move information encoded in one embedding into another, potentially one that is quite far away, and potentially information much richer than a single word.

The mystery novel, and why all of this has to happen before the last vector

This is the motivating case that makes the stakes concrete, and it is worth keeping in full.

Chapter 5 showed that after all the vectors flow through the network, including many different attention blocks, the computation that produces the prediction of the next token is entirely a function of the last vector in the sequence. Nothing else. So imagine the text you input is most of an entire mystery novel, all the way up to a point near the end, which reads "therefore the murderer was."

If the model is going to accurately predict the next word, that final vector in the sequence, which began its life simply embedding the word "was," will have to have been updated by all of the attention blocks to represent much, much more than any individual word. It has to somehow encode all of the information from the full context window that is relevant to predicting the next word. An entire novel's worth of plot, motive and misdirection, compressed into the vector that started out meaning nothing but "was."

That is the job. Attention is the only mechanism in the architecture that can do it, because it is the only one that moves information between positions.

The sentence we follow all the way through

To step through the computations he drops to something far simpler. The input includes the phrase:

a fluffy blue creature roamed the verdant forest

And for the moment, suppose the only type of update we care about is having the adjectives adjust the meanings of their corresponding nouns. What he is about to describe is a single head of attention. Later we see how a full attention block consists of many different heads run in parallel.

Two caveats he plants here, both of which matter.

First, on the embeddings. The initial embedding for each word is a high dimensional vector that only encodes the meaning of that particular word with no context. Then he corrects himself on camera: actually, that is not quite true. They also encode the position of the word. There is a lot more to say about the specific way positions are encoded, but all you need right now is that the entries of this vector are enough to tell you both what the word is and where it exists in the context. He denotes these embeddings with the letter e.

Second, on the example itself. He is explicit that he is making up the adjectives updating nouns story purely to illustrate the type of behavior you could imagine an attention head doing. As with so much deep learning, the true behavior is much harder to parse, because it is based on tweaking and tuning a huge number of parameters to minimize some cost function. The invented example is a handrail for the matrices, nothing more.

The goal, restated in the terms the rest of the video uses: produce a new refined set of embeddings where the ones corresponding to nouns have ingested the meaning from their corresponding adjectives. And playing the deep learning game, we want most of the computations involved to look like matrix vector products, where the matrices are full of tuneable weights that the model learns from data.

One piece of notation to carry: whenever he puts a matrix next to an arrow, it means multiplying that matrix by the vector at the arrow's start gives you the vector at the arrow's end.

Queries: the matrix that turns an embedding into a question

For the first step of the process, you might imagine each noun, like creature, asking the question: hey, are there any adjectives sitting in front of me? And for the words fluffy and blue to each be able to answer: yeah, I am an adjective and I am in that position.

That question is somehow encoded as yet another vector, another list of numbers, which we call the query for this word. The query vector has a much smaller dimension than the embedding vector, say 128.

Computing it looks like taking a certain matrix, labelled W_Q, and multiplying it by the embedding. Compressed, the query vector is written q. You multiply this matrix by all of the embeddings in the context, producing one query vector for each token. The entries of the matrix are parameters of the model, which means the true behavior is learned from data, and in practice what this matrix does in a particular attention head is challenging to parse.

But for the sake of the example, suppose the query matrix maps the embeddings of nouns to certain directions in this smaller query space that somehow encode the notion of looking for adjectives in preceding positions. As to what it does to other embeddings, who knows. Maybe it simultaneously tries to accomplish some other goal with those. Right now we are laser focused on the nouns.

Keys: the matrix that turns an embedding into an answer

At the same time, associated with this is a second matrix, the key matrix, W_K, which you also multiply by every one of the embeddings. This produces a second sequence of vectors that we call the keys.

Conceptually, think of the keys as potentially answering the queries. The key matrix is also full of tuneable parameters, and just like the query matrix it maps the embedding vectors into that same smaller dimensional space. That shared destination is the whole point: two vectors have to live in the same space before you can compare them.

You think of the keys as matching the queries whenever they closely align with each other. In the example, you would imagine the key matrix maps the adjectives like fluffy and blue to vectors that are closely aligned with the query produced by the word creature.

The dot product grid, and what "attends to" actually means

To measure how well each key matches each query, you compute a dot product between each possible key query pair. Sanderson visualizes this as a grid full of dots, where the bigger dots correspond to the larger dot products, the places where the keys and queries align.

For the adjective noun example, if the keys produced by fluffy and blue really do align closely with the query produced by creature, then the dot products in those two spots would be some large positive numbers. In the lingo, machine learning people would say this means the embeddings of fluffy and blue attend to the embedding of creature. That is the whole origin of the word. By contrast, the dot product between the key for some other word like "the" and the query for creature would be some small or negative value, reflecting that they are unrelated to each other.

So now we have a grid of values that can be any real number from negative infinity to infinity, giving us a score for how relevant each word is to updating the meaning of every other word.

RAW SCORES: EVERY KEY MEASURED AGAINST EVERY QUERY THE TWO PROJECTIONS e 12,288 q 128 k 128 W_Q W_K QUERIES, ONE PER TOKEN a q1 fluffy q2 blue q3 creature q4 KEYS, ONE PER TOKEN a k1 fluffy k2 blue k3 creature k4 fluffy and blue attend to creature large and positive Every other pair gives a small or negative value, because the words are unrelated. These are raw dot products, so they range anywhere from negative infinity to infinity. Masking and softmax come next, column by column.
Figure 1. One head, before normalization. The query matrix and the key matrix both compress a 12,288 dimensional embedding down to the same 128 dimensional space, which is what makes the dot product between a query and a key meaningful. The grid is every key measured against every query, and the only large entries in this invented example are the two adjectives answering the noun's question.

Softmax down each column, and the attention pattern

The way those scores get used is to take a weighted sum along each column, weighted by the relevance. So instead of values ranging from negative infinity to infinity, what we want is for the numbers in each column to be between 0 and 1, and for each column to add up to 1, as if they were a probability distribution.

If you came in from chapter 5 you already know the move: compute a softmax along each one of these columns to normalize the values. Fill the grid back in with those normalized numbers, and at that point you are safe to think about each column as giving weights according to how relevant the word on the left is to the corresponding value at the top.

This grid is what we call the attention pattern. It is the thing people are showing you when they publish those heatmaps of a model attending to words.

The one line version from the paper, and the square root

If you look at the original transformer paper, Attention Is All You Need, there is a really compact way they write all of this down. In that expression the variables Q and K represent the full arrays of query and key vectors, the little vectors you get by multiplying the embeddings by the query and the key matrices. The product in the numerator is a really compact way to represent the grid of all possible dot products between pairs of keys and queries.

Then two details sit on top of it.

A small technical one he had not mentioned: for numerical stability, it happens to be helpful to divide all of these values by the square root of the dimension in that key query space. With a key query dimension of 128, that is a division by a little over 11.3 before anything else happens.

And a reading convention: the softmax wrapped around the full expression is meant to be understood to apply column by column, not to the matrix as a whole. The V term in that same formula is the subject of the values section further down.

Masking: why a token is never allowed to read the future

Here is the other technical detail he had skipped, and it is the one that explains why these models are called causal.

During training, you run the model on a given text example and all of the weights get slightly adjusted and tuned to either reward or punish it based on how high a probability it assigned to the true next word in the passage. It turns out to make the whole training process a lot more efficient if you simultaneously have it predict every possible next token following each initial subsequence of tokens in that passage. With the phrase we have been focusing on, it would also be predicting what words follow creature, and what words follow the.

This is really nice, because it means what would otherwise be a single training example effectively acts as many.

But it only works under one condition. For the purposes of the attention pattern, it means you never want to allow later words to influence earlier words, since otherwise they could give away the answer for what comes next. All of the spots representing later tokens influencing earlier ones have to somehow be forced to be zero.

The simplest thing you might think to do is set them equal to zero. But if you did that, the columns would no longer add up to one. They would not be normalized. So instead, a common way to do this is that before applying softmax, you set all of those entries to negative infinity. After the softmax, all of those turn into zero, and the columns stay normalized.

This process is called masking. There are versions of attention where you do not apply it, but in the GPT example, even though this matters more during the training phase than it would when running the model as a chatbot, you do always apply masking to prevent later tokens from influencing earlier ones.

AFTER MASKING AND SOFTMAX: EACH COLUMN SUMS TO ONE keys \ queries a fluffy blue creature a fluffy blue creature 1.00 0.10 0.10 0.05 masked 0.90 0.20 0.45 masked masked 0.70 0.40 masked masked masked 0.10 Read a column as one token asking its question and receiving weighted answers. Everything below the diagonal is set to negative infinity first, so a token can never draw on a token that comes after it. Zeros from a plain assignment would break the normalization; zeros produced by softmax do not. the adjectives feed the noun, and the noun feeds nothing backwards
Figure 2. The same grid as Figure 1 after the two normalization steps. Illustrative weights, with the structure Sanderson describes: the creature column concentrates on fluffy and blue, and the masked half is exactly the set of entries where a key token sits later in the sequence than the query token asking the question.

Context size: the one term that grows as the square

Another fact worth reflecting on about the attention pattern is how its size is equal to the square of the context size.

This is why context size can be a really huge bottleneck for large language models, and why scaling it up is non trivial. Motivated by a desire for bigger and bigger context windows, recent years have seen variations to the attention mechanism aimed at making context more scalable. But right here, he says, we are staying focused on the basics.

Worth attaching a number he does not give: the GPT-3 paper puts that model's context window at 2,048 tokens, which means the attention pattern inside every one of its heads is a 2,048 by 2,048 grid. Four million entries, ninety six times per block, ninety six blocks deep.

Values: the matrix that actually moves the meaning

Computing the pattern lets the model deduce which words are relevant to which other words. It does not yet change anything. Now you need to actually update the embeddings, allowing words to pass information to whichever other words they are relevant to.

For example, you want the embedding of fluffy to somehow cause a change to creature that moves it to a different part of this 12,000 dimensional embedding space, one that more specifically encodes a fluffy creature.

He shows the most straightforward way first, flagging that there is a slight modification once you get to multi headed attention.

The straightforward way uses a third matrix, the value matrix, which you multiply by the embedding of that first word, for example fluffy. The result is a value vector, and this is something you add to the embedding of the second word, in this case something you add to the embedding of creature. So the value vector lives in the same very high dimensional space as the embeddings.

The intuition he gives for what the value matrix means: when you multiply it by the embedding of a word, you might think of it as saying, if this word is relevant to adjusting the meaning of something else, what exactly should be added to the embedding of that something else in order to reflect this?

Then the mechanics, step by step. Looking back at the diagram, set aside all of the keys and the queries, since after you compute the attention pattern you are done with those. Take the value matrix and multiply it by every one of the embeddings to produce a sequence of value vectors. You might think of these value vectors as being associated with the corresponding keys.

For each column in the diagram, multiply each of the value vectors by the corresponding weight in that column. Under the embedding of creature, for example, you would be adding large proportions of the value vectors for fluffy and blue, while all of the other value vectors get zeroed out, or at least nearly zeroed out.

And then, to actually update the embedding associated with that column, previously encoding some context free meaning of creature, you add together all of these rescaled values in the column. That produces a change he labels delta e, and you add that to the original embedding. What results, hopefully, is a more refined vector encoding the more contextually rich meaning, that of a fluffy blue creature.

Of course you do not just do this to one embedding. You apply the same weighted sum across all of the columns, producing a sequence of changes, and adding all of those changes to the corresponding embeddings produces a full sequence of more refined embeddings popping out of the attention block.

Zooming out, this whole process is what you would describe as a single head of attention.

StepWhat runsWhat comes out
Query projectionMultiply every embedding by W_QOne query vector per token, 128 dimensions
Key projectionMultiply every embedding by W_KOne key vector per token, 128 dimensions
ScoringDot product of every key query pair, divided by the square root of 128A grid of real numbers, size equal to the square of the context
MaskingSet every entry where the key is later than the query to negative infinityA grid that cannot leak the future
SoftmaxApplied column by columnThe attention pattern: each column a distribution summing to one
Value projectionMultiply every embedding by the value mapOne value vector per token, in the full 12,288 dimensional space
Weighted sumEach column's weights applied to the value vectors, then summedDelta e, the change for that position
Residual addAdd delta e to the original embeddingA refined embedding that now carries context

The parameter tally, part one: queries and keys against GPT-3

As described so far, the process is parameterized by three distinct matrices, all filled with tunable parameters: the key, the query, and the value. Sanderson now continues the scorekeeping he started in chapter 5, counting up the total number of model parameters using the numbers from GPT-3.

The key and query matrices each have 12,288 columns, matching the embedding dimension, and 128 rows, matching the dimension of that smaller key query space. That gives an additional 1.5 million or so parameters for each one.

Run the multiplication yourself and 12,288 times 128 is 1,572,864, so "1.5 million or so" is doing honest rounding.

The value matrix would blow the budget, so it gets factored

Now look at the value matrix. The way he has described things so far would suggest it is a square matrix with 12,288 columns and 12,288 rows, since both its inputs and its outputs live in that very large embedding space.

If true, that would mean about 150 million added parameters. And to be clear, he says, you could do that. You could devote orders of magnitude more parameters to the value map than to the key and the query.

But in practice it is much more efficient if instead you make the number of parameters devoted to the value map the same as the number devoted to the key and the query. This is especially relevant in the setting of running multiple attention heads in parallel, which is where the whole video is heading.

The way this looks is that the value map is factored as a product of two smaller matrices.

Conceptually he still encourages you to think about the overall linear map as one with inputs and outputs both in the larger embedding space, for example taking the embedding of blue to the blueness direction that you would add to nouns. The factorization is an implementation of that map, not a different map.

The first matrix on the right has a smaller number of rows, typically the same size as the key query space. You can think of it as mapping the large embedding vectors down to a much smaller space. He calls it the value down matrix, and says plainly that this is not the conventional naming.

The second matrix maps from that smaller space back up to the embedding space, producing the vectors you use to make the actual updates. He calls that one the value up matrix, which again is not conventional.

In linear algebra jargon, what this amounts to is constraining the overall value map to be a low rank transformation. The rank cannot exceed 128 no matter what the weights learn, because the information has to squeeze through a 128 dimensional bottleneck on the way across.

THE VALUE MAP: ONE SQUARE MATRIX, OR TWO SKINNY ONES THE OBVIOUS WAY e 12,288 one square matrix 12,288 x 12,288 ~150,000,000 params delta e 12,288 Possible, he says, just not what is done. Ninety six times the query matrix on its own. WHAT IS ACTUALLY DONE: A LOW RANK FACTORIZATION e 12,288 value down 128 x 12,288 128 value up 12,288 x 128 delta e 12,288 bottleneck 1,572,864 + 1,572,864 = 3,145,728 parameters, about 48 times smaller than the square matrix Conceptually it is still one linear map from the embedding space back into the embedding space, for example taking the embedding of blue to the blueness direction you add to nouns. The constraint is that the map has rank at most 128, because everything it carries has to pass through that narrow middle. The naming of the two halves is Sanderson's own.
Figure 3. Why the value matrix is never built as a square. The factored version has exactly the same parameter count as the query and key matrices put together, which is the whole reason the per head tally comes out tidy and the reason ninety six heads per block is affordable at all.

The parameter tally, part two: one head

He notes that the way you see this written in most papers looks a little different, that he will come back to it in a minute, and that in his opinion the usual presentation tends to make things a little more conceptually confusing.

Turning back to the parameter count: all four of these matrices have the same size, and adding them all up gives about 6.3 million parameters for one attention head.

The four are the query matrix, the key matrix, the value down matrix and the value up matrix. Four times 1,572,864 is 6,291,456.

A footnote on self attention and cross attention

As a quick side note, and to be a little more accurate, everything described so far is what people would call a self attention head, to distinguish it from a variation that comes up in other models called cross attention.

This is not relevant to the GPT example, but if you are curious: cross attention involves models that process two distinct types of data, like text in one language and text in another language that is part of an ongoing generation of a translation, or audio input of speech and an ongoing transcription.

A cross attention head looks almost identical. The only difference is that the key and query maps act on different data sets. In a model doing translation, the keys might come from one language while the queries come from another, and the attention pattern could describe which words from one language correspond to which words in another. And in this setting there would typically be no masking, since there is not really any notion of later tokens affecting earlier ones.

Self attentionCross attention
Where queries come fromThe sequence being processedOne of the two data streams
Where keys come fromThe same sequenceThe other data stream
MaskingAlways applied in the GPT caseTypically none
What the pattern meansWhich words in a passage are relevant to which othersWhich words in one language correspond to which in another
Example settingGPT style next token predictionTranslation, or speech audio against a running transcription

Staying focused on self attention, he makes the claim that justifies the structure of the whole video: if you understood everything so far, and if you were to stop here, you would come away with the essence of what attention really is. All that is really left is to lay out the sense in which you do this many, many different times.

Many heads, because context changes meaning in many ways

The central example focused on adjectives updating nouns. But of course there are lots of different ways that context can influence the meaning of a word, and this is where the video opens out.

If the words "they crashed the" preceded the word "car," that has implications for the shape and structure of that car. A lot of associations might be much less grammatical than that. If the word "wizard" is anywhere in the same passage as "Harry," it suggests this might be referring to Harry Potter, whereas if instead the words "Queen," "Sussex" and "William" were in that passage, then perhaps the embedding of Harry should instead be updated to refer to the prince. Same token, same lookup table entry, two completely different destinations in embedding space depending on words that might be dozens of positions away.

For every different type of contextual updating you might imagine, the parameters of the key and query matrices would be different, to capture the different attention patterns, and the parameters of the value map would be different, based on what should be added to the embeddings. And again, in practice the true behavior of these maps is much more difficult to interpret, where the weights are set to do whatever the model needs them to do to best accomplish its goal of predicting the next token.

Everything described so far is a single head of attention. A full attention block inside a transformer consists of what is called multi headed attention, where you run a lot of these operations in parallel, each with its own distinct key, query and value maps.

GPT-3 uses 96 attention heads inside each block. He acknowledges, reasonably, that considering each one is already a bit confusing, it is certainly a lot to hold in your head.

Spelled out very explicitly:

What this means is that for each position in the context, for each token, every one of these heads produces a proposed change to be added to the embedding in that position. So you sum together all of those proposed changes, one for each head, and you add the result to the original embedding of that position. That entire sum is one slice of what is output from the multi headed attention block: a single one of those refined embeddings that pops out the other end.

Again, he says, this is a lot to think about, so do not worry at all if it takes some time to sink in. The overall idea is that by running many distinct heads in parallel, you are giving the model the capacity to learn many distinct ways that context changes meaning.

The parameter tally, part three: one block

Pulling up the running tally, with 96 heads, each including its own variation of these four matrices, each block of multi headed attention ends up with around 600 million parameters.

Six point three million times ninety six is 603,979,776. One attention block. There are ninety six of them.

The output matrix, and the convention the papers actually use

Here is the slightly annoying thing he says he really has to mention for anyone who goes on to read more about transformers. It is the detail that trips people up when they move from a clean explanation to real code.

He framed the value map as factored into two distinct matrices, the value down and the value up. That framing would suggest you see this pair of matrices inside each attention head, and you could absolutely implement it that way. It would be a valid design.

But the way you see it written in papers, and the way it is implemented in practice, looks a little different. All of the value up matrices for each head appear stapled together in one giant matrix, called the output matrix, associated with the entire multi headed attention block. And when you see people refer to the value matrix for a given attention head, they are typically only referring to the first step, the one he was labeling as the value down projection into the smaller space.

So the vocabulary mismatch is specific and worth memorizing:

ComponentSanderson's name in this videoWhat papers and code call it
Projection down into the 128 dimensional spaceValue down matrixThe value matrix, W_V, per head
Projection back up into the 12,288 dimensional spaceValue up matrix, one per headA slice of the output matrix, W_O, one per block
Where it livesInside each head, as a pairDown matrices per head, up matrices stapled into one block level matrix
Parameter countIdentical either way. This is packaging, not a different model.

He adds that for the curious he left an on screen note about it, and that it is one of those details that runs the risk of distracting from the main conceptual points, but that he wanted to call it out so that you know what you are looking at if you read about this in other sources.

Going deeper: ninety six layers of this

Setting aside all the technical nuances, the preview from chapter 5 showed that data flowing through a transformer does not just flow through a single attention block. For one thing, it also goes through these other operations called multi layer perceptrons, the subject of the next chapter. And then it repeatedly goes through many, many copies of both of these operations.

What this means is that after a given word imbibes some of its context, there are many more chances for this more nuanced embedding to be influenced by its more nuanced surroundings. The further down the network you go, with each embedding taking in more and more meaning from all the other embeddings, which themselves are getting more and more nuanced, the hope is that there is the capacity to encode higher level and more abstract ideas about a given input, beyond just descriptors and grammatical structure. Things like sentiment and tone, and whether it is a poem, and what underlying scientific truths are relevant to the piece.

That is the honest statement of the hope. Not a claim about what provably happens, but the reason depth is there.

The parameter tally, part four: the whole model

Turning back one more time to the scorekeeping: GPT-3 includes 96 distinct layers, so the total number of key, query and value parameters is multiplied by another 96. That brings the total sum to just under 58 billion distinct parameters devoted to all of the attention heads.

That is a lot, to be sure. But it is only about a third of the 175 billion that are in the network in total.

Which gives him the line the chapter is built toward: even though attention gets all of the attention, the majority of parameters come from the blocks sitting in between these steps. The next chapter is about those other blocks, and about the training process.

THE RUNNING TALLY, EXACTLY AS HE BUILDS IT One matrix, 128 x 12,288 1,572,864 "1.5 million or so" x 4 matrices per head query, key, value down, value up One attention head 6,291,456 "about 6.3 million" x 96 heads per block One multi headed attention block 603,979,776 "around 600 million" x 96 layers All attention in GPT-3 57,982,058,496 "just under 58 billion" AGAINST THE WHOLE NETWORK 58 B 117 B attention everything in between, mostly MLP 175 BILLION PARAMETERS TOTAL Attention is roughly a third of GPT-3, which means two thirds of the model is the multi layer perceptron blocks that sit between the attention layers. That is the subject of chapter 7. The rounded figure beside each box is his, stated on screen. The exact integers are what those roundings resolve to if you do the multiplication yourself. one skinny matrix, multiplied by four, by ninety six, by ninety six
Figure 4. The scorekeeping chain in full. The reason this is the memorable part of the chapter is that it converts every abstraction into an amount of hardware: one 128 by 12,288 matrix is the atom, and four multiplications later you have most of a frontier model's attention budget.
GPT-3 quantityValueWhere it comes from in the video
Embedding dimension12,288The column count of the key and query matrices, carried over from chapter 5
Key query space dimension128The row count of the key and query matrices, and the size of the value bottleneck
Scaling divisor before softmaxThe square root of 128The numerical stability term in the paper's formula
Query matrix~1.5 million parameters128 rows by 12,288 columns
Key matrix~1.5 million parametersSame shape as the query matrix
Value map if built square~150 million parameters12,288 by 12,288, the version nobody builds
Value map as factored~3.1 million parametersValue down plus value up, each the same size as the query matrix
One attention head~6.3 million parametersFour matrices of identical size
Heads per block96Stated directly for GPT-3
One attention block~600 million parameters96 heads times 6.3 million
Layers96Stated directly for GPT-3
All attentionJust under 58 billion parameters96 blocks times 600 million
Whole network175 billion parametersAttention is about a third of it

Why this architecture won

The closing argument is the one worth carrying out of the chapter, and it is not about language at all.

A big part of the story for the success of the attention mechanism is not so much any specific kind of behavior that it enables, but the fact that it is extremely parallelizable, meaning you can run a huge number of computations in a short time using GPUs.

Given that one of the big lessons about deep learning in the last decade or two has been that scale alone seems to give huge qualitative improvements in model performance, there is a huge advantage to parallelizable architectures that let you do this.

Read Figure 1 again with that in mind. Every operation in the chapter is a matrix multiplication over the whole sequence at once. Nothing in the head waits for the previous token to finish. That is the property the earlier recurrent models did not have, and it is the property that lets you spend 58 billion parameters on attention and still train the thing.

Key takeaways

Chapters

Notable quotes

"The aim of a transformer is to progressively adjust these embeddings so that they don't merely encode an individual word, but instead they bake in some much, much richer contextual meaning." Grant Sanderson, 1:32

"I should say up front that a lot of people find the attention mechanism, this key piece in a transformer, very confusing, so don't worry if it takes some time for things to sink in." Grant Sanderson, 1:32

"The vector that's associated with mole would be the same in all of these cases, because this initial token embedding is effectively a lookup table with no reference to the context." Grant Sanderson, 2:04

"If the model is going to accurately predict the next word, that final vector in the sequence, which began its life simply embedding the word was, will have to have been updated by all of the attention blocks to represent much, much more than any individual word." Grant Sanderson, 3:43, on the mystery novel

"As with so much deep learning, the true behavior is much harder to parse because it's based on tweaking and tuning a huge number of parameters to minimize some cost function." Grant Sanderson, 5:52

"You might imagine each noun, like creature, asking the question, hey, are there any adjectives sitting in front of me?" Grant Sanderson, 6:22

"In the lingo, machine learning people would say that this means the embeddings of fluffy and blue attend to the embedding of creature." Grant Sanderson, 9:02

"So instead, a common way to do this is that before applying softmax, you set all of those entries to be negative infinity. If you do that, then after applying softmax, all of those get turned into zero, but the columns stay normalized. This process is called masking." Grant Sanderson, 12:11

"Another fact that's worth reflecting on about this attention pattern is how its size is equal to the square of the context size. So this is why context size can be a really huge bottleneck for large language models, and scaling it up is non-trivial." Grant Sanderson, 12:42

"When you multiply this value matrix by the embedding of a word, you might think of it as saying, if this word is relevant to adjusting the meaning of something else, what exactly should be added to the embedding of that something else in order to reflect this?" Grant Sanderson, 13:44

"This is not the conventional naming, but I'm going to call this the value down matrix." Grant Sanderson, 17:23

"To throw in linear algebra jargon here, what we're basically doing is constraining the overall value map to be a low rank transformation." Grant Sanderson, 17:55

"If you understood everything so far, and if you were to stop here, you would come away with the essence of what attention really is." Grant Sanderson, 18:57

"By running many distinct heads in parallel, you're giving the model the capacity to learn many distinct ways that context changes meaning." Grant Sanderson, 21:32

"All of these value up matrices for each head appear stapled together in one giant matrix that we call the output matrix, associated with the entire multi-headed attention block." Grant Sanderson, 22:34

"So even though attention gets all of the attention, the majority of parameters come from the blocks sitting in between these steps." Grant Sanderson, 24:41

"A big part of the story for the success of the attention mechanism is not so much any specific kind of behaviour that it enables, but the fact that it's extremely parallelizable, meaning that you can run a huge number of computations in a short time using GPUs." Grant Sanderson, 24:41

"Anything produced by Andrej Karpathy or Chris Olah tend to be pure gold." Grant Sanderson, 25:13

Two names the captions get wrong

Spoken names have no spelling, and YouTube's automatic captions guess. Two guesses in the closing credits are worth correcting once here so the links above resolve to real people.

The "Vivek" he calls his friend is vcubingx, Vivek Verma, and the videos in question start with What does it mean for computers to understand language?.

Where this sits in the LLM Learning track

This is Part 1, slot two, of the LLM Learning track, and it is the resolution pass.

Karpathy's deep dive before it treats the transformer as a machine that ingests tokens and emits a probability distribution over the next one. It tells you what the machine is for. This is the chapter that opens the machine and names every part inside one attention head, with the dimensions attached. Read the two together and the phrase "the model attends to earlier tokens" stops being a figure of speech and becomes a 2,048 by 2,048 grid of softmaxed dot products.

Read it before Sasha Rush's Five Formulas, which takes the attention equation as one of its five and spends its time on what that equation fails to explain. The two are complementary in the exact right way: Sanderson derives the formula, Rush stress tests it.

And read it before Part 2, where you type this same computation into a file. Let's build the GPT tokenizer handles the step that happens before any of this, turning text into the tokens whose embeddings get fed in here. Let's reproduce GPT-2 is where the four matrices in Figure 4 become lines of PyTorch, where the mask becomes a triangular buffer, and where the output matrix convention from the section above is the thing that makes the shapes in the code look different from the shapes in this video. That mismatch is the single most common place people get stuck, and this page exists partly so you are not surprised by it.

The natural follow on outside the track is chapter 7 of the same series, How might LLMs store facts, which takes apart the multi layer perceptron blocks holding the other two thirds of GPT-3's parameters.

Resources mentioned

The paper and the model

The series

The people he recommends in the closing

Concepts worth a definition

Full transcript
[00:00:00] In the last chapter, you and I started to step through the internal workings of a transformer. This is one of the key pieces of technology inside large language models, and a lot of other tools in the modern wave of AI. It first hit the scene in a now-famous 2017 paper called Attention is All You Need, and in this chapter you and I will dig into what this attention mechanism is, visualizing how it processes data. As a quick recap, here's the important context I want you to have in mind. [00:00:30] The goal of the model that you and I are studying is to take in a piece of text and predict what word comes next. The input text is broken up into little pieces that we call tokens, and these are very often words or pieces of words, but just to make the examples in this video easier for you and me to think about, let's simplify by pretending that tokens are always just words. The first step in a transformer is to associate each token with a high-dimensional vector, what we call its embedding. The most important idea I want you to have in mind is how directions in this [00:01:02] high-dimensional space of all possible embeddings can correspond with semantic meaning. In the last chapter we saw an example for how direction can correspond to gender, in the sense that adding a certain step in this space can take you from the embedding of a masculine noun to the embedding of the corresponding feminine noun. That's just one example you could imagine how many other directions in this high-dimensional space could correspond to numerous other aspects of a word's meaning. The aim of a transformer is to progressively adjust these embeddings [00:01:32] so that they don't merely encode an individual word, but instead they bake in some much, much richer contextual meaning. I should say up front that a lot of people find the attention mechanism, this key piece in a transformer, very confusing, so don't worry if it takes some time for things to sink in. I think that before we dive into the computational details and all the matrix multiplications, it's worth thinking about a couple examples for the kind of behavior that we want attention to enable. Consider the phrases American shrew mole, one mole of carbon dioxide, [00:02:04] and take a biopsy of the mole. You and I know that the word mole has different meanings in each one of these, based on the context. But after the first step of a transformer, the one that breaks up the text and associates each token with a vector, the vector that's associated with mole would be the same in all of these cases, because this initial token embedding is effectively a lookup table with no reference to the context. It's only in the next step of the transformer that the surrounding embeddings have the chance to pass information into this one. The picture you might have in mind is that there are multiple distinct directions in [00:02:38] this embedding space encoding the multiple distinct meanings of the word mole, and that a well-trained attention block calculates what you need to add to the generic embedding to move it to one of these specific directions, as a function of the context. To take another example, consider the embedding of the word tower. This is presumably some very generic, non-specific direction in the space, associated with lots of other large, tall nouns. If this word was immediately preceded by Eiffel, you could imagine wanting the mechanism to update this vector so that [00:03:10] it points in a direction that more specifically encodes the Eiffel tower, maybe correlated with vectors associated with Paris and France and things made of steel. If it was also preceded by the word miniature, then the vector should be updated even further, so that it no longer correlates with large, tall things. More generally than just refining the meaning of a word, the attention block allows the model to move information encoded in one embedding to that of another, potentially ones that are quite far away, and potentially with information that's much richer than just a single word. [00:03:43] What we saw in the last chapter was how after all of the vectors flow through the network, including many different attention blocks, the computation you perform to produce a prediction of the next token is entirely a function of the last vector in the sequence. Imagine, for example, that the text you input is most of an entire mystery novel, all the way up to a point near the end, which reads, therefore the murderer was. If the model is going to accurately predict the next word, that final vector in the sequence, which began its life simply embedding the word was, [00:04:16] will have to have been updated by all of the attention blocks to represent much, much more than any individual word, somehow encoding all of the information from the full context window that's relevant to predicting the next word. To step through the computations, though, let's take a much simpler example. Imagine that the input includes the phrase, a fluffy blue creature roamed the verdant forest. And for the moment, suppose that the only type of update that we care about is having the adjectives adjust the meanings of their corresponding nouns. [00:04:47] What I'm about to describe is what we would call a single head of attention, and later we will see how the attention block consists of many different heads run in parallel. Again, the initial embedding for each word is some high dimensional vector that only encodes the meaning of that particular word with no context. Actually, that's not quite true. They also encode the position of the word. There's a lot more to say about the specific way that positions are encoded, but right now, all you need to know is that the entries of this vector are enough to tell you both what the word is and where it exists in the context. [00:05:19] Let's go ahead and denote these embeddings with the letter e. The goal is to have a series of computations produce a new refined set of embeddings where, for example, those corresponding to the nouns have ingested the meaning from their corresponding adjectives. And playing the deep learning game, we want most of the computations involved to look like matrix-vector products, where the matrices are full of tuneable weights, things that the model will learn based on data. To be clear, I'm making up this example of adjectives updating nouns just to illustrate the type of behavior that you could imagine an attention head doing. [00:05:52] As with so much deep learning, the true behavior is much harder to parse because it's based on tweaking and tuning a huge number of parameters to minimize some cost function. It's just that as we step through all of different matrices filled with parameters that are involved in this process, I think it's really helpful to have an imagined example of something that it could be doing to help keep it all more concrete. For the first step of this process, you might imagine each noun, like creature, asking the question, hey, are there any adjectives sitting in front of me? [00:06:22] And for the words fluffy and blue, to each be able to answer, yeah, I'm an adjective and I'm in that position. That question is somehow encoded as yet another vector, another list of numbers, which we call the query for this word. This query vector though has a much smaller dimension than the embedding vector, say 128. Computing this query looks like taking a certain matrix, which I'll label wq, and multiplying it by the embedding. Compressing things a bit, let's write that query vector as q, [00:06:54] and then anytime you see me put a matrix next to an arrow like this one, it's meant to represent that multiplying this matrix by the vector at the arrow's start gives you the vector at the arrow's end. In this case, you multiply this matrix by all of the embeddings in the context, producing one query vector for each token. The entries of this matrix are parameters of the model, which means the true behavior is learned from data, and in practice, what this matrix does in a particular attention head is challenging to parse. But for our sake, imagining an example that we might hope that it would learn, [00:07:27] we'll suppose that this query matrix maps the embeddings of nouns to certain directions in this smaller query space that somehow encodes the notion of looking for adjectives in preceding positions. As to what it does to other embeddings, who knows? Maybe it simultaneously tries to accomplish some other goal with those. Right now, we're laser focused on the nouns. At the same time, associated with this is a second matrix called the key matrix, which you also multiply by every one of the embeddings. This produces a second sequence of vectors that we call the keys. [00:07:59] Conceptually, you want to think of the keys as potentially answering the queries. This key matrix is also full of tuneable parameters, and just like the query matrix, it maps the embedding vectors to that same smaller dimensional space. You think of the keys as matching the queries whenever they closely align with each other. In our example, you would imagine that the key matrix maps the adjectives like fluffy and blue to vectors that are closely aligned with the query produced by the word creature. To measure how well each key matches each query, [00:08:30] you compute a dot product between each possible key-query pair. I like to visualize a grid full of a bunch of dots, where the bigger dots correspond to the larger dot products, the places where the keys and queries align. For our adjective noun example, that would look a little more like this, where if the keys produced by fluffy and blue really do align closely with the query produced by creature, then the dot products in these two spots would be some large positive numbers. In the lingo, machine learning people would say that this means the [00:09:02] embeddings of fluffy and blue attend to the embedding of creature. By contrast to the dot product between the key for some other word like the and the query for creature would be some small or negative value that reflects that are unrelated to each other. So we have this grid of values that can be any real number from negative infinity to infinity, giving us a score for how relevant each word is to updating the meaning of every other word. The way we're about to use these scores is to take a certain [00:09:32] weighted sum along each column, weighted by the relevance. So instead of having values range from negative infinity to infinity, what we want is for the numbers in these columns to be between 0 and 1, and for each column to add up to 1, as if they were a probability distribution. If you're coming in from the last chapter, you know what we need to do then. We compute a softmax along each one of these columns to normalize the values. In our picture, after you apply softmax to all of the columns, [00:10:03] we'll fill in the grid with these normalized values. At this point you're safe to think about each column as giving weights according to how relevant the word on the left is to the corresponding value at the top. We call this grid an attention pattern. Now if you look at the original transformer paper, there's a really compact way that they write this all down. Here the variables q and k represent the full arrays of query and key vectors respectively, those little vectors you get by multiplying the embeddings by the query and the key matrices. [00:10:35] This expression up in the numerator is a really compact way to represent the grid of all possible dot products between pairs of keys and queries. A small technical detail that I didn't mention is that for numerical stability, it happens to be helpful to divide all of these values by the square root of the dimension in that key query space. Then this softmax that's wrapped around the full expression is meant to be understood to apply column by column. As to that v term, we'll talk about it in just a second. [00:11:05] Before that, there's one other technical detail that so far I've skipped. During the training process, when you run this model on a given text example, and all of the weights are slightly adjusted and tuned to either reward or punish it based on how high a probability it assigns to the true next word in the passage, it turns out to make the whole training process a lot more efficient if you simultaneously have it predict every possible next token following each initial subsequence of tokens in this passage. For example, with the phrase that we've been focusing on, it might also be predicting what words follow creature and what words follow the. [00:11:39] This is really nice, because it means what would otherwise be a single training example effectively acts as many. For the purposes of our attention pattern, it means that you never want to allow later words to influence earlier words, since otherwise they could kind of give away the answer for what comes next. What this means is that we want all of these spots here, the ones representing later tokens influencing earlier ones, to somehow be forced to be zero. The simplest thing you might think to do is to set them equal to zero, but if you did that the columns wouldn't add up to one anymore, [00:12:11] they wouldn't be normalized. So instead, a common way to do this is that before applying softmax, you set all of those entries to be negative infinity. If you do that, then after applying softmax, all of those get turned into zero, but the columns stay normalized. This process is called masking. There are versions of attention where you don't apply it, but in our GPT example, even though this is more relevant during the training phase than it would be, say, running it as a chatbot or something like that, you do always apply this masking to prevent later tokens from influencing earlier ones. [00:12:42] Another fact that's worth reflecting on about this attention pattern is how its size is equal to the square of the context size. So this is why context size can be a really huge bottleneck for large language models, and scaling it up is non-trivial. As you imagine, motivated by a desire for bigger and bigger context windows, recent years have seen some variations to the attention mechanism aimed at making context more scalable, but right here, you and I are staying focused on the basics. Okay, great, computing this pattern lets the model [00:13:12] deduce which words are relevant to which other words. Now you need to actually update the embeddings, allowing words to pass information to whichever other words they're relevant to. For example, you want the embedding of Fluffy to somehow cause a change to Creature that moves it to a different part of this 12,000-dimensional embedding space that more specifically encodes a Fluffy creature. What I'm going to do here is first show you the most straightforward way that you could do this, though there's a slight way that this gets modified in the context of multi-headed attention. [00:13:44] This most straightforward way would be to use a third matrix, what we call the value matrix, which you multiply by the embedding of that first word, for example Fluffy. The result of this is what you would call a value vector, and this is something that you add to the embedding of the second word, in this case something you add to the embedding of Creature. So this value vector lives in the same very high-dimensional space as the embeddings. When you multiply this value matrix by the embedding of a word, you might think of it as saying, if this word is relevant to adjusting the meaning of [00:14:15] something else, what exactly should be added to the embedding of that something else in order to reflect this? Looking back in our diagram, let's set aside all of the keys and the queries, since after you compute the attention pattern you're done with those, then you're going to take this value matrix and multiply it by every one of those embeddings to produce a sequence of value vectors. You might think of these value vectors as being kind of associated with the corresponding keys. For each column in this diagram, you multiply each of the [00:14:45] value vectors by the corresponding weight in that column. For example here, under the embedding of Creature, you would be adding large proportions of the value vectors for Fluffy and Blue, while all of the other value vectors get zeroed out, or at least nearly zeroed out. And then finally, the way to actually update the embedding associated with this column, previously encoding some context-free meaning of Creature, you add together all of these rescaled values in the column, producing a change that you want to add, that I'll label delta-e, [00:15:16] and then you add that to the original embedding. Hopefully what results is a more refined vector encoding the more contextually rich meaning, like that of a fluffy blue creature. And of course you don't just do this to one embedding, you apply the same weighted sum across all of the columns in this picture, producing a sequence of changes, adding all of those changes to the corresponding embeddings, produces a full sequence of more refined embeddings popping out of the attention block. Zooming out, this whole process is what you would describe as a single head of attention. [00:15:49] As I've described things so far, this process is parameterized by three distinct matrices, all filled with tunable parameters, the key, the query, and the value. I want to take a moment to continue what we started in the last chapter, with the scorekeeping where we count up the total number of model parameters using the numbers from GPT-3. These key and query matrices each have 12,288 columns, matching the embedding dimension, and 128 rows, matching the dimension of that smaller key query space. [00:16:20] This gives us an additional 1.5 million or so parameters for each one. If you look at that value matrix by contrast, the way I've described things so far would suggest that it's a square matrix that has 12,288 columns and 12,288 rows, since both its inputs and outputs live in this very large embedding space. If true, that would mean about 150 million added parameters. And to be clear, you could do that. You could devote orders of magnitude more parameters to the value map than to the key and query. [00:16:52] But in practice, it is much more efficient if instead you make it so that the number of parameters devoted to this value map is the same as the number devoted to the key and the query. This is especially relevant in the setting of running multiple attention heads in parallel. The way this looks is that the value map is factored as a product of two smaller matrices. Conceptually, I would still encourage you to think about the overall linear map, one with inputs and outputs, both in this larger embedding space, for example taking the embedding of blue to this blueness direction that you would [00:17:23] add to nouns. The first matrix on the right here has a smaller number of rows, typically the same size as the key-query space What this means is you can think of it as mapping the large embedding vectors down to a much smaller space. This is not the conventional naming, but I'm going to call this the value down matrix. The second matrix maps from this smaller space back up to the embedding space, producing the vectors that you use to make the actual updates. I'm going to call this one the value up matrix, which again is not conventional. [00:17:55] The way that you would see this written in most papers looks a little different. I'll talk about it in a minute. In my opinion, it tends to make things a little more conceptually confusing. To throw in linear algebra jargon here, what we're basically doing is constraining the overall value map to be a low rank transformation. Turning back to the parameter count, all four of these matrices have the same size, and adding them all up we get about 6.3 million parameters for one attention head. As a quick side note, to be a little more accurate, everything described so far is what people would call a self-attention head, [00:18:27] to distinguish it from a variation that comes up in other models that's called cross-attention. This isn't relevant to our GPT example, but if you're curious, cross-attention involves models that process two distinct types of data, like text in one language and text in another language that's part of an ongoing generation of a translation, or maybe audio input of speech and an ongoing transcription. A cross-attention head looks almost identical. The only difference is that the key and query maps act on different data sets. [00:18:57] In a model doing translation, for example, the keys might come from one language, while the queries come from another, and the attention pattern could describe which words from one language correspond to which words in another. And in this setting there would typically be no masking, since there's not really any notion of later tokens affecting earlier ones. Staying focused on self-attention though, if you understood everything so far, and if you were to stop here, you would come away with the essence of what attention really is. All that's really left to us is to lay out the sense [00:19:28] in which you do this many many different times. In our central example we focused on adjectives updating nouns, but of course there are lots of different ways that context can influence the meaning of a word. If the words they crashed the preceded the word car, it has implications for the shape and structure of that car. And a lot of associations might be less grammatical. If the word wizard is anywhere in the same passage as Harry, it suggests that this might be referring to Harry Potter, whereas if instead the words Queen, Sussex, and William were in that passage, [00:19:59] then perhaps the embedding of Harry should instead be updated to refer to the prince. For every different type of contextual updating that you might imagine, the parameters of these key and query matrices would be different to capture the different attention patterns, and the parameters of our value map would be different based on what should be added to the embeddings. And again, in practice the true behavior of these maps is much more difficult to interpret, where the weights are set to do whatever the model needs them to do to best accomplish its goal of predicting the next token. [00:20:31] As I said before, everything we described is a single head of attention, and a full attention block inside a transformer consists of what's called multi-headed attention, where you run a lot of these operations in parallel, each with its own distinct key query and value maps. GPT-3 for example uses 96 attention heads inside each block. Considering that each one is already a bit confusing, it's certainly a lot to hold in your head. Just to spell it all out very explicitly, this means you have 96 distinct [00:21:01] key and query matrices producing 96 distinct attention patterns. Then each head has its own distinct value matrices used to produce 96 sequences of value vectors. These are all added together using the corresponding attention patterns as weights. What this means is that for each position in the context, each token, every one of these heads produces a proposed change to be added to the embedding in that position. So what you do is you sum together all of those proposed changes, one for each head, [00:21:32] and you add the result to the original embedding of that position. This entire sum here would be one slice of what's outputted from this multi-headed attention block, a single one of those refined embeddings that pops out the other end of it. Again, this is a lot to think about, so don't worry at all if it takes some time to sink in. The overall idea is that by running many distinct heads in parallel, you're giving the model the capacity to learn many distinct ways that context changes meaning. [00:22:03] Pulling up our running tally for parameter count with 96 heads, each including its own variation of these four matrices, each block of multi-headed attention ends up with around 600 million parameters. There's one added slightly annoying thing that I should really mention for any of you who go on to read more about transformers. You remember how I said that the value map is factored out into these two distinct matrices, which I labeled as the value down and the value up matrices. The way that I framed things would suggest that you see this pair of matrices [00:22:34] inside each attention head, and you could absolutely implement it this way. That would be a valid design. But the way that you see this written in papers and the way that it's implemented in practice looks a little different. All of these value up matrices for each head appear stapled together in one giant matrix that we call the output matrix, associated with the entire multi-headed attention block. And when you see people refer to the value matrix for a given attention head, they're typically only referring to this first step, the one that I was labeling as the value down projection into the smaller space. [00:23:08] For the curious among you, I've left an on-screen note about it. It's one of those details that runs the risk of distracting from the main conceptual points, but I do want to call it out just so that you know if you read about this in other sources. Setting aside all the technical nuances, in the preview from the last chapter we saw how data flowing through a transformer doesn't just flow through a single attention block. For one thing, it also goes through these other operations called multi-layer perceptrons. We'll talk more about those in the next chapter. And then it repeatedly goes through many many copies of both of these operations. [00:23:39] What this means is that after a given word imbibes some of its context, there are many more chances for this more nuanced embedding to be influenced by its more nuanced surroundings. The further down the network you go, with each embedding taking in more and more meaning from all the other embeddings, which themselves are getting more and more nuanced, the hope is that there's the capacity to encode higher level and more abstract ideas about a given input beyond just descriptors and grammatical structure. Things like sentiment and tone and whether it's a poem and what underlying [00:24:11] scientific truths are relevant to the piece and things like that. Turning back one more time to our scorekeeping, GPT-3 includes 96 distinct layers, so the total number of key query and value parameters is multiplied by another 96, which brings the total sum to just under 58 billion distinct parameters devoted to all of the attention heads. That is a lot to be sure, but it's only about a third of the 175 billion that are in the network in total. [00:24:41] So even though attention gets all of the attention, the majority of parameters come from the blocks sitting in between these steps. In the next chapter, you and I will talk more about those other blocks and also a lot more about the training process. A big part of the story for the success of the attention mechanism is not so much any specific kind of behaviour that it enables, but the fact that it's extremely parallelizable, meaning that you can run a huge number of computations in a short time using GPUs. Given that one of the big lessons about deep learning in the last decade or two has [00:25:13] been that scale alone seems to give huge qualitative improvements in model performance, there's a huge advantage to parallelizable architectures that let you do this. If you want to learn more about this stuff, I've left lots of links in the description. In particular, anything produced by Andrej Karpathy or Chris Ola tend to be pure gold. In this video, I wanted to just jump into attention in its current form, but if you're curious about more of the history for how we got here and how you might reinvent this idea for yourself, my friend Vivek just put up a couple videos giving a lot more of that motivation. [00:25:43] Also, Britt Cruz from the channel The Art of the Problem has a really nice video about the history of large language models. Thank you.