Think of attention as an allocation
Imagine reading a sentence and trying to understand one word. That word asks the context for useful information; the surrounding words offer possible clues.
In attention, the Query asks, the Key identifies a match, and the Value carries the information to retrieve. A dot product between the Query and each Key gives a score. Softmax converts those scores into weights, which are then used to combine the Values.
A higher score usually receives more weight. That does not mean larger scores are always better.
Why does dimension affect the scores?
A dot product adds the products of corresponding components:
More dimensions mean more terms in the sum. Even when component sizes stay similar, the variation in the resulting dot products can grow.
Softmax exponentiates its inputs. Large differences between scores can concentrate almost all the weight on one position. Near saturation, some gradients become small, making those choices harder to adjust.
Where does the square root come from?
Scaled dot-product attention is:
Consider a simplified model in which all the Query and Key components are independent, with mean zero and variance one. Each component product then has variance one, and the sum of terms has variance :
Dividing by brings the score variance back to one under these assumptions. It keeps the scale more manageable across dimensions. Real training distributions need not satisfy the assumptions exactly.
Try three Keys
Let , with:
The dot-product scores are . Without scaling:
After division by :
The ranking stays the same. The allocation becomes less concentrated. Try changing the divisor below.
Same scores. Different temperature.
The dot-product scores stay fixed at [3, 1, −1]. Change the divisor to see how the weights respond.
A calculated example. Larger divisors produce a less concentrated distribution. Standard attention uses √dₖ.
These values are a calculated example, not a model-training result. Scaling alone does not make attention meaningful; the model still has to learn useful Queries and Keys.
Masking answers a different question
Scaling controls the size of the scores. Masking determines which positions may participate. A causal mask blocks future positions, while a padding mask excludes positions added to fill a sequence.
In an additive mask, allowed positions receive zero and blocked positions receive before softmax.
An entirely masked row cannot produce a valid probability distribution through ordinary softmax. An implementation needs an explicit policy, such as preventing that input or returning zero weights where its interface requires them. Boolean-mask conventions can also differ between interfaces, so check their meaning before use.
What can you verify next?
Hold the same , , and fixed and compare scores, weights, and outputs with and without scaling. Then introduce causal and padding masks and inspect the excluded positions.
Keep the derivation separate from the execution record. That makes it easier to see what the experiment actually established.
Original method: Attention Is All You Need.