All notes
AI / EXPLAINER / 4 min read

Why attention needs scaling

A three-dimensional example of dot products, softmax, and the square root in between.

Think of attention as an allocation

Imagine reading a sentence and trying to understand one word. That word asks the context for useful information; the surrounding words offer possible clues.

In attention, the Query asks, the Key identifies a match, and the Value carries the information to retrieve. A dot product between the Query and each Key gives a score. Softmax converts those scores into weights, which are then used to combine the Values.

A higher score usually receives more weight. That does not mean larger scores are always better.

ATTENTION / ALLOCATION02 DOT PRODUCTAFTER ÷ √3 → SOFTMAX 30.70710.223−10.070 ONE QUERY / THREE KEYS / CALCULATED EXAMPLE
One Query, three Keys. The same ranking, with a less concentrated allocation.

Why does dimension affect the scores?

A dot product adds the products of corresponding components:

q⊤k=∑i=1dkqikiq^\top k = \sum_{i=1}^{d_k} q_i k_i

More dimensions mean more terms in the sum. Even when component sizes stay similar, the variation in the resulting dot products can grow.

Softmax exponentiates its inputs. Large differences between scores can concentrate almost all the weight on one position. Near saturation, some gradients become small, making those choices harder to adjust.

Where does the square root come from?

Scaled dot-product attention is:

Attention⁡(Q,K,V)=softmax⁡ ⁣(QK⊤dk+M)V\operatorname{Attention}(Q,K,V) = \operatorname{softmax}\!\left( \frac{QK^\top}{\sqrt{d_k}} + M \right)V

Consider a simplified model in which all the Query and Key components are independent, with mean zero and variance one. Each component product then has variance one, and the sum of dkd_k terms has variance dkd_k:

Var⁡(q⊤k)=dk,Var⁡ ⁣(q⊤kdk)=1.\operatorname{Var}(q^\top k) = d_k, \qquad \operatorname{Var}\!\left(\frac{q^\top k}{\sqrt{d_k}}\right) = 1.

Dividing by dk\sqrt{d_k} brings the score variance back to one under these assumptions. It keeps the scale more manageable across dimensions. Real training distributions need not satisfy the assumptions exactly.

Try three Keys

Let q=(1,1,1)q=(1,1,1), with:

k1=(1,1,1),k2=(1,0,0),k3=(−1,0,0).\begin{aligned} k_1 &= (1,1,1),\\ k_2 &= (1,0,0),\\ k_3 &= (-1,0,0). \end{aligned}

The dot-product scores are (3,1,−1)(3,1,-1). Without scaling:

softmax⁡(3,1,−1)≈(0.867, 0.117, 0.016).\operatorname{softmax}(3,1,-1) \approx (0.867,\ 0.117,\ 0.016).

After division by 3\sqrt{3}:

softmax⁡ ⁣((3,1,−1)3)≈(0.707, 0.223, 0.070).\operatorname{softmax}\!\left(\frac{(3,1,-1)}{\sqrt{3}}\right) \approx (0.707,\ 0.223,\ 0.070).

The ranking stays the same. The allocation becomes less concentrated. Try changing the divisor below.

TRY THE IDEA

Same scores. Different temperature.

The dot-product scores stay fixed at [3, 1, −1]. Change the divisor to see how the weights respond.

Key 1
70.7%
Key 2
22.3%
Key 3
7.0%

A calculated example. Larger divisors produce a less concentrated distribution. Standard attention uses √dₖ.

These values are a calculated example, not a model-training result. Scaling alone does not make attention meaningful; the model still has to learn useful Queries and Keys.

Masking answers a different question

Scaling controls the size of the scores. Masking determines which positions may participate. A causal mask blocks future positions, while a padding mask excludes positions added to fill a sequence.

In an additive mask, allowed positions receive zero and blocked positions receive −∞-\infty before softmax.

An entirely masked row cannot produce a valid probability distribution through ordinary softmax. An implementation needs an explicit policy, such as preventing that input or returning zero weights where its interface requires them. Boolean-mask conventions can also differ between interfaces, so check their meaning before use.

What can you verify next?

Hold the same QQ, KK, and VV fixed and compare scores, weights, and outputs with and without scaling. Then introduce causal and padding masks and inspect the excluded positions.

Keep the derivation separate from the execution record. That makes it easier to see what the experiment actually established.

Original method: Attention Is All You Need.

CONTINUE EXPLORINGFinite automata, one transition at a time.