Attention Mechanism
Every position can look directly at every other position and dynamically weight how much attention to pay
DEFINITION
Attention lets every position in a sequence dynamically look at other positions and take their weighted sum according to relevance. It rests on three groups of vectors — Query, Key and Value: each Query is scored against all Keys, the scores are normalised with softmax into weights, and the Values are summed with those weights to form a new representation for the current position.
Intuition
Reading a comprehension passage, the moment you hit the pronoun "it" you automatically glance back to find which animal it refers to and let that sentence carry more weight in your mind. Attention is exactly this "look back and allocate focus on demand" behaviour, run simultaneously for every position. In a meeting, you (Query) decide whom to listen to by how relevant each person’s (Key) point is, and what stays with you is the weighted content (Value).
One attention head’s weight matrix: rows are Queries (current token), columns are Keys (attended tokens), each row sums to 1; darker means more attention
Self-attention over a real sentence: the "it" Query row shows a clear spike in the "animal" column — visible evidence of the model resolving coreference
How it works
- 01
Produce Q, K and V
Multiply each token vector by three learnable matrices W_Q, W_K, W_V to get a Query (what am I looking for), a Key (what can I offer) and a Value (the information I actually carry). In self-attention all three come from the same sequence; in cross-attention the Query comes from the decoder while Key and Value come from the encoder.
- 02
Score with scaled dot products
Measure similarity by taking the dot product of a Query with each Key, then scale by 1/√d_k. The scaling matters: larger dimensions inflate the variance of the dot products, pushing softmax into saturation where gradients nearly vanish.
- 03
Normalise into weights with softmax
For each Query, turn the row of scores into a weight distribution summing to 1 via softmax. In a decoder, scores for future positions are first set to −∞ so the model cannot peek at words it has not generated yet (causal masking).
- 04
Sum the Values with those weights
Take the weighted sum of the Values to obtain the output for that position. Several attention heads run this in parallel, each learning a different focus — some track syntactic dependencies, some coreference, some local adjacency.
Key formula
Attention(Q, K, V) = softmax(Q·Kᵀ / √d_k) · VWhere it is used
- Word alignment in translation: correspondences between source and target words are learned automatically
- Long-document modelling: linking distant coreference and dependencies without passing them word by word
- Cross-modal alignment: image patches and text tokens retrieve one another, as in CLIP
- Interpretability: visualising attention weights to observe which positions the model attends to
Common misconceptions
- A high attention weight is not a causal explanation. It reflects one weighting in one layer and head; it cannot be read directly as "the model decided this because it looked there". The real cause lies in the compounding across layers.
- Multiple heads are not about speed but about capturing several relations in parallel across subspaces. Pruning heads usually degrades language ability directly.
- Compute and memory grow quadratically with sequence length. Doubling the context quadruples the attention matrix — the root reason long context is expensive and why sparse or linear attention matters.
Key terms
- Query / Key / Value
- The three vector roles: what you seek, what is on offer, what is carried
- Self-attention
- Attention whose Q, K and V all come from one sequence
- Cross-attention
- Query from one sequence, Key/Value from another
- Multi-head attention
- Several attentions in parallel, each learning a different focus