Aller au contenu
Atlas de l'IA

Mécanisme d’attention

Chaque position peut regarder directement toutes les autres et répartir dynamiquement son attention selon la pertinence

04 TAL et grands modèles de langageIntermédiaireEntrée 3 de ce domaine

Le texte intégral est présenté en anglais ; le titre et le résumé sont localisés.

DÉFINITION

Attention lets every position in a sequence dynamically look at other positions and take their weighted sum according to relevance. It rests on three groups of vectors — Query, Key and Value: each Query is scored against all Keys, the scores are normalised with softmax into weights, and the Values are summed with those weights to form a new representation for the current position.

Intuition

Reading a comprehension passage, the moment you hit the pronoun "it" you automatically glance back to find which animal it refers to and let that sentence carry more weight in your mind. Attention is exactly this "look back and allocate focus on demand" behaviour, run simultaneously for every position. In a meeting, you (Query) decide whom to listen to by how relevant each person’s (Key) point is, and what stays with you is the weighted content (Value).

Fig. 1

One attention head’s weight matrix: rows are Queries (current token), columns are Keys (attended tokens), each row sums to 1; darker means more attention

0.350.20.10.20.150.150.350.250.120.130.080.30.340.160.120.10.10.250.30.250.060.080.140.320.4
Fig. 2

Self-attention over a real sentence: the "it" Query row shows a clear spike in the "animal" column — visible evidence of the model resolving coreference

Fonctionnement

  1. 01

    Produce Q, K and V

    Multiply each token vector by three learnable matrices W_Q, W_K, W_V to get a Query (what am I looking for), a Key (what can I offer) and a Value (the information I actually carry). In self-attention all three come from the same sequence; in cross-attention the Query comes from the decoder while Key and Value come from the encoder.

  2. 02

    Score with scaled dot products

    Measure similarity by taking the dot product of a Query with each Key, then scale by 1/√d_k. The scaling matters: larger dimensions inflate the variance of the dot products, pushing softmax into saturation where gradients nearly vanish.

  3. 03

    Normalise into weights with softmax

    For each Query, turn the row of scores into a weight distribution summing to 1 via softmax. In a decoder, scores for future positions are first set to −∞ so the model cannot peek at words it has not generated yet (causal masking).

  4. 04

    Sum the Values with those weights

    Take the weighted sum of the Values to obtain the output for that position. Several attention heads run this in parallel, each learning a different focus — some track syntactic dependencies, some coreference, some local adjacency.

Formule clé

Attention(Q, K, V) = softmax(Q·Kᵀ / √d_k) · V
Scaled dot-product attention: QKᵀ gives similarity scores, √d_k keeps gradients stable, and softmax weights the Values.

Où c'est utilisé

  • Word alignment in translation: correspondences between source and target words are learned automatically
  • Long-document modelling: linking distant coreference and dependencies without passing them word by word
  • Cross-modal alignment: image patches and text tokens retrieve one another, as in CLIP
  • Interpretability: visualising attention weights to observe which positions the model attends to

Idées fausses courantes

  • A high attention weight is not a causal explanation. It reflects one weighting in one layer and head; it cannot be read directly as "the model decided this because it looked there". The real cause lies in the compounding across layers.
  • Multiple heads are not about speed but about capturing several relations in parallel across subspaces. Pruning heads usually degrades language ability directly.
  • Compute and memory grow quadratically with sequence length. Doubling the context quadruples the attention matrix — the root reason long context is expensive and why sparse or linear attention matters.

Termes clés

Query / Key / Value
The three vector roles: what you seek, what is on offer, what is carried
Self-attention
Attention whose Q, K and V all come from one sequence
Cross-attention
Query from one sequence, Key/Value from another
Multi-head attention
Several attentions in parallel, each learning a different focus

Lectures complémentaires