어텐션 메커니즘
모든 위치가 다른 모든 위치를 직접 보고 관련도에 따라 주의를 동적으로 배분한다
이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.
정의
Attention lets every position in a sequence dynamically look at other positions and take their weighted sum according to relevance. It rests on three groups of vectors — Query, Key and Value: each Query is scored against all Keys, the scores are normalised with softmax into weights, and the Values are summed with those weights to form a new representation for the current position.
직관적 이해
Reading a comprehension passage, the moment you hit the pronoun "it" you automatically glance back to find which animal it refers to and let that sentence carry more weight in your mind. Attention is exactly this "look back and allocate focus on demand" behaviour, run simultaneously for every position. In a meeting, you (Query) decide whom to listen to by how relevant each person’s (Key) point is, and what stays with you is the weighted content (Value).
One attention head’s weight matrix: rows are Queries (current token), columns are Keys (attended tokens), each row sums to 1; darker means more attention
Self-attention over a real sentence: the "it" Query row shows a clear spike in the "animal" column — visible evidence of the model resolving coreference
작동 원리
- 01
Produce Q, K and V
Multiply each token vector by three learnable matrices W_Q, W_K, W_V to get a Query (what am I looking for), a Key (what can I offer) and a Value (the information I actually carry). In self-attention all three come from the same sequence; in cross-attention the Query comes from the decoder while Key and Value come from the encoder.
- 02
Score with scaled dot products
Measure similarity by taking the dot product of a Query with each Key, then scale by 1/√d_k. The scaling matters: larger dimensions inflate the variance of the dot products, pushing softmax into saturation where gradients nearly vanish.
- 03
Normalise into weights with softmax
For each Query, turn the row of scores into a weight distribution summing to 1 via softmax. In a decoder, scores for future positions are first set to −∞ so the model cannot peek at words it has not generated yet (causal masking).
- 04
Sum the Values with those weights
Take the weighted sum of the Values to obtain the output for that position. Several attention heads run this in parallel, each learning a different focus — some track syntactic dependencies, some coreference, some local adjacency.
핵심 수식
Attention(Q, K, V) = softmax(Q·Kᵀ / √d_k) · V응용 분야
- Word alignment in translation: correspondences between source and target words are learned automatically
- Long-document modelling: linking distant coreference and dependencies without passing them word by word
- Cross-modal alignment: image patches and text tokens retrieve one another, as in CLIP
- Interpretability: visualising attention weights to observe which positions the model attends to
흔한 오해
- A high attention weight is not a causal explanation. It reflects one weighting in one layer and head; it cannot be read directly as "the model decided this because it looked there". The real cause lies in the compounding across layers.
- Multiple heads are not about speed but about capturing several relations in parallel across subspaces. Pruning heads usually degrades language ability directly.
- Compute and memory grow quadratically with sequence length. Doubling the context quadruples the attention matrix — the root reason long context is expensive and why sparse or linear attention matters.
핵심 용어
- Query / Key / Value
- The three vector roles: what you seek, what is on offer, what is carried
- Self-attention
- Attention whose Q, K and V all come from one sequence
- Cross-attention
- Query from one sequence, Key/Value from another
- Multi-head attention
- Several attentions in parallel, each learning a different focus