Kiến trúc Transformer
Thay cách truyền từng từ bằng một phòng họp nơi mọi từ cùng lên tiếng, để phụ thuộc xa chỉ còn cách một bước
Toàn văn của trang này được trình bày bằng tiếng Anh; tiêu đề và phần dẫn đã được bản địa hóa.
ĐỊNH NGHĨA
The Transformer is a sequence-modelling architecture introduced in 2017 that relies entirely on attention and abandons recurrence. It stacks structurally identical layers, each holding multi-head attention and a feed-forward network, wrapped with residual connections and layer normalisation. Order is no longer implicit in time steps but injected explicitly through positional encodings.
Trực giác
A recurrent network plays telephone: information travels one word at a time, degrades over distance and cannot be parallelised. The Transformer holds a plenary meeting instead — every word can see every word, relevance is decided on the spot by attention, and word order is marked by giving each position a "seat number" (a positional encoding).
The Transformer data flow: embeddings plus positional encoding enter stacked attention–feed-forward blocks and finally yield a probability distribution over the next token
Ability profiles of three architectural variants (qualitative, 0–5): there is no all-rounder — choose by task focus
- Encoder-only (BERT)
- Decoder-only (GPT)
- Encoder-Decoder (T5)
Cách hoạt động
- 01
Embeddings plus positional encoding
Tokens become embedding vectors, to which positional information is added. The original paper used sine and cosine waves at different frequencies to encode absolute position; modern models largely switch to Rotary Position Embedding (RoPE), which encodes position as a relative rotation angle between vectors and extrapolates better to sequences longer than those seen in training.
- 02
Multi-head self-attention
Each position attends to all positions at once, and multiple heads capture different relation patterns in parallel. An encoder permits bidirectional attention, while a decoder adds a causal mask so a position may only see itself and the past.
- 03
Feed-forward network
Attention moves information between positions; the feed-forward network then applies a nonlinear transform at each position independently, typically widening (say fourfold) before projecting back. In large models the great majority of parameters live here.
- 04
Residuals and layer norm, then stack
Each sublayer is wrapped in a residual (x + sublayer(x)) and a layer norm so gradients flow cleanly through dozens or hundreds of layers. Modern designs place the norm before the sublayer (Pre-LN) for stability; stacking these layers dozens deep yields the body of a large model.
Two mainstream variants and their division of labour: encoder-only excels at understanding, decoder-only at generation
The growth of context windows: better positional encoding schemes expanded usable context by three orders of magnitude within a few years
Ứng dụng
- Decoder-only (GPT family): autoregressive text generation, dialogue and code completion
- Encoder-only (BERT family): text classification, extractive QA and semantic understanding
- Encoder-Decoder (T5, the original Transformer): sequence-to-sequence tasks such as translation and summarisation
- Cross-modal extension: ViT for image patches, Whisper for audio — all sharing one backbone
Hiểu lầm thường gặp
- Remove positional encoding and self-attention becomes permutation-invariant: the model degrades into a bag of words with no sense of order. Position must be injected explicitly.
- No single variant wins everywhere. Encoder-only barely generates, decoder-only is weaker on tasks needing bidirectional understanding, and encoder-decoder is structurally heavier and costlier to serve — the choice follows the task, not novelty.
- The same architecture does not imply the same ability. Differences from data, scale and alignment are usually far larger than those between architectural variants.
Thuật ngữ chính
- Positional encoding
- An explicit order signal, sinusoidal or RoPE
- RoPE
- Rotary Position Embedding: relative position with better extrapolation
- Residual connection
- Adding the input past a sublayer to ease vanishing gradients in depth
- Pre-LN
- Placing layer norm before each sublayer for stability