معمارية Transformer
يستبدل النقل كلمةً بكلمة بغرفة يتحدث فيها الجميع معاً، فتصبح التبعيات البعيدة على مسافة خطوة واحدة
يُعرض النص الكامل باللغة الإنجليزية؛ وقد تمت ترجمة العنوان والملخص.
التعريف
The Transformer is a sequence-modelling architecture introduced in 2017 that relies entirely on attention and abandons recurrence. It stacks structurally identical layers, each holding multi-head attention and a feed-forward network, wrapped with residual connections and layer normalisation. Order is no longer implicit in time steps but injected explicitly through positional encodings.
الحدس المباشر
A recurrent network plays telephone: information travels one word at a time, degrades over distance and cannot be parallelised. The Transformer holds a plenary meeting instead — every word can see every word, relevance is decided on the spot by attention, and word order is marked by giving each position a "seat number" (a positional encoding).
The Transformer data flow: embeddings plus positional encoding enter stacked attention–feed-forward blocks and finally yield a probability distribution over the next token
Ability profiles of three architectural variants (qualitative, 0–5): there is no all-rounder — choose by task focus
- Encoder-only (BERT)
- Decoder-only (GPT)
- Encoder-Decoder (T5)
طريقة العمل
- 01
Embeddings plus positional encoding
Tokens become embedding vectors, to which positional information is added. The original paper used sine and cosine waves at different frequencies to encode absolute position; modern models largely switch to Rotary Position Embedding (RoPE), which encodes position as a relative rotation angle between vectors and extrapolates better to sequences longer than those seen in training.
- 02
Multi-head self-attention
Each position attends to all positions at once, and multiple heads capture different relation patterns in parallel. An encoder permits bidirectional attention, while a decoder adds a causal mask so a position may only see itself and the past.
- 03
Feed-forward network
Attention moves information between positions; the feed-forward network then applies a nonlinear transform at each position independently, typically widening (say fourfold) before projecting back. In large models the great majority of parameters live here.
- 04
Residuals and layer norm, then stack
Each sublayer is wrapped in a residual (x + sublayer(x)) and a layer norm so gradients flow cleanly through dozens or hundreds of layers. Modern designs place the norm before the sublayer (Pre-LN) for stability; stacking these layers dozens deep yields the body of a large model.
Two mainstream variants and their division of labour: encoder-only excels at understanding, decoder-only at generation
The growth of context windows: better positional encoding schemes expanded usable context by three orders of magnitude within a few years
مجالات الاستخدام
- Decoder-only (GPT family): autoregressive text generation, dialogue and code completion
- Encoder-only (BERT family): text classification, extractive QA and semantic understanding
- Encoder-Decoder (T5, the original Transformer): sequence-to-sequence tasks such as translation and summarisation
- Cross-modal extension: ViT for image patches, Whisper for audio — all sharing one backbone
مفاهيم خاطئة شائعة
- Remove positional encoding and self-attention becomes permutation-invariant: the model degrades into a bag of words with no sense of order. Position must be injected explicitly.
- No single variant wins everywhere. Encoder-only barely generates, decoder-only is weaker on tasks needing bidirectional understanding, and encoder-decoder is structurally heavier and costlier to serve — the choice follows the task, not novelty.
- The same architecture does not imply the same ability. Differences from data, scale and alignment are usually far larger than those between architectural variants.
مصطلحات أساسية
- Positional encoding
- An explicit order signal, sinusoidal or RoPE
- RoPE
- Rotary Position Embedding: relative position with better extrapolation
- Residual connection
- Adding the input past a sublayer to ease vanishing gradients in depth
- Pre-LN
- Placing layer norm before each sublayer for stability