Transformer आर्किटेक्चर
शब्द-दर-शब्द relay की जगह वह कक्ष जहाँ सब एक साथ बोलते हैं, जिससे दूर की निर्भरताएँ एक कदम पर आ जाती हैं
यह पृष्ठ अंग्रेज़ी में प्रस्तुत है; शीर्षक और सारांश का स्थानीयकरण किया गया है।
परिभाषा
The Transformer is a sequence-modelling architecture introduced in 2017 that relies entirely on attention and abandons recurrence. It stacks structurally identical layers, each holding multi-head attention and a feed-forward network, wrapped with residual connections and layer normalisation. Order is no longer implicit in time steps but injected explicitly through positional encodings.
सहज समझ
A recurrent network plays telephone: information travels one word at a time, degrades over distance and cannot be parallelised. The Transformer holds a plenary meeting instead — every word can see every word, relevance is decided on the spot by attention, and word order is marked by giving each position a "seat number" (a positional encoding).
The Transformer data flow: embeddings plus positional encoding enter stacked attention–feed-forward blocks and finally yield a probability distribution over the next token
Ability profiles of three architectural variants (qualitative, 0–5): there is no all-rounder — choose by task focus
- Encoder-only (BERT)
- Decoder-only (GPT)
- Encoder-Decoder (T5)
कार्यप्रणाली
- 01
Embeddings plus positional encoding
Tokens become embedding vectors, to which positional information is added. The original paper used sine and cosine waves at different frequencies to encode absolute position; modern models largely switch to Rotary Position Embedding (RoPE), which encodes position as a relative rotation angle between vectors and extrapolates better to sequences longer than those seen in training.
- 02
Multi-head self-attention
Each position attends to all positions at once, and multiple heads capture different relation patterns in parallel. An encoder permits bidirectional attention, while a decoder adds a causal mask so a position may only see itself and the past.
- 03
Feed-forward network
Attention moves information between positions; the feed-forward network then applies a nonlinear transform at each position independently, typically widening (say fourfold) before projecting back. In large models the great majority of parameters live here.
- 04
Residuals and layer norm, then stack
Each sublayer is wrapped in a residual (x + sublayer(x)) and a layer norm so gradients flow cleanly through dozens or hundreds of layers. Modern designs place the norm before the sublayer (Pre-LN) for stability; stacking these layers dozens deep yields the body of a large model.
Two mainstream variants and their division of labour: encoder-only excels at understanding, decoder-only at generation
The growth of context windows: better positional encoding schemes expanded usable context by three orders of magnitude within a few years
उपयोग के क्षेत्र
- Decoder-only (GPT family): autoregressive text generation, dialogue and code completion
- Encoder-only (BERT family): text classification, extractive QA and semantic understanding
- Encoder-Decoder (T5, the original Transformer): sequence-to-sequence tasks such as translation and summarisation
- Cross-modal extension: ViT for image patches, Whisper for audio — all sharing one backbone
सामान्य भ्रांतियाँ
- Remove positional encoding and self-attention becomes permutation-invariant: the model degrades into a bag of words with no sense of order. Position must be injected explicitly.
- No single variant wins everywhere. Encoder-only barely generates, decoder-only is weaker on tasks needing bidirectional understanding, and encoder-decoder is structurally heavier and costlier to serve — the choice follows the task, not novelty.
- The same architecture does not imply the same ability. Differences from data, scale and alignment are usually far larger than those between architectural variants.
मुख्य शब्द
- Positional encoding
- An explicit order signal, sinusoidal or RoPE
- RoPE
- Rotary Position Embedding: relative position with better extrapolation
- Residual connection
- Adding the input past a sublayer to ease vanishing gradients in depth
- Pre-LN
- Placing layer norm before each sublayer for stability