Перейти к содержимому
Атлас ИИ
04 Обработка языка и большие языковые моделиСреднийСтатья 4 в этой области

Архитектура Transformer

Замена эстафеты по одному слову залом, где все говорят сразу, — и дальние зависимости оказываются в одном шаге

Полный текст статьи представлен на английском; заголовок и аннотация локализованы.

ОПРЕДЕЛЕНИЕ

The Transformer is a sequence-modelling architecture introduced in 2017 that relies entirely on attention and abandons recurrence. It stacks structurally identical layers, each holding multi-head attention and a feed-forward network, wrapped with residual connections and layer normalisation. Order is no longer implicit in time steps but injected explicitly through positional encodings.

Интуиция

A recurrent network plays telephone: information travels one word at a time, degrades over distance and cannot be parallelised. The Transformer holds a plenary meeting instead — every word can see every word, relevance is decided on the spot by attention, and word order is marked by giving each position a "seat number" (a positional encoding).

Рис. 1

The Transformer data flow: embeddings plus positional encoding enter stacked attention–feed-forward blocks and finally yield a probability distribution over the next token

Рис. 2

Ability profiles of three architectural variants (qualitative, 0–5): there is no all-rounder — choose by task focus

  • Encoder-only (BERT)
  • Decoder-only (GPT)
  • Encoder-Decoder (T5)

Как это работает

  1. 01

    Embeddings plus positional encoding

    Tokens become embedding vectors, to which positional information is added. The original paper used sine and cosine waves at different frequencies to encode absolute position; modern models largely switch to Rotary Position Embedding (RoPE), which encodes position as a relative rotation angle between vectors and extrapolates better to sequences longer than those seen in training.

  2. 02

    Multi-head self-attention

    Each position attends to all positions at once, and multiple heads capture different relation patterns in parallel. An encoder permits bidirectional attention, while a decoder adds a causal mask so a position may only see itself and the past.

  3. 03

    Feed-forward network

    Attention moves information between positions; the feed-forward network then applies a nonlinear transform at each position independently, typically widening (say fourfold) before projecting back. In large models the great majority of parameters live here.

  4. 04

    Residuals and layer norm, then stack

    Each sublayer is wrapped in a residual (x + sublayer(x)) and a layer norm so gradients flow cleanly through dozens or hundreds of layers. Modern designs place the norm before the sublayer (Pre-LN) for stability; stacking these layers dozens deep yields the body of a large model.

Рис. 3

Two mainstream variants and their division of labour: encoder-only excels at understanding, decoder-only at generation

Рис. 4

The growth of context windows: better positional encoding schemes expanded usable context by three orders of magnitude within a few years

Области применения

  • Decoder-only (GPT family): autoregressive text generation, dialogue and code completion
  • Encoder-only (BERT family): text classification, extractive QA and semantic understanding
  • Encoder-Decoder (T5, the original Transformer): sequence-to-sequence tasks such as translation and summarisation
  • Cross-modal extension: ViT for image patches, Whisper for audio — all sharing one backbone

Частые заблуждения

  • Remove positional encoding and self-attention becomes permutation-invariant: the model degrades into a bag of words with no sense of order. Position must be injected explicitly.
  • No single variant wins everywhere. Encoder-only barely generates, decoder-only is weaker on tasks needing bidirectional understanding, and encoder-decoder is structurally heavier and costlier to serve — the choice follows the task, not novelty.
  • The same architecture does not imply the same ability. Differences from data, scale and alignment are usually far larger than those between architectural variants.

Ключевые термины

Positional encoding
An explicit order signal, sinusoidal or RoPE
RoPE
Rotary Position Embedding: relative position with better extrapolation
Residual connection
Adding the input past a sublayer to ease vanishing gradients in depth
Pre-LN
Placing layer norm before each sublayer for stability

Дополнительная литература