本文へスキップ
AI図鑑
04 自然言語処理と大規模言語モデル中級この領域の第 4 項目

Transformer アーキテクチャ

「一語ずつの伝言」を「全員が同時に話す会議」に置き換え、長距離依存を一跳で届かせる

本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。

定義

The Transformer is a sequence-modelling architecture introduced in 2017 that relies entirely on attention and abandons recurrence. It stacks structurally identical layers, each holding multi-head attention and a feed-forward network, wrapped with residual connections and layer normalisation. Order is no longer implicit in time steps but injected explicitly through positional encodings.

直観的な理解

A recurrent network plays telephone: information travels one word at a time, degrades over distance and cannot be parallelised. The Transformer holds a plenary meeting instead — every word can see every word, relevance is decided on the spot by attention, and word order is marked by giving each position a "seat number" (a positional encoding).

図 1

The Transformer data flow: embeddings plus positional encoding enter stacked attention–feed-forward blocks and finally yield a probability distribution over the next token

図 2

Ability profiles of three architectural variants (qualitative, 0–5): there is no all-rounder — choose by task focus

  • Encoder-only (BERT)
  • Decoder-only (GPT)
  • Encoder-Decoder (T5)

仕組み

  1. 01

    Embeddings plus positional encoding

    Tokens become embedding vectors, to which positional information is added. The original paper used sine and cosine waves at different frequencies to encode absolute position; modern models largely switch to Rotary Position Embedding (RoPE), which encodes position as a relative rotation angle between vectors and extrapolates better to sequences longer than those seen in training.

  2. 02

    Multi-head self-attention

    Each position attends to all positions at once, and multiple heads capture different relation patterns in parallel. An encoder permits bidirectional attention, while a decoder adds a causal mask so a position may only see itself and the past.

  3. 03

    Feed-forward network

    Attention moves information between positions; the feed-forward network then applies a nonlinear transform at each position independently, typically widening (say fourfold) before projecting back. In large models the great majority of parameters live here.

  4. 04

    Residuals and layer norm, then stack

    Each sublayer is wrapped in a residual (x + sublayer(x)) and a layer norm so gradients flow cleanly through dozens or hundreds of layers. Modern designs place the norm before the sublayer (Pre-LN) for stability; stacking these layers dozens deep yields the body of a large model.

図 3

Two mainstream variants and their division of labour: encoder-only excels at understanding, decoder-only at generation

図 4

The growth of context windows: better positional encoding schemes expanded usable context by three orders of magnitude within a few years

応用場面

  • Decoder-only (GPT family): autoregressive text generation, dialogue and code completion
  • Encoder-only (BERT family): text classification, extractive QA and semantic understanding
  • Encoder-Decoder (T5, the original Transformer): sequence-to-sequence tasks such as translation and summarisation
  • Cross-modal extension: ViT for image patches, Whisper for audio — all sharing one backbone

よくある誤解

  • Remove positional encoding and self-attention becomes permutation-invariant: the model degrades into a bag of words with no sense of order. Position must be injected explicitly.
  • No single variant wins everywhere. Encoder-only barely generates, decoder-only is weaker on tasks needing bidirectional understanding, and encoder-decoder is structurally heavier and costlier to serve — the choice follows the task, not novelty.
  • The same architecture does not imply the same ability. Differences from data, scale and alignment are usually far larger than those between architectural variants.

重要用語

Positional encoding
An explicit order signal, sinusoidal or RoPE
RoPE
Rotary Position Embedding: relative position with better extrapolation
Residual connection
Adding the input past a sublayer to ease vanishing gradients in depth
Pre-LN
Placing layer norm before each sublayer for stability

参考文献