Saltar al contenido
Atlas de IA
04 PLN y grandes modelos de lenguajeIntermedioEntrada 4 de este dominio

Arquitectura Transformer

Sustituir el relé palabra a palabra por una sala donde todos hablan a la vez, para que las dependencias lejanas estén a un salto

El texto completo se presenta en inglés; el título y el resumen están traducidos.

DEFINICIÓN

The Transformer is a sequence-modelling architecture introduced in 2017 that relies entirely on attention and abandons recurrence. It stacks structurally identical layers, each holding multi-head attention and a feed-forward network, wrapped with residual connections and layer normalisation. Order is no longer implicit in time steps but injected explicitly through positional encodings.

Intuición

A recurrent network plays telephone: information travels one word at a time, degrades over distance and cannot be parallelised. The Transformer holds a plenary meeting instead — every word can see every word, relevance is decided on the spot by attention, and word order is marked by giving each position a "seat number" (a positional encoding).

Fig. 1

The Transformer data flow: embeddings plus positional encoding enter stacked attention–feed-forward blocks and finally yield a probability distribution over the next token

Fig. 2

Ability profiles of three architectural variants (qualitative, 0–5): there is no all-rounder — choose by task focus

  • Encoder-only (BERT)
  • Decoder-only (GPT)
  • Encoder-Decoder (T5)

Cómo funciona

  1. 01

    Embeddings plus positional encoding

    Tokens become embedding vectors, to which positional information is added. The original paper used sine and cosine waves at different frequencies to encode absolute position; modern models largely switch to Rotary Position Embedding (RoPE), which encodes position as a relative rotation angle between vectors and extrapolates better to sequences longer than those seen in training.

  2. 02

    Multi-head self-attention

    Each position attends to all positions at once, and multiple heads capture different relation patterns in parallel. An encoder permits bidirectional attention, while a decoder adds a causal mask so a position may only see itself and the past.

  3. 03

    Feed-forward network

    Attention moves information between positions; the feed-forward network then applies a nonlinear transform at each position independently, typically widening (say fourfold) before projecting back. In large models the great majority of parameters live here.

  4. 04

    Residuals and layer norm, then stack

    Each sublayer is wrapped in a residual (x + sublayer(x)) and a layer norm so gradients flow cleanly through dozens or hundreds of layers. Modern designs place the norm before the sublayer (Pre-LN) for stability; stacking these layers dozens deep yields the body of a large model.

Fig. 3

Two mainstream variants and their division of labour: encoder-only excels at understanding, decoder-only at generation

Fig. 4

The growth of context windows: better positional encoding schemes expanded usable context by three orders of magnitude within a few years

Dónde se usa

  • Decoder-only (GPT family): autoregressive text generation, dialogue and code completion
  • Encoder-only (BERT family): text classification, extractive QA and semantic understanding
  • Encoder-Decoder (T5, the original Transformer): sequence-to-sequence tasks such as translation and summarisation
  • Cross-modal extension: ViT for image patches, Whisper for audio — all sharing one backbone

Errores comunes

  • Remove positional encoding and self-attention becomes permutation-invariant: the model degrades into a bag of words with no sense of order. Position must be injected explicitly.
  • No single variant wins everywhere. Encoder-only barely generates, decoder-only is weaker on tasks needing bidirectional understanding, and encoder-decoder is structurally heavier and costlier to serve — the choice follows the task, not novelty.
  • The same architecture does not imply the same ability. Differences from data, scale and alignment are usually far larger than those between architectural variants.

Términos clave

Positional encoding
An explicit order signal, sinusoidal or RoPE
RoPE
Rotary Position Embedding: relative position with better extrapolation
Residual connection
Adding the input past a sublayer to ease vanishing gradients in depth
Pre-LN
Placing layer norm before each sublayer for stability

Lecturas complementarias