Ingeniería de prompts y alineación
Hacer que un modelo sea útil, honesto e inofensivo es más difícil que simplemente agrandarlo
El texto completo se presenta en inglés; el título y el resumen están traducidos.
DEFINICIÓN
Prompt engineering is the practice of designing the input context to steer an already-trained model towards a task; alignment is the process of steering a model’s outputs towards human preferences and safety requirements. A typical pipeline starts with instruction tuning (SFT), then pushes the model towards preferred behaviour with reinforcement learning from human feedback (RLHF) or direct preference optimisation (DPO).
Intuición
A pretrained model is like a highly knowledgeable but undirected intern: it answers whatever you ask, but may drift, ramble or miss the format. Instruction tuning hands it a work manual; alignment uses repeated feedback to teach it what a satisfying answer looks like. Prompt engineering is the specific brief you give on the spot — it can guide within existing ability but cannot conjure new ability out of nothing.
The RLHF alignment loop: generate responses, rank them, train a reward model, update the policy — repeated until outputs settle into the preferred behaviour
MMLU (57-subject knowledge benchmark) 5-shot accuracy: from base models to aligned assistants, knowledge and alignment improve together
Cómo funciona
- 01
In-context learning and few-shot examples
Placing a few input–output examples in the prompt lets the model imitate the pattern without updating any parameters — in-context learning. It replaces "change the weights" with "change the input", letting one model switch tasks from a passage of text, at the cost of examples consuming precious context length.
- 02
Chain-of-thought: making reasoning explicit
Asking the model to think step by step before answering markedly improves accuracy on multi-step arithmetic and logic. Each step’s output becomes the next step’s input, effectively trading more computation steps for steadier reasoning.
- 03
Instruction tuning (SFT)
Supervised fine-tuning on a large set of high-quality instruction–response pairs teaches the model to read intent, follow formats and decline out-of-scope requests. This step turns a text completer into an assistant that listens, and it is the base for later alignment.
- 04
Preference alignment: RLHF and DPO
RLHF first trains a reward model on human rankings of several responses, then uses reinforcement learning (usually PPO) to maximise that reward while a KL penalty keeps the policy from drifting too far from the base model. DPO skips the explicit reward model and writes "preferred responses should be more probable than rejected ones" directly into a differentiable, classification-style loss, making the pipeline simpler and steadier.
Dónde se usa
- General assistants and chat products: shaping a base model into an instruction-following assistant
- Reasoning boost: chain-of-thought and self-consistency for maths, code and logic
- Safety guardrails: training the model to refuse harmful requests and reduce hallucination and jailbreak risk
- Low-cost customisation: adapting to support, writing and other scenarios via system prompts and few-shot examples
Errores comunes
- A prompt is not a spell. It can only guide within existing ability; where the model lacks a capability, no cleverness in the prompt will invent it.
- RLHF invites reward hacking: the model learns to please the reward model rather than genuinely improve, and the reward model’s own annotator bias gets amplified into the model’s behaviour.
- Alignment carries a tax: the price paid for safety and compliance may erode some raw capability. And how to trade off helpfulness, honesty and harmlessness when they conflict is itself a value judgement that a bigger model alone cannot settle.
Términos clave
- In-context learning
- Solving a task from prompt examples without updating parameters
- Chain-of-thought (CoT)
- Making the model write out intermediate reasoning steps
- SFT
- Supervised fine-tuning on instruction–response pairs
- DPO
- Direct preference optimisation without an explicit reward model