본문으로 건너뛰기
AI 도감

프롬프트 엔지니어링과 정렬

모델을 유용하고 정직하며 무해하게 만드는 일은 단순히 더 크게 만드는 것보다 어렵다

04 자연어 처리와 대규모 언어 모델고급이 영역의 6번째 항목

이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.

정의

Prompt engineering is the practice of designing the input context to steer an already-trained model towards a task; alignment is the process of steering a model’s outputs towards human preferences and safety requirements. A typical pipeline starts with instruction tuning (SFT), then pushes the model towards preferred behaviour with reinforcement learning from human feedback (RLHF) or direct preference optimisation (DPO).

직관적 이해

A pretrained model is like a highly knowledgeable but undirected intern: it answers whatever you ask, but may drift, ramble or miss the format. Instruction tuning hands it a work manual; alignment uses repeated feedback to teach it what a satisfying answer looks like. Prompt engineering is the specific brief you give on the spot — it can guide within existing ability but cannot conjure new ability out of nothing.

그림 1

The RLHF alignment loop: generate responses, rank them, train a reward model, update the policy — repeated until outputs settle into the preferred behaviour

① Policy generates…② Humans rank them③ Train a reward m…④ PPO updates the …RLHF loop
그림 2

MMLU (57-subject knowledge benchmark) 5-shot accuracy: from base models to aligned assistants, knowledge and alignment improve together

작동 원리

  1. 01

    In-context learning and few-shot examples

    Placing a few input–output examples in the prompt lets the model imitate the pattern without updating any parameters — in-context learning. It replaces "change the weights" with "change the input", letting one model switch tasks from a passage of text, at the cost of examples consuming precious context length.

  2. 02

    Chain-of-thought: making reasoning explicit

    Asking the model to think step by step before answering markedly improves accuracy on multi-step arithmetic and logic. Each step’s output becomes the next step’s input, effectively trading more computation steps for steadier reasoning.

  3. 03

    Instruction tuning (SFT)

    Supervised fine-tuning on a large set of high-quality instruction–response pairs teaches the model to read intent, follow formats and decline out-of-scope requests. This step turns a text completer into an assistant that listens, and it is the base for later alignment.

  4. 04

    Preference alignment: RLHF and DPO

    RLHF first trains a reward model on human rankings of several responses, then uses reinforcement learning (usually PPO) to maximise that reward while a KL penalty keeps the policy from drifting too far from the base model. DPO skips the explicit reward model and writes "preferred responses should be more probable than rejected ones" directly into a differentiable, classification-style loss, making the pipeline simpler and steadier.

응용 분야

  • General assistants and chat products: shaping a base model into an instruction-following assistant
  • Reasoning boost: chain-of-thought and self-consistency for maths, code and logic
  • Safety guardrails: training the model to refuse harmful requests and reduce hallucination and jailbreak risk
  • Low-cost customisation: adapting to support, writing and other scenarios via system prompts and few-shot examples

흔한 오해

  • A prompt is not a spell. It can only guide within existing ability; where the model lacks a capability, no cleverness in the prompt will invent it.
  • RLHF invites reward hacking: the model learns to please the reward model rather than genuinely improve, and the reward model’s own annotator bias gets amplified into the model’s behaviour.
  • Alignment carries a tax: the price paid for safety and compliance may erode some raw capability. And how to trade off helpfulness, honesty and harmlessness when they conflict is itself a value judgement that a bigger model alone cannot settle.

핵심 용어

In-context learning
Solving a task from prompt examples without updating parameters
Chain-of-thought (CoT)
Making the model write out intermediate reasoning steps
SFT
Supervised fine-tuning on instruction–response pairs
DPO
Direct preference optimisation without an explicit reward model

참고문헌