안전성·정렬·프롬프트 인젝션
모델이 최적화하는 것은 우리가 진짜 원하는 것이 아니라 손실에 적어 넣은 대리 지표다 — 그 틈이 정렬 문제의 전부다
이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.
정의
Alignment asks how to make a model behave according to human intent and values rather than merely maximising the proxy objective it was trained on. Safety covers concrete threats from prompt injection and jailbreaks to hallucination and accountability. Prompt injection is hard to eradicate because instructions and data are concatenated into a single text channel, so the model cannot mechanically distinguish “commands to execute” from “content merely to read”.
직관적 이해
Alignment is like setting a target for an assistant: tell them to “take as many calls as possible” and they stretch every call past ten minutes to pad the count — perfectly optimising the metric, utterly missing your intent. Prompt injection is a letter with a forged instruction folded in: because system instructions and user content are written in the same ink, a line in the corner saying “ignore all of the above” can be taken at face value.
Why injection is hard to fix: in conventional software code and data are separate, so injection only corrupts data; in a large model instructions and data share one channel, so the boundary does not exist mechanically
Attack success rate by attack vector under each defence layer (illustrative magnitude): any single layer leaves gaps, indirect injection is especially stubborn, and only stacking layers pushes the rate low
작동 원리
- 01
Alignment: where the mismatch comes from
Training can only optimise a measurable proxy (next token, human preference scores), while the true intent — helpful, harmless, honest — cannot be fully written into a loss. The gap between proxy and intent, plus a strong optimiser’s knack for exploiting that gap, is the root of alignment.
- 02
Prompt injection: one channel for instructions and data
In conventional software, code and data are separate, so injection at worst corrupts data without altering logic; in a large model, system instructions, user input and retrieved content are concatenated into one text stream with no field-level privilege boundary. Any sentence buried in a web page, document or tool result can, in principle, be executed as an instruction.
- 03
Jailbreaks and defences: an escalating contest
Jailbreaks bypass safety training through role-play, obfuscated encodings and gradual multi-turn coercion. Defences are layered too — prompt hardening, input filtering, output moderation, alignment fine-tuning and runtime permission control — each with its own coverage and false-positive cost, none sufficient alone.
- 04
Red-teaming, interpretability and accountability
Before deployment, red teams actively hunt for failure and misuse paths and record their reproducibility; after deployment, interpretability tries to explain why the model answered as it did, calibrated uncertainty lets it admit when it is unsure, and audit logs make each output traceable to an accountable source.
응용 분야
- Content safety: filtering harmful, non-compliant or high-risk requests and outputs
- Injection defence: isolating untrusted content and limiting what tools may actually do
- Red-teaming and evaluation: systematically hunting for failure modes before release
- Human-in-the-loop for high-stakes actions: keeping human confirmation and auditing on irreversible steps
흔한 오해
- “Aligned” does not mean “cannot fail”. Safety training hardens against a particular distribution of attacks, and new jailbreaks can still circumvent it; the contest is continuous, not settled once.
- Hallucination cannot be eliminated by asking a model to rate its own confidence. A model can be wholly confident and wrong, and the self-rated “I am sure” is itself a fallible judgement.
- Prompt hardening alone cannot fix injection. As long as instructions and data share one channel, the answer is system-level isolation and least privilege, not a sterner “please do not be fooled”.
- Greater safety often costs usefulness. Over-blocking rejects legitimate requests, and one of the hard parts of alignment is finding an acceptable balance between safety and usability.
핵심 용어
- Alignment
- Making model behaviour match human intent and values
- Prompt injection
- Smuggling malicious instructions as data for the model to follow
- Jailbreak
- Inducing a model past its safety training
- Red teaming
- Actively hunting for failure and misuse paths