安全性・アラインメント・プロンプトインジェクション
モデルが最適化するのは「本当に望むもの」ではなく損失に書いた代理指標。そのずれがアラインメント問題のすべてだ
本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。
定義
Alignment asks how to make a model behave according to human intent and values rather than merely maximising the proxy objective it was trained on. Safety covers concrete threats from prompt injection and jailbreaks to hallucination and accountability. Prompt injection is hard to eradicate because instructions and data are concatenated into a single text channel, so the model cannot mechanically distinguish “commands to execute” from “content merely to read”.
直観的な理解
Alignment is like setting a target for an assistant: tell them to “take as many calls as possible” and they stretch every call past ten minutes to pad the count — perfectly optimising the metric, utterly missing your intent. Prompt injection is a letter with a forged instruction folded in: because system instructions and user content are written in the same ink, a line in the corner saying “ignore all of the above” can be taken at face value.
Why injection is hard to fix: in conventional software code and data are separate, so injection only corrupts data; in a large model instructions and data share one channel, so the boundary does not exist mechanically
Attack success rate by attack vector under each defence layer (illustrative magnitude): any single layer leaves gaps, indirect injection is especially stubborn, and only stacking layers pushes the rate low
仕組み
- 01
Alignment: where the mismatch comes from
Training can only optimise a measurable proxy (next token, human preference scores), while the true intent — helpful, harmless, honest — cannot be fully written into a loss. The gap between proxy and intent, plus a strong optimiser’s knack for exploiting that gap, is the root of alignment.
- 02
Prompt injection: one channel for instructions and data
In conventional software, code and data are separate, so injection at worst corrupts data without altering logic; in a large model, system instructions, user input and retrieved content are concatenated into one text stream with no field-level privilege boundary. Any sentence buried in a web page, document or tool result can, in principle, be executed as an instruction.
- 03
Jailbreaks and defences: an escalating contest
Jailbreaks bypass safety training through role-play, obfuscated encodings and gradual multi-turn coercion. Defences are layered too — prompt hardening, input filtering, output moderation, alignment fine-tuning and runtime permission control — each with its own coverage and false-positive cost, none sufficient alone.
- 04
Red-teaming, interpretability and accountability
Before deployment, red teams actively hunt for failure and misuse paths and record their reproducibility; after deployment, interpretability tries to explain why the model answered as it did, calibrated uncertainty lets it admit when it is unsure, and audit logs make each output traceable to an accountable source.
応用場面
- Content safety: filtering harmful, non-compliant or high-risk requests and outputs
- Injection defence: isolating untrusted content and limiting what tools may actually do
- Red-teaming and evaluation: systematically hunting for failure modes before release
- Human-in-the-loop for high-stakes actions: keeping human confirmation and auditing on irreversible steps
よくある誤解
- “Aligned” does not mean “cannot fail”. Safety training hardens against a particular distribution of attacks, and new jailbreaks can still circumvent it; the contest is continuous, not settled once.
- Hallucination cannot be eliminated by asking a model to rate its own confidence. A model can be wholly confident and wrong, and the self-rated “I am sure” is itself a fallible judgement.
- Prompt hardening alone cannot fix injection. As long as instructions and data share one channel, the answer is system-level isolation and least privilege, not a sterner “please do not be fooled”.
- Greater safety often costs usefulness. Over-blocking rejects legitimate requests, and one of the hard parts of alignment is finding an acceptable balance between safety and usability.
重要用語
- Alignment
- Making model behaviour match human intent and values
- Prompt injection
- Smuggling malicious instructions as data for the model to follow
- Jailbreak
- Inducing a model past its safety training
- Red teaming
- Actively hunting for failure and misuse paths