コード生成
自然言語の説明から動くコードを書く
本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。
この能力とは何か
Takes a description — write a function that reads a CSV into a list of dicts and skips blank lines — and outputs the corresponding source. Unlike code completion it targets a new feature or a whole piece of logic, usually written from scratch; unlike reasoning the deliverable is compilable, runnable code rather than a textual conclusion.
技術的にどう実現するか
The base is a language model pre-trained on large code corpora; strict syntax makes code easier to verify by execution than prose, so unit-test pass signals work well as reward for reinforcement learning or rejection-sampling fine-tuning. Training and evaluation use problems with test cases such as HumanEval and MBPP, and generation can sample several candidates and discard the ones that fail.
代表的な製品
11GPT-4o
2024ネイティブにマルチモーダルな汎用モデル。テキスト・画像・音声をひとつの入口で扱う
Claude
2023長い文脈と安全性の調整で知られる汎用対話モデル
DeepSeek-V3
2024総パラメータ 671B、1 トークンあたり 37B だけを活性化するオープンウェイト MoE
Qwen
2023多様な規模とマルチモーダル版を備えたオープンウェイトのモデル群
Gemini
2023ネイティブにマルチモーダルで、超長文脈を扱う汎用モデル
GitHub Copilot
2021エディタ上で文脈に沿ってコードを補完・改写する
o3
2025答える前に長い推論を重ね、推論時の計算で正答率を上げる
Claude Code
2025ターミナルから多段のコーディング作業をこなすエージェント
DeepSeek-R1
2025強化学習で推論の連鎖を訓練し、MIT ライセンスで重みを公開した推論モデル
Devin
2024シェル・エディタ・ブラウザを備えた自律型コーディングエージェント
Cursor
2023リポジトリ全体を文脈にするデスクトップコードエディタ
関連する組織
代表的な用途
- New features and utility scripts
- Data cleaning and transformation scripts
- Test cases and project scaffolding
- Porting code across languages and frameworks
どう評価するか
- pass@k
- Share of problems with at least one sample passing all tests out of k
- Benchmark pass rate
- Pass rate on fixed sets such as HumanEval
- Compile and lint pass rate
- Whether generated code passes the compiler and type checks as-is
限界と難しさ
- It invents libraries or methods, producing plausible interfaces that do not exist
- Code can pass unit tests yet leave holes in edge cases, error handling and security
- Multi-file changes are inconsistent, with call sites disagreeing with definitions