本文へスキップ
AI図鑑

コード生成

自然言語の説明から動くコードを書く

コードとソフトウェア初級 #33
入力テキストコード

本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。

この能力とは何か

Takes a description — write a function that reads a CSV into a list of dicts and skips blank lines — and outputs the corresponding source. Unlike code completion it targets a new feature or a whole piece of logic, usually written from scratch; unlike reasoning the deliverable is compilable, runnable code rather than a textual conclusion.

技術的にどう実現するか

The base is a language model pre-trained on large code corpora; strict syntax makes code easier to verify by execution than prose, so unit-test pass signals work well as reward for reinforcement learning or rejection-sampling fine-tuning. Training and evaluation use problems with test cases such as HumanEval and MBPP, and generation can sample several candidates and discard the ones that fail.

代表的な製品

11

GPT-4o

2024
OpenAI

ネイティブにマルチモーダルな汎用モデル。テキスト・画像・音声をひとつの入口で扱う

モデル クローズド
テキスト画像音声テキスト音声

Claude

2023
Anthropic

長い文脈と安全性の調整で知られる汎用対話モデル

モデル クローズド
テキスト画像テキスト

DeepSeek-V3

2024
DeepSeek

総パラメータ 671B、1 トークンあたり 37B だけを活性化するオープンウェイト MoE

モデル オープンウェイト
テキストテキスト

Qwen

2023
Alibaba (Qwen)

多様な規模とマルチモーダル版を備えたオープンウェイトのモデル群

モデル オープンウェイト
テキスト画像テキスト

Gemini

2023
Google DeepMind

ネイティブにマルチモーダルで、超長文脈を扱う汎用モデル

モデル クローズド
テキスト画像音声動画テキスト

GitHub Copilot

2021
Microsoft

エディタ上で文脈に沿ってコードを補完・改写する

ツール フリーミアム
テキストコードコード

o3

2025
OpenAI

答える前に長い推論を重ね、推論時の計算で正答率を上げる

モデル クローズド
テキスト画像テキスト

Claude Code

2025
Anthropic

ターミナルから多段のコーディング作業をこなすエージェント

ツール クローズド
テキストコードコード

DeepSeek-R1

2025
DeepSeek

強化学習で推論の連鎖を訓練し、MIT ライセンスで重みを公開した推論モデル

モデル オープンウェイト
テキストテキスト

Devin

2024
Cognition (Devin)

シェル・エディタ・ブラウザを備えた自律型コーディングエージェント

アプリ クローズド
テキストコードコード

Cursor

2023
Anysphere (Cursor)

リポジトリ全体を文脈にするデスクトップコードエディタ

アプリ フリーミアム
テキストコードコード

関連する組織

代表的な用途

  • New features and utility scripts
  • Data cleaning and transformation scripts
  • Test cases and project scaffolding
  • Porting code across languages and frameworks

どう評価するか

pass@k
Share of problems with at least one sample passing all tests out of k
Benchmark pass rate
Pass rate on fixed sets such as HumanEval
Compile and lint pass rate
Whether generated code passes the compiler and type checks as-is

限界と難しさ

  • It invents libraries or methods, producing plausible interfaces that do not exist
  • Code can pass unit tests yet leave holes in edge cases, error handling and security
  • Multi-file changes are inconsistent, with call sites disagreeing with definitions

背景にある概念