Code Generation
Write runnable code straight from a description
WHAT THIS CAPABILITY MEANS
Takes a description — write a function that reads a CSV into a list of dicts and skips blank lines — and outputs the corresponding source. Unlike code completion it targets a new feature or a whole piece of logic, usually written from scratch; unlike reasoning the deliverable is compilable, runnable code rather than a textual conclusion.
How it is done
The base is a language model pre-trained on large code corpora; strict syntax makes code easier to verify by execution than prose, so unit-test pass signals work well as reward for reinforcement learning or rejection-sampling fine-tuning. Training and evaluation use problems with test cases such as HumanEval and MBPP, and generation can sample several candidates and discard the ones that fail.
Representative products
11GPT-4o
2024A natively multimodal general model, with text, image and audio through one door
Claude
2023A general chat model known for long context and safety alignment
DeepSeek-V3
2024An open-weight MoE with 671B parameters, activating 37B per token
Qwen
2023An open-weight family spanning many sizes, with multimodal versions
Gemini
2023A natively multimodal general model built for very long context
GitHub Copilot
2021Completing and rewriting code in your editor from context
o3
2025It reasons at length before answering, trading inference-time compute for steadier accuracy
Claude Code
2025A coding agent that carries out multi-step tasks from the terminal
DeepSeek-R1
2025A reasoning model trained with RL on chains of thought, its weights open under the MIT licence
Devin
2024An autonomous coding agent with its own shell, editor and browser
Cursor
2023A desktop code editor that treats the whole repository as context
Organizations involved
Typical uses
- New features and utility scripts
- Data cleaning and transformation scripts
- Test cases and project scaffolding
- Porting code across languages and frameworks
How it is evaluated
- pass@k
- Share of problems with at least one sample passing all tests out of k
- Benchmark pass rate
- Pass rate on fixed sets such as HumanEval
- Compile and lint pass rate
- Whether generated code passes the compiler and type checks as-is
Limits and hard parts
- It invents libraries or methods, producing plausible interfaces that do not exist
- Code can pass unit tests yet leave holes in edge cases, error handling and security
- Multi-file changes are inconsistent, with call sites disagreeing with definitions
Concepts behind it
Transformer Architecture
Replacing word-by-word relay with a room where everyone speaks at once, so long-range dependencies are one hop away
Pretraining & Fine-tuning
Learn language first from vast unlabelled text, then specialise with little data — the most data-efficient paradigm in modern AI
Prompting & Alignment
Making a model helpful, honest and harmless is harder than simply making it bigger