Codegenerierung
Aus einer Beschreibung lauffähigen Code schreiben
Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.
WAS DIESE FÄHIGKEIT BEDEUTET
Takes a description — write a function that reads a CSV into a list of dicts and skips blank lines — and outputs the corresponding source. Unlike code completion it targets a new feature or a whole piece of logic, usually written from scratch; unlike reasoning the deliverable is compilable, runnable code rather than a textual conclusion.
Wie sie technisch umgesetzt wird
The base is a language model pre-trained on large code corpora; strict syntax makes code easier to verify by execution than prose, so unit-test pass signals work well as reward for reinforcement learning or rejection-sampling fine-tuning. Training and evaluation use problems with test cases such as HumanEval and MBPP, and generation can sample several candidates and discard the ones that fail.
Repräsentative Produkte
11GPT-4o
2024Ein von Grund auf multimodales Allzweckmodell: Text, Bild und Audio über einen Zugang
Claude
2023Ein allgemeines Dialogmodell, bekannt für langen Kontext und Sicherheitsausrichtung
DeepSeek-V3
2024Ein Open-Weights-MoE mit 671B Parametern, das pro Token nur 37B aktiviert
Qwen
2023Eine Open-Weights-Familie über viele Größen hinweg, mit multimodalen Versionen
Gemini
2023Ein von Grund auf multimodales Allzweckmodell für sehr langen Kontext
GitHub Copilot
2021Ergänzt und überarbeitet Code im Editor aus dem Kontext
o3
2025Es denkt lange nach, bevor es antwortet: Rechenzeit bei der Inferenz gegen mehr Treffsicherheit
Claude Code
2025Ein Coding-Agent, der mehrstufige Aufgaben im Terminal erledigt
DeepSeek-R1
2025Ein mit RL auf Gedankenketten trainiertes Schlussfolgerungsmodell, dessen Gewichte unter MIT-Lizenz offen sind
Devin
2024Ein autonomer Coding-Agent mit eigener Shell, Editor und Browser
Cursor
2023Ein Desktop-Code-Editor, der das ganze Repository als Kontext nutzt
Beteiligte Organisationen
Typische Verwendungen
- New features and utility scripts
- Data cleaning and transformation scripts
- Test cases and project scaffolding
- Porting code across languages and frameworks
Wie sie bewertet wird
- pass@k
- Share of problems with at least one sample passing all tests out of k
- Benchmark pass rate
- Pass rate on fixed sets such as HumanEval
- Compile and lint pass rate
- Whether generated code passes the compiler and type checks as-is
Grenzen und schwierige Punkte
- It invents libraries or methods, producing plausible interfaces that do not exist
- Code can pass unit tests yet leave holes in edge cases, error handling and security
- Multi-file changes are inconsistent, with call sites disagreeing with definitions
Konzepte dahinter
Transformer-Architektur
Statt Wort-für-Wort-Stafette ein Raum, in dem alle zugleich sprechen — und weite Abhängigkeiten sind nur einen Schritt entfernt
Vor- und Feintraining
Erst Sprache aus riesigem ungelabelten Text lernen, dann mit wenig Daten spezialisieren – das dateneffizienteste Paradigma der modernen KI
Prompt-Engineering und Alignment
Ein Modell hilfreich, ehrlich und harmlos zu machen ist schwieriger, als es einfach größer zu machen