Textgenerierung
Setzt einen Text Wort für Wort fort
Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.
WAS DIESE FÄHIGKEIT BEDEUTET
Given a passage as context, the model writes what comes next. Both input and output are text and no particular transformation is targeted: it can continue a story, finish an explanation or draft an email. It differs from conversation in that no instruction format is required, and from summarisation or translation in that neither compression nor a language switch is the goal — both ends stay in the same modality.
Wie sie technisch umgesetzt wird
The dominant route is an autoregressive Transformer language model: text is split into tokens, the model predicts the probability of the next token position by position, and training maximises likelihood over the corpus. GPT, Llama and Qwen all follow it, differing in scale, data and post-training recipe. At inference, temperature, top-k and top-p shape diversity, while long outputs reuse cached keys and values to cut cost.
Repräsentative Produkte
27GPT-4o
2024Ein von Grund auf multimodales Allzweckmodell: Text, Bild und Audio über einen Zugang
Llama
2023Die Modellfamilie, die offene Gewichte zum Mainstream machte
Qwen
2023Eine Open-Weights-Familie über viele Größen hinweg, mit multimodalen Versionen
DeepSeek-V3
2024Ein Open-Weights-MoE mit 671B Parametern, das pro Token nur 37B aktiviert
Mistral Large
2024Das kommerzielle Flaggschiffmodell eines europäischen Open-Weights-Labors
Gemini
2023Ein von Grund auf multimodales Allzweckmodell für sehr langen Kontext
Claude
2023Ein allgemeines Dialogmodell, bekannt für langen Kontext und Sicherheitsausrichtung
MiniMax-M
2025Ein Open-Weights-Schlussfolgerungsmodell mit hybridem Attention und einer Million Token Kontext
DeepSeek-R1
2025Ein mit RL auf Gedankenketten trainiertes Schlussfolgerungsmodell, dessen Gewichte unter MIT-Lizenz offen sind
Apple Intelligence
2024System-KI, die Arbeit zwischen Gerät und privater Cloud aufteilt
Command R
2024Ein kommerzielles Modell für Retrieval-Augmentierung und Werkzeugnutzung
Jamba
2024Ein Open-Weights-Modell, das ein Zustandsraummodell mit einem Transformer mischt
DBRX
2024Offenes MoE-Sprachmodell von Databricks
Gemini
2024Ein Chat-Eingang, der Suche, Office-Apps und ein multimodales Modell bündelt
Grok
2023Ein Dialogmodell, an die Daten einer Social-Plattform gekoppelt
Yi
2023Ein zweisprachiges (Chinesisch–Englisch) Open-Weights-Modell mit sehr langem Kontext
Kimi
2023Ein chinesischer Chat-Assistent, bekannt für langen Kontext
Hunyuan
2023Tencent Generalmodell-Familie mit Open-Weights-Versionen
Microsoft Copilot
2023Konversations-KI, eingebettet in Betriebssystem und Office-Apps
Doubao
2023ByteDances allgemeines Dialogmodell und App
Phi
2023Kleine, effiziente Modelle aus kuratierten Daten
Step
2023Eine Allzweck-Modellfamilie für Multimodalität und On-Device-Betrieb
Baichuan
2023Ein Open-Weights-Allzweckmodell für den chinesischen Sprachraum
GLM
2023Ein chinesisches Allzweckmodell, das mit autoregressivem Lückenfüllen begann
ChatGPT
2022Das Chatfenster, das ein großes Sprachmodell für alle zugänglich machte
Character.AI
2022Rollenspiel-Chat mit Figuren, die du selbst festlegst und lange begleitest
ERNIE
2019Ein chinesisches Modell, das mit wissensverstärktem Vortraining begann – frühe prägende Version
Beteiligte Organisationen
Typische Verwendungen
- First drafts and rewriting
- Drafting email, documents and copy
- Code comments and API docs
- Test data and synthetic corpora
Wie sie bewertet wird
- Perplexity
- Predictive uncertainty on held-out text; lower is better
- Human-preference Elo
- Preference ranking from pairwise comparison
- ROUGE/BLEU on constrained tasks
- Reference overlap when an answer key exists
Grenzen und schwierige Punkte
- Fluency is not truth: the model states wrong facts with confidence — hallucination
- Over long outputs, earlier setup drifts and the text contradicts itself
- Poor sampling settings can trap the model in repetitive or degenerate loops
Konzepte dahinter
Transformer-Architektur
Statt Wort-für-Wort-Stafette ein Raum, in dem alle zugleich sprechen — und weite Abhängigkeiten sind nur einen Schritt entfernt
Vor- und Feintraining
Erst Sprache aus riesigem ungelabelten Text lernen, dann mit wenig Daten spezialisieren – das dateneffizienteste Paradigma der modernen KI
Prompt-Engineering und Alignment
Ein Modell hilfreich, ehrlich und harmlos zu machen ist schwieriger, als es einfach größer zu machen