Text Generation
Continue a passage, one word at a time
WHAT THIS CAPABILITY MEANS
Given a passage as context, the model writes what comes next. Both input and output are text and no particular transformation is targeted: it can continue a story, finish an explanation or draft an email. It differs from conversation in that no instruction format is required, and from summarisation or translation in that neither compression nor a language switch is the goal — both ends stay in the same modality.
How it is done
The dominant route is an autoregressive Transformer language model: text is split into tokens, the model predicts the probability of the next token position by position, and training maximises likelihood over the corpus. GPT, Llama and Qwen all follow it, differing in scale, data and post-training recipe. At inference, temperature, top-k and top-p shape diversity, while long outputs reuse cached keys and values to cut cost.
Representative products
27GPT-4o
2024A natively multimodal general model, with text, image and audio through one door
Llama
2023The model family that made the open-weight route mainstream
Qwen
2023An open-weight family spanning many sizes, with multimodal versions
DeepSeek-V3
2024An open-weight MoE with 671B parameters, activating 37B per token
Mistral Large
2024The flagship commercial model from a European open-weights lab
Gemini
2023A natively multimodal general model built for very long context
Claude
2023A general chat model known for long context and safety alignment
MiniMax-M
2025An open-weight reasoning model with hybrid attention and a million-token context
DeepSeek-R1
2025A reasoning model trained with RL on chains of thought, its weights open under the MIT licence
Apple Intelligence
2024System-level AI that splits work between on-device and private cloud
Command R
2024A commercial model built for retrieval augmentation and tool use
Jamba
2024An open-weight model that mixes a state-space model with a Transformer
DBRX
2024Databricks’ open-weights MoE language model
Gemini
2024One chat entry point that gathers search, office apps and a multimodal model
Grok
2023A chat model tied to a social platform’s live data
Yi
2023A bilingual Chinese–English open-weight model with very-long-context versions
Kimi
2023A Chinese chat assistant known for long-context handling
Hunyuan
2023Tencent’s general model family, with open-weight versions
Microsoft Copilot
2023Conversational AI woven into the operating system and office apps
Doubao
2023ByteDance’s general chat model and application
Phi
2023Small, efficient models built from curated data
Step
2023A general model family aimed at multimodality and on-device use
Baichuan
2023An open-weight general model aimed at Chinese-language use
GLM
2023A Chinese general model that began with autoregressive blank-infilling pretraining
ChatGPT
2022The chat window that put a large language model in everyone’s hands
Character.AI
2022Role-play chat with characters you define and talk to over time
ERNIE
2019A Chinese model that began with knowledge-enhanced pretraining, an early landmark version
Organizations involved
Typical uses
- First drafts and rewriting
- Drafting email, documents and copy
- Code comments and API docs
- Test data and synthetic corpora
How it is evaluated
- Perplexity
- Predictive uncertainty on held-out text; lower is better
- Human-preference Elo
- Preference ranking from pairwise comparison
- ROUGE/BLEU on constrained tasks
- Reference overlap when an answer key exists
Limits and hard parts
- Fluency is not truth: the model states wrong facts with confidence — hallucination
- Over long outputs, earlier setup drifts and the text contradicts itself
- Poor sampling settings can trap the model in repetitive or degenerate loops
Concepts behind it
Transformer Architecture
Replacing word-by-word relay with a room where everyone speaks at once, so long-range dependencies are one hop away
Pretraining & Fine-tuning
Learn language first from vast unlabelled text, then specialise with little data — the most data-efficient paradigm in modern AI
Prompting & Alignment
Making a model helpful, honest and harmless is harder than simply making it bigger