Conversation & Instruction Following
Understand intent across turns and act on it
WHAT THIS CAPABILITY MEANS
Takes a multi-turn message sequence with roles (system, user, assistant) and outputs the next assistant turn. It is not merely continuation: the model must hold a persona across turns, remember earlier constraints, and obey format and tone instructions. It differs from plain text generation by its explicit role structure and the demand to follow instructions.
How it is done
The base is the same autoregressive language model; the difference lies in chat templates and post-training. System prompts and history are concatenated into a fixed sequence of special tokens, then instruction tuning teaches the model to answer rather than continue a web page. Preference data then shapes it through RLHF or a direct-preference method to favour helpful, harmless and honest behaviour. Tools and memory are injected into the prompt by external systems.
Representative products
30ChatGPT
2022The chat window that put a large language model in everyone’s hands
Claude
2023A general chat model known for long context and safety alignment
Gemini
2024One chat entry point that gathers search, office apps and a multimodal model
Kimi
2023A Chinese chat assistant known for long-context handling
Doubao
2023ByteDance’s general chat model and application
Character.AI
2022Role-play chat with characters you define and talk to over time
Microsoft Copilot
2023Conversational AI woven into the operating system and office apps
MiniMax-M
2025An open-weight reasoning model with hybrid attention and a million-token context
o3
2025It reasons at length before answering, trading inference-time compute for steadier accuracy
DeepSeek-R1
2025A reasoning model trained with RL on chains of thought, its weights open under the MIT licence
DeepSeek-V3
2024An open-weight MoE with 671B parameters, activating 37B per token
Apple Intelligence
2024System-level AI that splits work between on-device and private cloud
GPT-4o
2024A natively multimodal general model, with text, image and audio through one door
Command R
2024A commercial model built for retrieval augmentation and tool use
Jamba
2024An open-weight model that mixes a state-space model with a Transformer
DBRX
2024Databricks’ open-weights MoE language model
Mistral Large
2024The flagship commercial model from a European open-weights lab
Gemini
2023A natively multimodal general model built for very long context
Grok
2023A chat model tied to a social platform’s live data
Yi
2023A bilingual Chinese–English open-weight model with very-long-context versions
Hunyuan
2023Tencent’s general model family, with open-weight versions
Qwen
2023An open-weight family spanning many sizes, with multimodal versions
Phi
2023Small, efficient models built from curated data
Step
2023A general model family aimed at multimodality and on-device use
Baichuan
2023An open-weight general model aimed at Chinese-language use
GLM
2023A Chinese general model that began with autoregressive blank-infilling pretraining
Llama
2023The model family that made the open-weight route mainstream
Pangu
2020Huawei’s Pangu family of foundation models
ERNIE
2019A Chinese model that began with knowledge-enhanced pretraining, an early landmark version
Siri
2011The early voice assistant that brought spoken control to mainstream phones
Organizations involved
Typical uses
- General assistants and support bots
- Coding and study tutoring
- Role-play and companion apps
- Internal knowledge-helpdesk entry points
How it is evaluated
- Human-preference Elo
- Ranking by blind pairwise win rate
- Instruction-following rate
- Share of verifiable constraints (format, length, banned words) satisfied
- Safety violation rate
- Rate of violations under red-team prompts
Limits and hard parts
- Constraints from early turns, especially format rules, are forgotten or softened in long chats
- It tends to agree with the user even when the premise is wrong — sycophancy
- Carefully crafted prompts can bypass safety policy; jailbreaks are hard to eliminate
Concepts behind it
Prompting & Alignment
Making a model helpful, honest and harmless is harder than simply making it bigger
Reinforcement Learning from Human Feedback
When the good answer cannot be written as a formula, let humans stand in as the reward function
Transformer Architecture
Replacing word-by-word relay with a room where everyone speaks at once, so long-range dependencies are one hop away