Summarization
Compress a long text into a shorter, faithful one
WHAT THIS CAPABILITY MEANS
Takes a long document and returns a shorter version that keeps the key information. It splits into extractive (selecting sentences) and abstractive (rewriting in new words) styles, the latter now dominant. Unlike free generation it has an explicit compression target, and unlike question answering it is not aimed at one query but should cover the whole thread.
How it is done
The classic approach is a sequence-to-sequence attention model producing abstractive summaries, with rewriting ability coming from large-scale pre-training. Long documents are usually handled hierarchically or by a map-reduce scheme that summarises chunks and then combines them. Controlled summarisation passes length, angle or audience constraints through the prompt or light fine-tuning to steer style.
Representative products
8GPT-4o
2024A natively multimodal general model, with text, image and audio through one door
Claude
2023A general chat model known for long context and safety alignment
Gemini
2024One chat entry point that gathers search, office apps and a multimodal model
NotebookLM
2023Answers only from the sources you give it, with citations
Kimi
2023A Chinese chat assistant known for long-context handling
Apple Intelligence
2024System-level AI that splits work between on-device and private cloud
Microsoft Copilot
2023Conversational AI woven into the operating system and office apps
Perplexity
2022Answers built as it searches, each one backed by sources
Organizations involved
Typical uses
- Quick reads of news and reports
- Meeting and call minutes
- Literature triage and paper skims
- Rolling summaries of tickets and email
How it is evaluated
- ROUGE
- N-gram overlap with reference summaries
- BERTScore
- Semantic-embedding similarity, more tolerant than literal overlap
- Factual consistency
- Share of statements conflicting with the source, as in FactCC-style evaluation
Limits and hard parts
- Middle sections of long documents are often dropped; models favour the start and end
- Abstractive models splice facts from different sentences into claims the source never made
- On ambiguous or multi-sided texts, minority views are often reported as the majority
Concepts behind it
Attention Mechanism
Every position can look directly at every other position and dynamically weight how much attention to pay
Prompting & Alignment
Making a model helpful, honest and harmless is harder than simply making it bigger
Transformer Architecture
Replacing word-by-word relay with a room where everyone speaks at once, so long-range dependencies are one hop away