CAPABILITIES
What AI can do
Forty capabilities gathered into nine families by what they operate on. Each one spells out what it means, how it is done, where it is used, and which products already deliver it.
Click any cell to see the capabilities that fall on that input → output pair.
Language & knowledge
10 capabilitiesFrom writing the first word to reading a whole document: this family works on symbol sequences themselves — generating, rewriting, translating, summarising, extracting, retrieving and reasoning.
Text Generation
Continue a passage, one word at a time
Conversation & Instruction Following
Understand intent across turns and act on it
Machine Translation
Render text in one language as another
Summarization
Compress a long text into a shorter, faithful one
Text Classification & Sentiment
Assign a predefined label to a piece of text
Information Extraction & NER
Pick names, places and relations out of free text
Question Answering & RAG
Retrieve the evidence first, then answer from it
Reasoning & Chain-of-Thought
Break a hard problem into intermediate steps
Text Embeddings & Semantic Search
Turn sentences into vectors and fetch nearest by meaning
Reranking
Reorder recalled candidates by true relevance
Vision understanding
7 capabilitiesTeaching a model to read an image. From the simplest classification, to drawing boxes, segmenting every pixel, recognising text and faces, and answering questions about what it sees.
Image Classification
Decide what category the whole image is
Object Detection
Draw a box around each object and name it
Image Segmentation
Label every pixel with the object it belongs to
Pose & Keypoint Estimation
Locate joints and recover a skeleton
Optical Character Recognition
Read the text in an image as editable characters
Face Recognition
Decide whether two faces are the same person
Image Understanding & VQA
Look at an image and answer free-form questions
Image generation & editing
5 capabilitiesStarting from a sentence or a sketch and producing, editing or restoring an image. This is the family of generative AI that the public met first.
Text-to-Image
Turn a sentence into an image
Image-to-Image
Regenerate a similar image from an input one
Image Editing & Inpainting
Change a spot in the image by a written instruction
Super-Resolution & Restoration
Make small, blurry or old images sharp
Matting & Background Removal
Cut the subject cleanly out of its background
Video
4 capabilitiesVideo is a moving image plus time. This family handles space and time together: generating clips, understanding actions, editing shots and syncing lips.
Text-to-Video
Turn a sentence into a short video clip
Image-to-Video
Set a still image in motion
Video Understanding
Understand what happens in a video
Video Editing & Lip Sync
Recut a video or re-voice it so lips match audio
Speech & music
4 capabilitiesSound is a waveform and a time series. Speech recognition turns it into text, speech synthesis goes the other way, and music generation produces listenable material outright.
Speech Recognition (ASR)
Transcribe spoken audio into text
Text-to-Speech
Read text aloud as natural speech
Voice Cloning & Conversion
Reproduce a voice from a few sample clips
Music & Sound Generation
Generate a piece of music or a sound effect from a description
3D
2 capabilitiesLifting two-dimensional input into three-dimensional assets: meshes, point clouds and renderable scenes. It sits upstream of games, film and robot simulation.
Code & software
2 capabilitiesCode is a language with strict grammar. This family does more than complete the next line: it reads a whole repository, runs commands and tracks down errors.
Agents & tool use
3 capabilitiesOnce a model can call tools, read and write files and drive a browser, it stops answering questions and starts finishing tasks. This family is about stringing actions into reliable flows.
Data & documents
3 capabilitiesTurning messy tables, documents and layouts into structured fields. It is the first hurdle most organisations hit when they wire AI into a real process.