Machine Translation
Render text in one language as another
WHAT THIS CAPABILITY MEANS
Takes source-language text and outputs a translation. Unlike open-ended generation it has an objective reference — whether the translation is faithful — and unlike summarisation it should not drop content but rebuild it as evenly as possible. Beyond prose it must handle terminology lists, markup and placeholders so the output can be pasted back into the original system.
How it is done
When the Transformer was introduced in 2017, the encoder–decoder with attention was designed for translation, and neural machine translation has since largely replaced statistical systems. Source sentences are encoded and target sentences decoded, trained on parallel corpora. Recent multilingual models cover a hundred-plus languages with one set of parameters and treat translation as an instruction, while low-resource languages lean on transfer from high-resource ones and back-translation.
Representative products
6DeepL Translator
2017A neural machine translation service that bets on translation quality
GPT-4o
2024A natively multimodal general model, with text, image and audio through one door
Gemini
2023A natively multimodal general model built for very long context
Qwen
2023An open-weight family spanning many sizes, with multimodal versions
Hunyuan
2023Tencent’s general model family, with open-weight versions
Whisper
2022Transcribes speech in many languages and translates it into English
Organizations involved
Typical uses
- Cross-border documents and contracts
- Localisation of software and websites
- Real-time cross-language communication
- Paper abstracts and subtitles
How it is evaluated
- BLEU
- N-gram overlap with references; higher is better but not readability
- COMET
- A neural faithfulness score for translations
- chrF
- Character-level F-score, steadier for morphologically rich languages
Limits and hard parts
- Quality on low-resource languages and dialects trails high-resource ones by a wide margin
- Terminology and proper-noun consistency drifts across a long document
- Idioms, puns and culture-specific wording are often translated literally and lose meaning
Concepts behind it
Transformer Architecture
Replacing word-by-word relay with a room where everyone speaks at once, so long-range dependencies are one hop away
Attention Mechanism
Every position can look directly at every other position and dynamically weight how much attention to pay
Tokenization
Models do not read characters, they read tokens — and how you split text quietly sets both capability and cost