ERNIE
A Chinese model that began with knowledge-enhanced pretraining, an early landmark version
WHAT IT IS
ERNIE is Baidu’s large language model series; the earliest, ERNIE 1.0, was released in March 2019 and stands for Enhanced Representation through Knowledge Integration. Its key idea is knowledge-enhanced pretraining: entity- and phrase-level masking during pretraining forces the model to learn knowledge relations between words, not just surface co-occurrence of neighbouring characters. The ERNIE line later grew into a large family spanning retrieval, dialogue and multimodality, and underpins Baidu’s ERNIE Bot product. This entry records the ERNIE starting point as represented by the early 2019 versions.
Why it matters
It introduced entity- and phrase-level masking into pretraining so the model learns knowledge relations rather than mere character co-occurrence — an idea widely borrowed in early Chinese pretraining. These early versions have since been superseded by later ERNIE generations.
Key specs
- Architecture
- Knowledge-enhanced pretraining (knowledge masking)
- Released
- 2019-03 (ERNIE 1.0)
- Modality
- Text
- Open weights
- No
Capabilities
Related concepts
Pretraining & Fine-tuning
Learn language first from vast unlabelled text, then specialise with little data — the most data-efficient paradigm in modern AI
Transformer Architecture
Replacing word-by-word relay with a room where everyone speaks at once, so long-range dependencies are one hop away
Tokenization
Models do not read characters, they read tokens — and how you split text quietly sets both capability and cost
Comparable products
GLM
2023A Chinese general model that began with autoregressive blank-infilling pretraining
Qwen
2023An open-weight family spanning many sizes, with multimodal versions
Hunyuan
2023Tencent’s general model family, with open-weight versions
Baichuan
2023An open-weight general model aimed at Chinese-language use
Doubao
2023ByteDance’s general chat model and application
Pangu
2020Huawei’s Pangu family of foundation models