NLP & Large Language Models
From tokenisation to Transformers to models that can talk
OVERVIEW
- All Concepts
- 6
- Beginner
- 2
- Intermediate
- 3
- Expert
- 1
Language is cut into tokens, compressed into vectors, and then made to “look at” itself through attention — a loop so plain it seems excessive, yet it lifted next-token prediction all the way to a general assistant. This domain traces that path in order: bag of words, embeddings, RNNs, attention, Transformers, pre-training and alignment.
Questions this domain answers
- Q1
What exactly does a model see when it reads a word?
- Q2
Why does attention beat recurrence?
- Q3
Why does next-token prediction yield conversation?
Concepts in this domain
- 01TokenizationBeginnerModels do not read characters, they read tokens — and how you split text quietly sets both capability and cost
- 02Word EmbeddingsBeginnerTurning words into coordinates — synonyms land near each other, and meaning becomes something you can add and subtract
- 03Attention MechanismIntermediateEvery position can look directly at every other position and dynamically weight how much attention to pay
- 04Transformer ArchitectureIntermediateReplacing word-by-word relay with a room where everyone speaks at once, so long-range dependencies are one hop away
- 05Pretraining & Fine-tuningIntermediateLearn language first from vast unlabelled text, then specialise with little data — the most data-efficient paradigm in modern AI
- 06Prompting & AlignmentExpertMaking a model helpful, honest and harmless is harder than simply making it bigger