CONCEPTS
All Concepts
Forty-eight core concepts, filterable by domain or level.
48 concepts
Vectors & Vector Spaces
AI turns everything — words, images, sounds, users — into one thing: a list of numbers
Gradients & Gradient Descent
All of deep learning comes down to one thing: take a small step downhill
Matrix Operations & Linear Maps
Matrix multiplication is not a pile of multiply-and-adds; it rewrites an entire space in one stroke
Probability & Distributions
A model never hands you an answer; it hands you a degree of belief over every possible answer
Bayes’ Theorem
Believe a little, see the evidence, revise a little — that is Bayes
Entropy & Information Theory
The more surprising a sentence, the more information it carries — and a language model’s loss measures exactly that surprise
Supervised Learning
Pairs of questions and answers teach a model to answer on its own
Unsupervised Learning
With no answers given, structure must emerge from the data itself — and “good” has to be redefined
Loss Functions
A loss function defines what you actually penalise — swap it and you swap your entire notion of right and wrong
Overfitting & Regularization
The model memorises the training data word for word, then fails on anything new
Model Evaluation & Cross-Validation
Accuracy is the easiest metric to fool you — get evaluation wrong and everything else follows
Bias–Variance Tradeoff
Every prediction error splits into three parts: the model too simple, the model too jumpy, and the world’s own randomness
Neuron & Perceptron
The smallest part of a neural network: a weighted sum, a bias, and one twist of nonlinearity
Backpropagation
Turning “compute the gradient” from mathematical drudgery into a single function call — the moment deep learning took off
Activation Functions
Without it, even a very deep network is only a single linear map
Convolutional Neural Networks
Replacing full connections with “look locally, reuse the same filter everywhere” — the idea that made image recognition work
Recurrent Neural Networks
Giving networks a memory: one unit reused across time to handle sequences of any length
Normalization & Residual Connections
Making hundred-layer networks trainable: an identity shortcut plus a per-layer rescaling
Tokenization
Models do not read characters, they read tokens — and how you split text quietly sets both capability and cost
Word Embeddings
Turning words into coordinates — synonyms land near each other, and meaning becomes something you can add and subtract
Attention Mechanism
Every position can look directly at every other position and dynamically weight how much attention to pay
Transformer Architecture
Replacing word-by-word relay with a room where everyone speaks at once, so long-range dependencies are one hop away
Pretraining & Fine-tuning
Learn language first from vast unlabelled text, then specialise with little data — the most data-efficient paradigm in modern AI
Prompting & Alignment
Making a model helpful, honest and harmless is harder than simply making it bigger
Image Representation
To a machine, a photo is nothing but stacked grids of numbers
Convolution Operations
One small stencil swept across the image finds edges, textures and shapes
Image Classification
Don’t tell it “cats have whiskers” — show it enough cats and it works it out
Object Detection
From “what is in the image” to “what, where, and how many”
Semantic Segmentation
Colouring every pixel: not a box around the object, but a colouring book
Self-Supervised Vision & Contrastive Multimodal Learning
No labels needed: learning to see by working out which images are the same
Markov Decision Process
Write "deciding step by step" as five symbols — everything in reinforcement learning starts here
Value Functions & Q-Learning
Instead of guessing what to do, first estimate what each choice is worth
Policy Gradients
Adjust the policy itself, so that good actions occur more often
Exploration vs Exploitation
The best option right now is not necessarily the best one in the long run
Deep Reinforcement Learning
Let a neural network decide straight from pixels — then hold it steady with decades-old tricks
Reinforcement Learning from Human Feedback
When the good answer cannot be written as a formula, let humans stand in as the reward function
Generative Models: An Overview
Discriminative models answer "what is this"; generative models answer "what should this look like"
Autoencoders & VAE
Squeeze information through a bottleneck, then let it grow back
Generative Adversarial Networks
A forger versus an inspector: each pushes the other to its limit
Diffusion Models
Learn a thousand tiny denoising steps, and you can build an image from pure noise
Latent Diffusion & Conditional Control
Run diffusion not over pixels, but inside a compressed semantic space
Multimodal Generation
One model that learns to speak, to draw, to move — even to model the 3D world
Training & Inference Infrastructure
Memory decides how large a model you can train, communication how long it takes — raw compute is rarely the bottleneck
Model Compression
Make a model smaller, faster and cheaper with almost no accuracy loss — but you can usually have only two of the three at once
Inference Optimization & Serving
Training happens once; inference happens a billion times a day — and serving is torn between fast first tokens and high throughput, which usually pull against each other
Retrieval-Augmented Generation
Rather than cramming knowledge into parameters, leave it outside and look it up on demand — an open-book exam instead of a closed-book one
Agents & Tool Use
Let a model do more than answer: search, call APIs, run code — and decide the next step from what came back
Safety, Alignment & Prompt Injection
A model optimises the proxy we wrote into the loss, never the thing we actually want — the gap between them is the whole alignment problem