Aprendizagem supervisionada
Pares de pergunta e resposta ensinam o modelo a responder sozinho
O texto completo é apresentado em inglês; o título e o resumo estão traduzidos.
DEFINIÇÃO
Supervised learning is the paradigm of learning a mapping from labelled examples: given many input–label pairs (x → y), find within a candidate family of functions (the hypothesis space) the one that minimises the discrepancy between predictions and true labels. When labels are discrete categories the task is classification; when they are continuous values it is regression. It rests on two premises: that a learnable mapping exists, and that we hold enough clean labels to pin it down.
Intuição
It is like a student working through a thick workbook of exercises with the answer key attached: checking each attempt against the key, the student gradually internalises how such problems are solved rather than memorising any single answer. The real bottleneck is not solving problems but where the answers come from — grading one paper is cheap, but having doctors label millions of scans by hand can cost more than the project is worth. Labels are typically far scarcer than raw data.
The supervised pipeline: collect, label, fit inside a hypothesis space, and finally accept or reject on unseen data
Progress in supervised learning: on one ImageNet dataset with the same labels, single-crop top-1 accuracy across model families
Como funciona
- 01
Collect and label paired samples
Attach a correct label to every input, then split the data into training, validation and test sets. The split must happen before any tuning, and the three parts must never contaminate one another.
- 02
Fix the hypothesis space and the loss
Choose a model family — linear models, decision trees, neural networks — and specify how “being wrong” is measured, i.e. the loss function. The hypothesis space caps what can be expressed; the loss sets the direction of optimisation.
- 03
Minimise the empirical risk
Find the parameters that make the average loss over the training set as small as possible. This is usually carried out by gradient descent or one of its variants, and it is the only stage where real computation happens.
- 04
Test generalisation on unseen data
Low training loss does not mean the regularity has been captured. The real acceptance test is whether the model stays accurate on samples it has never seen — which leads straight to overfitting and model evaluation.
Fórmula-chave
min_θ (1/n) Σᵢ L( f_θ(xᵢ), yᵢ )Onde é usado
- Spam and content classification: map a piece of text to a category label
- Medical imaging: predict the presence of a lesion from X-rays or pathology slides
- Price and demand regression: predict house prices, sales or energy use from historical features
- Fine-tuning for speech recognition and translation: specialise on human transcripts or aligned bitexts
Equívocos comuns
- Label noise and annotation bias: a labeller’s carelessness or systematic preference becomes the model’s ceiling — it may learn the labeller’s bias instead of the truth.
- Distribution shift: training data came from one environment while deployment changes it, so the learned mapping no longer holds. Beautiful offline numbers can collapse on day one.
- Correlation is not causation: what the model learns is a statistical association, not necessarily the causal mechanism you have in mind. Acting on it can backfire.
Termos-chave
- Input x
- The feature vector fed to the model
- Label y
- The correct output for each sample; the source of supervision
- Hypothesis space
- The set of all functions the model can represent
- Empirical risk
- The model’s average loss on the training samples