본문으로 건너뛰기
AI 도감

지도 학습

‘질문과 정답’의 짝을 보여 주어 모델이 스스로 답하게 만든다

02 머신러닝입문이 영역의 1번째 항목

이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.

정의

Supervised learning is the paradigm of learning a mapping from labelled examples: given many input–label pairs (x → y), find within a candidate family of functions (the hypothesis space) the one that minimises the discrepancy between predictions and true labels. When labels are discrete categories the task is classification; when they are continuous values it is regression. It rests on two premises: that a learnable mapping exists, and that we hold enough clean labels to pin it down.

직관적 이해

It is like a student working through a thick workbook of exercises with the answer key attached: checking each attempt against the key, the student gradually internalises how such problems are solved rather than memorising any single answer. The real bottleneck is not solving problems but where the answers come from — grading one paper is cheap, but having doctors label millions of scans by hand can cost more than the project is worth. Labels are typically far scarcer than raw data.

그림 1

The supervised pipeline: collect, label, fit inside a hypothesis space, and finally accept or reject on unseen data

그림 2

Progress in supervised learning: on one ImageNet dataset with the same labels, single-crop top-1 accuracy across model families

작동 원리

  1. 01

    Collect and label paired samples

    Attach a correct label to every input, then split the data into training, validation and test sets. The split must happen before any tuning, and the three parts must never contaminate one another.

  2. 02

    Fix the hypothesis space and the loss

    Choose a model family — linear models, decision trees, neural networks — and specify how “being wrong” is measured, i.e. the loss function. The hypothesis space caps what can be expressed; the loss sets the direction of optimisation.

  3. 03

    Minimise the empirical risk

    Find the parameters that make the average loss over the training set as small as possible. This is usually carried out by gradient descent or one of its variants, and it is the only stage where real computation happens.

  4. 04

    Test generalisation on unseen data

    Low training loss does not mean the regularity has been captured. The real acceptance test is whether the model stays accurate on samples it has never seen — which leads straight to overfitting and model evaluation.

핵심 수식

min_θ (1/n) Σᵢ L( f_θ(xᵢ), yᵢ )
Empirical risk minimisation: choose θ to minimise the average loss over training samples — the generic objective behind every supervised training loop.

응용 분야

  • Spam and content classification: map a piece of text to a category label
  • Medical imaging: predict the presence of a lesion from X-rays or pathology slides
  • Price and demand regression: predict house prices, sales or energy use from historical features
  • Fine-tuning for speech recognition and translation: specialise on human transcripts or aligned bitexts

흔한 오해

  • Label noise and annotation bias: a labeller’s carelessness or systematic preference becomes the model’s ceiling — it may learn the labeller’s bias instead of the truth.
  • Distribution shift: training data came from one environment while deployment changes it, so the learned mapping no longer holds. Beautiful offline numbers can collapse on day one.
  • Correlation is not causation: what the model learns is a statistical association, not necessarily the causal mechanism you have in mind. Acting on it can backfire.

핵심 용어

Input x
The feature vector fed to the model
Label y
The correct output for each sample; the source of supervision
Hypothesis space
The set of all functions the model can represent
Empirical risk
The model’s average loss on the training samples

참고문헌