손실 함수
손실 함수는 ‘무엇을 벌할지’를 정한다 — 바꾸면 옳고 그름의 기준 자체가 바뀐다
이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.
정의
A loss function maps the gap between a prediction and the true label to a single non-negative number, and training simply tries to make its average as small as possible. Mean squared error penalises squared deviations and suits regression; cross-entropy penalises “assigning too little probability to the correct answer” and suits classification. Choosing a loss is declaring how much you are willing to pay for which kind of mistake.
직관적 이해
Think of a loss as a penalty rule. The same ten-minute lateness can be fined linearly per minute, or doubled the moment you are late — and each rule elicits completely different behaviour. A loss function is exactly that rule for a model: it decides which mistakes get corrected first and which get tolerated. Swap the loss and you swap the entire notion of right and wrong.
The shape of two losses (binary classification, true label positive): cross-entropy rises sharply when wrong, while mean squared error stays too flat
- Cross-entropy (true = positive)
- Mean squared error
Convergence of three losses on one classification task (schematic): cross-entropy descends fastest, mean squared error learns noticeably slower in the saturated regime
- Cross-entropy
- MSE (used for classification)
- Hinge loss
작동 원리
- 01
Measure the per-sample error
For each sample, quantify the gap between prediction and label. Regression uses squared difference; classification uses the negative log of the probability assigned to the correct class — how surprised you should be by the outcome.
- 02
Average into a total loss
Sum or average the per-sample errors into a single scalar over the training set. That scalar is what gradient descent actually minimises.
- 03
Let the gradient drive correction
Differentiate the loss to obtain the direction and magnitude for every parameter. Whether the loss is differentiable and its gradient smooth directly determines how stable optimisation will be.
- 04
Pick the loss for the task
Regression defaults to mean squared error, classification to cross-entropy; when margin matters or outliers are a concern, Hinge or contrastive losses come in. The criterion is always: which kind of mistake do you fear most?
핵심 수식
L_CE = − Σ_c y_c · log p̂_c응용 분야
- Regression: house prices, temperature and coordinates trained with mean squared error and its variants
- Classification: image and text classifiers almost universally use cross-entropy
- Metric learning: contrastive and triplet losses pull similar samples together and push dissimilar ones apart
- Class imbalance: weighted cross-entropy or Focal Loss keep the model from fixating on the majority class
흔한 오해
- Mean squared error paired with a Sigmoid vanishes. In classification, MSE on a saturating Sigmoid output yields gradients so small the weights barely move — exactly why cross-entropy wins in classification.
- Falling loss does not mean a better metric. While training loss keeps dropping, the accuracy or recall you care about may stall or decline, because the loss optimises a proxy rather than the end goal.
- Sensitivity to outliers varies by loss. Mean squared error is dominated by extreme values; mean absolute error or Huber is more robust, at the cost of non-differentiability or a kink near zero.
핵심 용어
- Mean squared error (MSE)
- The average squared difference between prediction and label; the default regression loss
- Cross-entropy
- Negative log-probability of the correct class; the default classification loss
- Hinge loss
- Requires the correct class to win by a margin; the heart of the SVM
- Contrastive loss
- A loss that pulls same-class embeddings together and pushes different-class ones apart