베이즈 정리
조금 믿고, 증거를 보고, 조금 수정한다 — 이것이 베이즈다
이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.
정의
Bayes’ theorem gives a precise update rule: multiply your degree of belief in a hypothesis before seeing data (the prior) by how likely that data is under the hypothesis (the likelihood), then normalise, and you obtain your degree of belief after seeing the data (the posterior). It reads probability as a belief that adjusts with evidence, not merely as a long-run frequency.
직관적 이해
A rare disease affects one person in a thousand. The test is accurate: 99% of the sick test positive, and only 1% of the healthy falsely do. You test positive — yet your chance of actually being sick is only about 9%. The reason is plain: there are so many healthy people that their 1% of false positives outnumbers the sick people’s true positives. Intuition is defeated by the base rate, and that is exactly the illusion Bayes’ theorem corrects.
Expanding 1000 people by "sick or not — test result": among all positives, most come from healthy people’s false alarms — the arithmetic behind the base-rate fallacy
The same 99%-accurate test gives wildly different verdicts under different base rates: the prior dominates the post-test probability
- Before test (prior)
- After positive (posterior)
작동 원리
- 01
Prior: what you believe before seeing the data
The prior P(H) is your degree of belief in hypothesis H before looking at data. It may come from historical statistics or domain knowledge, or simply be the modeller’s judgement — Bayes’ honesty is that it puts this subjective assumption on the table rather than hiding it.
- 02
Likelihood: how common the data is under each hypothesis
The likelihood P(E|H) asks "if H were true, how common would evidence E be". Note its direction is opposite to P(H|E): 99% of the sick test positive (likelihood) does not mean 99% of positives are sick (posterior). Confusing the two is the root of most faulty reasoning.
- 03
Posterior: the new belief after weighting by the evidence
The posterior P(H|E) = P(E|H)·P(H) / P(E), where the denominator P(E) is the total probability of the evidence across all hypotheses. It normalises the result so the updated belief remains a valid probability.
- 04
Sequential updating: today’s posterior is tomorrow’s prior
Multiple pieces of evidence can be absorbed one by one: treat this round’s posterior as the next round’s prior and the result matches using all evidence at once. This makes Bayes natural for online learning. The naive Bayes classifier plugs a simplifying "conditional independence" assumption into this frame to filter spam at very low cost.
핵심 수식
P(H | E) = P(E | H) · P(H) / P(E)Sequential belief updating under repeated positive tests: each additional positive multiplies the posterior odds by the same likelihood ratio, and certainty approaches fast
- Probability of disease
- Probability of health
응용 분야
- Spam filtering: naive Bayes updates the probability of "this is spam" word by word
- Medical diagnosis: combining prevalence (prior) with sensitivity and specificity (likelihood) to interpret a positive result
- A/B testing and Bayesian optimisation: using the posterior to decide where the next experiment’s resources go
- Online learning and calibration: nudging beliefs with each new data point while quantifying the remaining uncertainty
흔한 오해
- Ignoring the base rate. Looking only at the likelihood ("a 99% accurate test") and forgetting the prior turns a rare disease’s positive result into a spurious 99% — the most common and most costly reasoning error.
- Sensitivity to the prior. With little data, the choice of prior markedly shapes the posterior; the stronger the prior, the more evidence is needed to overturn it.
- A posterior is not causation. Bayesian updating revises "your beliefs", not the world itself; it says how to believe given current hypotheses, not why the world is the way it is.
핵심 용어
- Prior
- The degree of belief in a hypothesis before seeing data
- Likelihood
- The probability of observed data given that the hypothesis is true
- Posterior
- The updated degree of belief after incorporating the evidence
- Evidence
- The total probability of the data across all hypotheses; it normalises the result