Skip to content
AI Atlas

Bayes’ Theorem

Believe a little, see the evidence, revise a little — that is Bayes

01 Math & Statistics FoundationsIntermediateEntry 5 in this domain

DEFINITION

Bayes’ theorem gives a precise update rule: multiply your degree of belief in a hypothesis before seeing data (the prior) by how likely that data is under the hypothesis (the likelihood), then normalise, and you obtain your degree of belief after seeing the data (the posterior). It reads probability as a belief that adjusts with evidence, not merely as a long-run frequency.

Intuition

A rare disease affects one person in a thousand. The test is accurate: 99% of the sick test positive, and only 1% of the healthy falsely do. You test positive — yet your chance of actually being sick is only about 9%. The reason is plain: there are so many healthy people that their 1% of false positives outnumbers the sick people’s true positives. Intuition is defeated by the base rate, and that is exactly the illusion Bayes’ theorem corrects.

Fig. 1

Expanding 1000 people by "sick or not — test result": among all positives, most come from healthy people’s false alarms — the arithmetic behind the base-rate fallacy

1000 people1000 peopleSick 0.1% (1 person)Sick 0.1% (1 person)Positive 99% (≈0.99)Positive 99% (≈0.99)Negative 1% (≈0.01)Negative 1% (≈0.01)Healthy 99.9% (999 people)Healthy 99.9% (999 people)Positive 1% (≈9.99)Positive 1% (≈9.99)Negative 99% (≈989)Negative 99% (≈989)
Fig. 2

The same 99%-accurate test gives wildly different verdicts under different base rates: the prior dominates the post-test probability

  • Before test (prior)
  • After positive (posterior)

How it works

  1. 01

    Prior: what you believe before seeing the data

    The prior P(H) is your degree of belief in hypothesis H before looking at data. It may come from historical statistics or domain knowledge, or simply be the modeller’s judgement — Bayes’ honesty is that it puts this subjective assumption on the table rather than hiding it.

  2. 02

    Likelihood: how common the data is under each hypothesis

    The likelihood P(E|H) asks "if H were true, how common would evidence E be". Note its direction is opposite to P(H|E): 99% of the sick test positive (likelihood) does not mean 99% of positives are sick (posterior). Confusing the two is the root of most faulty reasoning.

  3. 03

    Posterior: the new belief after weighting by the evidence

    The posterior P(H|E) = P(E|H)·P(H) / P(E), where the denominator P(E) is the total probability of the evidence across all hypotheses. It normalises the result so the updated belief remains a valid probability.

  4. 04

    Sequential updating: today’s posterior is tomorrow’s prior

    Multiple pieces of evidence can be absorbed one by one: treat this round’s posterior as the next round’s prior and the result matches using all evidence at once. This makes Bayes natural for online learning. The naive Bayes classifier plugs a simplifying "conditional independence" assumption into this frame to filter spam at very low cost.

Key formula

P(H | E) = P(E | H) · P(H) / P(E)
H is the hypothesis and E the evidence. Multiply the prior P(H) by the likelihood P(E|H), then divide by the evidence P(E) to normalise, yielding the posterior.
Fig. 4

Sequential belief updating under repeated positive tests: each additional positive multiplies the posterior odds by the same likelihood ratio, and certainty approaches fast

  • Probability of disease
  • Probability of health

Where it is used

  • Spam filtering: naive Bayes updates the probability of "this is spam" word by word
  • Medical diagnosis: combining prevalence (prior) with sensitivity and specificity (likelihood) to interpret a positive result
  • A/B testing and Bayesian optimisation: using the posterior to decide where the next experiment’s resources go
  • Online learning and calibration: nudging beliefs with each new data point while quantifying the remaining uncertainty

Common misconceptions

  • Ignoring the base rate. Looking only at the likelihood ("a 99% accurate test") and forgetting the prior turns a rare disease’s positive result into a spurious 99% — the most common and most costly reasoning error.
  • Sensitivity to the prior. With little data, the choice of prior markedly shapes the posterior; the stronger the prior, the more evidence is needed to overturn it.
  • A posterior is not causation. Bayesian updating revises "your beliefs", not the world itself; it says how to believe given current hypotheses, not why the world is the way it is.

Key terms

Prior
The degree of belief in a hypothesis before seeing data
Likelihood
The probability of observed data given that the hypothesis is true
Posterior
The updated degree of belief after incorporating the evidence
Evidence
The total probability of the data across all hypotheses; it normalises the result

Further reading