Statistics › Probability rules › full formula sheet

P(A|B) = P(B|A)·P(A) / P(B)

Bayes’ theorem

Flip a conditional: update what you believe when new evidence arrives.

Notation on this page: P(A|B) is read “the probability of A given B” — how likely A is, once we know B happened.

Where it comes from

Real life constantly hands us probabilities facing the wrong way. A test manufacturer tells you P(positive | disease) — “our test catches 99% of infections.” But the patient does not want that number. The patient wants P(disease | positive) — “given that I tested positive, what are the odds I am actually sick?” The first is about the test; the second is about you. Bayes’ theorem is the machine that flips one into the other.

The naive approach is to treat the flip as free — to read P(A|B) and P(B|A) as the same thing:

P(disease | positive) = P(positive | disease)  ??the tempting — and wrong — shortcut: “the test is 99% accurate, so a positive means a 99% chance of disease”

Kill it with one concrete example. Suppose 1% of the population has a disease, and a test is 99% sensitive (catches 99% of the sick) and 99% specific (correctly clears 99% of the healthy). Run the numbers:

P(positive | disease)
=
0.99
Given by the test manufacturer. This is the direction we have.
P(disease | positive)
=
0.99 × 0.01 / (0.99 × 0.01 + 0.01 × 0.99) = 0.0099 / 0.0198 = 0.50
The naive shortcut says 0.99 — nearly double the truth. The true chance is a coin flip: 50%.

Why? Think in people, not probabilities. Out of 10,000: 100 are sick and the test catches 99 of them; 9,900 are healthy and the test wrongly flags 1% of them — another 99. So there are 198 positives on the table, and only 99 of them are actually sick: 99/198 = 50%. The false positives exactly match the true positives, because the healthy population is 99 times bigger. Ignoring that base rate is called base-rate neglect, and it is the single most common probability error in the wild.

Here is the intuition before the formalism. Read the formula as a belief update: you start with a prior P(A) — how likely A was before the evidence. Evidence B arrives. You multiply by how likely that evidence should be if A were true, P(B|A). And you divide by the total probability of seeing the evidence at all, P(B) — how likely B is across every hypothesis. Evidence that would be surprising under all alternatives is strong evidence; evidence that happens all the time anyway barely moves the needle. That division is the whole trick — it re-normalizes, so your updated beliefs still sum to 1.

Derivation

Bayes’ theorem needs no new axioms — it is two applications of the definition of conditional probability, back to back. The second line is the whole proof.

P(A|B)
=
P(A∩B) / P(B)
Step 1 — the definition. Conditional probability is defined as the probability of both events over the probability of the conditioning event. (Requires P(B) > 0 — you cannot condition on the impossible.)
P(A∩B)
=
P(B|A)·P(A)
Step 2 — the trick: apply the same definition the other way. Swap the roles: P(B|A) = P(B∩A)/P(A), so P(B∩A) = P(B|A)·P(A). And A∩B is the same event as B∩A — order never matters for an intersection.
P(A|B)
=
P(B|A)·P(A) / P(B)
Step 3 — substitute. Replace the numerator from Step 1 with the rewritten intersection from Step 2. That is Bayes’ theorem. ∎

In practice the denominator P(B) is never given — you build it from the hypotheses with the law of total probability:

P(B) = P(B|A)·P(A) + P(B|¬A)·P(¬A)the usable form: total probability of the evidence = evidence under A, weighted by A’s prior, plus evidence under not-A, weighted by not-A’s prior

Why this form? It makes the normalizing role of the denominator visible: the numerator is one hypothesis’s “share” of the evidence, and the denominator is the sum of every hypothesis’s share. That is why the posteriors over a full partition always sum to 1.

How to use it

The procedure, every time:

  1. Name A = the hypothesis, B = the evidence. A is what you want to know (“patient has the disease”); B is what you observed (“the test came back positive”). Write them down explicitly — most Bayes errors start with A and B swapped.
  2. List the three inputs. The prior P(A) (base rate, before the evidence); the likelihood P(B|A) (how the evidence behaves if the hypothesis is true); and the false-positive rate P(B|¬A) (how the evidence behaves if the hypothesis is false).
  3. Build P(B) with total probability. P(B) = P(B|A)·P(A) + P(B|¬A)·P(¬A). Never guess the denominator, never set it to 0.5 “by symmetry” — it is determined by your model.
  4. Divide. P(A|B) = P(B|A)·P(A) / P(B). That is your posterior: the updated belief.
  5. Sanity-check. The posterior must lie between 0 and 1, and posteriors over a full partition must sum to 1. Also check the direction: if the evidence is likelier under A than under not-A, the posterior should be above the prior. If it is not, recheck your inputs.

Chain it: yesterday’s posterior is today’s prior

Bayes is meant to be applied repeatedly. Got a second, independent test? Do not start over — the posterior from the first test becomes the prior for the second. Each round of evidence multiplies in, and beliefs converge toward the truth. See Example 3 below.

When the prior is shaky

Sometimes nobody knows the base rate. That does not break the method — it exposes the subjectivity debate: different priors give different posteriors. Two defenses: (1) use the best data you have and say so; (2) with enough evidence the likelihood overwhelms the prior and honest analysts converge anyway. One hard rule: never use a prior of exactly 0 or 1 (Cromwell’s rule) — 0 × likelihood stays 0 forever, so no evidence could ever move you.

Common mistake — base-rate neglect: plugging in the test’s accuracy and forgetting the prior entirely. A 99%-accurate test on a 1%-prevalent disease gives 50%, not 99%. Whenever your posterior barely mentions the base rate, you have probably dropped it.

Worked examples

Four problems, easiest first. In each one, read every step — the why of each move is the lesson.

Example 1 — the medical test: 99% accurate, 50% answer

1% of people have a disease. The test is 99% sensitive and 99% specific. You test positive. What is P(disease | positive)?

  1. Name the parts. A = “has the disease”, B = “tests positive”. (Why this way? We want P(disease | positive), so the disease is A, the evidence is B.)
  2. List the inputs. Prior P(A) = 0.01; likelihood P(B|A) = 0.99; false-positive rate P(B|¬A) = 1 − 0.99 = 0.01. (Why 1 − 0.99? Specificity is P(negative | healthy); the false-positive rate is its complement.)
  3. Build P(B). = 0.99·0.01 + 0.01·0.99 = 0.0099 + 0.0099 = 0.0198.
  4. Divide. P(A|B) = 0.0099 / 0.0198 = 0.50.
  5. Check with a frequency table (the no-algebra safety net). Start with 10,000 people: 100 are sick, 9,900 healthy. True positives: 99% of 100 = 99. False positives: 1% of 9,900 = 99. Total positives = 198; sick among them = 99/198 = 50%. Matches ✓
Common mistake: answering 99% — quoting P(positive | disease) as if it were P(disease | positive). The frequency table is the antidote: it forces you to count the healthy population’s false positives.

Example 2 — the spam filter: P(spam | “free”)

20% of mail is spam. The word “free” appears in 40% of spam but only 5% of legitimate mail (ham). A message contains “free”. What is P(spam | “free”)?

  1. Name the parts. A = “is spam”, B = “contains ‘free’”.
  2. List the inputs. Prior P(A) = 0.2; likelihood P(B|A) = 0.4; P(B|¬A) = 0.05. (Why is the prior 0.2 and not 0.5? The base rate of spam is a fact about your inbox — never assume “50/50” by symmetry.)
  3. Build P(B). = 0.4·0.2 + 0.05·0.8 = 0.08 + 0.04 = 0.12. (Why two terms? “Free” can arrive via spam or via ham — total probability adds both paths.)
  4. Divide. P(A|B) = 0.08 / 0.12 = 2/3 ≈ 66.7%.
  5. Sanity-check the direction. “Free” is 8× likelier under spam (0.4 vs 0.05), so the posterior should exceed the 20% prior — and 66.7% does. ✓
Common mistake: answering 40% — the likelihood P(“free” | spam), quoted as if it were the posterior. The likelihood is an input, not the answer.

Example 3 — chaining: a second positive test

Same test as Example 1 (99% sensitive, 99% specific). The patient from Example 1 takes a second, independent test and it is also positive. What is P(disease | two positives)?

  1. Reuse the posterior as the new prior. After the first positive, P(disease) = 0.5. (Why? Bayes is iterative — yesterday’s conclusion is today’s starting point.)
  2. List the inputs. New prior P(A) = 0.5; likelihood P(B|A) = 0.99; false-positive rate P(B|¬A) = 0.01.
  3. Build P(B). = 0.99·0.5 + 0.01·0.5 = 0.495 + 0.005 = 0.50.
  4. Divide. P(A|B) = 0.495 / 0.50 = 0.99.
  5. Read the lesson. One positive moved belief from 1% to 50%; a second, independent positive moves it to 99%. (Why does independence matter? If the two tests shared the same failure mode, the second test would add less — possibly nothing. Only independent evidence chains this cleanly.)
Common mistake: starting the second test from the 1% base rate again, as if the first test never happened. That double-counts the prior and throws away everything learned.

Example 4 — judgment call: the prosecutor’s fallacy

A crime-scene DNA sample matches the defendant. An expert testifies: “the chance of this match if the defendant were innocent is 1 in a million.” The prosecutor concludes the defendant is guilty beyond doubt. What went wrong?

  1. Identify the two conditionals. The expert gave P(match | innocent) = 1/1,000,000. The prosecutor wants P(innocent | match) — the probability the defendant is innocent given the match.
  2. Spot the flip. The prosecutor equated P(match | innocent) with P(innocent | match). That is the transposed conditional — the same error as reading P(positive | healthy) as P(healthy | positive).
  3. See why Bayes is mandatory here. P(innocent | match) = P(match | innocent)·P(innocent) / P(match). The prior P(innocent) — the chance a random person in the suspect pool is innocent — is enormous, and the denominator P(match) includes matches against every innocent person tested. Without them, the 1-in-a-million number says nothing about guilt.
  4. Intuition check. If the police tested a million innocent people, they would expect about one match by chance. A match from a million-person sweep barely moves the needle — the posterior depends on how big the pool was, i.e., on the base rate.

The lesson: whenever someone quotes P(evidence | hypothesis) as though it were P(hypothesis | evidence), you are looking at the prosecutor’s fallacy. Bayes’ theorem is the correction — and the prior is the part they are hoping you forget.

Common mistake: thinking this fallacy only happens in courtrooms. It is the same error as the 99%-test paradox and the spam-filter slip: P(B|A) quoted as P(A|B), with the base rate quietly dropped.

Memorization tips

  • Say it aloud: “posterior equals likelihood times prior, over evidence.” Three named parts, one division — the rhythm matches the formula.
  • Name the parts every time: P(A) = prior, P(B|A) = likelihood, P(B) = evidence. A problem half-solved is a problem with A and B labeled.
  • Draw the tree: first branches = hypotheses (priors), second branches = evidence (likelihoods). Bayes just re-weights the paths that end at B. If you can draw the tree, you can build P(B) — just add the B-paths.
  • P(A|B) ≠ P(B|A) — the whole point of the theorem is that they differ. Confusing them is the prosecutor’s fallacy, and it costs marks, cases, and diagnoses.
  • Respect the base rate: rare things stay unlikely even after positive evidence. If your posterior ignores the prior, you have committed base-rate neglect.
  • The 50% test (your 5-second self-check): 1% base rate, 99%/99% test, one positive → exactly 50%. Run it in your head whenever a Bayes answer feels off — if your method cannot reproduce 50%, the method is wrong.

Final challenge

Five mixed questions — basics, applications, and the traps, all in one. Score 5/5 and Bayes’ theorem is yours.

← Back to the Statistics formula sheet

How to learn a formula here

  1. Read each section in order — every section ends with a short quiz. Take it before moving on; the questions test exactly what you just read.
  2. Work the examples with the answers covered, then uncover one step at a time and compare.
  3. Finish with the final challenge — five mixed questions including the classic traps.
  4. Retake what you miss — every quiz reshuffles each attempt, and every answer explains itself.

Frequently asked questions

What is Bayes’ theorem?

Bayes’ theorem, P(A|B) = P(B|A)P(A)/P(B), flips a conditional probability: it tells you the probability of a hypothesis given evidence, from the probability of the evidence given the hypothesis. It updates a prior belief into a posterior belief.

Why can a 99% accurate test be misleading?

Because of the base rate. With 1% disease prevalence, a 99%-sensitive, 99%-specific test gives only a 50% chance you’re sick after a positive result — the false positives from the healthy majority equal the true positives from the tiny sick group.

Is P(A|B) the same as P(B|A)?

No — confusing them is the prosecutor’s fallacy. P(DNA match|innocent) being 1 in a million does not mean P(innocent|match) is 1 in a million. Bayes’ theorem exists precisely because these two quantities differ.

How do I compute P(B) in the denominator?

Use the law of total probability: P(B) = sum over all hypotheses of P(B|Ai)P(Ai). It just normalizes the posteriors so they sum to 1 — never guess it.

Where does Bayes’ theorem come from?

Rev. Thomas Bayes’ 1763 posthumous essay (Laplace developed it independently). It is a direct rearrangement of the definition of conditional probability.

More from the codex