Statistics › Probability rules › full formula sheet
Bayes’ theorem
Flip a conditional: update what you believe when new evidence arrives.
Notation on this page: P(A|B) is read “the probability of A given B” — how likely A is, once we know B happened.
Where it comes from
Real life constantly hands us probabilities facing the wrong way. A test manufacturer tells you P(positive | disease) — “our test catches 99% of infections.” But the patient does not want that number. The patient wants P(disease | positive) — “given that I tested positive, what are the odds I am actually sick?” The first is about the test; the second is about you. Bayes’ theorem is the machine that flips one into the other.
The naive approach is to treat the flip as free — to read P(A|B) and P(B|A) as the same thing:
Kill it with one concrete example. Suppose 1% of the population has a disease, and a test is 99% sensitive (catches 99% of the sick) and 99% specific (correctly clears 99% of the healthy). Run the numbers:
Why? Think in people, not probabilities. Out of 10,000: 100 are sick and the test catches 99 of them; 9,900 are healthy and the test wrongly flags 1% of them — another 99. So there are 198 positives on the table, and only 99 of them are actually sick: 99/198 = 50%. The false positives exactly match the true positives, because the healthy population is 99 times bigger. Ignoring that base rate is called base-rate neglect, and it is the single most common probability error in the wild.
Here is the intuition before the formalism. Read the formula as a belief update: you start with a prior P(A) — how likely A was before the evidence. Evidence B arrives. You multiply by how likely that evidence should be if A were true, P(B|A). And you divide by the total probability of seeing the evidence at all, P(B) — how likely B is across every hypothesis. Evidence that would be surprising under all alternatives is strong evidence; evidence that happens all the time anyway barely moves the needle. That division is the whole trick — it re-normalizes, so your updated beliefs still sum to 1.
Derivation
Bayes’ theorem needs no new axioms — it is two applications of the definition of conditional probability, back to back. The second line is the whole proof.
In practice the denominator P(B) is never given — you build it from the hypotheses with the law of total probability:
Why this form? It makes the normalizing role of the denominator visible: the numerator is one hypothesis’s “share” of the evidence, and the denominator is the sum of every hypothesis’s share. That is why the posteriors over a full partition always sum to 1.
How to use it
The procedure, every time:
- Name A = the hypothesis, B = the evidence. A is what you want to know (“patient has the disease”); B is what you observed (“the test came back positive”). Write them down explicitly — most Bayes errors start with A and B swapped.
- List the three inputs. The prior P(A) (base rate, before the evidence); the likelihood P(B|A) (how the evidence behaves if the hypothesis is true); and the false-positive rate P(B|¬A) (how the evidence behaves if the hypothesis is false).
- Build P(B) with total probability. P(B) = P(B|A)·P(A) + P(B|¬A)·P(¬A). Never guess the denominator, never set it to 0.5 “by symmetry” — it is determined by your model.
- Divide. P(A|B) = P(B|A)·P(A) / P(B). That is your posterior: the updated belief.
- Sanity-check. The posterior must lie between 0 and 1, and posteriors over a full partition must sum to 1. Also check the direction: if the evidence is likelier under A than under not-A, the posterior should be above the prior. If it is not, recheck your inputs.
Chain it: yesterday’s posterior is today’s prior
Bayes is meant to be applied repeatedly. Got a second, independent test? Do not start over — the posterior from the first test becomes the prior for the second. Each round of evidence multiplies in, and beliefs converge toward the truth. See Example 3 below.
When the prior is shaky
Sometimes nobody knows the base rate. That does not break the method — it exposes the subjectivity debate: different priors give different posteriors. Two defenses: (1) use the best data you have and say so; (2) with enough evidence the likelihood overwhelms the prior and honest analysts converge anyway. One hard rule: never use a prior of exactly 0 or 1 (Cromwell’s rule) — 0 × likelihood stays 0 forever, so no evidence could ever move you.
Worked examples
Four problems, easiest first. In each one, read every step — the why of each move is the lesson.
Example 1 — the medical test: 99% accurate, 50% answer
1% of people have a disease. The test is 99% sensitive and 99% specific. You test positive. What is P(disease | positive)?
- Name the parts. A = “has the disease”, B = “tests positive”. (Why this way? We want P(disease | positive), so the disease is A, the evidence is B.)
- List the inputs. Prior P(A) = 0.01; likelihood P(B|A) = 0.99; false-positive rate P(B|¬A) = 1 − 0.99 = 0.01. (Why 1 − 0.99? Specificity is P(negative | healthy); the false-positive rate is its complement.)
- Build P(B). = 0.99·0.01 + 0.01·0.99 = 0.0099 + 0.0099 = 0.0198.
- Divide. P(A|B) = 0.0099 / 0.0198 = 0.50.
- Check with a frequency table (the no-algebra safety net). Start with 10,000 people: 100 are sick, 9,900 healthy. True positives: 99% of 100 = 99. False positives: 1% of 9,900 = 99. Total positives = 198; sick among them = 99/198 = 50%. Matches ✓
Example 2 — the spam filter: P(spam | “free”)
20% of mail is spam. The word “free” appears in 40% of spam but only 5% of legitimate mail (ham). A message contains “free”. What is P(spam | “free”)?
- Name the parts. A = “is spam”, B = “contains ‘free’”.
- List the inputs. Prior P(A) = 0.2; likelihood P(B|A) = 0.4; P(B|¬A) = 0.05. (Why is the prior 0.2 and not 0.5? The base rate of spam is a fact about your inbox — never assume “50/50” by symmetry.)
- Build P(B). = 0.4·0.2 + 0.05·0.8 = 0.08 + 0.04 = 0.12. (Why two terms? “Free” can arrive via spam or via ham — total probability adds both paths.)
- Divide. P(A|B) = 0.08 / 0.12 = 2/3 ≈ 66.7%.
- Sanity-check the direction. “Free” is 8× likelier under spam (0.4 vs 0.05), so the posterior should exceed the 20% prior — and 66.7% does. ✓
Example 3 — chaining: a second positive test
Same test as Example 1 (99% sensitive, 99% specific). The patient from Example 1 takes a second, independent test and it is also positive. What is P(disease | two positives)?
- Reuse the posterior as the new prior. After the first positive, P(disease) = 0.5. (Why? Bayes is iterative — yesterday’s conclusion is today’s starting point.)
- List the inputs. New prior P(A) = 0.5; likelihood P(B|A) = 0.99; false-positive rate P(B|¬A) = 0.01.
- Build P(B). = 0.99·0.5 + 0.01·0.5 = 0.495 + 0.005 = 0.50.
- Divide. P(A|B) = 0.495 / 0.50 = 0.99.
- Read the lesson. One positive moved belief from 1% to 50%; a second, independent positive moves it to 99%. (Why does independence matter? If the two tests shared the same failure mode, the second test would add less — possibly nothing. Only independent evidence chains this cleanly.)
Example 4 — judgment call: the prosecutor’s fallacy
A crime-scene DNA sample matches the defendant. An expert testifies: “the chance of this match if the defendant were innocent is 1 in a million.” The prosecutor concludes the defendant is guilty beyond doubt. What went wrong?
- Identify the two conditionals. The expert gave P(match | innocent) = 1/1,000,000. The prosecutor wants P(innocent | match) — the probability the defendant is innocent given the match.
- Spot the flip. The prosecutor equated P(match | innocent) with P(innocent | match). That is the transposed conditional — the same error as reading P(positive | healthy) as P(healthy | positive).
- See why Bayes is mandatory here. P(innocent | match) = P(match | innocent)·P(innocent) / P(match). The prior P(innocent) — the chance a random person in the suspect pool is innocent — is enormous, and the denominator P(match) includes matches against every innocent person tested. Without them, the 1-in-a-million number says nothing about guilt.
- Intuition check. If the police tested a million innocent people, they would expect about one match by chance. A match from a million-person sweep barely moves the needle — the posterior depends on how big the pool was, i.e., on the base rate.
The lesson: whenever someone quotes P(evidence | hypothesis) as though it were P(hypothesis | evidence), you are looking at the prosecutor’s fallacy. Bayes’ theorem is the correction — and the prior is the part they are hoping you forget.
Memorization tips
- Say it aloud: “posterior equals likelihood times prior, over evidence.” Three named parts, one division — the rhythm matches the formula.
- Name the parts every time: P(A) = prior, P(B|A) = likelihood, P(B) = evidence. A problem half-solved is a problem with A and B labeled.
- Draw the tree: first branches = hypotheses (priors), second branches = evidence (likelihoods). Bayes just re-weights the paths that end at B. If you can draw the tree, you can build P(B) — just add the B-paths.
- P(A|B) ≠ P(B|A) — the whole point of the theorem is that they differ. Confusing them is the prosecutor’s fallacy, and it costs marks, cases, and diagnoses.
- Respect the base rate: rare things stay unlikely even after positive evidence. If your posterior ignores the prior, you have committed base-rate neglect.
- The 50% test (your 5-second self-check): 1% base rate, 99%/99% test, one positive → exactly 50%. Run it in your head whenever a Bayes answer feels off — if your method cannot reproduce 50%, the method is wrong.
Final challenge
Five mixed questions — basics, applications, and the traps, all in one. Score 5/5 and Bayes’ theorem is yours.
← Back to the Statistics formula sheet
How to learn a formula here
- Read each section in order — every section ends with a short quiz. Take it before moving on; the questions test exactly what you just read.
- Work the examples with the answers covered, then uncover one step at a time and compare.
- Finish with the final challenge — five mixed questions including the classic traps.
- Retake what you miss — every quiz reshuffles each attempt, and every answer explains itself.
Frequently asked questions
What is Bayes’ theorem?
Bayes’ theorem, P(A|B) = P(B|A)P(A)/P(B), flips a conditional probability: it tells you the probability of a hypothesis given evidence, from the probability of the evidence given the hypothesis. It updates a prior belief into a posterior belief.
Why can a 99% accurate test be misleading?
Because of the base rate. With 1% disease prevalence, a 99%-sensitive, 99%-specific test gives only a 50% chance you’re sick after a positive result — the false positives from the healthy majority equal the true positives from the tiny sick group.
Is P(A|B) the same as P(B|A)?
No — confusing them is the prosecutor’s fallacy. P(DNA match|innocent) being 1 in a million does not mean P(innocent|match) is 1 in a million. Bayes’ theorem exists precisely because these two quantities differ.
How do I compute P(B) in the denominator?
Use the law of total probability: P(B) = sum over all hypotheses of P(B|Ai)P(Ai). It just normalizes the posteriors so they sum to 1 — never guess it.
Where does Bayes’ theorem come from?
Rev. Thomas Bayes’ 1763 posthumous essay (Laplace developed it independently). It is a direct rearrangement of the definition of conditional probability.
More from the codex
Statistics formula sheet
All 49 formulas — distributions, inference, probability — printable and quiz-ready.
Open sheet → LiveFormula Sheet Builder
Mix and match any sections into your own printable sheet.
Open tool → LivePrompt Simulator
Practice prompt engineering with deterministic scoring.
Open tool →Support the codex
This page is free, with no account and no ads. If it helped you learn, consider supporting the indie dev behind it.
Questions or a bug to report? Email [email protected].