Skip to content
Family Table Math
Auto

Bayes' Theorem

Often the probability you know runs the “wrong way”. A factory knows how often each machine makes a faulty item, but when a customer returns a faulty item, the question is which machine probably made it. A doctor knows how accurate a test is for people who have a disease, but the patient wants to know the chance they have the disease given a positive result. Bayes’ theorem turns P(B∣A)P(B \mid A) into P(A∣B)P(A \mid B). It’s one formula, but it changes how you think about evidence.

From the definition of conditional probability, P(A∩B)=P(A)P(B∣A)P(A \cap B) = P(A)P(B \mid A) and also P(A∩B)=P(B)P(A∣B)P(A \cap B) = P(B)P(A \mid B). Setting these equal and dividing by P(B)P(B):

P(A∣B)=P(A) P(B∣A)P(B)P(A \mid B) = \frac{P(A)\,P(B \mid A)}{P(B)}

That’s the heart of Bayes’ theorem: to reverse the condition, you need the “forward” probability P(B∣A)P(B \mid A) and the overall probability P(B)P(B).

Usually P(B)P(B) isn’t given directly. Event AA either happens or it doesn’t, so BB can happen in two ways, and

P(B)=P(A)P(B∣A)+P(A′)P(B∣A′)P(B) = P(A)P(B \mid A) + P(A')P(B \mid A')

Putting this in the denominator gives Bayes’ theorem:

P(A∣B)=P(A) P(B∣A)P(A) P(B∣A)+P(A′) P(B∣A′)P(A \mid B) = \frac{P(A)\,P(B \mid A)}{P(A)\,P(B \mid A) + P(A')\,P(B \mid A')}

If exactly one of A1A_1, A2A_2, A3A_3 must happen (they’re mutually exclusive and their probabilities add to 11), the same idea gives

P(Ai∣B)=P(Ai) P(B∣Ai)P(A1)P(B∣A1)+P(A2)P(B∣A2)+P(A3)P(B∣A3)P(A_i \mid B) = \frac{P(A_i)\,P(B \mid A_i)}{P(A_1)P(B \mid A_1) + P(A_2)P(B \mid A_2) + P(A_3)P(B \mid A_3)}

The IB course uses Bayes’ theorem for at most three events like this.

You don’t have to memorize the formula if you can draw the tree:

  1. First branches: the possible causes (AA and A′A', or A1A_1, A2A_2, A3A_3), with their probabilities.
  2. Second branches: whether BB happens, with the conditional probabilities P(B∣…)P(B \mid \ldots).
  3. Multiply along each path that ends in BB. Their sum is P(B)P(B) (the denominator).
  4. P(Ai∣B)P(A_i \mid B) is the path through AiA_i divided by that sum.

In words:

P(cause∣evidence)=the path through that causeall the paths that give the evidenceP(\text{cause} \mid \text{evidence}) = \frac{\text{the path through that cause}}{\text{all the paths that give the evidence}}

P(A)P(A) is sometimes called the prior probability (before you know about BB), and P(A∣B)P(A \mid B) the updated (posterior) probability. If new evidence arrives, the updated probability can become the prior for the next step. This only works if the pieces of evidence are independent of each other once you know whether AA happened.

On any school day, the probability of rain is 0.30.3. If it rains, the probability that Maya is late is 0.250.25; if it doesn’t rain, it’s 0.080.08. Given that Maya is late, find the probability that it was raining.

Solution. Let RR = rain and LL = late. The two paths that end in “late”:

  • P(R∩L)=0.3×0.25=0.075P(R \cap L) = 0.3 \times 0.25 = 0.075
  • P(R′∩L)=0.7×0.08=0.056P(R' \cap L) = 0.7 \times 0.08 = 0.056

So P(L)=0.075+0.056=0.131P(L) = 0.075 + 0.056 = 0.131 and

P(R∣L)=0.0750.131=0.5725…≈0.573 (3 s.f.)P(R \mid L) = \frac{0.075}{0.131} = 0.5725\ldots \approx 0.573 \text{ (3 s.f.)}

Knowing Maya is late raises the probability of rain from 0.30.3 to about 0.5730.573.

A factory has three machines. Machine A makes 50%50\% of the items, machine B makes 30%30\% and machine C makes 20%20\%. The proportions of defective items are 2%2\% for A, 3%3\% for B and 5%5\% for C. An item is chosen at random and is defective. Find the probability that it was made by each machine.

Tree diagram: machines A, B and C make 0.5, 0.3 and 0.2 of the items; the defective rates are 0.02, 0.03 and 0.05. 0.5 A 0.02 D 0.5 × 0.02 = 0.010 0.98 D′ 0.3 B 0.03 D 0.3 × 0.03 = 0.009 0.97 D′ 0.2 C 0.05 D 0.2 × 0.05 = 0.010 0.95 D′ machine defective?
Multiply along each branch that ends in D. The three products add to P(D)=0.029P(D) = 0.029.

Solution. Let DD = defective. From the tree:

P(D)=0.5(0.02)+0.3(0.03)+0.2(0.05)=0.010+0.009+0.010=0.029P(D) = 0.5(0.02) + 0.3(0.03) + 0.2(0.05) = 0.010 + 0.009 + 0.010 = 0.029

Then divide each path by 0.0290.029:

P(A∣D)=0.0100.029=1029≈0.345P(B∣D)=0.0090.029=929≈0.310P(C∣D)=0.0100.029=1029≈0.345P(A \mid D) = \frac{0.010}{0.029} = \frac{10}{29} \approx 0.345 \qquad P(B \mid D) = \frac{0.009}{0.029} = \frac{9}{29} \approx 0.310 \qquad P(C \mid D) = \frac{0.010}{0.029} = \frac{10}{29} \approx 0.345

Check: 10+9+1029=1\dfrac{10 + 9 + 10}{29} = 1. ✓

Machine C makes only 20%20\% of the items, but it’s just as likely as machine A to be the source of a defective one, because its defect rate is so much higher.

A disease affects 1%1\% of a population. A test gives a positive result for 98%98\% of people who have the disease and for 4%4\% of people who don’t.

  • (a) A person tests positive. Find the probability that they have the disease.
  • (b) The person takes the test again and again tests positive. Assuming the two results are independent given whether the person has the disease, find the new probability that they have the disease.

Solution.

(a) Let DD = has the disease and ++ = positive.

P(D∣+)=0.01(0.98)0.01(0.98)+0.99(0.04)=0.00980.0098+0.0396=0.00980.0494=0.19838…P(D \mid +) = \frac{0.01(0.98)}{0.01(0.98) + 0.99(0.04)} = \frac{0.0098}{0.0098 + 0.0396} = \frac{0.0098}{0.0494} = 0.19838\ldots

Only about 0.1980.198: most positives come from the large healthy group.

(b) Now the prior is P(D)=0.19838P(D) = 0.19838 (keep it unrounded):

P(D∣second +)=0.19838(0.98)0.19838(0.98)+0.80162(0.04)=0.194410.19441+0.03206=0.858 (3 s.f.)P(D \mid \text{second } +) = \frac{0.19838(0.98)}{0.19838(0.98) + 0.80162(0.04)} = \frac{0.19441}{0.19441 + 0.03206} = 0.858 \text{ (3 s.f.)}

A second positive result raises the probability from about 20%20\% to about 86%86\%. That’s why positive screening results are confirmed with a repeat test.

A bag holds three coins: a fair coin, a coin with heads on both sides, and a biased coin with P(head)=14P(\text{head}) = \tfrac{1}{4}. A coin is chosen at random and tossed, and it lands heads. Find the probability that it was the two-headed coin.

Solution. Each coin is chosen with probability 13\tfrac{1}{3}. The paths ending in heads:

P(H)=13⋅12+13⋅1+13⋅14=13(24+44+14)=13⋅74=712P(H) = \frac{1}{3}\cdot\frac{1}{2} + \frac{1}{3}\cdot 1 + \frac{1}{3}\cdot\frac{1}{4} = \frac{1}{3}\left(\frac{2}{4} + \frac{4}{4} + \frac{1}{4}\right) = \frac{1}{3}\cdot\frac{7}{4} = \frac{7}{12} P(two-headed∣H)=13⋅1712=13×127=47P(\text{two-headed} \mid H) = \frac{\tfrac{1}{3} \cdot 1}{\tfrac{7}{12}} = \frac{1}{3} \times \frac{12}{7} = \frac{4}{7}

Mixing up P(A∣B)P(A \mid B) and P(B∣A)P(B \mid A). ”98%98\% of people with the disease test positive” is P(+∣D)=0.98P(+ \mid D) = 0.98. It does not mean that 98%98\% of positive results have the disease; Example 3 shows that’s only about 20%20\%. Write each given probability in P(…∣…)P(\ldots \mid \ldots) notation before you start.

Forgetting a path in the denominator. P(B)P(B) is the sum of all the paths that end in BB. With three causes, there are three paths. Leaving one out makes the answer too big.

Using the wrong second-branch probability. On the “not defective” branches you need 1−0.02=0.981 - 0.02 = 0.98, and so on. If the question is about a negative test or a non-defective item, use those complementary branches in both the numerator and the denominator.

Ignoring the base rate. A very accurate test for a rare condition can still produce mostly false positives. The prior probability P(A)P(A) matters as much as the test accuracy.

Rounding the updated probability before reusing it. In Example 3(b), using 0.20.2 instead of 0.198380.19838 changes the answer to 0.8600.860. Keep full accuracy until the end.

1. (Warm-up) P(A)=0.4P(A) = 0.4, P(B∣A)=0.3P(B \mid A) = 0.3 and P(B∣A′)=0.6P(B \mid A') = 0.6. Find P(A∣B)P(A \mid B).

SolutionP(B)=0.4(0.3)+0.6(0.6)=0.12+0.36=0.48P(B) = 0.4(0.3) + 0.6(0.6) = 0.12 + 0.36 = 0.48P(A∣B)=0.120.48=0.25P(A \mid B) = \frac{0.12}{0.48} = 0.25

2. (Warm-up) P(A)=0.3P(A) = 0.3, P(B)=0.2P(B) = 0.2 and P(A∣B)=0.6P(A \mid B) = 0.6. Find P(B∣A)P(B \mid A).

Solution

P(A∩B)=P(B)P(A∣B)=0.2(0.6)=0.12P(A \cap B) = P(B)P(A \mid B) = 0.2(0.6) = 0.12. Then

P(B∣A)=P(A∩B)P(A)=0.120.3=0.4P(B \mid A) = \frac{P(A \cap B)}{P(A)} = \frac{0.12}{0.3} = 0.4

3. (Core) A phone maker buys a part from three suppliers: 45%45\% from supplier 1, 35%35\% from supplier 2 and 20%20\% from supplier 3. The parts are faulty 1%1\%, 2%2\% and 4%4\% of the time. A part is found to be faulty.

  • (a) Find the probability that it came from supplier 3.
  • (b) Which supplier is it most likely to have come from?
Solution

(a) The faulty paths: 0.45(0.01)=0.00450.45(0.01) = 0.0045, 0.35(0.02)=0.0070.35(0.02) = 0.007, 0.2(0.04)=0.0080.2(0.04) = 0.008. So P(F)=0.0195P(F) = 0.0195.

P(S3∣F)=0.0080.0195=0.410 (3 s.f.)P(S_3 \mid F) = \frac{0.008}{0.0195} = 0.410 \text{ (3 s.f.)}

(b) P(S1∣F)=0.00450.0195=0.231P(S_1 \mid F) = \dfrac{0.0045}{0.0195} = 0.231 and P(S2∣F)=0.0070.0195=0.359P(S_2 \mid F) = \dfrac{0.007}{0.0195} = 0.359 (3 s.f.). Supplier 3 is the most likely source, even though it supplies the fewest parts.

4. (Core) 30%30\% of the emails a company receives are spam. A filter flags 95%95\% of spam emails and 2%2\% of genuine emails.

  • (a) An email is flagged. Find the probability that it is spam.
  • (b) An email is not flagged. Find the probability that it is spam.
Solution

(a) Flagged paths: spam 0.3(0.95)=0.2850.3(0.95) = 0.285; genuine 0.7(0.02)=0.0140.7(0.02) = 0.014.

P(S∣flagged)=0.2850.285+0.014=0.2850.299=0.953 (3 s.f.)P(S \mid \text{flagged}) = \frac{0.285}{0.285 + 0.014} = \frac{0.285}{0.299} = 0.953 \text{ (3 s.f.)}

(b) Not-flagged paths: spam 0.3(0.05)=0.0150.3(0.05) = 0.015; genuine 0.7(0.98)=0.6860.7(0.98) = 0.686.

P(S∣not flagged)=0.0150.015+0.686=0.0150.701=0.0214 (3 s.f.)P(S \mid \text{not flagged}) = \frac{0.015}{0.015 + 0.686} = \frac{0.015}{0.701} = 0.0214 \text{ (3 s.f.)}

5. (Core) Students at a school travel by bus (50%50\%), by car (30%30\%) or on foot (20%20\%). The probabilities of being late are 0.10.1 by bus, 0.050.05 by car and 0.150.15 on foot. A student chosen at random arrived on time. Find the probability that they walked.

Solution

Use the “on time” branches: 0.90.9, 0.950.95 and 0.850.85.

P(on time)=0.5(0.9)+0.3(0.95)+0.2(0.85)=0.45+0.285+0.17=0.905P(\text{on time}) = 0.5(0.9) + 0.3(0.95) + 0.2(0.85) = 0.45 + 0.285 + 0.17 = 0.905P(walk∣on time)=0.170.905=0.188 (3 s.f.)P(\text{walk} \mid \text{on time}) = \frac{0.17}{0.905} = 0.188 \text{ (3 s.f.)}

6. (Core) A fair six-sided die is rolled. If it shows 11 or 22, a ball is taken from box A, which holds 44 green and 22 yellow balls. Otherwise, a ball is taken from box B, which holds 11 green and 55 yellow balls. Given that the ball is green, find the exact probability that it came from box A.

Solution

P(A)=26=13P(A) = \dfrac{2}{6} = \dfrac{1}{3} and P(B)=23P(B) = \dfrac{2}{3}. P(G∣A)=46=23P(G \mid A) = \dfrac{4}{6} = \dfrac{2}{3} and P(G∣B)=16P(G \mid B) = \dfrac{1}{6}.

P(G)=13⋅23+23⋅16=29+19=39=13P(G) = \frac{1}{3}\cdot\frac{2}{3} + \frac{2}{3}\cdot\frac{1}{6} = \frac{2}{9} + \frac{1}{9} = \frac{3}{9} = \frac{1}{3}P(A∣G)=2/91/3=29×3=23P(A \mid G) = \frac{2/9}{1/3} = \frac{2}{9} \times 3 = \frac{2}{3}

7. (Core) For the factory in Example 2, an item is chosen at random and is not defective. Find the probability that it was made by machine C.

Solution

P(D′)=1−0.029=0.971P(D') = 1 - 0.029 = 0.971. The path through C is 0.2(0.95)=0.190.2(0.95) = 0.19.

P(C∣D′)=0.190.971=0.196 (3 s.f.)P(C \mid D') = \frac{0.19}{0.971} = 0.196 \text{ (3 s.f.)}

That’s slightly less than C’s share of 0.20.2, since C’s items are more likely to be defective.

8. (Challenge) A test for a condition gives a positive result for 90%90\% of people who have it and for 10%10\% of people who don’t. In a certain population, half of the people who test positive actually have the condition. Find the proportion of the population that has the condition.

Solution

Let p=P(C)p = P(C). Then

0.9p0.9p+0.1(1−p)=0.50.9p=0.45p+0.05(1−p)0.9p=0.45p+0.05−0.05p0.5p=0.05p=0.1\begin{aligned} \frac{0.9p}{0.9p + 0.1(1 - p)} &= 0.5 \\ 0.9p &= 0.45p + 0.05(1 - p) \\ 0.9p &= 0.45p + 0.05 - 0.05p \\ 0.5p &= 0.05 \\ p &= 0.1 \end{aligned}

10%10\% of the population has the condition.

Check: 0.090.09+0.09=0.5\dfrac{0.09}{0.09 + 0.09} = 0.5. ✓

9. (Challenge) In a multiple-choice quiz, each question has 44 options. A student knows the answer to a question with probability pp; otherwise she guesses at random. If she knows the answer, she gets it right.

  • (a) Show that the probability she knew the answer, given that she got it right, is 4p3p+1\dfrac{4p}{3p + 1}.
  • (b) Find this probability when p=0.6p = 0.6.
  • (c) Find pp if this probability is 0.90.9.
Solution

(a) Let KK = knows and RR = right. P(R∣K)=1P(R \mid K) = 1 and P(R∣K′)=14P(R \mid K') = \dfrac{1}{4}.

P(K∣R)=p⋅1p⋅1+(1−p)⋅14=4p4p+1−p=4p3p+1P(K \mid R) = \frac{p \cdot 1}{p \cdot 1 + (1 - p)\cdot\frac{1}{4}} = \frac{4p}{4p + 1 - p} = \frac{4p}{3p + 1}

(b) 4(0.6)3(0.6)+1=2.42.8=67≈0.857\dfrac{4(0.6)}{3(0.6) + 1} = \dfrac{2.4}{2.8} = \dfrac{6}{7} \approx 0.857.

(c)

4p3p+1=0.9⇒4p=2.7p+0.9⇒1.3p=0.9⇒p=913≈0.692\frac{4p}{3p + 1} = 0.9 \quad\Rightarrow\quad 4p = 2.7p + 0.9 \quad\Rightarrow\quad 1.3p = 0.9 \quad\Rightarrow\quad p = \frac{9}{13} \approx 0.692