Mistake Master
Student view — seeing the site as a student does
Home Unit 2 · Probability, Random Variables, and Probability Distributions 2.1·2.2·2.3·2.4·2.5·2.6·2.7·2.8·2.9·2.10·2.11·2.12 Lesson
Skill Check 0 / 10 complete

The condition sets the denominator

Conditioning is not a new kind of probability; it is the same counting done inside a smaller world. Once you are told that B happened, everything outside B is gone, and the question becomes what share of what remains is also A. Reverse the two events and you have changed which world you are standing in, which is why a test that catches 95% of a disease can still be wrong most of the time when it says yes.

§1

Conditioning throws away everything outside the condition.

The conditional probability of $A$ given $B$ is the probability that $A$ occurs when it is already known that $B$ did:

$$P(A \mid B) = \frac{P(A \text{ and } B)}{P(B)}, \qquad P(B) > 0.$$

The formula is a restatement of one idea. Learning that $B$ happened deletes every outcome outside $B$, so $B$ becomes the new sample space. The outcomes still available that also satisfy $A$ are exactly the ones in $A$ and $B$, and they are measured against the size of $B$ rather than against the whole space.

In a two-way table this is a row or a column and nothing more. Among 200 students, 60 are seniors and 30 of them hold a job, so $P(\text{job} \mid \text{senior}) = \frac{30}{60} = 0.50$: cover every row but the seniors and read the fraction across it. That is why the condition, not the event, decides which total goes underneath.

§2

The direction of the bar changes the question.

$P(A \mid B)$ and $P(B \mid A)$ share a numerator and differ in denominator, so they are usually different numbers answering different questions:

$$P(A \mid B) = \frac{P(A \text{ and } B)}{P(B)}, \qquad P(B \mid A) = \frac{P(A \text{ and } B)}{P(A)}.$$

Almost every NBA player is over six feet tall, and almost nobody over six feet tall is an NBA player. Both statements are true, and swapping them turns a fact into an absurdity. In the students table, $P(\text{job} \mid \text{senior}) = \frac{30}{60} = 0.50$ while $P(\text{senior} \mid \text{job}) = \frac{30}{72} \approx 0.417$: same 30 students, two populations.

The reliable habit is to read the condition first and write down its total before doing anything else. The phrase after "given", or after "of the", or after "among", names the world; everything else is the event being counted inside it.

§3

A tree diagram multiplies along a branch and adds across paths.

Rearranging the definition gives the multiplication rule, $P(A \text{ and } B) = P(B) \cdot P(A \mid B)$, which is what a tree diagram draws. The first split carries unconditional probabilities; each later branch carries a probability conditional on the path so far; a path's probability is the product along it; and the probability of an event is the sum over every path that produces it.

A disease affects 1% of a population. A test correctly returns positive for 95% of people who have it, and correctly returns negative for 90% of people who do not. In a population of 10,000:

  1. 100 have the disease; 95 of them test positive.
  2. 9,900 do not; 10% of them, or 990, test positive anyway.
  3. Total positives: $95 + 990 = 1{,}085$.

So $P(\text{positive} \mid \text{disease}) = 0.95$, while

$$P(\text{disease} \mid \text{positive}) = \frac{95}{1085} \approx 0.088.$$

Fewer than one positive result in ten comes from a person who has the disease. Nothing is wrong with the test: the healthy group is 99 times larger, so its 10% error rate produces ten times more positives than the sick group's 95% success rate does. This is what a base rate does, and it is the reason the reversed conditional is a distinct question rather than a rephrasing.

§4

Conditional probabilities are read back in the language of the group.

An interpretation names the restricted group, the event, and the context: "among students who ate at the cafeteria, about 8% reported illness". Dropping the group turns a conditional into a claim about everyone, which is a different and usually false statement.

Three checks catch most errors before they leave the page:

  1. Which total went underneath? It should be the size of the condition, not the grand total and not the size of the event.
  2. Does the conditional distribution sum to 1? $P(A \mid B) + P(A^c \mid B) = 1$, since within $B$ the event either happens or it does not. Note that $P(A \mid B) + P(A \mid B^c)$ has no reason to sum to anything.
  3. Would the reversed version sound different? If so, make sure the one written is the one asked for.

One more consequence worth keeping: conditioning can raise or lower a probability. If $P(A \mid B) > P(A)$, then $B$ makes $A$ more likely, and the relationship is symmetric, so $P(B \mid A) > P(B)$ as well. When conditioning changes nothing, the events are independent, which is the next topic.

§5

Skill Check.

Ten scenarios. Pick the chips that match your answer, then check. A scenario marks complete the first time every part is right. Progress saves on this device.

0 of 10 scenarios complete