Mistake Master
Student view — seeing the site as a student does

Inference for Categorical Data: Proportions

Fifteen topics on the move from a sample to a claim about a population, done for categorical data. Estimators and the sampling distribution that makes them trustworthy, confidence intervals for one proportion and for the difference between two, what a confidence level is actually a property of, significance tests and what a p-value is and is not, the two kinds of error and what each one costs, and chi-square tests for two or more categories. This unit is the hinge of the course: the procedures here repeat, with new arithmetic, in Unit 4.

Exam weight 15-25%15 topics
Topics
Key forms For every problem in this unit
Statistic vs parameter
p̂ is computed from the sample; p is the unknown being estimated
Unbiased estimator
its sampling distribution is CENTERED at the parameter, not "right every time"
Spread of p̂
√(p(1 − p) / n). Quadrupling n halves it
Three conditions
randomization; 10% (sampling without replacement); large counts
Large counts, INTERVAL
n·p̂ and n(1 − p̂) both at least 10
One-proportion interval
p̂ ± z* · √(p̂(1 − p̂) / n)
Two-proportion interval
(p̂1 − p̂2) ± z* · √(p̂1(1−p̂1)/n1 + p̂2(1−p̂2)/n2). NOT pooled
Interpreting an interval
plausible values for the PARAMETER, in context. Not for the data, not for future samples
Interpreting the LEVEL
a property of the METHOD over many samples: about 95% of such intervals capture p
Judging a claim
values inside are plausible; a claim entirely outside is contradicted by the data
Hypotheses
always about the PARAMETER, stated before the data: H₀: p = p₀ against a one- or two-sided Hₑ
One-proportion z
z = (p̂ − p₀) / √(p₀(1 − p₀) / n): the null value p₀ goes in the standard error, not p̂
Large counts, TEST
n·p₀ and n(1 − p₀) both at least 10
p-value
P(a statistic at least this extreme GIVEN H₀ is true). Not P(H₀ true), not "the chance it was random"
Decision
p-value ≤ α: reject H₀. Otherwise FAIL TO REJECT: never "accept", never "prove"
Two-proportion test
the null says the proportions are equal, so POOL: p̂c = (x₁ + x₂) / (n₁ + n₂)
Type I error
rejecting a TRUE null. Its probability is α
Type II error
failing to reject a FALSE null. Its probability is β
Power
1 − β: the chance of detecting a real effect. Rises with n, with α, and with a bigger true effect
Chi-square statistic
Σ (observed − expected)² / expected, over every cell
Expected count
(row total × column total) / grand total. Condition: every EXPECTED count at least 5
Chi-square df
(rows − 1)(columns − 1). Never n − 1
Homogeneity vs independence
several populations compared on one variable, versus ONE population classified two ways
Unit 3 tools
Challenge bank
1 / 60

60 open-ended problems.

Read the question, work it out, then flip the card to compare your reasoning to the worked solution. Mark each card so you can return to the ones that still bite.

0 mastered · 0 to revisit · 60 total
Question
Tap card to reveal explanation
Worked solution
Tap card to return to question
Cumulative assessment

Test the unit.

Twenty mixed items drawn from across all 15 topics, with guaranteed misconception-code coverage. Identifies which misconceptions still bite when you cannot see which topic the question came from.

20questions
15topics
15codes covered
Begin assessment →
Course so far

Check what stuck.

Units 1 through 3, drawn evenly so earlier units get the same share as this one. Twenty questions or a full 42-question section, your choice. Even coverage means this is a retention check rather than a score estimate.

20 or 42questions
40topics
44codes covered
Begin assessment →