Mistake Master
Home Unit 1 · Exploring One-Variable Data and Collecting Data 1.1·1.2·1.3·1.4·1.5·1.6·1.7·1.8·1.9·1.10·1.11·1.12·1.13 Lesson
Skill Check 0 / 10 complete

Where samples go wrong

A biased sampling method does not miss the truth at random: it misses in the same direction every time. That is what makes bias different from bad luck, and it is why the standard rescue, collect more data, fails completely. More data from a tilted method is just a sharper picture of the wrong number.

§1

Bias is a property of the method, not of one sample.

Bias in a sampling method is a systematic error in the procedure that makes the resulting statistic consistently larger or consistently smaller than the parameter it estimates. The word systematic is doing all the work. A fair method can produce an unlucky sample; run it again and the errors scatter in both directions and average away. A biased method tilts the same way on every run, so repeating it, or enlarging it, only piles up estimates around the wrong value.

That is the whole case against "just ask more people". Sample size shrinks the scatter around whatever the method targets. If the method targets the wrong value, a huge sample delivers the wrong value with great precision. Ten thousand responses to a tilted poll are not ten thousand pieces of evidence; they are one mistake, measured very carefully.

Because bias has a direction, a full description of it always states one: the estimate is likely too high, or too low, because of who was missed or how answers were bent. On the AP exam, naming the direction and the reason is the difference between a scored answer and a vocabulary word.

§2

Undercoverage: part of the population never had a chance.

Undercoverage occurs when the sampling method leaves out part of the population, or gives part of it a reduced chance of selection. The randomness can be flawless within the frame; the trouble is who never made it into the frame at all. A telephone survey using only landlines random-dials perfectly, and still cannot reach the mostly younger adults who have no landline.

Undercoverage matters exactly when the missed group differs on the question asked. Landline-only sampling barely distorts a question about garden pests; it badly distorts a question about streaming habits, because the excluded group streams the most, and the estimate lands too low.

Convenience sampling, selecting whoever is easy to reach, is undercoverage at its bluntest: everyone not at the mall, the gym, or the front of the line has selection probability zero. However many easy people are measured, the hard-to-reach remain unmeasured, and they were different in whatever way made them hard to reach.

§3

Voluntary response and nonresponse: who opts in, who drops out.

Voluntary response bias arises when the sample consists of people who chose themselves: call-in polls, open web polls, clip-out ballots. Opting in costs effort, and effort is spent by people with strong feelings, usually the aggrieved. The sample fills with intensity, and the estimate slides toward the passionate side of the issue.

Nonresponse bias is the mirror problem. The sample was properly selected, but some of the chosen individuals cannot be reached or refuse to answer, and the respondents differ from the nonrespondents in ways that matter to the study. A mailed survey about a park returned by 30 percent of recipients hears from the park's devoted users; the indifferent majority threw it away, and measured enthusiasm runs high.

Keep the two apart by asking who assembled the sample. In voluntary response, the volunteers did: nobody was selected. In nonresponse, the researcher selected properly and then lost people. The AP exam rewards using the right name, and punishes claims like "nonresponse bias, because people chose to respond", which mixes the two mechanisms into neither.

§4

Response bias: the answers themselves are bent.

Even with a perfect frame and full participation, the recorded values can drift from the truth in one direction. Response bias covers systematic distortion in the answers themselves. Its two staple sources:

  1. Question wording. A leading or loaded question pushes respondents toward the phrasing. "Do you support wasting taxpayer money on a new stadium" and "Do you support investing in a new stadium" measure two different quantities, neither of them plain support.
  2. Self-report. People shade answers toward what is flattering or expected: screen time and calories get underreported, exercise and reading get inflated, sensitive behaviors vanish when the interviewer is watching.

The same direction-and-reason discipline applies: say which way the wording or the self-interest pushes, and why. And note what randomness cannot do here: random selection decides who is asked, not how truthfully they answer. A flawless SRS asked a loaded question returns a beautifully precise distortion.

§5

Skill Check.

Ten scenarios. Pick the chips that match your answer, then check. A scenario marks complete the first time every part is right. Progress saves on this device.

0 of 10 scenarios complete