Mistake Master
Home Unit 1 · Exploring One-Variable Data and Collecting Data 1.1·1.2·1.3·1.4·1.5·1.6·1.7·1.8·1.9·1.10·1.11·1.12·1.13 Lesson
Skill Check 0 / 10 complete

From question to data plan

Before any data are collected, the investigative question has already made three promises: what will be measured, how it will be analyzed, and who the answer will be about. The design you choose either keeps those promises or quietly breaks them, and the two uses of randomness, selection and assignment, are what separate a claim about a population from a claim about cause.

§1

The question comes first, and it names the plan.

An investigative question worth collecting data for has three working parts. First, it is phrased in terms of the variables of interest, which tells you exactly what to measure. "Do students get enough sleep?" measures nothing; "What is the mean nightly sleep time of students at this school?" tells you to record hours per night. Second, it points at the analysis: a question about estimating a value calls for an interval around a parameter, while a question asking whether one group differs from another calls for a test with a stated direction. Third, it names the population the conclusion should reach, and the kind of conclusion, descriptive or cause-and-effect, that the study is being built to support.

Read the three parts back and the data collection plan almost writes itself: what to record, from whom, and by what selection mechanism. When a study goes wrong at the design stage, it is usually because one of these parts was never pinned down.

§2

Census, survey, observational study, experiment: four ways to collect.

A census records information from every item or individual in the population. It answers the question outright, with no inference needed, and it is usually impractical: measuring all 2,000 students is a census, measuring 50 of them is a sample, and those are different claims with different strengths.

An observational study records the values of variables of interest without imposing anything on anyone. A sample survey is the human version: an observational study that collects data from people using a standard set of questions. Observational studies come in two time directions: a prospective study selects units now and gathers data going forward, while a retrospective study selects units now and gathers data from the past.

An experiment is the one design where the researcher intervenes: conditions, called treatments, are assigned to experimental units (called subjects or participants when they are people). The variable whose levels are imposed is the explanatory variable, and the outcome measured afterward on each unit is the response variable. The assignment is the defining act: choosing whom to study is not an experiment, imposing a condition on them is.

§3

Two randomnesses license two different conclusions.

Randomness enters a study in two places, and they do different jobs:

  1. Random selection decides who gets into the sample, using a chance mechanism such as a random number generator. It makes the sample representative of the population it was drawn from, so the results can be generalized to that population.
  2. Random assignment decides which treatment each unit receives. It balances the treatment groups, so a difference in the response can be attributed to the treatment: a cause-and-effect conclusion.

A study can have either, both, or neither. Volunteers randomly assigned to treatments support a causal claim, but only about people similar to those volunteers. A random sample that is merely observed supports a generalization, but not a causal claim. Citing the wrong randomness for a conclusion is one of the most reliably penalized errors on the AP exam.

When units are not randomly selected, because they were deliberately chosen or volunteered themselves, generalization shrinks to a population of individuals similar to those studied. And the number computed from a sample is a statistic: an estimate of the population's parameter, never the parameter itself. "42 percent of our sample" and "42 percent of the school" are different sentences, and only one of them was measured.

§4

Confounding is why observation cannot establish cause.

In an observational study, the groups being compared assembled themselves, and whatever made them different in the first place travels with them. A confounding variable is a variable associated with both the explanatory variable and the response variable, which lets it provide an alternative explanation for the observed relationship.

Both links are required, and naming them in the specific study is what earns credit. "Maybe other stuff matters" identifies nothing. "Children who eat breakfast tend to have more involved parents, and parental involvement also tends to raise test scores" names one variable and traces both of its connections, so the breakfast-and-scores association no longer forces a causal reading.

This is not a defect to be patched with a bigger sample. However many units you observe, the self-assembled groups still differ in the same systematic ways. The only design that breaks the link between the explanatory variable and everything tangled with it is random assignment, which is the subject of Topic 1.13.

§5

Skill Check.

Ten scenarios. Pick the chips that match your answer, then check. A scenario marks complete the first time every part is right. Progress saves on this device.

0 of 10 scenarios complete