AP-STATS-3.7

U3.7 Inference and Generalizability

Learn how random sampling and random assignment determine the scope of inference in AP Statistics — when you can generalize and when you can claim causation.

What you'll do in this lesson

A voice-first session with the Crimsora tutor on U3.7 Inference and Generalizability, then targeted practice and FRQs — with the tutor adapting to where you get stuck.

What this lesson covers

Every statistics study ends with the same crucial question: what can we actually conclude? The answer never depends on how big the effect looks or how excited the researcher is. It depends on two design choices — whether subjects were randomly selected and whether they were randomly assigned to groups. This lesson shows you how to read a study's design and instantly determine its scope of inference. Master this and you will nail one of the most predictable question types on the exam: the "can we conclude cause?" and "can we generalize?" pair. Two independent switches, four possible verdicts — that is the whole game.

The Two Independent Switches

Scope of inference is controlled by two separate features of a study, and they answer two different questions.

Random sampling (how subjects were selected from a population) determines whether you can generalize results to that larger population. If subjects were randomly chosen, the sample is representative, so conclusions extend to the population it was drawn from. Without random sampling, conclusions apply only to the subjects actually studied.

Random assignment (how subjects were sorted into treatment groups) determines whether you can claim causation. Randomly assigning treatments balances out confounding variables on average, so any significant difference in outcomes can be attributed to the treatment itself. Without random assignment, lurking variables could explain the difference, so you can only claim association, not cause.

The key insight the exam rewards: these switches are independent. One controls generalizing, the other controls cause-and-effect. A study can have either, both, or neither. Do not let a well-run experiment trick you into generalizing to a population it never sampled from, and do not let a huge random survey trick you into claiming causation. Always evaluate the two questions separately, then combine the answers.

The Four Combinations

Because each switch is on or off, there are exactly four scenarios. Memorizing this table lets you answer almost any 3.7 question in seconds.
Random sample?Random assignment?Can generalize?Can claim cause?
YesYesYes, to populationYes
YesNoYes, to populationNo, association only
NoYesNo, only these subjectsYes, for these subjects
NoNoNo, only these subjectsNo, association only
The top-left cell is the gold standard: a randomized experiment on a randomly selected sample lets you generalize a cause-and-effect conclusion to the whole population. This is rare in practice because random sampling of people who will then submit to random assignment is logistically hard.

Most real experiments live in the bottom-left row: volunteers randomly assigned to treatments. You get causation but only for subjects like those studied. Most surveys live in the top-right row: random samples with no assignment, so you generalize a relationship but cannot say what causes it. Read the stem carefully to place each study in the correct cell.

How the Exam Phrases It

AP questions rarely use the words "random sampling" and "random assignment" for you. You must detect them from context. Phrases like "a random sample of 200 adults," "selected at random from the registry," or "randomly chosen from all students" signal random sampling, which unlocks generalization. Phrases like "subjects were randomly assigned to receive," "randomly divided into two groups," or "the treatment was randomly allocated" signal random assignment, which unlocks causation.

Watch for traps. A study that says "we surveyed 5,000 volunteers who signed up online" has a large sample but no random sampling — no generalization beyond those volunteers. A study that "compared people who chose to exercise with those who did not" is observational with no random assignment — association only, even if the sample was random.

When writing free responses, always justify with the mechanism, not just the label. For causation, say random assignment balances confounding variables. For generalization, say the random sample is representative of the population. Naming the population precisely matters too: generalize to "adults in this city," not "all people," if that is what was sampled. Vague answers lose points even when the direction is correct.

Common Misconceptions to Avoid

The single biggest error students make is thinking a large sample size fixes everything. Sample size affects the precision of estimates and the power of tests, but it does nothing for scope of inference. A million-person voluntary poll still cannot generalize to a population, and a huge observational study still cannot prove causation.

A second misconception is confusing the two switches. Students often write "random assignment means we can generalize" — this is backwards. Assignment is about causation; sampling is about generalization. Keep them straight by asking two separate questions in order: first "were groups formed randomly?" (cause), then "were subjects selected randomly?" (generalize).

A third error is over-generalizing an experiment. Randomized experiments on volunteers are extremely common. They earn causation but not generalization. The correct conclusion is a cause-and-effect statement limited to subjects similar to those in the study.

Finally, do not confuse a statistically significant result with proof of cause. Significance only tells you the difference is unlikely due to chance. Whether that difference reflects the treatment or a confounder depends entirely on random assignment. Significance and scope of inference are answered by different tools — the p-value versus the study design.

Key terms

Random sampling.
Selecting subjects from a population using chance so every unit has a known, nonzero probability of selection; it makes the sample representative and permits generalization.
Random assignment.
Using chance to allocate subjects to treatment groups; it balances confounding variables on average and permits cause-and-effect conclusions.
Scope of inference.
The set of conclusions justified by a study's design — specifically whether results generalize to a population and whether a causal claim is warranted.
Generalization.
Extending conclusions from a sample to the larger population it was drawn from; requires random sampling.
Causation.
A conclusion that a treatment produces a change in the response; requires random assignment to rule out confounding.
Confounding variable.
A variable associated with both the explanatory and response variables, making it impossible to isolate the treatment's effect in observational studies.
Observational study.
A study that measures subjects without imposing treatments; can show association but not causation because groups are not randomly assigned.

Worked example

A researcher emails a survey to a random sample of 800 employees at a large company. Among respondents, those who reported drinking coffee daily had higher self-rated productivity than non-coffee-drinkers. The researcher concludes that coffee causes higher productivity for employees nationwide. Evaluate this conclusion.
First, ask about random assignment to judge causation. The employees chose whether to drink coffee; the researcher did not randomly assign coffee versus no coffee. This is an observational study, so confounding variables — such as workload, sleep, or job type — could explain the difference. Therefore no cause-and-effect conclusion is warranted; only an association between coffee drinking and higher self-rated productivity can be claimed.

Second, ask about random sampling to judge generalization. The 800 employees were randomly sampled, but only from one large company, not from the national workforce. So results may generalize to employees at that company, but not to employees nationwide.

Combining both: the correct conclusion is that among employees at this company there is an association between daily coffee drinking and higher self-rated productivity, but we cannot conclude coffee causes it, and we cannot extend the finding to all employees nationwide. The researcher's conclusion overreaches on both counts — claiming causation without random assignment and generalizing beyond the sampled population.

Practice questions

Researchers randomly assigned 60 volunteer college students to either a standing desk or a sitting desk for two weeks and measured back pain. The standing group reported significantly less pain. Which conclusion is best supported?
  1. Standing desks reduce back pain for all college students
  2. Standing desks reduce back pain for students similar to these volunteers
  3. Standing desks are associated with less back pain, but no cause can be claimed
  4. Standing desks reduce back pain for all adults nationwide

Answer: Standing desks reduce back pain for students similar to these volunteers

Random assignment was used, so a cause-and-effect conclusion is justified — the standing desk reduced pain. But the subjects were volunteers, not a random sample, so the causal claim cannot be generalized beyond students similar to those studied. Options generalizing to all college students or all adults overreach, and the association-only option ignores that random assignment permits causation.
A polling organization takes a random sample of 1,500 registered voters in a state and finds that 62% support a new transportation bill. Explain what scope of inference is justified and why.

Answer: You can generalize the 62% support estimate to all registered voters in that state, but you cannot make any cause-and-effect claim.

Random sampling from the state's registered voters means the sample is representative, so the result generalizes to that population — registered voters in the state, not the whole country. Because this is a survey with no treatment and no random assignment, there is no causal question at all; it simply estimates a population proportion. A full-credit response names the specific population and notes that no causation is involved.
A gym recruits members who volunteer for a study and randomly assigns half to a new workout plan and half to their usual routine, then compares fitness gains. Why can the researcher claim causation but not generalize broadly?

Answer: Random assignment justifies causation, but the volunteer sample prevents generalization beyond members like those studied.

Random assignment balances confounding variables across the two groups, so a significant difference in fitness gains can be attributed to the workout plan — causation is valid. However, the subjects were self-selected gym volunteers, not a random sample of any larger population, so the conclusion applies only to people similar to those volunteers, not to all gym members or the general public.

FAQ

Does random assignment let me generalize to a population?
No. Random assignment only justifies cause-and-effect conclusions by balancing confounding variables. Generalizing to a population requires random sampling. These are two independent design features that answer two different questions.
If a study has a very large sample size, can I generalize even without random sampling?
No. Sample size affects precision and power, not scope of inference. A huge voluntary or convenience sample can still be badly unrepresentative, so its results cannot be generalized to a population. Only random sampling justifies generalization.
Can an observational study ever prove causation?
Not on its own. Without random assignment, confounding variables may explain any observed difference, so observational studies show association only. You can strengthen causal arguments by controlling for confounders, but the AP exam expects you to say observational studies cannot establish causation.
How do I state the scope of inference on a free-response question?
Answer two separate questions. For causation, check for random assignment and mention that it balances confounders. For generalization, check for random sampling and name the specific population it represents. Combine both into one precise sentence limited to what the design supports.

Learn this with a teacher, not a page

The Crimsora tutor teaches U3.7 Inference and Generalizability live — explaining on a whiteboard, asking you questions, and adapting to where you get stuck.