U5.5 Sampling Distribution of p̂
Master the sampling distribution of p̂: find its mean and SD, check Large Counts and 10% conditions, and compute normal probabilities for p̂ and p̂₁−p̂₂.
What you'll do in this lesson
A voice-first session with the Crimsora tutor on U5.5 Sampling Distribution of p̂, then targeted practice and FRQs — with the tutor adapting to where you get stuck.
What this lesson covers
When you take a sample and calculate a proportion — the fraction of voters favoring a candidate, the share of defective parts — that sample proportion changes from sample to sample. The sampling distribution of describes exactly how it varies, and once you know its center, spread, and shape, you can answer questions like "How likely is a sample proportion this far from the truth?"
This lesson gives you the three-step toolkit the exam rewards: state the mean and standard deviation, verify the conditions that let you treat the distribution as approximately Normal, and then compute a probability with a -score. We finish by extending everything to the difference of two proportions .
This lesson gives you the three-step toolkit the exam rewards: state the mean and standard deviation, verify the conditions that let you treat the distribution as approximately Normal, and then compute a probability with a -score. We finish by extending everything to the difference of two proportions .
Mean and Standard Deviation of p̂
Suppose a population has true proportion of successes, and you draw a simple random sample of size . The count of successes follows a binomial model, and the sample proportion is . Its sampling distribution has a predictable center and spread.
The mean is . This says is an unbiased estimator of — on average, across all possible samples, it hits the true value.
The standard deviation is . Notice two things. First, larger shrinks the standard deviation: bigger samples give more precise estimates. Second, spread depends on itself, and is largest when .
A common exam trap: this formula uses the true , not . In this unit you are given , so plug it in directly. (Later, in inference, when is unknown you will substitute and call it the standard error — but that is a different setting.)
Always state the mean and SD together before doing any probability work. Full-credit FRQ responses name the distribution and show both values with the formula substituted, not just the final number.
The mean is . This says is an unbiased estimator of — on average, across all possible samples, it hits the true value.
The standard deviation is . Notice two things. First, larger shrinks the standard deviation: bigger samples give more precise estimates. Second, spread depends on itself, and is largest when .
A common exam trap: this formula uses the true , not . In this unit you are given , so plug it in directly. (Later, in inference, when is unknown you will substitute and call it the standard error — but that is a different setting.)
Always state the mean and SD together before doing any probability work. Full-credit FRQ responses name the distribution and show both values with the formula substituted, not just the final number.
Checking the Conditions for Normality
The mean and SD are always correct, but you may only treat the shape as approximately Normal after checking two conditions.
Large Counts condition: both and . You need at least 10 expected successes AND at least 10 expected failures. If either count falls short, the distribution is too skewed to model with the Normal curve.
10% condition: when sampling without replacement, the sample size must be no more than 10% of the population, . This keeps the observations approximately independent so the SD formula stays valid.
On the exam, show the arithmetic, not just the words. Write " and " to earn the point. A frequent error is checking only one part of Large Counts — you must verify both successes and failures.
Large Counts condition: both and . You need at least 10 expected successes AND at least 10 expected failures. If either count falls short, the distribution is too skewed to model with the Normal curve.
10% condition: when sampling without replacement, the sample size must be no more than 10% of the population, . This keeps the observations approximately independent so the SD formula stays valid.
| Condition | Requirement | What it protects |
|---|---|---|
| Large Counts | and | Normal shape |
| 10% | Independence / SD formula |
Computing Probabilities for p̂
Once the conditions hold, is approximately . Finding a probability is then a standard Normal calculation.
Standardize with , then find the area under the standard Normal curve. The -score answers "how many standard deviations is this sample proportion from the true proportion?"
A clean four-step process earns full credit: state the shape, mean, and SD; check both conditions; compute the -score; find and interpret the probability. Draw a quick sketch of the Normal curve with the region shaded — graders and your own accuracy both benefit.
Watch the direction of the inequality. is the upper tail, so subtract the table value from 1, or use the calculator with the correct bounds. Also keep more decimal places in during intermediate steps; rounding early can shift the -score enough to change the answer.
Standardize with , then find the area under the standard Normal curve. The -score answers "how many standard deviations is this sample proportion from the true proportion?"
A clean four-step process earns full credit: state the shape, mean, and SD; check both conditions; compute the -score; find and interpret the probability. Draw a quick sketch of the Normal curve with the region shaded — graders and your own accuracy both benefit.
Watch the direction of the inequality. is the upper tail, so subtract the table value from 1, or use the calculator with the correct bounds. Also keep more decimal places in during intermediate steps; rounding early can shift the -score enough to change the answer.
The Difference of Two Proportions p̂₁ − p̂₂
When you compare two independent groups, the statistic is . Its sampling distribution combines the two individual ones.
The mean is the difference of the true proportions: .
Because the samples are independent, variances add. So . A key insight students miss: you add the variances (the terms under the root), never the standard deviations directly.
Check Large Counts for each sample separately — four inequalities total: , , , all . Check the 10% condition for each sample too. When all hold, is approximately Normal.
To find a probability such as , standardize with . The value 0 corresponds to "the two sample proportions are equal," a question that appears often.
The mean is the difference of the true proportions: .
Because the samples are independent, variances add. So . A key insight students miss: you add the variances (the terms under the root), never the standard deviations directly.
Check Large Counts for each sample separately — four inequalities total: , , , all . Check the 10% condition for each sample too. When all hold, is approximately Normal.
To find a probability such as , standardize with . The value 0 corresponds to "the two sample proportions are equal," a question that appears often.
Key terms
- Sample proportion .
- The fraction of successes in a sample, ; a statistic that varies from sample to sample.
- Sampling distribution of .
- The distribution of values takes across all possible samples of size from a population.
- Unbiased estimator.
- A statistic whose sampling-distribution mean equals the parameter it estimates; .
- Large Counts condition.
- Requirement that and , ensuring the sampling distribution of is approximately Normal.
- 10% condition.
- Requirement that when sampling without replacement, keeping observations approximately independent.
- Standard deviation of .
- The spread of the sampling distribution, , computed with the true proportion .
- z-score.
- The standardized distance used to find Normal probabilities.
- .
- The difference of two independent sample proportions, with mean and variance equal to the sum of the two individual variances.
Worked example
A city claims 40% of its residents recycle regularly. A researcher takes a simple random sample of 150 residents. Assuming the claim is true, find the probability that the sample proportion of recyclers is greater than 0.47. The city has 80,000 residents.
Step 1 — State the distribution. Here and . The mean is . The standard deviation is .
Step 2 — Check conditions. Large Counts: and , both satisfied. 10% condition: , satisfied. So is approximately Normal.
Step 3 — Standardize. .
Step 4 — Find the probability. .
Conclusion: If the true recycling rate is 40%, there is about a 0.040 probability that a random sample of 150 residents yields a sample proportion above 0.47. This is fairly unlikely, so such a sample would be mild evidence against the city's claim.
Step 2 — Check conditions. Large Counts: and , both satisfied. 10% condition: , satisfied. So is approximately Normal.
Step 3 — Standardize. .
Step 4 — Find the probability. .
Conclusion: If the true recycling rate is 40%, there is about a 0.040 probability that a random sample of 150 residents yields a sample proportion above 0.47. This is fairly unlikely, so such a sample would be mild evidence against the city's claim.
Practice questions
A population has proportion . Which sample size makes the standard deviation of smallest?
Answer:
The standard deviation is , which decreases as increases. The largest sample size, , gives the smallest spread. Larger samples produce more precise, less variable estimates of .
In a large batch of manufactured chips, 8% are defective. A quality inspector randomly samples 60 chips. Explain whether it is appropriate to use a Normal approximation for the sampling distribution of , and justify with the relevant condition.
Answer: No, the Normal approximation is not appropriate because the Large Counts condition fails.
Check Large Counts: , which is less than 10. Although , both parts must be at least 10. Since the expected number of defectives is only 4.8, the distribution of is right-skewed and should not be modeled as Normal.
Two independent SRSs are taken. In group 1, with ; in group 2, with . Find the mean and standard deviation of the sampling distribution of .
Answer: Mean ; standard deviation .
The mean is . Variances add: . The standard deviation is . Note you add variances, then take the square root — never add the two standard deviations directly.
FAQ
- When do I use p versus p̂ in the standard deviation formula?
- In this unit (sampling distributions) you are given the true population proportion , so use it: . Only later, in inference when is unknown, do you substitute and call the result the standard error.
- Why do I add variances but not standard deviations for the difference of two proportions?
- Variances of independent random variables add, but standard deviations do not. So you compute each variance, sum them, and only then take the square root: .
- What happens if the Large Counts condition fails?
- If or , the sampling distribution of is too skewed to model with a Normal curve. You should not compute probabilities using -scores; the binomial model or a note that the shape is skewed is the honest answer.
- Do I always have to check the 10% condition?
- Check it whenever you sample without replacement, which is nearly always in practice. It confirms so observations stay approximately independent and the SD formula remains valid. On the exam, showing this arithmetic is required for full credit.
Learn this with a teacher, not a page
The Crimsora tutor teaches U5.5 Sampling Distribution of p̂ live — explaining on a whiteboard, asking you questions, and adapting to where you get stuck.