Random Sampling & Inferences
Learn why random samples represent a population, how to scale a sample proportion up to estimate a population total, and why bigger samples give steadier estimates.
What you'll do in this lesson
A voice-first session with the Crimsora tutor on Random Sampling & Inferences, then targeted practice and FRQs — with the tutor adapting to where you get stuck.
What this lesson covers
In this lesson you will learn what makes a sample generalizable, how to turn a sample proportion into an estimate of a population total, and why a sample of 100 gives a steadier estimate than a sample of 10. You will also see why two honest random samples from the same population can give slightly different answers — and why that is normal, not a mistake.
Why Random Selection Makes a Sample Generalizable
When selection is not random, the sample carries bias — a built-in tilt in one direction. Bias does not come from bad math; it comes from who got asked. Notice that bias is not fixed by asking more people. Surveying 300 students in the cafeteria line still misses everyone who brings lunch from home.
| Sampling method | Random? | Problem |
|---|---|---|
| Draw 40 ID numbers from a hat containing all students | Yes | None; every student equally likely |
| Use a random number generator on the full student roster | Yes | None |
| Survey the first 40 students off bus 12 | No | Only bus riders represented |
| Post an online poll and use whoever answers | No | Only students who choose to reply answer |
| Survey 40 students in the band room | No | Only music students represented |
From Sample Proportion to Population Estimate
You can also set it up as a proportion equation, which many students find easier to keep straight:Cross-multiplying gives , so .
Two places students go wrong here. First, flipping the fraction — writing — which produces an answer larger than the population and should immediately look wrong. Always check that your estimate is between 0 and the population size. Second, reporting the decimal as if it were exact. You cannot have 302.4 students, and more importantly, the estimate is an approximation, so say "about 300 students" rather than "302.4 students." The word about is part of a complete answer, because a different random sample of 50 would likely have produced a slightly different number.
How Sample Size Affects Reliability
Imagine a jar with a known mix of 60 percent red beads. Repeatedly scooping samples might look like this:
| Sample size | Typical range of sample percent red |
|---|---|
| 10 | 30% to 90% |
| 25 | 44% to 76% |
| 100 | 51% to 69% |
| 400 | 55% to 65% |
When you compare two estimates, ask both questions. "Ana surveyed 15 random students, Ben surveyed 90 random students" — Ben's estimate is more reliable because sampling variability is smaller. But "Ana surveyed 15 random students, Ben surveyed 90 students on the soccer team" — now Ana's is better, because Ben's sample cannot represent non-athletes at all.
The practical takeaway: to improve an estimate, either take a bigger random sample or take several random samples and look at how much their proportions differ. If several samples cluster tightly, your estimate is trustworthy.
Using Multiple Samples to Judge an Estimate
Suppose five students each randomly sample 20 books from a library of 3,000 and count the fiction titles: 11, 13, 9, 12, 15. The proportions are , , , , and . Scaling each up to 3,000 books gives estimates of 1,650, 1,950, 1,350, 1,800, and 2,250. The spread from 1,350 to 2,250 is wide, which honestly reflects how little information 20 books carry.
A smarter move is to pool the samples. Together the students examined books and found fiction titles. That gives and an estimate of books — based on 100 observations instead of 20, so it is far more reliable.
This is exactly why pollsters report sample sizes. A result from 30 people and a result from 1,000 people are both estimates, but only one is precise enough to act on. When you write up an inference, name three things: the sample size, the fact that selection was random, and the word "about" in front of your number. Those three pieces are what make the claim defensible to someone who did not collect the data.
Key terms
- Population.
- The entire group you want to learn about, such as all 840 students in a school.
- Random sample.
- A sample chosen so that every member of the population has an equal chance of being selected.
- Generalizable.
- Describes a sample whose results can reasonably be applied to the whole population.
- Bias.
- A systematic tilt in a sample caused by how members were selected, which does not go away when the sample gets larger.
- Sample proportion.
- The fraction of the sample having a trait, found by dividing the count with the trait by the sample size.
- Inference.
- A conclusion about a population drawn from data in a sample, always stated as an approximation.
- Sampling variability.
- The natural differences among proportions from different random samples of the same population.
- Sample size.
- How many members are in the sample; larger sizes reduce sampling variability and make estimates more precise.
Worked example
Step 2: Check that the sample is random. The researcher used a random number generator on the complete voter roll, so every voter had an equal chance of selection. The sample is generalizable to the town's voters.
Step 3: Find the sample proportion.So 35 percent of the sampled voters support the park.
Step 4: Scale up to the population.Equivalently, solve , which gives and .
Step 5: State the inference correctly. About 1,575 of the town's 4,500 voters support the new park. Check reasonableness: 1,575 is between 0 and 4,500, and it is a bit more than a third of the population, matching the 35 percent found in the sample.
Step 6: Answer the sample-size question. With only 20 voters, the estimate would still be centered near the true value because selection is random, but sampling variability would be much larger. One voter changing an answer would shift the proportion by , which moves the population estimate by 225 voters. With 120 voters, one changed answer shifts the proportion by only about , moving the estimate by about 38 voters. The larger sample gives a far more precise estimate.
Practice questions
A middle school has 600 students. Which sampling plan would give the most generalizable estimate of the fraction of students who walk to school?
- Survey 200 students as they leave the gym after a basketball game
- Survey the 150 students who arrive before 7:30 a.m.
- Survey 60 students chosen using a random number generator applied to the full student roster
- Survey 250 students who volunteer to answer an online poll
Answer: Survey 60 students chosen using a random number generator applied to the full student roster
In a random sample of 80 cars passing a checkpoint, 28 were red. There are 3,200 cars registered in the area. Estimate the number of red cars, and state one thing that would make your estimate more reliable.
Answer: About 1,120 red cars; taking a larger random sample (or combining several random samples) would make the estimate more reliable.
Two students each take a random sample from the same bag of 500 marbles to estimate how many are blue. Maya samples 25 marbles and finds 10 blue. Devon samples 100 marbles and finds 36 blue. Their estimates differ. Explain why this happens and whose estimate you would trust more.
Answer: The difference is ordinary sampling variability, not a mistake; Devon's estimate is more trustworthy because his larger sample has less variability.
FAQ
- Does a bigger sample fix a biased sample?
- No. Bias comes from how members were selected, not from how many were selected. If you only survey students in the band room, surveying 300 of them instead of 30 still tells you nothing about students who are not in band. Fixing bias requires changing the selection method to give everyone in the population an equal chance.
- How big does a sample need to be?
- There is no single magic number, and this lesson does not ask you to compute one. The rule to remember is directional: larger random samples give estimates with less variability, so they are more precise. In class problems, if you are asked to compare two random samples, the larger one is the more reliable estimate.
- Why do two random samples from the same population give different answers?
- Because each sample happens to catch a different set of individuals. This is called sampling variability and it is expected, not an error. Both estimates are centered near the true value; they just land in slightly different spots. That is exactly why inferences are always stated with the word "about" instead of as exact counts.
- What is the difference between the sample proportion and the population total?
- The sample proportion is a fraction or percent, like . The population total is a count of individuals, found by multiplying that proportion by the population size, like students. If your answer to a "how many" question is less than 1, you probably stopped at the proportion and forgot to multiply.
Learn this with a teacher, not a page
The Crimsora tutor teaches Random Sampling & Inferences live — explaining on a whiteboard, asking you questions, and adapting to where you get stuck.