M7MATH-9.2

Random Sampling & Inferences

Learn why random samples represent a population, how to scale a sample proportion up to estimate a population total, and why bigger samples give steadier estimates.

What you'll do in this lesson

A voice-first session with the Crimsora tutor on Random Sampling & Inferences, then targeted practice and FRQs — with the tutor adapting to where you get stuck.

What this lesson covers

Suppose your school has 840 students and the principal wants to know how many would ride a late bus. Asking all 840 takes forever. Asking the 30 students standing at the bus stop is fast — but those students already ride buses, so the answer will be way too high. The fix is random selection: give every student an equal chance of being picked, then scale up what the sample tells you.

In this lesson you will learn what makes a sample generalizable, how to turn a sample proportion into an estimate of a population total, and why a sample of 100 gives a steadier estimate than a sample of 10. You will also see why two honest random samples from the same population can give slightly different answers — and why that is normal, not a mistake.

Why Random Selection Makes a Sample Generalizable

A sample is generalizable when what is true in the sample is probably close to what is true in the whole population. Random selection is what earns you that permission. In a random sample, every member of the population has an equal chance of being chosen, so no group gets systematically over-represented.

When selection is not random, the sample carries bias — a built-in tilt in one direction. Bias does not come from bad math; it comes from who got asked. Notice that bias is not fixed by asking more people. Surveying 300 students in the cafeteria line still misses everyone who brings lunch from home.
Sampling methodRandom?Problem
Draw 40 ID numbers from a hat containing all studentsYesNone; every student equally likely
Use a random number generator on the full student rosterYesNone
Survey the first 40 students off bus 12NoOnly bus riders represented
Post an online poll and use whoever answersNoOnly students who choose to reply answer
Survey 40 students in the band roomNoOnly music students represented
A common mistake is thinking "random" means "whatever happens to be convenient." Standing in one hallway and grabbing whoever walks by feels unplanned, but it is not random — students who never use that hallway had zero chance of being selected. Real randomness requires a list of the whole population and a fair chance device: a spinner, dice, drawn slips, or a random number generator.

From Sample Proportion to Population Estimate

Once you trust the sample, the arithmetic is short. First find the sample proportion:p=number in sample with the traitsample sizep = \frac{\text{number in sample with the trait}}{\text{sample size}}Then assume the population has roughly the same proportion and multiply by the population size:estimated population total≈p×N\text{estimated population total} \approx p \times NExample: in a random sample of 50 students from a school of 840, 18 say they would ride a late bus. Then p=1850=0.36p = \frac{18}{50} = 0.36, and the estimate is 0.36×840=302.40.36 \times 840 = 302.4, so about 302 students.

You can also set it up as a proportion equation, which many students find easier to keep straight:1850=x840\frac{18}{50} = \frac{x}{840}Cross-multiplying gives 50x=1512050x = 15120, so x=302.4x = 302.4.

Two places students go wrong here. First, flipping the fraction — writing 5018\frac{50}{18} — which produces an answer larger than the population and should immediately look wrong. Always check that your estimate is between 0 and the population size. Second, reporting the decimal as if it were exact. You cannot have 302.4 students, and more importantly, the estimate is an approximation, so say "about 300 students" rather than "302.4 students." The word about is part of a complete answer, because a different random sample of 50 would likely have produced a slightly different number.

How Sample Size Affects Reliability

Every random sample gives a slightly different proportion. That wobble is called sampling variability, and it is not an error — it is what happens when you look at part of a whole. The key idea is that larger samples wobble less.

Imagine a jar with a known mix of 60 percent red beads. Repeatedly scooping samples might look like this:
Sample sizeTypical range of sample percent red
1030% to 90%
2544% to 76%
10051% to 69%
40055% to 65%
The center stays near the true 60 percent no matter the sample size — that is what randomness buys you. What shrinks is the spread. So sample size controls precision, while random selection controls accuracy. A huge biased sample is still wrong, just confidently wrong. A small random sample points the right direction but loosely.

When you compare two estimates, ask both questions. "Ana surveyed 15 random students, Ben surveyed 90 random students" — Ben's estimate is more reliable because sampling variability is smaller. But "Ana surveyed 15 random students, Ben surveyed 90 students on the soccer team" — now Ana's is better, because Ben's sample cannot represent non-athletes at all.

The practical takeaway: to improve an estimate, either take a bigger random sample or take several random samples and look at how much their proportions differ. If several samples cluster tightly, your estimate is trustworthy.

Using Multiple Samples to Judge an Estimate

Classrooms often gather many samples at once — each student takes their own random sample of 20 and the class compares results. This turns an invisible idea into something you can see.

Suppose five students each randomly sample 20 books from a library of 3,000 and count the fiction titles: 11, 13, 9, 12, 15. The proportions are 0.550.55, 0.650.65, 0.450.45, 0.600.60, and 0.750.75. Scaling each up to 3,000 books gives estimates of 1,650, 1,950, 1,350, 1,800, and 2,250. The spread from 1,350 to 2,250 is wide, which honestly reflects how little information 20 books carry.

A smarter move is to pool the samples. Together the students examined 5×20=1005 \times 20 = 100 books and found 11+13+9+12+15=6011 + 13 + 9 + 12 + 15 = 60 fiction titles. That gives p=60100=0.60p = \frac{60}{100} = 0.60 and an estimate of 0.60×3000=18000.60 \times 3000 = 1800 books — based on 100 observations instead of 20, so it is far more reliable.

This is exactly why pollsters report sample sizes. A result from 30 people and a result from 1,000 people are both estimates, but only one is precise enough to act on. When you write up an inference, name three things: the sample size, the fact that selection was random, and the word "about" in front of your number. Those three pieces are what make the claim defensible to someone who did not collect the data.

Key terms

Population.
The entire group you want to learn about, such as all 840 students in a school.
Random sample.
A sample chosen so that every member of the population has an equal chance of being selected.
Generalizable.
Describes a sample whose results can reasonably be applied to the whole population.
Bias.
A systematic tilt in a sample caused by how members were selected, which does not go away when the sample gets larger.
Sample proportion.
The fraction of the sample having a trait, found by dividing the count with the trait by the sample size.
Inference.
A conclusion about a population drawn from data in a sample, always stated as an approximation.
Sampling variability.
The natural differences among proportions from different random samples of the same population.
Sample size.
How many members are in the sample; larger sizes reduce sampling variability and make estimates more precise.

Worked example

A town has 4,500 registered voters. A researcher uses a random number generator on the voter roll to select 120 voters and finds that 42 of them support building a new park. Estimate how many voters in the town support the park. Then explain what would change if the researcher had sampled only 20 voters instead.
Step 1: Identify the population size and the sample. The population is N=4500N = 4500 voters. The sample size is 120, and 42 of them support the park.

Step 2: Check that the sample is random. The researcher used a random number generator on the complete voter roll, so every voter had an equal chance of selection. The sample is generalizable to the town's voters.

Step 3: Find the sample proportion.p=42120=0.35p = \frac{42}{120} = 0.35So 35 percent of the sampled voters support the park.

Step 4: Scale up to the population.0.35×4500=15750.35 \times 4500 = 1575Equivalently, solve 42120=x4500\frac{42}{120} = \frac{x}{4500}, which gives 120x=189000120x = 189000 and x=1575x = 1575.

Step 5: State the inference correctly. About 1,575 of the town's 4,500 voters support the new park. Check reasonableness: 1,575 is between 0 and 4,500, and it is a bit more than a third of the population, matching the 35 percent found in the sample.

Step 6: Answer the sample-size question. With only 20 voters, the estimate would still be centered near the true value because selection is random, but sampling variability would be much larger. One voter changing an answer would shift the proportion by 120=0.05\frac{1}{20} = 0.05, which moves the population estimate by 225 voters. With 120 voters, one changed answer shifts the proportion by only about 0.0080.008, moving the estimate by about 38 voters. The larger sample gives a far more precise estimate.

Practice questions

A middle school has 600 students. Which sampling plan would give the most generalizable estimate of the fraction of students who walk to school?
  1. Survey 200 students as they leave the gym after a basketball game
  2. Survey the 150 students who arrive before 7:30 a.m.
  3. Survey 60 students chosen using a random number generator applied to the full student roster
  4. Survey 250 students who volunteer to answer an online poll

Answer: Survey 60 students chosen using a random number generator applied to the full student roster

Only the random number generator plan gives every one of the 600 students an equal chance of being selected. The other three plans are biased no matter how many students they include: gym-exit and early-arrival groups over-represent particular kinds of students, and a volunteer poll only captures students who choose to respond. Notice that the biased samples are all larger than 60 — this is the key point that bigger does not fix bias. Sample size improves precision; random selection is what makes an estimate point at the right value in the first place.
In a random sample of 80 cars passing a checkpoint, 28 were red. There are 3,200 cars registered in the area. Estimate the number of red cars, and state one thing that would make your estimate more reliable.

Answer: About 1,120 red cars; taking a larger random sample (or combining several random samples) would make the estimate more reliable.

The sample proportion is p=2880=0.35p = \frac{28}{80} = 0.35. Multiplying by the population, 0.35×3200=11200.35 \times 3200 = 1120, so about 1,120 cars are red. You could also solve 2880=x3200\frac{28}{80} = \frac{x}{3200}. To improve reliability, increase the sample size, because sampling variability shrinks as the sample grows — a sample of 400 cars would produce proportions clustered much more tightly around the true value than samples of 80 do. Combining results from several random samples of 80 accomplishes the same thing, since pooling 5 such samples gives you 400 observations.
Two students each take a random sample from the same bag of 500 marbles to estimate how many are blue. Maya samples 25 marbles and finds 10 blue. Devon samples 100 marbles and finds 36 blue. Their estimates differ. Explain why this happens and whose estimate you would trust more.

Answer: The difference is ordinary sampling variability, not a mistake; Devon's estimate is more trustworthy because his larger sample has less variability.

Maya's proportion is 1025=0.40\frac{10}{25} = 0.40, giving an estimate of 0.40×500=2000.40 \times 500 = 200 blue marbles. Devon's is 36100=0.36\frac{36}{100} = 0.36, giving 0.36×500=1800.36 \times 500 = 180. Both used random selection, so both estimates are aimed at the true value — the gap of 20 marbles is just sampling variability, the natural difference between random samples. Devon's estimate is more reliable because with 100 marbles each individual marble changes the proportion by only 0.010.01, while in Maya's sample each marble changes it by 0.040.04. Larger samples produce proportions that cluster more tightly around the truth.

FAQ

Does a bigger sample fix a biased sample?
No. Bias comes from how members were selected, not from how many were selected. If you only survey students in the band room, surveying 300 of them instead of 30 still tells you nothing about students who are not in band. Fixing bias requires changing the selection method to give everyone in the population an equal chance.
How big does a sample need to be?
There is no single magic number, and this lesson does not ask you to compute one. The rule to remember is directional: larger random samples give estimates with less variability, so they are more precise. In class problems, if you are asked to compare two random samples, the larger one is the more reliable estimate.
Why do two random samples from the same population give different answers?
Because each sample happens to catch a different set of individuals. This is called sampling variability and it is expected, not an error. Both estimates are centered near the true value; they just land in slightly different spots. That is exactly why inferences are always stated with the word "about" instead of as exact counts.
What is the difference between the sample proportion and the population total?
The sample proportion is a fraction or percent, like 1850=0.36\frac{18}{50} = 0.36. The population total is a count of individuals, found by multiplying that proportion by the population size, like 0.36×840≈3020.36 \times 840 \approx 302 students. If your answer to a "how many" question is less than 1, you probably stopped at the proportion and forgot to multiply.

Learn this with a teacher, not a page

The Crimsora tutor teaches Random Sampling & Inferences live — explaining on a whiteboard, asking you questions, and adapting to where you get stuck.