Populations & Samples
Learn to name the population and sample in a statistical question, see why surveying everyone is rarely possible, and spot biased sampling methods fast.
What you'll do in this lesson
A voice-first session with the Crimsora tutor on Populations & Samples, then targeted practice and FRQs — with the tutor adapting to where you get stuck.
What this lesson covers
This lesson teaches you to read a statistical question and name the population and the sample precisely, explain why samples are used instead of counting everybody, and decide whether a sampling method fairly reflects the group or tilts the results. That last skill matters most. A sample of 50 students grabbed from the after-school robotics club will not tell you about the whole school, no matter how carefully you compute the average. Getting the sample right comes before any arithmetic.
Population, Sample, and the Question Behind Them
Compare these two questions about the same school. "What is the average height of students at Kennedy Middle School?" has a population of all Kennedy students. "What is the average height of seventh-grade students at Kennedy Middle School?" has a population of only the seventh-grade class. The same 40 measured students could be a sample of either population, depending on which question you asked.
A useful habit: write the population as a complete phrase starting with "all." All registered voters in the county. All boxes of cereal produced by the factory on Tuesday. All bass in Lake Monroe. Vague answers like "students" or "the school" are where students lose accuracy, because a population must have clear boundaries — you should be able to tell whether any given individual is in it or not.
The sample is described the same way, but with the how included: "the 40 students whose names were drawn from a hat" or "the 12 boxes pulled from the end of the assembly line." Naming how the sample was chosen is not extra decoration; it is the information you need to judge whether the sample can be trusted.
| Term | What it is | Example |
|---|---|---|
| Population | Whole group of interest | All 840 students at the school |
| Sample | Part actually studied | 50 students who were surveyed |
| Census | Data from every member | Surveying all 840 students |
Why Sample Instead of Taking a Census?
Cost and time: a company that wants to know how long its batteries last cannot test millions of batteries. Physical impossibility: no one can catch and measure every fish in a lake, and populations change as new fish hatch. Destruction: crash-testing every car, or opening every bag of chips to weigh the contents, leaves nothing to sell. Access: you cannot reach every adult in a country, and many would refuse to answer anyway.
Because of these limits, statisticians use a well-chosen sample and then make an inference — a reasoned conclusion about the population based on the sample. If 18 of your 50 surveyed students walk to school, that is 36 percent of the sample, so a reasonable estimate is that about 36 percent of all 840 students walk, roughly 302 students.
The trade-off is honest to state: a sample gives you an estimate, not a certainty. A different random sample of 50 would probably give a slightly different percentage. That normal wobble is called sampling variability, and it is the price of not counting everyone. What a good sampling method promises is not a perfect answer but an answer that is close, with errors that go both directions rather than always leaning one way.
Representative Samples and Biased Samples
The key idea is that bias comes from the method, not from bad luck or from a small size. A sample of 200 people collected outside a gym is biased about exercise habits; a random sample of 30 people from the same town is not biased, just less precise. Ask yourself: does every member of the population have a fair chance of being selected? If some group is shut out or oversampled, bias is present.
Watch for three common sources. Convenience sampling takes whoever is easiest to reach — the people already in the lunch line, or your own friends. Voluntary response lets people choose to participate, and those with strong opinions respond far more often. Location or timing bias hides in wording like "surveyed at the skate park" or "called at 10 a.m. on a weekday."
| Method described | Representative? | Why |
|---|---|---|
| Names drawn from a list of all students | Yes | Everyone has an equal chance |
| Every 10th student off the bus, all buses | Usually yes | No group is systematically skipped |
| Students at the basketball game | No | Sports fans are overrepresented |
| Online poll anyone may answer | No | Only strongly opinionated people reply |
Where Students Go Wrong
Swapping population and sample. The population is always the larger group you want to describe; the sample is the smaller group you actually measured. If a problem says "a scientist tags 60 turtles to estimate the number of turtles in the marsh," the 60 turtles are the sample and all marsh turtles are the population — even though only the 60 were ever counted.
Thinking bigger always means better. Size improves precision but cannot repair a flawed method. Doubling a survey taken only at the mall gives you twice as much lopsided information. Fix the selection process first, then worry about size.
Blaming the question instead of the selection. Confusing wording ("Don't you agree that recess is too short?") is a real problem, but it is separate from sampling bias. This lesson asks who got chosen; leading questions are a different flaw worth mentioning separately.
Refusing to make any inference. Some students write "you can't know anything from a sample." That is too strong. If the sample was chosen fairly, you can and should estimate — carefully, with words like "about" and "approximately." Say "about 36 percent of all students walk to school," not "exactly 302 students walk."
One more habit worth building: when a survey result surprises you, ask how the data were collected before you argue about what the number means. Claims in news stories and advertisements often rest on convenience or voluntary-response samples, and spotting that is the most useful thing this topic gives you outside of math class.
Key terms
- Population.
- The entire group of people or objects you want information about in a statistical question.
- Sample.
- The part of the population that data is actually collected from.
- Census.
- Collecting data from every single member of the population rather than from a sample.
- Representative sample.
- A sample whose characteristics closely mirror the population, so conclusions about it apply to the whole group.
- Biased sample.
- A sample chosen by a method that systematically favors some members, pushing results consistently too high or too low.
- Convenience sample.
- A sample made up of whoever is easiest to reach, which usually leaves out part of the population.
- Voluntary response sample.
- A sample of people who choose to participate; those with strong opinions are overrepresented.
- Inference.
- A conclusion or estimate about a population that is based on data from a sample.
Worked example
Step 2: Identify the sample. Data was collected only from the 45 students waiting in the front office, so that group of 45 is the sample.
Step 3: Compute the sample result. The fraction saying yes is .
Step 4: Judge the method. Ask whether every student had a fair chance of being chosen. They did not. Students in the front office in the morning are often arriving late, being signed in, or handling attendance issues — a group unusually likely to have transportation problems. This is a convenience sample, and it over-represents students who already struggle to get to and from school on time.
Step 5: Predict the direction of the error. Because that group is more likely than average to want a late bus, the 67 percent figure probably overestimates the true interest across all 840 students.
Step 6: State a complete conclusion. The 67 percent should not be extended to the whole school. A better method would be to assign every student a number and randomly draw 45 numbers, or to survey a few students from every homeroom. Then an estimate such as "about two-thirds of all students would use the late bus" would be reasonable.
Practice questions
A factory produces 12,000 light bulbs a day. To check quality, a worker tests 60 bulbs pulled at random times throughout the day. Which statement is correct?
- The population is the 60 tested bulbs and the sample is the 12,000 bulbs.
- The population is the 12,000 bulbs made that day and the sample is the 60 tested bulbs.
- Both the population and the sample are the 60 tested bulbs.
- The factory should test all 12,000 bulbs to avoid bias.
Answer: The population is the 12,000 bulbs made that day and the sample is the 60 tested bulbs.
A student wants to know the average number of hours per week that students at her school play sports. She posts a survey link on the school's sports-team message board and gets 120 responses. Explain two separate problems with her sampling method and describe a better method.
Answer: Problem 1: the sample is drawn only from a sports message board, so athletes are heavily over-represented and students who play no sports are mostly excluded. Problem 2: it is a voluntary response sample — people choose whether to answer, and students who care a lot about sports are the most likely to reply. Both problems push the estimated average hours too high. A better method: get a list of all students, assign each a number, and randomly select 120 numbers to survey, following up so most selected students actually respond.
A wildlife biologist says, "I cannot possibly count every deer in the state forest, so I surveyed twelve randomly chosen square-mile plots." Why is a census impossible here, and what does the biologist gain by choosing the plots randomly?
Answer: A census is impossible because deer move, hide, and are born or die during the counting, and no team could search every acre of the forest at once — it would cost too much and take too long. Choosing the plots randomly means no part of the forest is systematically favored, so plots with dense deer populations and plots with few deer both have a fair chance of being included. That makes the estimate representative of the whole forest instead of leaning too high or too low.
FAQ
- How do I tell the population from the sample when a problem gives two numbers?
- Ask which group the question wants to describe. That group is the population, and it is always the larger one. The group that data was actually collected from is the sample. In "a manager surveys 80 of the 2,000 members of a gym," the 2,000 members are the population and the 80 surveyed are the sample.
- Does a bigger sample automatically mean better results?
- No. A large sample makes your estimate more precise only if the sampling method is fair. If the method leaves out part of the population — like surveying only people at a gym about exercise — a bigger sample just gives you more of the same lopsided information. Fix the method first, then increase the size.
- Is a sample ever a bad idea? Should we just count everybody?
- Count everybody when the population is small and easy to reach, like your own class. For large populations a census usually costs too much, takes too long, is physically impossible, or would destroy the items being tested. A well-chosen sample gives a close estimate for a tiny fraction of the effort.
- What makes a sampling method biased rather than just imperfect?
- Bias comes from the selection process consistently favoring or excluding certain members, so results lean the same direction every time you repeat it. Ordinary variation from a fair random sample is not bias — it just means different samples give slightly different estimates that center on the true value.
Learn this with a teacher, not a page
The Crimsora tutor teaches Populations & Samples live — explaining on a whiteboard, asking you questions, and adapting to where you get stuck.