AP-STATS-1.9

U1.9 Comparing Distributions

Learn to compare two or more distributions on shape, outliers, center, and spread using parallel boxplots and back-to-back stemplots, with the comparative language AP graders reward.

What you'll do in this lesson

A voice-first session with the Crimsora tutor on U1.9 Comparing Distributions, then targeted practice and FRQs — with the tutor adapting to where you get stuck.

What this lesson covers

When you have data from two groups — test scores from two classes, rainfall in two cities, wait times at two clinics — the interesting question is rarely "describe one distribution." It's "how do they differ, and what does that mean?" Topic 1.9 takes the single-distribution skills you built with SOCS and boxplots and turns them outward toward comparison.

The key is that comparing is not the same as describing twice. AP readers look for explicit comparative statements that connect the two groups in context. In this lesson you'll learn the tools (parallel boxplots, back-to-back stemplots), the four features to address (shape, outliers, center, spread), and the exact sentence structure that earns credit.

Why comparison needs comparative language

The single most common way students lose points on a comparison FRQ is by describing each distribution separately instead of comparing them. Writing "Group A has a median of 40. Group B has a median of 55." states two facts but never links them. The comparison the reader wants is: "Group B's median (55) is higher than Group A's median (40)," or better, "the typical value for Group B is about 15 units higher than for Group A."

Comparative language uses words like higher, lower, greater, more spread out, less variable, more symmetric, more skewed. Every claim about center, spread, or shape should be phrased as a relationship between the groups, not a solo statement.
Weak (parallel description)Strong (true comparison)
A is skewed right. B is symmetric.A is skewed right while B is roughly symmetric.
A's IQR is 12. B's IQR is 20.B is more variable than A (IQR 20 vs. 12).
A median 40. B median 55.B's median is higher than A's (55 vs. 40).
Always tie the comparison back to context — what the numbers represent — so a reader who doesn't see your graph understands the real-world meaning.

The four features: shape, outliers, center, spread

Comparing distributions extends the SOCS framework from Topic 1.6. You address the same four features, but for each you make a between-group statement.

Shape: Compare skewness and symmetry, and note modality if relevant. Example: "Distribution X is skewed right, whereas Y is approximately symmetric."

Outliers: Point out unusual values in either group and whether one group has them and the other doesn't. If you use the 1.5×IQR1.5 \times IQR rule to justify an outlier, state it, but for comparison you can also note apparent gaps.

Center: Compare medians (or means, if appropriate). Skewed data are usually compared by median. Say which group's center is higher and by roughly how much.

Spread: Compare range, IQR, or standard deviation. State which group is more variable. For boxplots, IQR (box width) is the natural spread measure.

A complete answer touches all four features that the graph lets you see. On boxplots you cannot judge shape precisely (you can't see modality or clustering), so you comment on skewness from the relative position of the median inside the box and the whisker lengths, and you always address center, spread, and outliers.

Parallel boxplots and back-to-back stemplots

Two graph types dominate this topic. Parallel (side-by-side) boxplots draw multiple boxplots on the same axis, making center and spread comparisons visual: a box shifted right has a higher center; a wider box has greater IQR. Skewness shows up as an off-center median line and unequal whisker lengths — a long right whisker suggests right skew. Boxplots hide gaps, clusters, and multiple peaks, so never claim a boxplot is "bimodal."

Back-to-back stemplots share a single stem in the middle, with one group's leaves extending left and the other's right. Because they keep individual data values, they reveal shape, clusters, gaps, and outliers that boxplots cannot. They work best for small-to-moderate data sets of two groups.
FeatureParallel boxplotsBack-to-back stemplot
Best forMany groups, quick center/spreadTwo groups, small n
Shows exact valuesNoYes
Shows shape detailLimitedYes
Shows outliersYes (by rule)Yes (visually)
Match your reading of the graph to what it actually shows. Read the axis carefully, and remember the box spans Q1 to Q3 with the median line inside.

How the exam tests this and common traps

On the AP exam, comparing distributions appears both as multiple-choice items (identifying which statement is a valid comparison, or reading a boxplot) and as free-response parts that ask you to "compare the distributions of [variable] for the two groups."

To earn full credit on the FRQ version, address multiple features with explicit comparative language and include context. A reliable template: "The center of ___ is (higher/lower) than ___ (give values). ___ is (more/less) variable than ___ (give spread values). ___ is (skewed/symmetric) while ___ is ___. ___ (does/does not) contain outlier(s)."

Common traps: (1) describing each group separately without connecting words; (2) forgetting context; (3) claiming a boxplot shows modality or clusters; (4) comparing means for clearly skewed data when median is more appropriate; (5) saying "the ranges are different" without saying which is larger. Also avoid vague words — "different" is not a comparison unless you specify direction and, ideally, magnitude.

Key terms

Comparative statement.
A sentence that relates two groups using directional words (higher, lower, more variable), rather than describing each separately.
Parallel boxplots.
Multiple boxplots drawn on a common axis to visually compare center, spread, and outliers across groups.
Back-to-back stemplot.
A stemplot sharing one central stem, with one group's leaves to the left and another's to the right, preserving individual values.
SOCS.
Shape, Outliers, Center, Spread — the four features to address when describing or comparing distributions.
IQR.
Interquartile range, IQR=Q3Q1IQR = Q_3 - Q_1; the width of a boxplot's box and a resistant measure of spread.
1.5 × IQR rule.
A value is an outlier if it is below Q11.5×IQRQ_1 - 1.5 \times IQR or above Q3+1.5×IQRQ_3 + 1.5 \times IQR.
Skewness (from a boxplot).
Inferred from the median's position in the box and whisker lengths; a longer right whisker suggests right skew.
Resistant measure.
A statistic like the median or IQR that is not strongly affected by outliers, preferred for skewed data.

Worked example

Parallel boxplots show daily commute times (minutes) for two routes. Route A: min 10, Q1 20, median 25, Q3 32, max 45. Route B: min 15, Q1 30, median 45, Q3 52, max 90, with a point plotted at 90 flagged as an outlier. Compare the distributions of commute time for the two routes.
Start with center. Route B's median commute time (45 min) is higher than Route A's (25 min) — about 20 minutes longer on a typical day.

Next, spread. Compute IQRs: Route A IQR=3220=12IQR = 32 - 20 = 12 minutes; Route B IQR=5230=22IQR = 52 - 30 = 22 minutes. Route B is more variable than Route A. The ranges agree: A's range is 4510=3545 - 10 = 35, B's is 9015=7590 - 15 = 75, so B's commute times are more spread out.

Now shape. In Route A the median (25) sits near the middle of the box and whiskers are fairly balanced, so A is roughly symmetric. In Route B the upper whisker is long and there's a high outlier, so B is skewed right.

Finally outliers. Route B has a high outlier at 90 minutes (an unusually long commute), while Route A shows no outliers.

Full-credit summary: "Route B has a higher center and greater spread than Route A, Route B is skewed right while A is roughly symmetric, and Route B contains a high outlier (90 min) that A lacks — overall, Route B's commutes are typically longer and less predictable."

Practice questions

Which of the following is the best example of a valid comparative statement about two distributions of exam scores?
  1. Class 1 has a median of 78 and Class 2 has a median of 85.
  2. The two classes have different medians.
  3. Class 2's median score (85) is higher than Class 1's median score (78).
  4. Class 1 is skewed left.

Answer: Class 2's median score (85) is higher than Class 1's median score (78).

A valid comparison must relate the two groups with directional language. The first choice states two separate facts without connecting them, the second gives direction-less "different," and the last describes only one class. Only the third explicitly says which median is higher and by how much, which is what earns credit.
A back-to-back stemplot displays reaction times for a caffeine group and a placebo group. The caffeine side is tightly clustered with a peak in the low values, while the placebo side is spread out with a long tail toward high values. Describe how you would compare these two distributions, and name one thing this graph shows that parallel boxplots could not.

Answer: Compare center (caffeine faster/lower typical reaction time), spread (placebo more variable), shape (caffeine more symmetric/clustered, placebo right-skewed), and outliers, all in context. The stemplot reveals clustering and exact values / modality that a boxplot hides.

A strong response addresses all four SOCS features with comparative wording and context (reaction time), e.g., "the caffeine group's typical reaction time is lower and less variable than the placebo group's, and the placebo distribution is skewed right." The distinguishing advantage of a back-to-back stemplot is that it preserves individual data values, so it can show clusters, gaps, and the shape/peaks that a boxplot's five-number summary conceals.
On parallel boxplots, Group X has a box from 40 to 60 with median 55, and Group Y has a box from 40 to 60 with median 45. What can you conclude about spread and shape?
  1. X and Y have the same IQR; X appears skewed left and Y appears skewed right.
  2. X has a larger IQR than Y.
  3. Y is more variable than X.
  4. Both groups are perfectly symmetric.

Answer: X and Y have the same IQR; X appears skewed left and Y appears skewed right.

Both boxes span 40 to 60, so both have IQR=20IQR = 20 — equal spread, eliminating the two IQR-difference choices. The median position signals shape: X's median (55) sits near the top of the box, suggesting left skew, while Y's median (45) sits near the bottom, suggesting right skew, so they are not symmetric.

FAQ

Do I have to mention all four features every time I compare distributions?
Address every feature the graph lets you see and that is relevant. For center and spread you should almost always give a comparative statement. Comment on shape and outliers when they are visible and meaningful. On boxplots you can't judge modality, so you focus on skewness, center, spread, and outliers.
Should I compare means or medians?
Use the median (and IQR) when a distribution is skewed or has outliers, because these measures are resistant. Use the mean (and standard deviation) for roughly symmetric distributions without outliers. When comparing two groups, use the same measure for both so the comparison is fair.
Why do I lose points if I describe each distribution separately?
The objective is comparison, not two descriptions. Stating "A's median is 40" and "B's median is 55" as separate facts doesn't demonstrate the relationship. Graders look for linking language such as "higher than" or "more variable than" that directly compares the groups in context.
Can a boxplot show whether a distribution is bimodal?
No. A boxplot only shows the five-number summary, so it cannot reveal peaks, clusters, or gaps. Never claim a boxplot is bimodal. If you need to see shape in detail, use a stemplot, dotplot, or histogram instead.

Learn this with a teacher, not a page

The Crimsora tutor teaches U1.9 Comparing Distributions live — explaining on a whiteboard, asking you questions, and adapting to where you get stuck.