One-Variable Statistics: Center & Spread
Mean, median, mode, range, and IQR: how to compute each, read dot plots and box plots, and pick the measure of center that survives outliers and skew.
What you'll do in this lesson
A voice-first session with the Crimsora tutor on One-Variable Statistics: Center & Spread, then targeted practice and FRQs — with the tutor adapting to where you get stuck.
What this lesson covers
In this lesson you will compute the mean, median, and mode of a one-variable data set, find the range and interquartile range, and pull those same numbers out of a dot plot or a box plot without doing any arithmetic. The judgment call at the end is the part that matters most: when a data set has an extreme value or a long tail, the mean and the range get dragged around, while the median and IQR stay put. Knowing which pair to report is a real decision that scientists, coaches, and city planners make all the time.
Three Measures of Center
The mean is the balance point: add every value and divide by how many there are, . Every data point pulls on the mean, so one huge or tiny value moves it noticeably.
The median is the positional middle. Order the data from least to greatest. If is odd, the median is the single middle value; if is even, it is the mean of the two middle values. Only the position of an extreme value matters to the median, not how extreme it is.
The mode is the value that occurs most often. A data set can have one mode, several modes, or none at all if every value appears exactly once. The mode is the only measure of center that works for categorical data such as favorite color.
| Measure | How to find it | Sensitive to outliers? |
|---|---|---|
| Mean | Sum divided by | Yes, strongly |
| Median | Middle of the ordered list | No |
| Mode | Most frequent value | No |
Range, Quartiles, and the IQR
The range is the simplest measure: . It uses only two numbers, so a single unusual value can inflate it enormously.
The interquartile range describes the spread of the middle half of the data. First find the median, which splits the ordered list into a lower half and an upper half. The median of the lower half is the first quartile ; the median of the upper half is the third quartile . ThenWhen is odd, do not include the overall median in either half. For 4, 7, 9, 11, 20 the median is 9, the lower half is 4 and 7, so , and the upper half is 11 and 20, so , giving .
The five-number summary is minimum, , median, , maximum. It is exactly the information a box plot displays.
A common way to flag an outlier is the 1.5 IQR rule: any value below or above is treated as an outlier. Because the IQR depends only on the middle of each half of the data, pushing an extreme value farther out barely moves it — which is precisely why the IQR is the spread measure you report whenever outliers exist.
Reading Dot Plots and Box Plots
A box plot compresses the data into the five-number summary. The left whisker starts at the minimum, the box spans from to with a vertical segment at the median, and the right whisker ends at the maximum. The box therefore contains the middle 50 percent of the data and its width is the IQR. Each whisker covers about 25 percent of the data, and the two halves of the box cover about 25 percent each.
A long right whisker with the median sitting toward the left side of the box signals a distribution skewed right; the mirror image signals skewed left. If the median sits near the center and the whiskers are about equal, the distribution is roughly symmetric.
The biggest misconception here: a wide section of a box plot does not mean more data points are there. It means the same 25 percent of the data is spread over a wider interval. A long whisker signals sparse, stretched-out values, not a crowd of them. Also, you cannot find the mean or the mode from a box plot — the individual values are gone.
Choosing the Measure That Fits the Data
| Shape of distribution | Report for center | Report for spread |
|---|---|---|
| Roughly symmetric, no outliers | Mean | Standard deviation or range |
| Skewed or contains outliers | Median | IQR |
In a right-skewed distribution the mean is pulled toward the long tail, so the mean sits above the median. In a left-skewed distribution the mean sits below the median. In a symmetric distribution they are approximately equal. That relationship lets you predict skew from two numbers alone.
When you justify a choice in writing, name the shape, name the measure, and connect them: "The distribution is skewed right because of the 120 minute value, so the median of 30 minutes better represents a typical day than the mean of about 42 minutes." Answers that only state a measure without the reason are incomplete. Also resist deleting an outlier just because it is inconvenient — an outlier is legitimate data unless you know it was a recording error.
Comparing Two Data Sets
Work through it in a fixed order. First compare centers: which group is typically higher? Then compare spreads: which group is more consistent? Smaller IQR means more consistent. Finally comment on shape and any outliers.
For example, if Team A has median 12 points with and Team B has median 12 points with , the two teams are typically equal, but Team A is far more predictable. A coach who needs a dependable scorer prefers Team A; a coach who needs occasional huge games might accept Team B's variability.
When comparing box plots stacked on the same axis, look at whether the boxes overlap. If Team B's entire box sits above Team A's box, more than half of B's values exceed nearly all of A's, which is strong visual evidence of a real difference. If the boxes overlap heavily, the groups are similar even when their medians differ slightly.
One caution: never compare a mean from one group to a median from another. Comparisons only make sense between the same statistic, computed the same way, in the same units.
Key terms
- Mean.
- The sum of all data values divided by the number of values, . It is the balance point of the distribution and is sensitive to outliers.
- Median.
- The middle value of an ordered data set; with an even number of values it is the mean of the two middle values. It is resistant to outliers.
- Mode.
- The value that appears most frequently. A data set may have no mode, one mode, or several.
- Range.
- Maximum minus minimum. A quick measure of total spread that depends entirely on the two most extreme values.
- Interquartile range (IQR).
- , the spread of the middle 50 percent of the data. It is unaffected by outliers.
- Five-number summary.
- Minimum, , median, , and maximum — the five values a box plot displays.
- Outlier.
- A value far from the rest of the data, commonly identified as any value below or above .
- Skew.
- Asymmetry in a distribution. Skewed right means a long tail of high values pulling the mean above the median; skewed left is the reverse.
Worked example
Mean. The sum is , so minutes.
Median. With , the median is the 4th value: 30 minutes.
Mode. The value 25 appears twice and every other value appears once, so the mode is 25 minutes.
Range. minutes.
Quartiles. Exclude the median from both halves. The lower half is 20, 25, 25, so . The upper half is 35, 40, 120, so . Therefore minutes.
Outlier check. . The upper fence is , and the lower fence is . Since , the value 120 is an outlier.
Choice of measures. The single long practice session pulls the mean up to about 42 minutes, which is higher than six of the seven actual days — no typical day looks like that. The distribution is skewed right, so report the median of 30 minutes for center and the IQR of 15 minutes for spread. Notice how much the outlier matters to the non-resistant statistics: dropping the 120 would move the mean from about 42.1 to about 29.2 and the range from 100 to 20, while the median moves only from 30 to 27.5 and the IQR from 15 to 10.
Practice questions
A data set has the five-number summary 4, 9, 11, 15, 40. Which statement is best supported?
- The IQR is 36, so the data are very consistent.
- The mean is 11 because the median is 11.
- The distribution is skewed right, so the median and IQR are the better summary.
- Exactly half of the data values lie between 11 and 15.
Answer: The distribution is skewed right, so the median and IQR are the better summary.
A dot plot of the number of pets owned by 10 students shows: 0, 0, 1, 1, 1, 2, 2, 3, 4, 6. Compute the mean and the median, then explain what the relationship between them tells you about the shape of the distribution.
Answer: Mean = 2, median = 1.5; the mean exceeds the median, indicating a right-skewed distribution.
Class A's test scores have a median of 82 with an IQR of 4. Class B's scores have a median of 82 with an IQR of 18. Which class is more consistent, and why does comparing only the medians hide this?
Answer: Class A is more consistent, because its smaller IQR means the middle 50 percent of its scores span only 4 points compared with 18 points for Class B.
FAQ
- When should I use the mean instead of the median?
- Use the mean when the distribution is roughly symmetric and has no outliers, because then it uses all the data and is the most informative single number. Switch to the median whenever the data are skewed or contain an outlier, since the mean gets pulled toward extreme values and can end up describing no actual data point.
- Do I include the median when finding the quartiles?
- No. When the number of values is odd, leave the overall median out of both halves and find the median of the values strictly below it and strictly above it. When the number of values is even, the list already splits cleanly into two equal halves, so nothing is excluded.
- Can I find the mean from a box plot?
- No. A box plot displays only the minimum, , median, , and maximum. The individual data values are not shown, and many different data sets share the same five-number summary, so the mean cannot be recovered. If you need the mean, you need the raw data or a dot plot.
- Does a wider box mean more data points are in that region?
- No. Each of the four sections of a box plot — lower whisker, lower box, upper box, upper whisker — contains about 25 percent of the data no matter how long it is. A wide section means those values are spread far apart; a narrow section means they are packed close together.
Learn this with a teacher, not a page
The Crimsora tutor teaches One-Variable Statistics: Center & Spread live — explaining on a whiteboard, asking you questions, and adapting to where you get stuck.