M7MATH-9.3

Measures of Center & Variability

Learn to compute mean, median, mode, range, and mean absolute deviation of small data sets, and pick the best measure of center when an outlier appears.

What you'll do in this lesson

A voice-first session with the Crimsora tutor on Measures of Center & Variability, then targeted practice and FRQs — with the tutor adapting to where you get stuck.

What this lesson covers

Suppose six classmates report how many minutes they spent on homework last night, and one of them says 100 minutes while everyone else says about half an hour. What single number honestly describes that group? This lesson gives you five tools — mean, median, mode, range, and mean absolute deviation (MAD) — and, just as importantly, teaches you when each one tells the truth about the data.

By the end you will be able to compute all five measures for a small data set by hand, explain the difference between a measure of center (where the data clusters) and a measure of variability (how spread out it is), and defend a choice between mean and median when an extreme value is pulling the average around. These skills are the foundation for the work later in this unit, where you compare two different groups using their centers and their spread together.

Mean, Median, and Mode: Three Ways to Find the Center

A measure of center is a single number that stands in for a whole data set. The three you need are the mean, the median, and the mode.

The mean is the arithmetic average: add every value, then divide by the number of values, xˉ=sum of valuesn\bar{x} = \frac{\text{sum of values}}{n}. The median is the middle value once the data is written in order from least to greatest. If there is an even number of values, the median is the mean of the two middle numbers. The mode is the value that appears most often; a data set can have one mode, several modes, or no mode at all if every value appears once.
MeasureHow to find itDescribes the data well when
MeanAdd all values, divide by nnValues are fairly close together, no extremes
MedianOrder the values, take the middleOne or two values are far from the rest
ModeCount which value repeats mostYou want the most common response or category
Two mistakes cause most wrong answers here. First, students find the median without ordering the data — the middle number of an unordered list means nothing. Second, with an even count they pick one of the two middle values instead of averaging them. For the ordered set 3,5,8,143, 5, 8, 14, the median is 5+82=6.5\frac{5+8}{2} = 6.5, not 55 and not 88. Also remember that the median does not have to be a value in the data set, but the mode always is.

Range and Mean Absolute Deviation: Measuring Spread

Two data sets can share the same mean and still look nothing alike. The set 49,50,5149, 50, 51 and the set 10,50,9010, 50, 90 both have a mean of 5050, but the second is wildly more spread out. Measures of variability capture that difference.

The range is the simplest: subtract the least value from the greatest. It uses only two numbers, so a single extreme value can make the range huge even when most of the data is tightly packed.

The mean absolute deviation (MAD) is more informative because every value contributes. It answers the question: on average, how far is a typical data value from the mean? The formula isMAD=∑∣x−xˉ∣nMAD = \frac{\sum |x - \bar{x}|}{n}To compute it, follow three steps in order. Find the mean. Find the distance from each value to the mean, always taking the absolute value so every distance is positive or zero. Then average those distances.

The biggest error is dropping the absolute value bars. Without them the positive and negative deviations always cancel, and you get a sum of exactly zero every single time — if your deviations add to zero, that is a signal you skipped the absolute value. A second common slip is dividing by the wrong number, such as dividing by the number of nonzero deviations instead of by nn. A third is computing deviations from the median instead of the mean; MAD is defined using the mean.

Interpret MAD in the units of the data. A MAD of 44 minutes means values sit about 44 minutes away from the mean on a typical day.

Outliers: When the Mean Stops Telling the Truth

An outlier is a data value that lies far away from the rest of the data. Outliers matter because the mean and the median react to them very differently.

The mean is computed from every value, so one enormous number drags it upward and one tiny number drags it down. The median only cares about position in the ordered list, so an outlier moves it slightly at most. Statisticians say the median is resistant to outliers while the mean is sensitive to them.

Here is the comparison in numbers. Take the ages 11,12,12,13,1411, 12, 12, 13, 14. The mean is 625=12.4\frac{62}{5} = 12.4 and the median is 1212 — the two agree closely. Now replace the last value with 5454: the data becomes 11,12,12,13,5411, 12, 12, 13, 54. The mean jumps to 1025=20.4\frac{102}{5} = 20.4, but the median is still 1212. A mean of 20.420.4 is a poor description, because four of the five people are younger than 1414.

The rule of thumb for the class: when a data set contains an outlier or is clearly lopsided, report the median as the measure of center; when the values are bunched together without extremes, the mean is fine and is usually preferred because it uses all the information.

Where students go wrong is answering "median" or "mean" with no justification. A complete answer names the outlier, shows what it does to the mean, and then explains why the other measure describes the typical value better. Notice too that an outlier inflates both the range and the MAD, so a large MAD compared with the data values is often a hint that an outlier is hiding in the set.

Reading Center and Spread Together

Center and spread are a package deal. Reporting only a center hides how reliable that center is, and reporting only spread tells you nothing about where the data sits.

Compare two students' quiz scores out of 2020. Ana scores 14,15,15,1614, 15, 15, 16; Ben scores 9,13,17,219, 13, 17, 21. Both have a mean of 1515. Ana's deviations are 1,0,0,11, 0, 0, 1, so her MAD=24=0.5MAD = \frac{2}{4} = 0.5. Ben's deviations are 6,2,2,66, 2, 2, 6, so his MAD=164=4MAD = \frac{16}{4} = 4. Same center, but Ana's MAD is eight times smaller, which means she is far more consistent. If you had to predict the next quiz, you would trust the number 1515 much more for Ana.
Data setMeanMADWhat it says
Ana: 14, 15, 15, 16150.5Very consistent, tightly clustered
Ben: 9, 13, 17, 21154Same average, much less predictable
A useful way to describe a difference between two groups is to ask how many MADs apart their centers are. If two groups have means of 1515 and 1919 and each has a MAD of about 11, the four-point gap is large compared with the ordinary variation inside each group. If each MAD were 55 instead, that same gap would be small and unremarkable.

That reasoning is exactly the tool used later in this unit when comparing two populations, so practicing it now on small sets pays off. For any data set you meet, get in the habit of stating three things: the center you chose, why you chose it, and the spread that goes with it.

Key terms

Mean.
The sum of all data values divided by the number of values, written xˉ\bar{x}; the arithmetic average.
Median.
The middle value of a data set arranged in order; with an even number of values, the mean of the two middle values.
Mode.
The value that occurs most often in a data set. A set may have no mode, one mode, or more than one.
Range.
The greatest value minus the least value; a quick but crude measure of variability that uses only two data points.
Mean absolute deviation (MAD).
The average distance from the mean, MAD=∑∣x−xˉ∣nMAD = \frac{\sum |x - \bar{x}|}{n}, reported in the same units as the data.
Outlier.
A data value that lies far from the rest of the data; it pulls the mean toward itself but barely moves the median.
Measure of center.
A single number summarizing where the data clusters — mean, median, or mode.
Measure of variability.
A number describing how spread out the data is, such as the range or the mean absolute deviation.

Worked example

Six students reported the minutes they spent on homework last night: 20, 25, 30, 30, 35, 100. Find the mean, median, mode, range, and mean absolute deviation. Then decide which measure of center best describes a typical student, and explain why.
Step 1: Order the data. It is already in order: 20,25,30,30,35,10020, 25, 30, 30, 35, 100. There are n=6n = 6 values.

Step 2: Mean. The sum is 20+25+30+30+35+100=24020 + 25 + 30 + 30 + 35 + 100 = 240, so xˉ=2406=40\bar{x} = \frac{240}{6} = 40 minutes.

Step 3: Median. With six values, average the third and fourth: 30+302=30\frac{30 + 30}{2} = 30 minutes.

Step 4: Mode. The value 3030 appears twice and every other value appears once, so the mode is 3030 minutes.

Step 5: Range. 100−20=80100 - 20 = 80 minutes.

Step 6: MAD. Find each distance from the mean of 4040: ∣20−40∣=20|20-40| = 20, ∣25−40∣=15|25-40| = 15, ∣30−40∣=10|30-40| = 10, ∣30−40∣=10|30-40| = 10, ∣35−40∣=5|35-40| = 5, ∣100−40∣=60|100-40| = 60. The distances sum to 20+15+10+10+5+60=12020+15+10+10+5+60 = 120, so MAD=1206=20MAD = \frac{120}{6} = 20 minutes.

Step 7: Choose a center. The value 100100 is an outlier — it is 6565 minutes above the next largest value. It pushes the mean up to 4040, even though five of the six students worked 3535 minutes or less, so no one is actually near 4040. The median of 3030 minutes sits right in the cluster and is the better description of a typical student. Notice also that the MAD of 2020 minutes is huge compared with how close most values are to each other; that inflated spread is another sign the outlier is distorting the mean.

Practice questions

What is the mean absolute deviation of the data set 4,4,6,9,124, 4, 6, 9, 12?
  1. 2.52.5
  2. 2.82.8
  3. 3.53.5
  4. 1414

Answer: 2.82.8

First find the mean: 4+4+6+9+125=355=7\frac{4+4+6+9+12}{5} = \frac{35}{5} = 7. Now find each absolute deviation from 77: 3,3,1,2,53, 3, 1, 2, 5. Their sum is 1414, and dividing by n=5n = 5 gives MAD=145=2.8MAD = \frac{14}{5} = 2.8. The answer 1414 is what you get if you stop at the sum and forget to divide, and 3.53.5 comes from dividing 1414 by 44 instead of 55.
A店 owner records the number of customers for seven days: 18,21,22,22,24,25,9618, 21, 22, 22, 24, 25, 96. Find the mean and the median, state which one better describes a typical day, and justify your choice.

Answer: Mean =32.57= 32.57 (about 3333) customers; median =22= 22 customers. The median is the better description because 9696 is an outlier that inflates the mean.

The sum is 18+21+22+22+24+25+96=22818+21+22+22+24+25+96 = 228, so the mean is 2287≈32.6\frac{228}{7} \approx 32.6. With seven ordered values the median is the fourth one, which is 2222. Six of the seven days had between 1818 and 2525 customers, so a mean of about 3333 is higher than every ordinary day — the single value of 9696 pulled it up. The median of 2222 sits inside the cluster where the data actually lives, so it describes a typical day. A complete answer names 9696 as the outlier and explains its effect on the mean, not just which number is chosen.
Two groups of four students each have a mean height of 150150 cm. Group A has a MAD of 22 cm and Group B has a MAD of 99 cm. What does this tell you about the two groups?

Answer: Both groups center at the same height, but Group B's heights are much more spread out; Group A's members are typically only about 22 cm from 150150 cm while Group B's are about 99 cm away.

MAD measures the average distance of the data values from the mean, in the same units as the data. Equal means tell you the two groups balance at the same place, so the mean alone cannot distinguish them. The much larger MAD for Group B shows its heights vary widely, so 150150 cm is a far less reliable prediction for a randomly chosen member of Group B than for one from Group A.

FAQ

When should I use the median instead of the mean?
Use the median whenever the data contains an outlier or is clearly lopsided, because the median barely moves when one value is extreme. Use the mean when the values are bunched together with no extremes, since it uses every data point. Always justify your choice by pointing to the outlier and what it does to the mean.
Why does the mean absolute deviation use absolute value?
Because the signed deviations from the mean always add to zero — the mean is the balance point of the data. Taking the absolute value turns every deviation into a positive distance, so the sum measures real spread instead of collapsing to zero. If your deviations ever sum to zero, you forgot the absolute value bars.
Can a data set have more than one mode, or none at all?
Yes to both. If two or more values tie for the most appearances, every one of them is a mode. If every value appears exactly once, the set has no mode. That is why the mode is the least useful center for numerical data and is most helpful for the most common response or category.
Is the range or the MAD the better measure of variability?
The MAD is usually better because it uses every value, while the range depends only on the largest and smallest. One extreme value can make the range enormous even when nearly all the data is tightly packed. The range is still handy as a fast first look at how wide a data set is.

Learn this with a teacher, not a page

The Crimsora tutor teaches Measures of Center & Variability live — explaining on a whiteboard, asking you questions, and adapting to where you get stuck.