Measures of Center & Variability
Learn to compute mean, median, mode, range, and mean absolute deviation of small data sets, and pick the best measure of center when an outlier appears.
What you'll do in this lesson
A voice-first session with the Crimsora tutor on Measures of Center & Variability, then targeted practice and FRQs — with the tutor adapting to where you get stuck.
What this lesson covers
By the end you will be able to compute all five measures for a small data set by hand, explain the difference between a measure of center (where the data clusters) and a measure of variability (how spread out it is), and defend a choice between mean and median when an extreme value is pulling the average around. These skills are the foundation for the work later in this unit, where you compare two different groups using their centers and their spread together.
Mean, Median, and Mode: Three Ways to Find the Center
The mean is the arithmetic average: add every value, then divide by the number of values, . The median is the middle value once the data is written in order from least to greatest. If there is an even number of values, the median is the mean of the two middle numbers. The mode is the value that appears most often; a data set can have one mode, several modes, or no mode at all if every value appears once.
| Measure | How to find it | Describes the data well when |
|---|---|---|
| Mean | Add all values, divide by | Values are fairly close together, no extremes |
| Median | Order the values, take the middle | One or two values are far from the rest |
| Mode | Count which value repeats most | You want the most common response or category |
Range and Mean Absolute Deviation: Measuring Spread
The range is the simplest: subtract the least value from the greatest. It uses only two numbers, so a single extreme value can make the range huge even when most of the data is tightly packed.
The mean absolute deviation (MAD) is more informative because every value contributes. It answers the question: on average, how far is a typical data value from the mean? The formula isTo compute it, follow three steps in order. Find the mean. Find the distance from each value to the mean, always taking the absolute value so every distance is positive or zero. Then average those distances.
The biggest error is dropping the absolute value bars. Without them the positive and negative deviations always cancel, and you get a sum of exactly zero every single time — if your deviations add to zero, that is a signal you skipped the absolute value. A second common slip is dividing by the wrong number, such as dividing by the number of nonzero deviations instead of by . A third is computing deviations from the median instead of the mean; MAD is defined using the mean.
Interpret MAD in the units of the data. A MAD of minutes means values sit about minutes away from the mean on a typical day.
Outliers: When the Mean Stops Telling the Truth
The mean is computed from every value, so one enormous number drags it upward and one tiny number drags it down. The median only cares about position in the ordered list, so an outlier moves it slightly at most. Statisticians say the median is resistant to outliers while the mean is sensitive to them.
Here is the comparison in numbers. Take the ages . The mean is and the median is — the two agree closely. Now replace the last value with : the data becomes . The mean jumps to , but the median is still . A mean of is a poor description, because four of the five people are younger than .
The rule of thumb for the class: when a data set contains an outlier or is clearly lopsided, report the median as the measure of center; when the values are bunched together without extremes, the mean is fine and is usually preferred because it uses all the information.
Where students go wrong is answering "median" or "mean" with no justification. A complete answer names the outlier, shows what it does to the mean, and then explains why the other measure describes the typical value better. Notice too that an outlier inflates both the range and the MAD, so a large MAD compared with the data values is often a hint that an outlier is hiding in the set.
Reading Center and Spread Together
Compare two students' quiz scores out of . Ana scores ; Ben scores . Both have a mean of . Ana's deviations are , so her . Ben's deviations are , so his . Same center, but Ana's MAD is eight times smaller, which means she is far more consistent. If you had to predict the next quiz, you would trust the number much more for Ana.
| Data set | Mean | MAD | What it says |
|---|---|---|---|
| Ana: 14, 15, 15, 16 | 15 | 0.5 | Very consistent, tightly clustered |
| Ben: 9, 13, 17, 21 | 15 | 4 | Same average, much less predictable |
That reasoning is exactly the tool used later in this unit when comparing two populations, so practicing it now on small sets pays off. For any data set you meet, get in the habit of stating three things: the center you chose, why you chose it, and the spread that goes with it.
Key terms
- Mean.
- The sum of all data values divided by the number of values, written ; the arithmetic average.
- Median.
- The middle value of a data set arranged in order; with an even number of values, the mean of the two middle values.
- Mode.
- The value that occurs most often in a data set. A set may have no mode, one mode, or more than one.
- Range.
- The greatest value minus the least value; a quick but crude measure of variability that uses only two data points.
- Mean absolute deviation (MAD).
- The average distance from the mean, , reported in the same units as the data.
- Outlier.
- A data value that lies far from the rest of the data; it pulls the mean toward itself but barely moves the median.
- Measure of center.
- A single number summarizing where the data clusters — mean, median, or mode.
- Measure of variability.
- A number describing how spread out the data is, such as the range or the mean absolute deviation.
Worked example
Step 2: Mean. The sum is , so minutes.
Step 3: Median. With six values, average the third and fourth: minutes.
Step 4: Mode. The value appears twice and every other value appears once, so the mode is minutes.
Step 5: Range. minutes.
Step 6: MAD. Find each distance from the mean of : , , , , , . The distances sum to , so minutes.
Step 7: Choose a center. The value is an outlier — it is minutes above the next largest value. It pushes the mean up to , even though five of the six students worked minutes or less, so no one is actually near . The median of minutes sits right in the cluster and is the better description of a typical student. Notice also that the MAD of minutes is huge compared with how close most values are to each other; that inflated spread is another sign the outlier is distorting the mean.
Practice questions
What is the mean absolute deviation of the data set ?
Answer:
A店 owner records the number of customers for seven days: . Find the mean and the median, state which one better describes a typical day, and justify your choice.
Answer: Mean (about ) customers; median customers. The median is the better description because is an outlier that inflates the mean.
Two groups of four students each have a mean height of cm. Group A has a MAD of cm and Group B has a MAD of cm. What does this tell you about the two groups?
Answer: Both groups center at the same height, but Group B's heights are much more spread out; Group A's members are typically only about cm from cm while Group B's are about cm away.
FAQ
- When should I use the median instead of the mean?
- Use the median whenever the data contains an outlier or is clearly lopsided, because the median barely moves when one value is extreme. Use the mean when the values are bunched together with no extremes, since it uses every data point. Always justify your choice by pointing to the outlier and what it does to the mean.
- Why does the mean absolute deviation use absolute value?
- Because the signed deviations from the mean always add to zero — the mean is the balance point of the data. Taking the absolute value turns every deviation into a positive distance, so the sum measures real spread instead of collapsing to zero. If your deviations ever sum to zero, you forgot the absolute value bars.
- Can a data set have more than one mode, or none at all?
- Yes to both. If two or more values tie for the most appearances, every one of them is a mode. If every value appears exactly once, the set has no mode. That is why the mode is the least useful center for numerical data and is most helpful for the most common response or category.
- Is the range or the MAD the better measure of variability?
- The MAD is usually better because it uses every value, while the range depends only on the largest and smallest. One extreme value can make the range enormous even when nearly all the data is tightly packed. The range is still handy as a fast first look at how wide a data set is.
Learn this with a teacher, not a page
The Crimsora tutor teaches Measures of Center & Variability live — explaining on a whiteboard, asking you questions, and adapting to where you get stuck.