AP-STATS-1.6

U1.6 Describing a Distribution (SOCS)

Master SOCS for AP Statistics: describe a quantitative distribution's shape, outliers, center, and spread with exam-ready language and context.

What you'll do in this lesson

A voice-first session with the Crimsora tutor on U1.6 Describing a Distribution (SOCS), then targeted practice and FRQs — with the tutor adapting to where you get stuck.

What this lesson covers

When an AP Statistics question shows you a dotplot, stemplot, or histogram and asks you to "describe the distribution," it is testing one specific skill: SOCS. That acronym stands for Shape, Outliers, Center, and Spread — the four features you must comment on to earn full credit.

This lesson teaches you how to read a graph and translate it into precise, context-rich sentences. You already know how to build these graphs from U1.5; now you will learn what to say about them. Getting the vocabulary exactly right here pays off all year, because comparing distributions (U1.9) and analyzing normal models (U1.10) both rest on the SOCS foundation.

What SOCS Means and Why All Four Matter

SOCS is a checklist. A complete description of a quantitative distribution addresses every letter, always in the context of the variable being measured.

Shape describes the overall pattern: is it symmetric or skewed, how many peaks (modes) does it have? Outliers are individual values that fall far from the rest of the data; you note whether any appear and where. Center is a single number summarizing a typical value — the mean or median. Spread describes how much the values vary, using the range, interquartile range (IQR), or standard deviation.
LetterWhat you reportExample phrasing
Shapesymmetry, skew, modes"roughly symmetric with a single peak"
Outliersunusual values"a possible high outlier near 95"
Centermean or median"a median of about 40 points"
Spreadrange, IQR, or SD"values range from 10 to 60"
On the AP exam, leaving out a component is the most common way students lose points. Even if a distribution has no outliers, you should state that explicitly: "there do not appear to be any outliers." Never describe these features with bare numbers — always tie them to what the variable measures and its units.

Describing Shape Correctly

Shape has two main dimensions: symmetry and modality.

A distribution is symmetric if the left and right halves are approximate mirror images. It is skewed right if it has a long tail stretching toward larger values, and skewed left if the long tail stretches toward smaller values. A memory trick: the skew direction is the direction the tail points, not where the bulk of the data sits.

Modality counts the peaks. One clear peak is unimodal, two peaks is bimodal, and a roughly flat shape with no dominant peak is uniform. A bimodal shape often signals two groups mixed together (for example, heights of a co-ed class).

A critical misconception: students say "the data is skewed" without a direction, or confuse the tail with the mound. Look at where the sparse, stretched-out values are — that tail names the skew.

Shape also drives your choice of center and spread. When a distribution is skewed or has outliers, the median and IQR are resistant and better summaries. When it is roughly symmetric with no outliers, the mean and standard deviation work well. The exam frequently rewards students who justify this connection explicitly.

Outliers, Center, and Spread in Practice

For outliers in a describe-the-distribution question, you can often identify them visually as gaps separating extreme points from the cluster. (A formal rule, the 1.5×IQR1.5 \times IQR criterion, is developed in U1.7; here a visual note is usually acceptable.) Say where they are: "one unusually large value near 200."

For center, report the median when the shape is skewed or has outliers, and the mean when it is symmetric — but any reasonable estimate of a typical value earns credit as long as you name it. Give an approximate number: "the center is around 52 seconds."

For spread, describe variability. The simplest is the range (maximum minus minimum). The IQR (Q3Q1Q_3 - Q_1) captures the middle 50% and resists outliers. Standard deviation measures typical distance from the mean. For a graph without exact values, stating the range from the visible minimum to maximum is usually enough: "data spread from about 5 to 48."
FeatureResistant choiceNon-resistant choice
Centermedianmean
SpreadIQRstandard deviation, range
Remember to always attach units and context. "Center about 40" is weaker than "a typical commute was about 40 minutes."

How the AP Exam Tests SOCS

On free-response questions, a prompt like "Describe the distribution of [variable]" expects a short paragraph hitting all four SOCS components in context. Scoring guidelines typically require at least shape, center, and spread mentioned correctly with context to earn full credit; outliers should be addressed when present.

High-scoring answers read like this: "The distribution of daily rainfall is skewed right, with most days recording under 0.5 inches and a long tail toward higher amounts. There appears to be a high outlier near 3 inches. The center is around 0.3 inches, and values range from 0 to about 3 inches."

Common point-losing errors include describing a categorical variable with SOCS (SOCS only applies to quantitative data), reporting numbers with no context or units, forgetting to mention outliers, and using vague words like "normal" when you mean "symmetric." The word normal refers to a specific bell-shaped model from U1.10 — reserve it for that.

A final tip: if the question gives a comparison across two groups, you still use SOCS but with explicit comparative language (higher, lower, more spread out). That extension is the focus of U1.9, but the vocabulary starts here.

Key terms

SOCS.
An acronym for the four features used to describe a quantitative distribution: Shape, Outliers, Center, and Spread — always reported in context.
Skewed right.
A distribution with a long tail extending toward larger values; most data cluster at the lower end.
Skewed left.
A distribution with a long tail extending toward smaller values; most data cluster at the higher end.
Unimodal / Bimodal.
Describes the number of peaks: unimodal has one clear peak, bimodal has two, often signaling two mixed groups.
Outlier.
An individual value that falls notably far from the rest of the data, seen as a gap in a graph or flagged by the 1.5×IQR1.5 \times IQR rule.
Resistant statistic.
A summary measure, like the median or IQR, that is not strongly affected by outliers or skew.
Spread.
The variability in the data, described by range, interquartile range (IQR), or standard deviation.
Center.
A single number representing a typical value in the data, usually the mean or the median.

Worked example

A histogram shows the number of text messages 40 students sent in one day. The bars rise sharply from 0, peak between 20 and 40 messages, and trail off with a few students in the 120–140 range. The median is 35 messages. Write a complete SOCS description.
Start with Shape. The bars pile up at the low end and stretch out toward high values, so the distribution is skewed right and unimodal, with its single peak in the 20–40 range.

Next, Outliers. The students in the 120–140 range sit far from the main cluster and appear as separated bars, so we note a few possible high outliers near 120 to 140 messages.

Now Center. Because the distribution is skewed with high outliers, the median is the better summary. The center is about 35 messages per day — a typical student sent roughly 35 texts.

Finally, Spread. The values run from 0 to about 140 messages, giving a range near 140. Since the data is skewed, the IQR would be the preferred numerical spread if quartiles were given.

Put it together: "The distribution of daily text messages is skewed right and unimodal, peaking around 20 to 40 messages. A few students appear to be high outliers near 120 to 140 messages. The center is about 35 messages (median), and values range from 0 to roughly 140 messages." This addresses all four SOCS components in context with units.

Practice questions

A dotplot of the ages (in years) of cars in a parking lot is strongly skewed left with no gaps. Which combination of statistics best summarizes this distribution?
  1. Mean and standard deviation
  2. Median and IQR
  3. Mean and range
  4. Mode and standard deviation

Answer: Median and IQR

Because the distribution is skewed, the mean and standard deviation are pulled toward the tail and are not resistant. The median and IQR resist the effect of skew and extreme values, making them the appropriate summaries of center and spread for a skewed distribution.
Explain why stating "the distribution is skewed" would not earn full credit on an AP free-response question, and rewrite it as a complete phrase.

Answer: You must name the direction of the skew and add context.

Saying only "skewed" is incomplete because skew has a direction determined by the tail. A full-credit phrase names that direction and ties it to the variable, for example: "The distribution of household incomes is skewed right, with a long tail toward higher incomes." Graders look for direction plus context, not just the word skewed.
A histogram of quiz scores is roughly symmetric and unimodal with no unusual values. A student writes: "The center is 82 and the spread is 15." Identify two ways to improve this description.

Answer: Add context/units and name which statistics are used, and state the shape and that there are no outliers.

The description omits shape and the absence of outliers, both required parts of SOCS. It also lacks context: "center is 82" should read "the mean score is about 82 points" and "spread is 15" should specify "a standard deviation of about 15 points." A complete answer might be: "The distribution of quiz scores is roughly symmetric and unimodal with no apparent outliers; the mean is about 82 points with a standard deviation of about 15 points."

FAQ

What does SOCS stand for in AP Statistics?
SOCS stands for Shape, Outliers, Center, and Spread. These are the four features you must describe whenever an AP question asks you to describe the distribution of a quantitative variable, and you should always describe them in the context of the variable and its units.
Do I always have to mention outliers even if there are none?
Yes. If a distribution has no outliers, state that explicitly, for example "there do not appear to be any outliers." Skipping the outlier component is a common way to lose credit, so address it either way.
When should I use the mean versus the median to describe center?
Use the median when the distribution is skewed or has outliers, because the median is resistant to those extremes. Use the mean when the distribution is roughly symmetric with no outliers. Justifying your choice based on shape strengthens your answer.
Can I use SOCS to describe a categorical variable?
No. SOCS applies only to quantitative (numerical) variables. Categorical variables are described using counts, proportions, or the most common category, since concepts like center, spread, and skew do not apply to categories.

Learn this with a teacher, not a page

The Crimsora tutor teaches U1.6 Describing a Distribution (SOCS) live — explaining on a whiteboard, asking you questions, and adapting to where you get stuck.