U2.5 Correlation
Master the correlation coefficient r for AP Statistics: how to compute it, interpret its value, and recognize its key properties and limitations.
What you'll do in this lesson
A voice-first session with the Crimsora tutor on U2.5 Correlation, then targeted practice and FRQs — with the tutor adapting to where you get stuck.
What this lesson covers
After describing a scatterplot's direction, unusual features, form, and strength in words (U2.4), you need a single number that pins down how strong and in what direction a linear relationship is. That number is the correlation coefficient . This lesson shows you how is built from standardized scores, what its value actually tells you, and — just as important on the AP exam — what it does NOT tell you. Get comfortable with its bounds, its unit-free nature, and why one stray point can wreck it. These properties show up constantly in multiple-choice traps and free-response justification.
What r Measures and How It's Built
The correlation coefficient measures the strength and direction of a linear association between two quantitative variables. Its formula isNotice what is inside the sum: each variable is converted to a standardized value (a z-score) before multiplying. This is the key idea. You are asking, for each point, whether and are on the same side of their means (both above, or both below) or opposite sides.
When a point is above average in and above average in , both z-scores are positive, so the product is positive. Same for both below average (negative times negative). These points push toward . When a point is above average in one variable but below average in the other, the product is negative, pulling toward .
Because the formula uses z-scores, has no units. It does not matter whether you measure height in inches or centimeters — comes out the same. On the AP exam you will rarely compute by hand for many points; instead you read it from calculator output or a computer regression table. But understanding the z-score structure explains every property below.
When a point is above average in and above average in , both z-scores are positive, so the product is positive. Same for both below average (negative times negative). These points push toward . When a point is above average in one variable but below average in the other, the product is negative, pulling toward .
Because the formula uses z-scores, has no units. It does not matter whether you measure height in inches or centimeters — comes out the same. On the AP exam you will rarely compute by hand for many points; instead you read it from calculator output or a computer regression table. But understanding the z-score structure explains every property below.
The Core Properties of r
The AP exam tests these properties relentlessly, often in multiple-choice questions and in FRQ justifications.
A value of or occurs only when all points lie exactly on a straight line. A common misconception is that means "60% linear" or that it is twice as strong as . It does not work that scale; is not a percentage. (That interpretation belongs to , coming in U2.6.)
Also remember: describes only the LINEAR component. A perfect U-shaped parabola can have even though and are strongly related — just not linearly.
| Property | What it means |
|---|---|
| Bounds | always |
| Sign | Positive = positive (upward) linear trend; negative = downward trend |
| Magnitude | Closer to = stronger linear pattern; near = weak or no linear pattern |
| Unit-free | Changing units of or does not change |
| Symmetric | The correlation of with equals the correlation of with |
| Requires two quantitative variables | is undefined for categorical data |
Also remember: describes only the LINEAR component. A perfect U-shaped parabola can have even though and are strongly related — just not linearly.
Limitations: Outliers, Linearity, and Causation
Three limitations generate most exam mistakes.
First, is not resistant — it is strongly affected by outliers. Because a single point far from the pattern can have large z-scores, its product dominates the sum and drags up or down. If a question shows a scatterplot with one influential point and asks what happens to when it is removed, expect a noticeable change.
Second, measures only linear association. A high does not confirm the relationship is linear, and a low does not mean there is no relationship — it means no strong LINEAR one. Always look at the scatterplot. This is why the DUFS description (U2.4) and work together.
Third, and heavily tested: correlation does not imply causation. A strong between ice cream sales and drownings does not mean ice cream causes drowning; a lurking variable (hot weather) drives both. On the FRQ, if you claim one variable causes another based only on correlation, you lose the point. The safe language is "associated with" or "tends to be higher when."
Finally, says nothing about the steepness of a line. A very steep and a very shallow trend can share the same ; slope and correlation are different quantities.
First, is not resistant — it is strongly affected by outliers. Because a single point far from the pattern can have large z-scores, its product dominates the sum and drags up or down. If a question shows a scatterplot with one influential point and asks what happens to when it is removed, expect a noticeable change.
Second, measures only linear association. A high does not confirm the relationship is linear, and a low does not mean there is no relationship — it means no strong LINEAR one. Always look at the scatterplot. This is why the DUFS description (U2.4) and work together.
Third, and heavily tested: correlation does not imply causation. A strong between ice cream sales and drownings does not mean ice cream causes drowning; a lurking variable (hot weather) drives both. On the FRQ, if you claim one variable causes another based only on correlation, you lose the point. The safe language is "associated with" or "tends to be higher when."
Finally, says nothing about the steepness of a line. A very steep and a very shallow trend can share the same ; slope and correlation are different quantities.
How the Exam Tests Correlation
Expect three question styles.
Interpretation in context: given for hours of TV and test score, you should say something like, "There is a strong, negative, linear association between hours of TV watched and test score." Always include strength, direction, linear, and the variable names in context — a bare "strong negative correlation" often misses full credit.
Property reasoning: multiple-choice items ask what happens to if you add a constant, multiply by a constant, swap the axes, or change units. Because is unit-free and symmetric, all of these leave unchanged. Adding an outlier, however, changes it.
Matching scatterplots to values: you may be shown four plots and asked which has . Judge both the direction (sign) and how tightly points cluster around a line. Watch for a plot with an obvious curve — its can be misleadingly small or near zero.
A subtle trap: a question describes a clear curved relationship and reports , then asks if there is "no relationship." The correct response is that there is no strong LINEAR relationship, but there may be a strong nonlinear one. Never let alone override what the scatterplot shows.
Interpretation in context: given for hours of TV and test score, you should say something like, "There is a strong, negative, linear association between hours of TV watched and test score." Always include strength, direction, linear, and the variable names in context — a bare "strong negative correlation" often misses full credit.
Property reasoning: multiple-choice items ask what happens to if you add a constant, multiply by a constant, swap the axes, or change units. Because is unit-free and symmetric, all of these leave unchanged. Adding an outlier, however, changes it.
Matching scatterplots to values: you may be shown four plots and asked which has . Judge both the direction (sign) and how tightly points cluster around a line. Watch for a plot with an obvious curve — its can be misleadingly small or near zero.
A subtle trap: a question describes a clear curved relationship and reports , then asks if there is "no relationship." The correct response is that there is no strong LINEAR relationship, but there may be a strong nonlinear one. Never let alone override what the scatterplot shows.
Key terms
- Correlation coefficient (r).
- A number between and that measures the strength and direction of a linear association between two quantitative variables.
- Standardized value (z-score).
- How many standard deviations a value lies from its mean, ; correlation is built from products of z-scores.
- Unit-free.
- A property meaning the value of does not depend on the units of measurement of either variable.
- Resistant.
- A statistic is resistant if it is not greatly affected by outliers; is NOT resistant.
- Linear association.
- A relationship whose points cluster around a straight line; measures only this type of association.
- Lurking variable.
- A variable not included in the analysis that influences both variables studied, often explaining a correlation without causation.
- Direction.
- Whether a linear relationship trends upward (positive ) or downward (negative ).
Worked example
A researcher records the number of study hours () and quiz score () for 5 students. The summary statistics are , , , . For the five students the sum of the products of standardized scores . Compute and interpret it. Then state what happens to if study hours are recorded in minutes instead of hours.
Start with the correlation formula written in terms of z-scores:Here , so , and the sum of z-score products is given as .Interpretation in context: There is a strong, positive, linear association between the number of hours studied and quiz score. Students who studied more hours tended to score higher.
Note we do NOT say studying more hours causes higher scores — this is observational data, so we only describe association.
Now, converting hours to minutes multiplies every value by 60. This changes and by the same factor, so each standardized value is unchanged. Because depends only on z-scores, stays exactly . This demonstrates that is unit-free.
Note we do NOT say studying more hours causes higher scores — this is observational data, so we only describe association.
Now, converting hours to minutes multiplies every value by 60. This changes and by the same factor, so each standardized value is unchanged. Because depends only on z-scores, stays exactly . This demonstrates that is unit-free.
Practice questions
A scatterplot of two quantitative variables shows a clear U-shaped (parabolic) pattern, and the computed correlation is . Which statement is the best interpretation?
- There is no relationship between the two variables
- There is a strong positive linear relationship between the variables
- There is no strong linear relationship, but there may be a strong nonlinear relationship
- The correlation must have been computed incorrectly because the points show a clear pattern
Answer: There is no strong linear relationship, but there may be a strong nonlinear relationship
Correlation measures only linear association. A U-shaped pattern is a strong relationship, but it is nonlinear, so can be near zero. This does not mean the variables are unrelated, nor that the computation is wrong — it correctly reflects the weak LINEAR component. This is exactly the trap the exam sets.
The correlation between two variables is . A researcher then multiplies every -value by 3 and adds 10, and separately measures in different units. Describe the effect on the correlation, and explain why.
Answer: The correlation remains .
Correlation is built from standardized scores (z-scores). Multiplying a variable by a positive constant and adding a constant is a linear transformation that changes the mean and standard deviation proportionally, leaving every z-score unchanged. Changing units of does the same. Since depends only on z-scores, none of these operations affects it — this reflects the unit-free property. (Multiplying by a negative constant would flip the sign, but that was not done here.)
Two datasets each contain 20 points. Dataset A has points tightly clustered along an upward line. Dataset B has the same 20 points plus one additional point far below and to the right of the pattern. Which dataset likely has the correlation closer to 1, and why?
Answer: Dataset A, because r is not resistant to outliers.
Because is not resistant, the single outlier in Dataset B produces large-magnitude z-score products with the opposite sign of the main trend, dragging the correlation away from 1. Dataset A, with tightly clustered upward points and no outlier, keeps closer to 1. This illustrates that a single influential point can substantially change the correlation.
FAQ
- What is the difference between r and r-squared?
- is the correlation coefficient, ranging from to , and shows direction and strength of a linear association. (the coefficient of determination, covered in U2.6) is the square of that value and represents the proportion of variation in explained by the linear model. Because it is squared, is always between 0 and 1 and loses the sign information.
- Does a high correlation mean one variable causes the other?
- No. Correlation never establishes causation on its own. A strong can be produced by a lurking variable influencing both, or by coincidence. Only a well-designed randomized experiment can support a cause-and-effect claim. On the AP exam, describe relationships using words like 'associated with,' not 'causes.'
- What counts as a strong correlation?
- There is no official cutoff, and the AP exam does not expect a fixed number. Loosely, values near or beyond are often called strong, around moderate, and near weak, but always judge in context and look at the scatterplot rather than relying on a threshold.
- Can correlation be used with categorical variables?
- No. The correlation coefficient requires two quantitative variables because it is computed from means and standard deviations. Relationships between categorical variables are analyzed with two-way tables and related tools (see U2.1), not with .
Learn this with a teacher, not a page
The Crimsora tutor teaches U2.5 Correlation live — explaining on a whiteboard, asking you questions, and adapting to where you get stuck.