Scatterplots, Trend Lines & Correlation
Learn to read scatterplots, describe direction and strength, fit a trend line, write and interpret its equation, predict values, and tell correlation from causation.
What you'll do in this lesson
A voice-first session with the Crimsora tutor on Scatterplots, Trend Lines & Correlation, then targeted practice and FRQs — with the tutor adapting to where you get stuck.
What this lesson covers
In this lesson you will describe an association by its direction (positive, negative, or none), its form (linear or nonlinear), and its strength (how tightly the points hug a line). Then you will fit a trend line, interpret its slope and its -intercept in the actual units of the problem, predict with it, and — just as important — recognize when a strong association still does not mean one variable causes the other.
Reading the Association: Direction, Form, and Strength
Describe every scatterplot with three words plus context.
Direction. A positive association means that as increases, tends to increase — the cloud of points rises left to right. A negative association means tends to decrease as increases. If there is no consistent rise or fall, there is no association.
Form. Is the pattern roughly a straight line, or does it curve? Only linear patterns should get a straight trend line. A curved pattern with a line forced through it will produce badly wrong predictions at the ends.
Strength. How closely do the points cluster around the pattern? Tight clustering is a strong association; a wide, fuzzy cloud is weak.
| Description | What you see |
|---|---|
| Strong positive linear | Points rise steadily, nearly on one line |
| Weak negative linear | Points drift downward but scatter widely |
| Nonlinear | Clear pattern, but it bends |
| No association | Shapeless blob, no rise or fall |
Fitting a Trend Line and Writing Its Equation
To get the equation, pick two convenient points on your line (grid intersections are easiest), then compute slope and use point-slope or slope-intercept form:Suppose your line passes through and . Then . Substituting into with gives , so and the equation is .
A graphing calculator or spreadsheet can compute the least-squares regression line, the line that makes the total squared vertical distance from the points as small as possible. Different students will draw slightly different trend lines by hand, and that is expected; least squares gives one agreed-upon answer. Technology also reports , the correlation coefficient, a number from to . Its sign matches the direction and its size measures linear strength: near is strong, near is weak. Note that means no linear pattern — a perfect U-shape can have near zero.
Interpreting Slope and Intercept in Context
Slope is the predicted change in for each one-unit increase in . If is hours studied and is quiz score out of 100, then means: for each additional hour of study, the model predicts a quiz score about 3 points higher. Two habits make interpretations complete — include the units of both variables, and use hedging words like "predicted" or "on average," because the line describes a tendency, not a guarantee for every student.
The -intercept is the predicted when . Here predicts a score of 8 for a student who studies zero hours. Sometimes that is meaningful; sometimes it is nonsense. If is the height of a tree in meters and is its age, an intercept at describes a tree with no height, which is outside any sensible range. Always ask whether falls inside the data you actually have.
That question leads to extrapolation: predicting far outside the range of the observed -values. A model built from students studying 0 to 6 hours says nothing reliable about 40 hours — the line would predict a score of 128, which is impossible. Predicting inside the data range is interpolation and is generally trustworthy. Also remember the residual, the difference ; a positive residual means the real point sits above your line.
Correlation Is Not Causation
When a strong association appears, at least four explanations are possible:
| Explanation | Example |
|---|---|
| causes | More fertilizer causes more plant growth |
| causes | Better health causes more exercise, not only the reverse |
| A lurking variable causes both | Hot weather drives ice cream sales and swimming |
| Coincidence in the sample | Two unrelated trends both rise over a decade |
When a question asks whether the data prove that one variable causes the other, a complete answer names a plausible lurking variable and says that only a randomized experiment could settle the question. Writing "no, correlation is not causation" and stopping there leaves out the reasoning that shows you understand why. Where students most often go wrong is sliding causal language into an interpretation: "each extra hour of study raises the score by 3 points" claims cause, while "is associated with a 3-point higher predicted score" states exactly what the data support.
Key terms
- Scatterplot.
- A graph of paired numerical data in which each point represents one individual, plotted as (explanatory variable, response variable).
- Positive association.
- A relationship in which tends to increase as increases; the points rise from left to right.
- Trend line (line of best fit).
- A straight line drawn or computed to summarize a linear pattern in a scatterplot, passing through the middle of the point cloud.
- Correlation coefficient .
- A number between and whose sign gives the direction and whose absolute value gives the strength of a linear association.
- Residual.
- The difference between an observed value and the value predicted by the trend line: .
- Extrapolation.
- Using a model to predict for -values outside the range of the collected data, where the pattern may no longer hold.
- Lurking variable.
- A variable not included in the study that influences both of the plotted variables and can create an association without causation.
- Least-squares regression line.
- The unique trend line that minimizes the sum of the squared vertical distances between the data points and the line.
Worked example
(b) The slope means that for each additional year of age, the model predicts the selling price drops by about 14 hundred dollars, that is, about 1,400 dollars. The intercept means a brand-new car (age 0) is predicted to sell for about 176 hundred dollars, or about 17,600 dollars. Since the data only start at age 1, an age of 0 sits just outside the observed range, so treat that figure as an estimate rather than a measured fact.
(c) Substitute :The predicted price is about 92 hundred dollars, roughly 9,200 dollars. Age 6 is inside the data range of 1 to 10 years, so this interpolation is reasonable.
(d) No. Age 25 is far outside the range of 1 to 10 years. The model would predict , a negative price, which is impossible. This shows why extrapolating far beyond the data breaks down.
Practice questions
A scatterplot of daily high temperature (, in degrees Fahrenheit) versus number of hot chocolates sold () shows points falling steadily from upper left to lower right, tightly clustered around a line. Which description is best?
- Strong negative linear association
- Weak negative linear association
- Strong positive linear association
- No association
Answer: Strong negative linear association
A researcher finds that towns with more firefighters at a blaze tend to have more fire damage, with a strong positive correlation. Does this show that sending more firefighters causes more damage? Explain fully.
Answer: No. The size of the fire is a lurking variable: large fires both attract more firefighters and cause more damage. Only a randomized experiment could establish causation, and such correlational data cannot.
A trend line for plant height (cm) versus days since planting is , based on data from day 5 through day 30. Interpret the slope and the intercept, and state one prediction you should not make with this model.
Answer: Slope: each additional day is associated with a predicted increase of about 0.8 cm in height. Intercept: at day 0 the model predicts a height of 3 cm. You should not predict the height at, say, day 200, because that is far outside the range of days 5 to 30.
FAQ
- What is the difference between the strength of a correlation and the steepness of the trend line?
- They measure completely different things. Strength is about how closely the points hug the line — measured by and visible as tight versus fuzzy scatter. Steepness is the slope, which tells how much changes per unit of . A line with slope can have if every point sits right on it, and a line with slope can have if the points are widely scattered.
- Does a trend line have to pass through data points?
- No. A trend line summarizes the overall pattern, so it may pass through several points, one point, or none at all. What matters is that it runs through the center of the cloud with roughly equal numbers of points above and below and follows the data's direction. Connecting the first and last points is a common mistake because it lets two values, possibly outliers, determine the whole line.
- When is it okay to use a trend line to predict?
- Predict for -values inside the range of your data (interpolation) and only when the pattern really is linear. Predicting far outside the range (extrapolation) assumes the pattern continues, which often fails — it can yield negative prices, impossible test scores, or heights no plant reaches. Always check that your predicted value makes physical sense.
- How can I tell whether an association is causal?
- From a scatterplot alone, you cannot. Association only shows that two variables move together. Causation requires a well-designed experiment with random assignment to treatment groups, which spreads lurking variables evenly so that outcome differences can be attributed to the treatment. With observational data, always ask what third variable might be influencing both quantities.
Learn this with a teacher, not a page
The Crimsora tutor teaches Scatterplots, Trend Lines & Correlation live — explaining on a whiteboard, asking you questions, and adapting to where you get stuck.