Two-Variable Data: Scatterplots & Lines of Best Fit
Master Digital SAT scatterplots: describe association, read lines of best fit, interpret slope and intercept in context, and judge residuals and extrapolation.
What you'll do in this lesson
A voice-first session with the Crimsora tutor on Two-Variable Data: Scatterplots & Lines of Best Fit, then targeted practice and FRQs — with the tutor adapting to where you get stuck.
What this lesson covers
This lesson walks through describing association, using the fitted line to make predictions, interpreting slope and intercept in context, and evaluating residuals and the danger of extrapolation. By the end, you'll know exactly what the test is asking when it shows you a cloud of dots and a straight line running through them.
Describing Association in a Scatterplot
Three features matter: direction, form, and strength. Direction is positive (dots rise left to right), negative (dots fall), or none. Form is linear (dots follow a straight-line pattern) or nonlinear (a curve, such as quadratic or exponential). Strength describes how tightly the dots cluster around the pattern — strong means little scatter, weak means lots of scatter.
| Feature | What to look for |
|---|---|
| Direction | Do dots rise (positive) or fall (negative)? |
| Form | Straight line or a curve? |
| Strength | Tightly clustered (strong) or spread out (weak)? |
Reading and Using the Line of Best Fit
To predict a value, substitute a known variable and solve. If the line is and you want the predicted when , compute . To find the that produces a given , set equal and solve for .
Read carefully whether the question wants a value predicted by the line or an actual data point read from the graph. These often differ. A question might say "according to the line of best fit" — use the equation. If it says "based on the data" or points to a specific dot, read the scatterplot directly.
Watch units and scale. Axis labels frequently include units like "thousands of dollars" or "minutes," and gridlines may count by 2, 5, or 25. Misreading the scale is the most common avoidable error. Always check what one gridline represents before you estimate a coordinate.
Interpreting Slope and Intercept in Context
The slope is the predicted change in for each one-unit increase in . Always state it with units: "For each additional hour studied, the predicted score increases by points." A negative slope means decreases as increases.
The -intercept is the predicted value of when . In context this is a starting value or baseline — but only meaningful if falls within a sensible range. Sometimes is impossible (a person of height 0), so the intercept is just a mathematical anchor, not a real quantity.
| Symbol | Meaning | Context template |
|---|---|---|
| rate of change | "per one-unit increase in , changes by " | |
| value at | "when is 0, predicted is " |
Residuals and the Limits of Prediction
To compute a residual, read the actual from the data point, plug that point's into the line's equation to get the predicted , and subtract. For example, if a point is at and the line predicts , the residual is — the point is above the line.
Reliability depends on where you predict. Interpolation (predicting within the range of the data) is generally reliable if the association is strong and linear. Extrapolation (predicting far outside the data range) is risky: the linear pattern may not continue, so those predictions can be badly off. If a question asks whether a prediction for is trustworthy when data only span to , the answer is that it is unreliable because it extrapolates well beyond the observed data. Recognizing this distinction is a favorite Digital SAT concept.
Key terms
- Association.
- The relationship between two variables in a scatterplot, described by direction (positive/negative), form (linear/nonlinear), and strength (strong/weak).
- Line of best fit.
- The straight line that minimizes the total squared vertical distance to the data points; used to model the trend and make predictions, written as .
- Slope.
- The rate of change in a linear model; the predicted change in for each one-unit increase in .
- y-intercept.
- The value ; the predicted value of when , meaningful only if is realistic for the data.
- Residual.
- The observed value minus the predicted value: ; positive means the point is above the line, negative means below.
- Interpolation.
- Predicting a value within the range of the observed data, generally reliable for a strong linear pattern.
- Extrapolation.
- Predicting a value outside the range of the observed data, which is unreliable because the trend may not continue.
Worked example
Part (b): Substitute into the line: . The predicted temperature at 15 minutes is 65 degrees Celsius.
Part (c): First find the predicted value at : . The actual value is 78. The residual is . The positive residual means this point lies above the line, so the line underestimated the temperature here.
Part (d): The data only cover 0 to 25 minutes, but is far outside that range. This is extrapolation, so the prediction is unreliable — the cooling likely slows and levels off near room temperature rather than continuing to drop 1.8 degrees per minute, which would give an unrealistic negative temperature.
Practice questions
A scatterplot relating a car's age (in years) to its resale value (in thousands of dollars) has line of best fit . Which statement best interpprets the slope?
- For each additional year of age, the predicted resale value decreases by 1.4 thousand dollars.
- For each additional year of age, the predicted resale value increases by 1.4 thousand dollars.
- When the car is new, its predicted resale value is 1.4 thousand dollars.
- For each additional 1.4 years, the resale value decreases by 1 thousand dollars.
Answer: For each additional year of age, the predicted resale value decreases by 1.4 thousand dollars.
Using the same line , a 5-year-old car actually sold for 17 thousand dollars. Find the residual and state whether the point lies above or below the line of best fit.
Answer: The residual is thousand dollars, and the point lies above the line.
The data used to build the line covered cars from 1 to 9 years old. Explain why using the line to predict the resale value of a 30-year-old car may be unreliable.
Answer: Because is far outside the 1-to-9-year range of the data, the prediction is an extrapolation and the linear trend may not hold that far out.
FAQ
- How do I tell the difference between the line of best fit prediction and an actual data point?
- Read the question wording. Phrases like "according to the line of best fit" or "predicted by the model" mean you plug into the equation . Phrases like "based on the actual data" or references to a specific plotted point mean you read the dot's coordinates directly from the graph. The two values often differ, and that difference is the residual.
- Does a strong association mean one variable causes the other?
- No. Association is not causation. Even a tight linear pattern only shows that two variables move together; a third factor could drive both, or the link could be coincidental. The Digital SAT will not ask you to claim causation from a scatterplot alone.
- What exactly is a residual and how do I compute it?
- A residual is the actual observed -value minus the value the line predicts at that same : . Compute the predicted value by substituting the point's into the line's equation, then subtract. Positive residuals sit above the line; negative ones sit below.
- When is a prediction from the line trustworthy?
- Predictions are most reliable when the association is strong and linear and when the -value falls within the range of the collected data (interpolation). Predicting far outside that range (extrapolation) is risky because the trend may change, sometimes producing impossible values like negative prices or temperatures.
Learn this with a teacher, not a page
The Crimsora tutor teaches Two-Variable Data: Scatterplots & Lines of Best Fit live — explaining on a whiteboard, asking you questions, and adapting to where you get stuck.