U2.1 Two Categorical Variables
Master AP Statistics 2.1-2.3: build two-way tables, compute marginal and conditional distributions, read segmented bar graphs, and decide when two categorical variables are associated.
What you'll do in this lesson
A voice-first session with the Crimsora tutor on U2.1 Two Categorical Variables, then targeted practice and FRQs — with the tutor adapting to where you get stuck.
What this lesson covers
Two-Way Tables and Marginal Distributions
A marginal distribution describes just one of the variables, ignoring the other. You compute it by dividing each row total (or each column total) by the grand total. Marginal distributions answer questions like "what percent of all people surveyed are seniors?" They tell you nothing about the relationship between the two variables — only about one variable on its own.
| Coffee | Tea | Total | |
|---|---|---|---|
| Morning | 40 | 10 | 50 |
| Evening | 15 | 35 | 50 |
| Total | 55 | 45 | 100 |
Conditional Distributions
Using the table above, the conditional distribution of drink given morning is coffee and tea. The conditional distribution of drink given evening is coffee and tea. Because these two conditional distributions are very different, the two variables appear associated.
The biggest error students make is dividing by the wrong total. The phrase "given" or "among" tells you which subgroup to isolate, and that subgroup's total becomes your denominator. "What percent of morning drinkers chose coffee?" uses 50 (morning total), while "what percent of all people were morning coffee drinkers?" uses 100 (grand total). On the AP exam, conditional distributions are the tool you almost always use to argue whether an association exists, so practice computing them quickly and stating the condition in words.
Segmented Bar Graphs and Comparing Distributions
The key reading skill: if the segments line up at the same heights across every bar, the conditional distributions are (nearly) identical, suggesting no association. If the segment heights clearly shift from bar to bar, the conditional distributions differ, suggesting an association.
| Graph type | What each bar shows | Best for |
|---|---|---|
| Segmented (stacked) | One group's full conditional distribution, totaling 100% | Seeing proportions within groups |
| Side-by-side | Separate categories placed adjacently | Comparing individual category counts or percents |
Deciding Whether an Association Exists
Write your conclusion in context and back it with numbers. A strong AP answer sounds like: "Among morning drinkers, chose coffee, but among evening drinkers only chose coffee. Because these conditional distributions differ substantially, there is an association between time of day and drink choice." Notice it cites specific conditional percents and names both variables.
A subtle point: association does not mean causation. Even a strong association could be explained by a lurking variable, so avoid causal language unless the data came from a randomized experiment. Also remember that "association" is symmetric — if drink depends on time, then time depends on drink; the exam may ask you to compare in either direction, and either set of conditional distributions can reveal the same association. Finally, small differences may just be sampling variation; formal testing (chi-square) comes later, so at this stage describe differences as apparent rather than proven.
Key terms
- Two-way table.
- A table of counts showing how individuals are distributed across the combinations of two categorical variables, with row and column totals in the margins.
- Marginal distribution.
- The distribution of a single categorical variable alone, found by dividing each row total or column total by the grand total.
- Joint distribution.
- The distribution across combinations of both variables, found by dividing each interior cell count by the grand total.
- Conditional distribution.
- The distribution of one variable within a fixed category of the other, found by dividing each cell by its row or column total.
- Segmented bar graph.
- A bar graph in which each bar reaches 100% and is split into segments representing a conditional distribution.
- Association.
- A relationship in which the conditional distributions of one variable differ across the categories of the other.
- Independence.
- The condition in which conditional distributions are the same across categories, so knowing one variable gives no information about the other.
Worked example
| AP | No AP | Total | |
|---|---|---|---|
| Athlete | 90 | 30 | 120 |
| Non-athlete | 40 | 40 | 80 |
| Total | 130 | 70 | 200 |
Now compare. Athletes take AP courses at , while non-athletes take them at only . Because these conditional distributions differ substantially, knowing whether a student is an athlete changes the likelihood they take an AP course.
Conclusion in context: There appears to be an association between playing a sport and taking an AP course, since a higher proportion of athletes () than non-athletes () enroll in AP courses. Because this is observational data, we should not claim that playing a sport causes AP enrollment.
Practice questions
In a two-way table of 300 people classified by region (North, South) and phone type (Android, iPhone), 60% of Northerners use iPhones and 60% of Southerners use iPhones. Based on conditional distributions, which statement is best supported?
- There is a strong association between region and phone type
- There is no apparent association between region and phone type
- Region causes phone choice
- iPhones are more popular in the North
Answer: There is no apparent association between region and phone type
A researcher records 150 people by exercise habit (Regular, None) and sleep quality (Good, Poor). Among 90 regular exercisers, 63 report good sleep. Among 60 with no exercise, 24 report good sleep. Build the relevant conditional distributions and state, with justification in context, whether an association exists.
Answer: Among regular exercisers, report good sleep; among non-exercisers, report good sleep. Because these conditional distributions differ substantially, there appears to be an association between exercise habit and sleep quality.
FAQ
- How do I know whether to divide by a row total, a column total, or the grand total?
- Read what the question conditions on. If it says "among" or "given" a subgroup, divide each cell by that subgroup's row or column total to get a conditional distribution. If it asks about one variable overall, divide the margin total by the grand total for a marginal distribution. If it asks about a combination out of everyone, divide the cell by the grand total for a joint value.
- What is the difference between association and independence?
- They are opposites. Two categorical variables are associated when their conditional distributions differ across groups, meaning knowing one variable changes the percents for the other. They are independent when the conditional distributions are the same, so knowing one variable gives no information about the other.
- Can I claim one variable causes the other if I find an association?
- Not from observational data. An association only shows the variables tend to occur together; a lurking variable could explain the pattern. You can only argue causation from a well-designed randomized experiment. On the AP exam, describe associations without causal language unless randomization is present.
- How different do conditional distributions have to be to say there is an association?
- At this level you judge it informally: if the conditional percents are clearly and consistently different, describe an apparent association; if they are nearly identical, say there is little or no association. A formal test for whether the difference is beyond sampling variation is the chi-square test, which comes in a later unit.
Learn this with a teacher, not a page
The Crimsora tutor teaches U2.1 Two Categorical Variables live — explaining on a whiteboard, asking you questions, and adapting to where you get stuck.