AP-STATS-2.1-2.3

U2.1 Two Categorical Variables

Master AP Statistics 2.1-2.3: build two-way tables, compute marginal and conditional distributions, read segmented bar graphs, and decide when two categorical variables are associated.

What you'll do in this lesson

A voice-first session with the Crimsora tutor on U2.1 Two Categorical Variables, then targeted practice and FRQs — with the tutor adapting to where you get stuck.

What this lesson covers

When you survey people about two categorical traits at once — say, grade level and preferred lunch, or smoking status and disease — you need a way to see whether the two traits travel together. That is the heart of association between categorical variables. In this lesson you will organize counts into a two-way table, turn raw counts into percents using marginal and conditional distributions, and display those distributions with segmented bar graphs. Most importantly, you will learn the single decision the AP exam asks over and over: do the conditional distributions differ enough to claim an association, or are the variables independent? Get comfortable moving fluidly between counts and percents, because that translation is where most points are won or lost.

Two-Way Tables and Marginal Distributions

A two-way table (also called a contingency table) records the counts of individuals who fall into each combination of two categorical variables. One variable labels the rows, the other labels the columns, and each interior cell is a joint count. The margins — the row totals and column totals — give the totals for each single category, and the grand total sits in the bottom-right corner.

A marginal distribution describes just one of the variables, ignoring the other. You compute it by dividing each row total (or each column total) by the grand total. Marginal distributions answer questions like "what percent of all people surveyed are seniors?" They tell you nothing about the relationship between the two variables — only about one variable on its own.
CoffeeTeaTotal
Morning401050
Evening153550
Total5545100
Here the marginal distribution of drink is 55%55\% coffee and 45%45\% tea. A common misconception is to report a joint count (like 40) as if it answered a marginal question. Always identify whether the question is about one variable alone (marginal), a combination (joint), or one variable given the other (conditional). Reading the question carefully determines which denominator you use.

Conditional Distributions

A conditional distribution describes the distribution of one categorical variable for a specific value (a subgroup) of the other variable. It answers "given that a person drinks in the morning, what percent choose coffee?" To compute it, you restrict attention to one row or one column, then divide each cell in that row or column by that row or column total — not by the grand total.

Using the table above, the conditional distribution of drink given morning is 40/50=80%40/50 = 80\% coffee and 10/50=20%10/50 = 20\% tea. The conditional distribution of drink given evening is 15/50=30%15/50 = 30\% coffee and 35/50=70%35/50 = 70\% tea. Because these two conditional distributions are very different, the two variables appear associated.

The biggest error students make is dividing by the wrong total. The phrase "given" or "among" tells you which subgroup to isolate, and that subgroup's total becomes your denominator. "What percent of morning drinkers chose coffee?" uses 50 (morning total), while "what percent of all people were morning coffee drinkers?" uses 100 (grand total). On the AP exam, conditional distributions are the tool you almost always use to argue whether an association exists, so practice computing them quickly and stating the condition in words.

Segmented Bar Graphs and Comparing Distributions

A segmented bar graph (also called a stacked bar graph) displays conditional distributions visually. Each bar represents one subgroup and stands at a total height of 100%100\%; the bar is divided into segments whose sizes show the conditional percents for the response variable. A side-by-side (clustered) bar graph puts separate bars next to each other for the same comparison. Both let you compare conditional distributions across groups at a glance.

The key reading skill: if the segments line up at the same heights across every bar, the conditional distributions are (nearly) identical, suggesting no association. If the segment heights clearly shift from bar to bar, the conditional distributions differ, suggesting an association.
Graph typeWhat each bar showsBest for
Segmented (stacked)One group's full conditional distribution, totaling 100%Seeing proportions within groups
Side-by-sideSeparate categories placed adjacentlyComparing individual category counts or percents
A frequent misconception is comparing raw-count bars of different total heights and concluding association from the height difference alone. Because groups can have different sizes, always convert to percents (conditional distributions) before comparing. When the AP exam gives you a graph, describe what the segments reveal in context, then connect that description to the association question.

Deciding Whether an Association Exists

Two categorical variables are associated if knowing the value of one variable changes the probabilities (percents) for the other. Operationally: compare the conditional distributions of one variable across the categories of the other. If those conditional distributions are notably different, there is an association. If they are essentially the same, the variables are independent (no association).

Write your conclusion in context and back it with numbers. A strong AP answer sounds like: "Among morning drinkers, 80%80\% chose coffee, but among evening drinkers only 30%30\% chose coffee. Because these conditional distributions differ substantially, there is an association between time of day and drink choice." Notice it cites specific conditional percents and names both variables.

A subtle point: association does not mean causation. Even a strong association could be explained by a lurking variable, so avoid causal language unless the data came from a randomized experiment. Also remember that "association" is symmetric — if drink depends on time, then time depends on drink; the exam may ask you to compare in either direction, and either set of conditional distributions can reveal the same association. Finally, small differences may just be sampling variation; formal testing (chi-square) comes later, so at this stage describe differences as apparent rather than proven.

Key terms

Two-way table.
A table of counts showing how individuals are distributed across the combinations of two categorical variables, with row and column totals in the margins.
Marginal distribution.
The distribution of a single categorical variable alone, found by dividing each row total or column total by the grand total.
Joint distribution.
The distribution across combinations of both variables, found by dividing each interior cell count by the grand total.
Conditional distribution.
The distribution of one variable within a fixed category of the other, found by dividing each cell by its row or column total.
Segmented bar graph.
A bar graph in which each bar reaches 100% and is split into segments representing a conditional distribution.
Association.
A relationship in which the conditional distributions of one variable differ across the categories of the other.
Independence.
The condition in which conditional distributions are the same across categories, so knowing one variable gives no information about the other.

Worked example

A school surveys 200 students, recording whether they play a sport and whether they take an AP course. Of 120 athletes, 90 take an AP course; of 80 non-athletes, 40 take an AP course. Determine whether there is an association between playing a sport and taking an AP course.
First organize the counts into a two-way table. Athletes: 90 take AP, so 12090=30120-90=30 do not. Non-athletes: 40 take AP, so 8040=4080-40=40 do not.
APNo APTotal
Athlete9030120
Non-athlete404080
Total13070200
To test for association, compute the conditional distribution of AP status given sport status. Among athletes, the percent taking AP is 90/120=0.75=75%90/120 = 0.75 = 75\%. Among non-athletes, the percent taking AP is 40/80=0.50=50%40/80 = 0.50 = 50\%.

Now compare. Athletes take AP courses at 75%75\%, while non-athletes take them at only 50%50\%. Because these conditional distributions differ substantially, knowing whether a student is an athlete changes the likelihood they take an AP course.

Conclusion in context: There appears to be an association between playing a sport and taking an AP course, since a higher proportion of athletes (75%75\%) than non-athletes (50%50\%) enroll in AP courses. Because this is observational data, we should not claim that playing a sport causes AP enrollment.

Practice questions

In a two-way table of 300 people classified by region (North, South) and phone type (Android, iPhone), 60% of Northerners use iPhones and 60% of Southerners use iPhones. Based on conditional distributions, which statement is best supported?
  1. There is a strong association between region and phone type
  2. There is no apparent association between region and phone type
  3. Region causes phone choice
  4. iPhones are more popular in the North

Answer: There is no apparent association between region and phone type

Association is judged by comparing conditional distributions. The conditional distribution of phone type given region is identical for both regions (60% iPhone in each), so knowing region tells you nothing about phone type. That means the variables appear independent, so there is no apparent association. Causation cannot be claimed from observational data, and the popularity claim is false because both regions match.
A researcher records 150 people by exercise habit (Regular, None) and sleep quality (Good, Poor). Among 90 regular exercisers, 63 report good sleep. Among 60 with no exercise, 24 report good sleep. Build the relevant conditional distributions and state, with justification in context, whether an association exists.

Answer: Among regular exercisers, 63/90=70%63/90 = 70\% report good sleep; among non-exercisers, 24/60=40%24/60 = 40\% report good sleep. Because these conditional distributions differ substantially, there appears to be an association between exercise habit and sleep quality.

The correct comparison is the conditional distribution of sleep quality given exercise habit, so you divide by each group's own total (90 and 60), not by 150. The resulting percents, 70% versus 40%, are far apart, indicating that knowing exercise habit changes the likelihood of good sleep. A complete answer names both variables in context and, since the data are observational, avoids claiming exercise causes better sleep.

FAQ

How do I know whether to divide by a row total, a column total, or the grand total?
Read what the question conditions on. If it says "among" or "given" a subgroup, divide each cell by that subgroup's row or column total to get a conditional distribution. If it asks about one variable overall, divide the margin total by the grand total for a marginal distribution. If it asks about a combination out of everyone, divide the cell by the grand total for a joint value.
What is the difference between association and independence?
They are opposites. Two categorical variables are associated when their conditional distributions differ across groups, meaning knowing one variable changes the percents for the other. They are independent when the conditional distributions are the same, so knowing one variable gives no information about the other.
Can I claim one variable causes the other if I find an association?
Not from observational data. An association only shows the variables tend to occur together; a lurking variable could explain the pattern. You can only argue causation from a well-designed randomized experiment. On the AP exam, describe associations without causal language unless randomization is present.
How different do conditional distributions have to be to say there is an association?
At this level you judge it informally: if the conditional percents are clearly and consistently different, describe an apparent association; if they are nearly identical, say there is little or no association. A formal test for whether the difference is beyond sampling variation is the chi-square test, which comes in a later unit.

Learn this with a teacher, not a page

The Crimsora tutor teaches U2.1 Two Categorical Variables live — explaining on a whiteboard, asking you questions, and adapting to where you get stuck.