U1.1 Variables and Categorical Data
Master AP Stats 1.1-1.4: tell categorical from quantitative variables and build frequency, relative frequency, and two-way tables plus bar and pie charts.
What you'll do in this lesson
A voice-first session with the Crimsora tutor on U1.1 Variables and Categorical Data, then targeted practice and FRQs — with the tutor adapting to where you get stuck.
What this lesson covers
These skills anchor the entire course. The AP exam constantly asks you to identify variable type before choosing a graph or test, and two-way tables reappear in probability (Unit 4) and inference for categorical data. Get comfortable now, and later units feel far easier.
Categorical vs. Quantitative Variables
A categorical variable places each individual into a group or category. Examples include eye color, favorite sport, zip code, and grade level (freshman, sophomore, etc.). A quantitative variable takes numerical values for which arithmetic like averaging makes sense — height, test score, number of siblings, temperature.
The classic trap is a number that is really a category. A phone area code or a jersey number is stored as digits, but averaging them is meaningless, so they are categorical. Ask yourself: does it make sense to compute a mean? If not, it is categorical.
| Feature | Categorical | Quantitative |
|---|---|---|
| Values | Labels/groups | Numbers with meaning |
| Summary | Counts, proportions | Mean, median, spread |
| Displays | Bar, pie | Dot, stem, histogram, box |
| Example | Blood type | Weight in kg |
Frequency and Relative Frequency Tables
A relative frequency table converts each count to a proportion or percent by dividing by the total: . Relative frequencies always sum to (or , allowing for tiny rounding). They matter because they let you compare groups of different sizes fairly.
Suppose 12 of 40 students chose pizza. The relative frequency is , or . If a second, larger class also had pizza fans, relative frequency reveals the similarity that raw counts would hide.
A common misconception is treating relative frequency as a count. Always label whether a value is a count or a proportion. On the AP exam, expect to be asked to compute a proportion from a table, or to explain why relative frequencies are preferred when comparing groups of unequal size. Show the fraction you divide, not just the final decimal, so a reader can follow your reasoning.
Two-Way Tables: Joint, Marginal, and Conditional
| Coffee | Tea | Total | |
|---|---|---|---|
| Under 30 | 40 | 20 | 60 |
| 30+ | 15 | 25 | 40 |
| Total | 55 | 45 | 100 |
Conditional distributions are how you check for an association between the two variables. If the conditional distribution of drink preference differs across age groups, the variables are associated; if the conditionals are essentially identical, there is no association. The exam frequently asks you to compute a conditional proportion and then state, in context, whether an association exists. Read carefully which total the question conditions on — that choice changes the denominator and the answer.
Bar Charts and Pie Charts
A pie chart shows each category as a slice of a circle, with slice size proportional to its relative frequency. To find a slice's angle, multiply its proportion by : a category spans .
Bar charts are usually more useful because the human eye compares bar heights more accurately than pie-slice areas, and bar charts can display counts, percentages, or side-by-side and segmented comparisons of two variables. A segmented (stacked) bar chart displays conditional distributions and is ideal for spotting association.
Watch for misleading graphs: a vertical axis that does not start at zero exaggerates differences, and reordering categories can hide patterns. On free-response questions, always label axes, include a scale, and title the graph. When asked to compare two groups, use relative frequencies so unequal group sizes do not distort the picture. Never use a pie chart for two variables at once — reach for a segmented or side-by-side bar chart instead.
Key terms
- Categorical variable.
- A variable that assigns each individual to a group or label, such as gender, color, or region; summarized with counts and proportions.
- Quantitative variable.
- A numerical variable for which arithmetic operations like averaging are meaningful, such as height, age, or score.
- Relative frequency.
- A category's count divided by the total number of individuals, expressed as a proportion or percent; relative frequencies sum to 1.
- Two-way table.
- A table displaying counts for two categorical variables simultaneously, with rows for one variable and columns for the other.
- Marginal relative frequency.
- A row or column total divided by the grand total, giving the distribution of a single variable ignoring the other.
- Conditional relative frequency.
- A cell divided by its row or column total, giving the distribution of one variable within a fixed category of the other.
- Association.
- A relationship in which the conditional distributions of one variable differ across categories of another variable.
- Segmented bar chart.
- A stacked bar chart in which each bar is divided into parts showing a conditional distribution, used to compare groups.
Worked example
Long commuters total . The marginal relative frequency of long commutes is , about .
For the conditional relative frequency of a long commute given biking, restrict to bike users: , or .
Now compare conditionals across modes. Car: . Bus: . Bike: . Because these conditional proportions differ substantially ( vs. vs. ), commute length is associated with transportation mode: car users are far more likely to have long commutes than bikers. A segmented bar chart of the three conditional distributions would show this pattern visually.
Practice questions
A researcher records each student's number on their sports jersey. What type of variable is jersey number, and why?
- Quantitative, because it is written using digits
- Quantitative, because averaging jersey numbers is meaningful
- Categorical, because the numbers serve as labels and averaging them is meaningless
- Categorical, because there are only a few possible values
Answer: Categorical, because the numbers serve as labels and averaging them is meaningless
In a two-way table of 300 people classified by gender (male, female) and pet preference (dog, cat), 80 males prefer dogs. The table shows 150 males total and 180 dog-lovers total. Compute (a) the joint relative frequency of male dog-lovers, (b) the conditional relative frequency of preferring dogs given male, and explain what each proportion describes.
Answer: (a) Joint relative frequency = 80/300 ≈ 0.267; (b) conditional relative frequency = 80/150 ≈ 0.533.
Explain why a segmented bar chart is generally a better choice than two separate pie charts for deciding whether two categorical variables are associated.
Answer: A segmented bar chart lets you directly compare conditional distributions across groups, making differences (and therefore association) easy to see, while pie charts force less accurate area comparisons.
FAQ
- How do I quickly tell if a variable is categorical or quantitative?
- Ask whether averaging the values makes sense. If computing a mean is meaningful (heights, scores, ages), it is quantitative. If the values are labels or groups even when written as numbers (zip codes, area codes, jersey numbers), it is categorical.
- What is the difference between marginal and conditional relative frequency?
- A marginal relative frequency uses a row or column total over the grand total and describes one variable alone. A conditional relative frequency uses a cell over a single row or column total, describing one variable within a fixed category of the other. Conditionals are what you compare to check for association.
- When should I use a bar chart versus a histogram?
- Use a bar chart for categorical data; its bars have gaps because categories are distinct. Use a histogram for quantitative data, where bars touch to show continuous intervals. Choosing the wrong one signals a misunderstanding of variable type on the exam.
- How do I show that two variables are associated using a table?
- Compute the conditional distribution of one variable for each category of the other. If these conditional distributions differ noticeably, the variables are associated. If they are essentially the same across categories, there is no association. Always state your conclusion in context.
Learn this with a teacher, not a page
The Crimsora tutor teaches U1.1 Variables and Categorical Data live — explaining on a whiteboard, asking you questions, and adapting to where you get stuck.