U4.1 Probability Foundations
Master AP Statistics probability foundations: long-run relative frequency, the basic probability rules, complements, addition for mutually exclusive events, and simulation.
What you'll do in this lesson
A voice-first session with the Crimsora tutor on U4.1 Probability Foundations, then targeted practice and FRQs — with the tutor adapting to where you get stuck.
What this lesson covers
Every big idea in AP Statistics Unit 4 rests on one question: what does it actually mean to say the probability of something is ? In this lesson you'll learn that probability describes the long-run behavior of a random process, not what happens on any single trial. You'll pin down the axioms every probability must obey, use the complement rule to work smarter, add probabilities for events that can't happen together, and design simulations to estimate probabilities when exact math is hard. These tools are the grammar of the whole unit, so getting them precise now pays off through conditional probability, random variables, and the binomial and geometric distributions later.
Probability as Long-Run Relative Frequency
The AP definition of probability is the long-run relative frequency of an outcome: if you repeated a random process an enormous number of times, the proportion of times an event occurs settles down to a fixed number, and that number is .
The key word is long-run. In the short run, results are unpredictable and can look lopsided. Flip a fair coin four times and you might get four heads; that does not make . As the number of trials grows, the cumulative proportion of heads converges toward . This convergence is the law of large numbers.
A common misconception the exam loves to test is the gambler's fallacy — the false belief that a run of heads makes tails "due" on the next flip. Independent trials have no memory. The long-run proportion self-corrects only because early results get diluted by a growing number of future trials, not because the process compensates.
Probability applies to random processes: procedures whose individual outcomes are uncertain but whose long-run pattern is predictable. On free-response questions, a strong answer explicitly references "in the long run" or "over many repetitions" when interpreting a probability. Saying "it will happen 70% of the time" for a single trial is imprecise and can cost you.
The key word is long-run. In the short run, results are unpredictable and can look lopsided. Flip a fair coin four times and you might get four heads; that does not make . As the number of trials grows, the cumulative proportion of heads converges toward . This convergence is the law of large numbers.
A common misconception the exam loves to test is the gambler's fallacy — the false belief that a run of heads makes tails "due" on the next flip. Independent trials have no memory. The long-run proportion self-corrects only because early results get diluted by a growing number of future trials, not because the process compensates.
Probability applies to random processes: procedures whose individual outcomes are uncertain but whose long-run pattern is predictable. On free-response questions, a strong answer explicitly references "in the long run" or "over many repetitions" when interpreting a probability. Saying "it will happen 70% of the time" for a single trial is imprecise and can cost you.
The Basic Probability Rules
Every legitimate probability model obeys a small set of rules. Learn them as a checklist.
The sample space is the set of all possible outcomes. Because one of those outcomes must occur, the probabilities of all distinct outcomes sum to exactly .
The complement rule is a workhorse. When a problem asks for "at least one," it is almost always faster to compute the probability of "none" and subtract from . For example, .
The addition rule for mutually exclusive (disjoint) events applies only when and share no outcomes, so . If events can overlap, you must subtract the overlap: . The general form appears later, but you should recognize when disjointness is or isn't a safe assumption.
| Rule | Statement | Meaning |
|---|---|---|
| Bounds | No probability is negative or above 1 | |
| Sample space | Something in must happen | |
| Complement | "Not A" fills the rest | |
| Addition (disjoint) | Only if A, B cannot both occur |
The complement rule is a workhorse. When a problem asks for "at least one," it is almost always faster to compute the probability of "none" and subtract from . For example, .
The addition rule for mutually exclusive (disjoint) events applies only when and share no outcomes, so . If events can overlap, you must subtract the overlap: . The general form appears later, but you should recognize when disjointness is or isn't a safe assumption.
Disjoint vs. Independent — Don't Confuse Them
Students routinely mix up mutually exclusive (disjoint) and independent, and the AP exam deliberately probes this.
Disjoint means the two events cannot both happen in the same trial: their intersection is empty, . Drawing a single card that is both a King and a Queen is impossible, so those events are disjoint.
Independent means knowing one event occurred does not change the probability of the other. Independence is about information, not overlap, and it is developed fully in the next lesson (U4.4).
Here is the trap: two events with positive probability that are disjoint are actually not independent. If happens, then definitely did not (since they can't co-occur), so knowing about changes 's probability from positive to zero. So "mutually exclusive" does not mean "independent" — in fact for nonzero-probability events it guarantees dependence.
In this lesson focus on the addition rule and recognizing disjointness. Just carry away the warning: never assume events are disjoint or independent unless the problem's context justifies it.
Disjoint means the two events cannot both happen in the same trial: their intersection is empty, . Drawing a single card that is both a King and a Queen is impossible, so those events are disjoint.
Independent means knowing one event occurred does not change the probability of the other. Independence is about information, not overlap, and it is developed fully in the next lesson (U4.4).
Here is the trap: two events with positive probability that are disjoint are actually not independent. If happens, then definitely did not (since they can't co-occur), so knowing about changes 's probability from positive to zero. So "mutually exclusive" does not mean "independent" — in fact for nonzero-probability events it guarantees dependence.
| Concept | Condition | Formula |
|---|---|---|
| Disjoint | Can't co-occur | |
| Independent | No info gained |
Estimating Probabilities with Simulation
When a probability is hard to compute directly, you can estimate it with a simulation: an imitation of the random process using a chance device such as a random number table, calculator, or coins and dice.
A clear simulation design specifies four things. First, state how one trial models the real situation and assign digits or outcomes (for example, digits 0–6 = "makes the free throw," 7–9 = "misses," to model a 70% shooter). Second, define what a single repetition consists of. Third, state the response variable you record each repetition. Fourth, run many repetitions and compute the proportion meeting your condition — that proportion is your estimated probability.
The estimate improves as the number of repetitions increases, echoing the law of large numbers. A simulation with only ten trials gives a rough estimate; thousands of trials gives a precise one.
On the exam, free-response questions may ask you to describe a simulation using a supplied line of random digits. Common errors include ignoring the stated probabilities when assigning digits, failing to say how you handle repeated or unused digits, and not clearly defining a stopping rule or what counts as a success. Write your digit assignment explicitly, carry out the requested number of repetitions from the given digits, and report the resulting relative frequency as the answer.
A clear simulation design specifies four things. First, state how one trial models the real situation and assign digits or outcomes (for example, digits 0–6 = "makes the free throw," 7–9 = "misses," to model a 70% shooter). Second, define what a single repetition consists of. Third, state the response variable you record each repetition. Fourth, run many repetitions and compute the proportion meeting your condition — that proportion is your estimated probability.
The estimate improves as the number of repetitions increases, echoing the law of large numbers. A simulation with only ten trials gives a rough estimate; thousands of trials gives a precise one.
On the exam, free-response questions may ask you to describe a simulation using a supplied line of random digits. Common errors include ignoring the stated probabilities when assigning digits, failing to say how you handle repeated or unused digits, and not clearly defining a stopping rule or what counts as a success. Write your digit assignment explicitly, carry out the requested number of repetitions from the given digits, and report the resulting relative frequency as the answer.
Key terms
- Probability.
- A number between 0 and 1 giving the long-run relative frequency with which an outcome occurs over many repetitions of a random process.
- Law of Large Numbers.
- The principle that as the number of trials increases, the observed proportion of an event converges to its true probability.
- Sample Space (S).
- The set of all possible outcomes of a random process; the total probability across it equals 1.
- Complement.
- The event that A does not occur, denoted , with .
- Mutually Exclusive (Disjoint).
- Two events that cannot occur on the same trial, so and .
- Simulation.
- An imitation of a random process using a chance device to estimate a probability from the resulting relative frequency.
- Gambler's Fallacy.
- The mistaken belief that past independent outcomes make certain future outcomes more or less 'due.'
Worked example
A basketball player makes 80% of her free throws. In a game she attempts 3 free throws, and each attempt is independent. Use the complement rule to find the probability she makes at least one, then describe how you would estimate this with a simulation.
Direct computation of "at least one make" would require adding the cases of exactly one, exactly two, and exactly three makes. The complement rule is far faster.
The complement of "at least one make" is "zero makes," meaning she misses all three. She misses each attempt with probability . Since attempts are independent, .
Apply the complement rule: .
To estimate this by simulation, assign digits so that 0–7 represent a made free throw (that is 8 of 10 digits, matching 80%) and 8–9 represent a miss. One repetition is reading three consecutive random digits, one per free-throw attempt. The response variable is whether at least one of the three digits is 0–7. Run many repetitions, count how many contain at least one make, and divide by the number of repetitions. With enough repetitions this proportion should land near , confirming the calculation.
The complement of "at least one make" is "zero makes," meaning she misses all three. She misses each attempt with probability . Since attempts are independent, .
Apply the complement rule: .
To estimate this by simulation, assign digits so that 0–7 represent a made free throw (that is 8 of 10 digits, matching 80%) and 8–9 represent a miss. One repetition is reading three consecutive random digits, one per free-throw attempt. The response variable is whether at least one of the three digits is 0–7. Run many repetitions, count how many contain at least one make, and divide by the number of repetitions. With enough repetitions this proportion should land near , confirming the calculation.
Practice questions
Events and are mutually exclusive with and . What is ?
- 0.135
- 0.75
- 0.85
- Cannot be determined without P(A and B)
Answer: 0.75
Because the events are mutually exclusive, they cannot occur together, so . The addition rule for disjoint events gives . You do not need extra information — disjointness supplies the missing overlap value as zero. The choice 0.135 wrongly multiplies the probabilities.
A fair coin has landed heads five times in a row. A student claims tails is now more likely on the sixth flip because it is 'due.' Explain why this reasoning is incorrect, referencing the correct interpretation of probability.
Answer: The reasoning commits the gambler's fallacy; the sixth flip still has probability 0.5 of tails.
Coin flips are independent, so the coin has no memory of previous results — the probability of tails on the next flip remains regardless of the streak. The law of large numbers does not force short-run balancing; instead, as trials pile up, early imbalances become a smaller fraction of the total, so the long-run proportion drifts toward without any single flip being 'corrected.' Probability describes long-run relative frequency, not guarantees about the next trial.
In a raffle, the probability of winning a prize is 0.02. Using the complement rule, find the probability of winning at least one prize if you buy 4 independent tickets.
Answer: About 0.0776
The complement of winning at least one prize is winning nothing on all four tickets. Each ticket loses with probability , and tickets are independent, so . Then . Using the complement avoids summing the separate cases of exactly one, two, three, and four wins.
FAQ
- What is the difference between mutually exclusive and independent events?
- Mutually exclusive means two events cannot happen on the same trial, so their intersection has probability zero. Independent means one event's occurrence gives no information about the other. They are different ideas: two events with positive probability that are mutually exclusive are necessarily dependent, because knowing one occurred tells you the other did not.
- When should I use the complement rule?
- Reach for the complement rule whenever a problem asks for the probability of 'at least one' of something, or when computing the event directly requires adding many cases. Finding the probability of the opposite ('none' or 'not A') and subtracting from 1 is usually much faster and less error-prone.
- How many trials does a simulation need to give a good estimate?
- More trials give more accurate estimates because of the law of large numbers. A handful of trials produces a rough, unstable estimate, while hundreds or thousands drive the estimated relative frequency close to the true probability. On the AP exam, carry out exactly the number of repetitions the question specifies using the given random digits.
- Does a probability of 0.7 mean an event happens 70% of the time in the short run?
- No. It means that over a very large number of repetitions, the event occurs in about 70% of them. In the short run the proportion can vary widely. Always phrase interpretations in terms of the long run or many repetitions to earn full credit on free-response questions.
Learn this with a teacher, not a page
The Crimsora tutor teaches U4.1 Probability Foundations live — explaining on a whiteboard, asking you questions, and adapting to where you get stuck.