U3.8 Operant Conditioning
Master AP Psychology operant conditioning: positive vs negative reinforcement and punishment, reinforcement schedules, and shaping — with worked examples and practice questions.
What you'll do in this lesson
A voice-first session with the Crimsora tutor on U3.8 Operant Conditioning, then targeted practice and FRQs — with the tutor adapting to where you get stuck.
What this lesson covers
The AP exam loves to test whether you can correctly label a scenario as positive reinforcement, negative reinforcement, positive punishment, or negative punishment, and whether you can match a reward pattern to the right schedule. This lesson gives you a reliable system for sorting these terms, plus the logic behind shaping and why consequences drive learning.
The Core Logic: Consequences Shape Behavior
Skinner tested this using an operant chamber (the "Skinner box"), where an animal's lever press or key peck could be automatically rewarded or punished. The behavior here is voluntary and emitted by the organism — this is the crucial contrast with classical conditioning, where the response is automatic and elicited by a stimulus.
Every consequence does one of two things: it either increases a behavior (reinforcement) or decreases a behavior (punishment). To keep these straight, always ask two separate questions. First: does the consequence make the behavior go up or down? Up means reinforcement; down means punishment. Second: is something being added or removed? Added means positive; removed means negative.
Here the words positive and negative do NOT mean good and bad. Positive means "add a stimulus" and negative means "remove a stimulus." This is the single most common misconception the AP exam exploits, so train yourself to read positive as a plus sign and negative as a minus sign.
The Four Combinations
| Consequence | Add stimulus (positive) | Remove stimulus (negative) |
|---|---|---|
| Increases behavior (reinforcement) | Positive reinforcement: give a reward (praise, candy) | Negative reinforcement: take away something unpleasant (turn off loud alarm) |
| Decreases behavior (punishment) | Positive punishment: apply something unpleasant (scolding, extra chores) | Negative punishment: take away something pleasant (lose phone privileges) |
A primary reinforcer satisfies a biological need (food, water, warmth) and needs no learning. A secondary (conditioned) reinforcer gains value through association, like money or grades. The AP exam often asks you to identify which type a scenario involves.
Watch for escape and avoidance learning, both driven by negative reinforcement: an organism learns to escape or avoid an unpleasant condition. Also note that punishment tends to suppress behavior without teaching a desired alternative, which is why reinforcement is generally considered more effective for building new behaviors.
Reinforcement Schedules
Partial schedules are defined by two dimensions. Ratio schedules depend on the NUMBER of responses; interval schedules depend on TIME passing. Fixed schedules are predictable; variable schedules are unpredictable.
| Schedule | Reinforcer delivered | Example | Response pattern |
|---|---|---|---|
| Fixed-ratio | After a set number of responses | Free coffee after 10 purchases | High rate, brief pause after reward |
| Variable-ratio | After an unpredictable number of responses | Slot machines | Highest, steadiest rate; very resistant to extinction |
| Fixed-interval | First response after a set time | Weekly paycheck; studying before a scheduled exam | Scalloped: activity rises near the deadline |
| Variable-interval | First response after an unpredictable time | Checking for a text reply | Slow, steady rate |
Shaping and How Behaviors Are Built
Shaping shows that operant conditioning is not just about waiting for behavior to happen; it is an active molding process. Animal trainers, therapists teaching new skills, and teachers scaffolding complex tasks all use shaping.
Several related concepts appear on the exam. A discriminative stimulus is a signal indicating that a response will be reinforced (a lit light meaning lever presses now pay off). Extinction occurs when reinforcement stops and the behavior gradually declines. A token economy uses secondary reinforcers (tokens) that can later be exchanged for rewards, common in classrooms and treatment settings.
Finally, know the limits. Instinctive drift describes how animals tend to revert to instinctive behaviors that interfere with conditioned ones — a raccoon trained to deposit coins may start rubbing them together instead. This, along with research on latent learning and cognitive maps from the next lesson, shows that biology and cognition constrain pure consequence-based learning.
Key terms
- Reinforcement.
- Any consequence that increases the likelihood a behavior will be repeated.
- Positive reinforcement.
- Adding a pleasant stimulus after a behavior to increase that behavior.
- Negative reinforcement.
- Removing an aversive stimulus after a behavior to increase that behavior; not punishment.
- Positive punishment.
- Adding an unpleasant stimulus after a behavior to decrease that behavior.
- Negative punishment.
- Removing a pleasant stimulus after a behavior to decrease that behavior.
- Shaping.
- Reinforcing successive approximations that gradually build toward a target behavior.
- Variable-ratio schedule.
- Reinforcement after an unpredictable number of responses; yields the highest, most extinction-resistant response rate.
- Primary vs secondary reinforcer.
- Primary reinforcers satisfy biological needs innately; secondary reinforcers gain reinforcing value through learned association.
Worked example
Next analyze the praise. Praise is a pleasant stimulus being added (positive), and it increases hand-raising (reinforcement), so praise is positive reinforcement. Praise here is a secondary reinforcer because its value is learned, not biological.
Now analyze the timing of the praise. It is given for hand-raising responses, so reinforcement depends on responses, not time — that makes it a ratio schedule. Because it is given unpredictably rather than after a set number, it is a variable-ratio schedule. This predicts a high, steady rate of hand-raising that resists extinction even if praise later becomes rare.
The combined strategy strengthens the same target behavior through two reinforcement routes, and the variable-ratio praise is especially effective for maintaining the behavior over time.
Practice questions
A child stops whining once a parent hands over a candy bar. Because the whining ended, the parent gives in more quickly the next time. From the parent's perspective, giving in is an example of which principle?
- Positive reinforcement
- Negative reinforcement
- Positive punishment
- Negative punishment
Answer: Negative reinforcement
Explain why a variable-ratio reinforcement schedule produces behavior that is more resistant to extinction than a continuous reinforcement schedule, and give one real-world example.
Answer: On a variable-ratio schedule, reinforcement arrives after an unpredictable number of responses, so the learner cannot tell when reward stops. On continuous reinforcement, a single missed reward signals a change, so behavior extinguishes quickly.
A dog trainer gives a treat first when the dog sits halfway, then only when it sits fully, then only when it sits and stays. What learning process is this, and what is the term for the intermediate steps being rewarded?
Answer: This is shaping, and the rewarded intermediate steps are called successive approximations.
FAQ
- What is the difference between negative reinforcement and punishment?
- Negative reinforcement increases a behavior by removing something unpleasant, while punishment decreases a behavior. They are opposites in effect. Buckling a seatbelt to end an annoying beep is negative reinforcement; losing your phone for texting in class is punishment.
- How is operant conditioning different from classical conditioning?
- Classical conditioning links two stimuli to produce an automatic, involuntary response. Operant conditioning links a voluntary behavior to a consequence, so the behavior becomes more or less likely. Simply put, classical conditioning is about associations before behavior, and operant conditioning is about consequences after behavior.
- Which reinforcement schedule is hardest to extinguish?
- The variable-ratio schedule produces the highest response rate and is the most resistant to extinction because rewards come after an unpredictable number of responses. This is why gambling behaviors are so persistent.
- Do the words positive and negative mean good and bad in operant conditioning?
- No. Positive means a stimulus is added and negative means a stimulus is removed. This is the most common exam trap, so read positive as a plus sign and negative as a minus sign, regardless of whether the outcome feels good or bad.
Learn this with a teacher, not a page
The Crimsora tutor teaches U3.8 Operant Conditioning live — explaining on a whiteboard, asking you questions, and adapting to where you get stuck.