AP-PSYCH-3.8

U3.8 Operant Conditioning

Master AP Psychology operant conditioning: positive vs negative reinforcement and punishment, reinforcement schedules, and shaping — with worked examples and practice questions.

What you'll do in this lesson

A voice-first session with the Crimsora tutor on U3.8 Operant Conditioning, then targeted practice and FRQs — with the tutor adapting to where you get stuck.

What this lesson covers

You already know classical conditioning links stimuli to automatic responses. Operant conditioning is different: it explains how the consequences of a behavior make that behavior more or less likely in the future. This is B.F. Skinner's territory — rats pressing levers, pigeons pecking keys, and students studying for grades.

The AP exam loves to test whether you can correctly label a scenario as positive reinforcement, negative reinforcement, positive punishment, or negative punishment, and whether you can match a reward pattern to the right schedule. This lesson gives you a reliable system for sorting these terms, plus the logic behind shaping and why consequences drive learning.

The Core Logic: Consequences Shape Behavior

Operant conditioning, studied systematically by B.F. Skinner, is learning in which behavior is influenced by its consequences. The key idea from Edward Thorndike's law of effect is that behaviors followed by satisfying consequences tend to be repeated, while behaviors followed by unpleasant consequences tend to fade.

Skinner tested this using an operant chamber (the "Skinner box"), where an animal's lever press or key peck could be automatically rewarded or punished. The behavior here is voluntary and emitted by the organism — this is the crucial contrast with classical conditioning, where the response is automatic and elicited by a stimulus.

Every consequence does one of two things: it either increases a behavior (reinforcement) or decreases a behavior (punishment). To keep these straight, always ask two separate questions. First: does the consequence make the behavior go up or down? Up means reinforcement; down means punishment. Second: is something being added or removed? Added means positive; removed means negative.

Here the words positive and negative do NOT mean good and bad. Positive means "add a stimulus" and negative means "remove a stimulus." This is the single most common misconception the AP exam exploits, so train yourself to read positive as a plus sign and negative as a minus sign.

The Four Combinations

Because you can add or remove something, and the goal can be to increase or decrease behavior, there are exactly four possibilities. The table below shows them.
ConsequenceAdd stimulus (positive)Remove stimulus (negative)
Increases behavior (reinforcement)Positive reinforcement: give a reward (praise, candy)Negative reinforcement: take away something unpleasant (turn off loud alarm)
Decreases behavior (punishment)Positive punishment: apply something unpleasant (scolding, extra chores)Negative punishment: take away something pleasant (lose phone privileges)
Negative reinforcement is the trickiest term. It is NOT punishment. It strengthens a behavior by removing an aversive stimulus. Buckling your seatbelt to stop the annoying beeping is negative reinforcement — the behavior (buckling) increases because it removes the beeping.

A primary reinforcer satisfies a biological need (food, water, warmth) and needs no learning. A secondary (conditioned) reinforcer gains value through association, like money or grades. The AP exam often asks you to identify which type a scenario involves.

Watch for escape and avoidance learning, both driven by negative reinforcement: an organism learns to escape or avoid an unpleasant condition. Also note that punishment tends to suppress behavior without teaching a desired alternative, which is why reinforcement is generally considered more effective for building new behaviors.

Reinforcement Schedules

Once a behavior exists, how often you reinforce it dramatically affects how fast it is learned and how resistant it is to extinction. Continuous reinforcement rewards every correct response — learning is fast but extinction is also fast once rewards stop. Partial (intermittent) reinforcement rewards only some responses and produces behavior that is much more resistant to extinction.

Partial schedules are defined by two dimensions. Ratio schedules depend on the NUMBER of responses; interval schedules depend on TIME passing. Fixed schedules are predictable; variable schedules are unpredictable.
ScheduleReinforcer deliveredExampleResponse pattern
Fixed-ratioAfter a set number of responsesFree coffee after 10 purchasesHigh rate, brief pause after reward
Variable-ratioAfter an unpredictable number of responsesSlot machinesHighest, steadiest rate; very resistant to extinction
Fixed-intervalFirst response after a set timeWeekly paycheck; studying before a scheduled examScalloped: activity rises near the deadline
Variable-intervalFirst response after an unpredictable timeChecking for a text replySlow, steady rate
A memory tip: variable schedules produce the steadiest responding, and ratio schedules produce faster responding than interval schedules because reward depends on how much you do. Variable-ratio is the champion for both response rate and resistance to extinction — that is why gambling is so hard to quit.

Shaping and How Behaviors Are Built

You cannot reinforce a behavior that never occurs. Shaping solves this by reinforcing successive approximations — steps that get progressively closer to the target behavior. To teach a rat to press a lever, you might first reward it for facing the lever, then for approaching it, then for touching it, then finally for pressing it.

Shaping shows that operant conditioning is not just about waiting for behavior to happen; it is an active molding process. Animal trainers, therapists teaching new skills, and teachers scaffolding complex tasks all use shaping.

Several related concepts appear on the exam. A discriminative stimulus is a signal indicating that a response will be reinforced (a lit light meaning lever presses now pay off). Extinction occurs when reinforcement stops and the behavior gradually declines. A token economy uses secondary reinforcers (tokens) that can later be exchanged for rewards, common in classrooms and treatment settings.

Finally, know the limits. Instinctive drift describes how animals tend to revert to instinctive behaviors that interfere with conditioned ones — a raccoon trained to deposit coins may start rubbing them together instead. This, along with research on latent learning and cognitive maps from the next lesson, shows that biology and cognition constrain pure consequence-based learning.

Key terms

Reinforcement.
Any consequence that increases the likelihood a behavior will be repeated.
Positive reinforcement.
Adding a pleasant stimulus after a behavior to increase that behavior.
Negative reinforcement.
Removing an aversive stimulus after a behavior to increase that behavior; not punishment.
Positive punishment.
Adding an unpleasant stimulus after a behavior to decrease that behavior.
Negative punishment.
Removing a pleasant stimulus after a behavior to decrease that behavior.
Shaping.
Reinforcing successive approximations that gradually build toward a target behavior.
Variable-ratio schedule.
Reinforcement after an unpredictable number of responses; yields the highest, most extinction-resistant response rate.
Primary vs secondary reinforcer.
Primary reinforcers satisfy biological needs innately; secondary reinforcers gain reinforcing value through learned association.

Worked example

A teacher wants students to raise their hands before speaking. She stops giving pop quizzes on days when the whole class raises hands consistently. She also praises individual students each time they raise a hand, but only occasionally and unpredictably. Identify the operant principles at work.
First analyze the removal of pop quizzes. A pop quiz is an aversive stimulus for most students. By removing it when hand-raising occurs, the teacher makes hand-raising more likely. Something is being removed (negative) and behavior increases (reinforcement), so this is negative reinforcement.

Next analyze the praise. Praise is a pleasant stimulus being added (positive), and it increases hand-raising (reinforcement), so praise is positive reinforcement. Praise here is a secondary reinforcer because its value is learned, not biological.

Now analyze the timing of the praise. It is given for hand-raising responses, so reinforcement depends on responses, not time — that makes it a ratio schedule. Because it is given unpredictably rather than after a set number, it is a variable-ratio schedule. This predicts a high, steady rate of hand-raising that resists extinction even if praise later becomes rare.

The combined strategy strengthens the same target behavior through two reinforcement routes, and the variable-ratio praise is especially effective for maintaining the behavior over time.

Practice questions

A child stops whining once a parent hands over a candy bar. Because the whining ended, the parent gives in more quickly the next time. From the parent's perspective, giving in is an example of which principle?
  1. Positive reinforcement
  2. Negative reinforcement
  3. Positive punishment
  4. Negative punishment

Answer: Negative reinforcement

Focus on the parent's behavior of giving in. The whining (an aversive stimulus) is removed when the parent gives candy, and this makes giving-in more likely in the future. Removing an unpleasant stimulus to increase a behavior is negative reinforcement. Students often mislabel this as punishment because a child is involved, but the parent's giving-in behavior is being strengthened, not weakened.
Explain why a variable-ratio reinforcement schedule produces behavior that is more resistant to extinction than a continuous reinforcement schedule, and give one real-world example.

Answer: On a variable-ratio schedule, reinforcement arrives after an unpredictable number of responses, so the learner cannot tell when reward stops. On continuous reinforcement, a single missed reward signals a change, so behavior extinguishes quickly.

Under continuous reinforcement, the organism expects a reward every time, so as soon as rewards stop the absence is obvious and the behavior fades fast. Under a variable-ratio schedule, long unrewarded stretches are normal, so the organism keeps responding, expecting the next reward could come at any moment — this delays extinction. A strong example is a slot machine or any form of gambling, where unpredictable payouts sustain persistent play.
A dog trainer gives a treat first when the dog sits halfway, then only when it sits fully, then only when it sits and stays. What learning process is this, and what is the term for the intermediate steps being rewarded?

Answer: This is shaping, and the rewarded intermediate steps are called successive approximations.

Shaping is used to build a behavior that does not yet occur on its own by reinforcing steps that get progressively closer to the goal. Each rewarded step — sitting halfway, sitting fully, sitting and staying — is a successive approximation. This demonstrates that operant conditioning can actively construct complex behaviors, not just strengthen existing ones.

FAQ

What is the difference between negative reinforcement and punishment?
Negative reinforcement increases a behavior by removing something unpleasant, while punishment decreases a behavior. They are opposites in effect. Buckling a seatbelt to end an annoying beep is negative reinforcement; losing your phone for texting in class is punishment.
How is operant conditioning different from classical conditioning?
Classical conditioning links two stimuli to produce an automatic, involuntary response. Operant conditioning links a voluntary behavior to a consequence, so the behavior becomes more or less likely. Simply put, classical conditioning is about associations before behavior, and operant conditioning is about consequences after behavior.
Which reinforcement schedule is hardest to extinguish?
The variable-ratio schedule produces the highest response rate and is the most resistant to extinction because rewards come after an unpredictable number of responses. This is why gambling behaviors are so persistent.
Do the words positive and negative mean good and bad in operant conditioning?
No. Positive means a stimulus is added and negative means a stimulus is removed. This is the most common exam trap, so read positive as a plus sign and negative as a minus sign, regardless of whether the outcome feels good or bad.

Learn this with a teacher, not a page

The Crimsora tutor teaches U3.8 Operant Conditioning live — explaining on a whiteboard, asking you questions, and adapting to where you get stuck.