Psychology
Operant Conditioning: Skinner's 4 Cases in One Table
Reinforcement or punishment, positive or negative: the four cases in one table with examples. Plus 9 solved exercises and all five reinforcement schedules.
Recommended for: Grade 11 · Grade 12
Operant conditioning is the form of learning in which a voluntary behavior becomes more or less frequent depending on the consequences it produces. It is Burrhus F. Skinner’s model, and it is called operant because the subject does not simply undergo the environment: it operates on it, and learns from the environment’s response what is worth repeating.
The table of the four cases
The whole topic reduces to two questions, always in this order. First: does the consequence add a stimulus or remove one? Second: after that consequence, does the behavior increase or decrease? Cross the two answers and you get four cases, and no others.
| The consequence adds a stimulus (positive) | The consequence removes a stimulus (negative) | |
|---|---|---|
| Behavior increases → reinforcement | Positive reinforcement: the teacher praises a student who contributes, and that student raises their hand more often | Negative reinforcement: you buckle up and the chime stops, so you buckle up straight away every time |
| Behavior decreases → punishment | Positive punishment: a detention follows every late arrival, and lateness drops | Negative punishment: your phone is confiscated for a week, and the shouting at home drops |
Two reading rules are worth fixing immediately, because almost every mistake in this topic comes from ignoring them.
Positive and negative do not mean pleasant and unpleasant. They mean added and removed, exactly like the plus and minus signs. Negative reinforcement is therefore not a punishment: it is reinforcement, and like all reinforcement it makes behavior increase.
Reinforcement and punishment are defined by their effect, not by intention. If a behavior does not decrease after a telling-off, that telling-off was not a punishment, whatever the person delivering it had in mind. And if the child was seeking attention, it may even have acted as positive reinforcement.
From the law of effect to the Skinner box
The starting point is Edward Thorndike, who in the late nineteenth century shut cats in a puzzle box that could be opened by operating a latch. Timing each escape, he observed gradual improvement rather than sudden insight, and derived the law of effect: behaviors followed by satisfying consequences tend to be repeated, behaviors followed by annoying consequences tend to disappear.
In the 1930s Skinner made that principle measurable with the operant conditioning chamber, the so-called Skinner box: an isolated cage in which a rat presses a lever or a pigeon pecks a disc, with a food dispenser and a cumulative recorder tracing responses over time. The dependent variable is no longer the time taken, but the frequency of the behavior.
Negative reinforcement: escape and avoidance
Negative reinforcement comes in two forms worth keeping apart. In escape, the behavior terminates an aversive stimulus that is already present: you take a painkiller and the headache goes away. In avoidance, the behavior prevents the stimulus from appearing at all: you leave half an hour early so you never hit the traffic.
Avoidance is the clinically interesting case, because it is almost impossible to extinguish. Someone who avoids never faces the feared situation, so they never collect the evidence that the feared outcome would not have happened: the behavior confirms itself every time it works. This is the mechanism that keeps phobias alive, where operant avoidance is layered on top of the learned association described in classical conditioning.
Primary and secondary reinforcers
Primary reinforcers work without any learning because they satisfy biological needs: food, water, warmth, contact. Secondary or conditioned reinforcers, by contrast, have acquired their value through association with primary ones or with other already effective reinforcers: money, grades, the tokens of a token economy, applause.
The Premack principle is also useful: a high-frequency activity can reinforce a low-frequency one. No external prize is needed, only the right order. Homework first, video games after: reverse the order and you reinforce not doing the homework.
Reinforcement schedules
Reinforcing every response is the fastest way to raise a behavior, and also the fastest way to lose it. The schedule, that is the rule by which reinforcement is delivered, changes both the rhythm of responding and how long the behavior survives once reinforcement stops.
| Schedule | Rule | Example | Response pattern | Resistance to extinction |
|---|---|---|---|---|
| Continuous (CRF) | every response is reinforced | a vending machine delivering for every coin | very fast learning | very low |
| Fixed ratio (FR) | one reinforcer every n responses | piece rate: a bonus every 50 units | high, with a pause after each reinforcer | medium |
| Variable ratio (VR) | on average one every n, unpredictable | slot machine | the highest and steadiest of all | the highest |
| Fixed interval (FI) | the first response after time t | monthly salary, a scheduled test | scalloped: rises near the deadline | low |
| Variable interval (VI) | the first response after an average time | checking notifications, fishing | moderate and very steady | high |
Two regularities explain the whole table. Ratio schedules produce higher rates than interval schedules, because responding faster genuinely brings the next reinforcer closer. Variable schedules resist extinction better than fixed ones, because a long unreinforced run falls within the schedule’s normal behavior and never signals that it is time to stop.
Shaping, chaining and discriminative stimuli
Reinforcement can only act on a behavior that has already been emitted: if you wait for the dog to ring the bell by itself, you will never get to reinforce it. Shaping solves this by reinforcing successive approximations, behaviors increasingly similar to the target, and dropping reinforcement for steps already mastered. Chaining then links several responses into a sequence in which each link signals the next.
A behavior also does not appear at random: it is emitted in the presence of a discriminative stimulus, the signal indicating that reinforcement is available under those conditions. Asking a favour when the other person is in a good mood is stimulus control, not luck.
Extinction and the extinction burst
When the reinforcer that maintained a behavior stops arriving, the behavior weakens until it disappears: this is extinction. Before dropping, however, it almost always rises: more frequent, more intense, often accompanied by frustration or anger. This is the extinction burst, and it is the moment when the procedure looks like it is failing while it is in fact working.
Here lies the most expensive trap in the whole topic. Giving in during the peak means reinforcing the most intense version of the behavior and, above all, shifting it from continuous reinforcement to a variable ratio, the most extinction-resistant schedule there is. It is the most effective way to make permanent a behavior you were trying to eliminate.
Superstitious behavior
In 1948 Skinner set the feeder to release food every 15 seconds regardless of what the pigeons were doing. Six out of eight birds developed stereotyped movements: turning in circles, bobbing the head, thrusting the beak into a corner. The mechanism is accidental reinforcement: whatever the animal happened to be doing at the moment of delivery gets reinforced by coincidence, becomes more probable, and is therefore more likely to be under way at the next delivery.
It is worth adding that the interpretation is disputed: in 1971 Staddon and Simmelhag showed that many of those movements were species-typical behaviors linked to food anticipation rather than random responses. As with the biological constraints revealed by the Brelands’ instinctive drift, not everything is equally conditionable: a species’ inherited repertoire, with the transmission rules studied in Mendelian genetics, sets limits on what reinforcement can build.
Operant or classical? How to tell in ten seconds
| Classical conditioning | Operant conditioning | |
|---|---|---|
| What gets associated | two stimuli (neutral and unconditioned) | a behavior and its consequence |
| Role of the subject | passive: the response is elicited | active: the response is emitted |
| Type of response | reflexive, involuntary (salivating, flinching) | voluntary (pressing, asking, studying) |
| Temporal order | the stimulus precedes the response | the consequence follows the behavior |
| Key figures | Pavlov, Watson | Thorndike, Skinner |
The quickest test: ask whether the subject has to do something for what follows to happen. If yes it is operant, if no it is classical.
The most common mistakes
- Reading negative as unpleasant. Positive and negative only indicate whether a stimulus is added or removed.
- Calling negative reinforcement a punishment. Reinforcement always increases behavior, in both of its forms.
- Confusing negative punishment with negative reinforcement. In the first, something desirable is removed and the behavior falls; in the second, something aversive is removed and the behavior rises.
- Classifying by intention. The label depends on what happens to the frequency of the behavior after the consequence, not on what was meant to happen.
- Mistaking the stimulus for the behavior. The grade is not the behavior: the behavior is studying.
- Giving in during the extinction burst. It is the surest way to install a variable ratio on a behavior you were trying to switch off.
How to use the exercises
The nine exercises below follow a real progression: first classifying the four cases and dismantling the usual confusions, then distinguishing escape from avoidance, identifying reinforcement schedules and ranking them by resistance, calculating reinforcers and response rates from numerical data, and finally designing a shaping procedure, analysing an extinction burst, explaining why a punitive strategy fails, and accounting for superstitious behavior. Exercise 5 requires a percentage change: if that calculation is not automatic yet, revise it with the solved exercises on percentages.
Solved exercises
1. Classify each situation as positive reinforcement, negative reinforcement, positive punishment or negative punishment. a) Maya puts on sunscreen and stops getting burned, so now she always applies it. b) Ryan gets a detention for being late and starts arriving on time. c) Students who tidy up get a sticker, and tidying becomes more frequent. d) Sophia loses her car keys for a month after speeding, and she slows down. base
Show solution
- For each case ask two questions in this order. First: does the consequence ADD something (positive) or REMOVE something (negative)? Second: after the consequence, does the behavior INCREASE (reinforcement) or DECREASE (punishment)?
- a) The sunscreen removes the sunburn, an unpleasant stimulus, so it is negative. Applying sunscreen increases, so it is reinforcement. Negative reinforcement.
- b) The detention adds an unpleasant stimulus, so it is positive. Lateness decreases, so it is punishment. Positive punishment.
- c) The sticker adds a pleasant stimulus, so it is positive. Tidying increases, so it is reinforcement. Positive reinforcement.
- d) Losing the keys removes something desirable, so it is negative. Speeding decreases, so it is punishment. Negative punishment.
- Final check: in a) and c) the behavior increases even though one case adds and the other removes. Positive and negative never mean pleasant and unpleasant, only added and removed.
Answer: a) negative reinforcement; b) positive punishment; c) positive reinforcement; d) negative punishment
2. True or false? Correct the false statements. a) Negative reinforcement is another name for punishment. b) If the behavior does not decrease, the consequence was not a punishment. c) Money is a primary reinforcer. d) In operant conditioning the response is an involuntary reflex. base
Show solution
- a) False. This is the single most common error in the topic. Reinforcement, positive or negative, always INCREASES behavior; punishment always DECREASES it. Negative reinforcement increases behavior by removing something aversive, like taking a painkiller because the headache stops.
- b) True. Reinforcement and punishment are defined by their effect, not by the intention behind them. A telling-off that does not reduce the behavior is not a punishment, and if the student was seeking attention it may even have worked as positive reinforcement.
- c) False. Primary reinforcers satisfy biological needs: food, water, warmth. Money is a secondary or conditioned reinforcer, and its value is learned through association with what it can buy.
- d) False. That describes classical conditioning. In operant conditioning the response is voluntary and emitted by the subject, who acts on the environment.
Answer: a) False (reinforcement always increases behavior); b) True; c) False (it is a secondary reinforcer); d) False (that is classical conditioning)
3. Distinguish escape from avoidance in these three cases and explain why the third is the hardest to extinguish. a) The seatbelt chime is beeping and Anna buckles up to make it stop. b) Anna buckles up before turning the key, so the chime never starts. c) Liam is afraid of freezing during oral exams, so he stays home on presentation days and has avoided every unannounced presentation for years. base
Show solution
- In escape the behavior terminates an aversive stimulus that is ALREADY present. In a) the chime is already sounding, so this is escape.
- In avoidance the behavior prevents the aversive stimulus from appearing at all. In b) the chime never starts, so this is avoidance.
- In c) staying home prevents the presentation, so this is also avoidance, driven by fear.
- All three are negative reinforcement: in every case the behavior increases and what follows is the removal, or non-appearance, of something unpleasant.
- Case c) resists extinction for a precise reason: Liam never faces the situation, so he never discovers that the feared outcome might not happen. Avoidance confirms itself, because every time it works it appears to justify itself.
- Check: if a behavior prevents you from collecting the evidence that would disprove it, no disconfirming evidence can ever arrive. This is why avoidance keeps phobias alive.
Answer: a) escape; b) avoidance; c) avoidance. All three are negative reinforcement. The third is the most resistant because avoidance prevents the person from testing whether the feared consequence would occur, so the association is never disconfirmed.
4. Identify the reinforcement schedule in each situation (continuous, fixed ratio, variable ratio, fixed interval, variable interval), then rank them from most to least resistant to extinction. a) A factory worker earns a bonus every 50 units assembled. b) A slot machine pays out on average once every 30 plays, but you never know when. c) A salary arrives on the last working day of every month. d) A vending machine delivers a can for every coin. e) An angler watches the line and now and then, at unpredictable moments, a fish bites. intermedio
Show solution
- First sort the ratio/interval axis: does the reinforcer depend on the NUMBER of responses (ratio) or on the TIME elapsed (interval)? Then the fixed/variable axis: is the rule constant or does it vary around an average?
- a) The bonus depends on the number of units and it is always 50: fixed ratio (FR 50).
- b) It depends on the number of plays, but the number varies around 30: variable ratio (VR 30).
- c) It depends only on time, always the same amount: fixed interval (FI one month).
- d) Every response is reinforced: continuous reinforcement (CRF).
- e) It depends on unpredictable waiting time: variable interval (VI).
- For the ranking, use two rules. VARIABLE schedules resist better than fixed ones, because the subject cannot tell when reinforcement has stopped and keeps responding. RATIO schedules produce higher rates than interval schedules, because responding faster genuinely brings the next reinforcer closer.
- Resulting order, most to least resistant: b (VR), e (VI), a (FR), c (FI), d (CRF). Continuous reinforcement extinguishes fastest, because a handful of unreinforced responses is already an obvious change.
Answer: a) fixed ratio; b) variable ratio; c) fixed interval; d) continuous; e) variable interval. Resistance to extinction, most to least: b, e, a, c, d.
5. In a 30-minute session, pigeon A on a fixed ratio FR 20 schedule emits 1840 pecks; pigeon B on a variable ratio VR 20 schedule emits 2100. Calculate for each the number of reinforcers obtained and the response rate per minute, work out by what percentage B responded more than A, and predict which bird will keep pecking longer once food is withheld. intermedio
Show solution
- Reinforcers for A: on FR 20 one reinforcer arrives every 20 pecks, so 1840 divided by 20 = 92 reinforcers.
- Reinforcers for B: on VR 20 the ratio is 20 on average, so the expected number is 2100 divided by 20 = 105 reinforcers.
- Rate for A: 1840 pecks divided by 30 minutes = 61.3 pecks per minute.
- Rate for B: 2100 pecks divided by 30 minutes = 70 pecks per minute.
- Percentage difference: (2100 minus 1840) divided by 1840 = 260 divided by 1840 = 0.1413, which is about 14.1 per cent more.
- Extinction prediction: B lasts longer. On FR 20, twenty or thirty pecks with no food is a clear signal that something has changed; on VR 20, long unreinforced runs are a normal part of the schedule, so the bird has no immediate way of detecting that food has stopped.
- Consistency check: both rates are high, as expected for ratio schedules, and B is higher partly because the variable schedule lacks the post-reinforcement pause typical of fixed ratio.
Answer: A: 92 reinforcers and 61.3 pecks per minute. B: 105 reinforcers and 70 pecks per minute. B responds about 14.1 per cent more. B resists extinction longer, because on a variable ratio the unreinforced pauses are indistinguishable from the normal pauses of the schedule.
6. You have to teach a dog to ring a bell hanging on the door when it wants to go out. The dog has never touched the bell, so you cannot wait for the behavior to appear on its own. Design the shaping procedure with at least four successive approximations, and explain two technical points: when to deliver the reinforcer and which schedule to use before and after learning. intermedio
Show solution
- The constraint is that reinforcement can only act on a behavior that has already been emitted. If I wait for the final behavior I will never see it, so I reinforce behaviors that resemble it more and more closely: this is shaping by successive approximations.
- First approximation: reinforce every time the dog approaches the door. Once this happens often, stop reinforcing mere approach.
- Second approximation: reinforce only contact with the area around the bell, for example sniffing it.
- Third approximation: reinforce only touching the bell with the nose or a paw, even without any sound.
- Fourth approximation: reinforce only when the bell actually rings. Fifth and last: reinforce only when it rings and the dog then heads for the door.
- First technical point, timing: the reinforcer must arrive within a couple of seconds. Any delay risks reinforcing whatever behavior occurred IN THE MEANTIME, such as turning away, rather than the target one.
- Second technical point, schedule: use continuous reinforcement during acquisition, because it raises the behavior fastest. Once it is stable, switch to a variable ratio, because continuous reinforcement extinguishes quickly as soon as the treats stop.
Answer: Approximations: approaching the door, touching the area around the bell, touching the bell, making it ring, ringing it and then going to the door. The reinforcer must arrive within a few seconds or a different behavior gets reinforced. Use continuous reinforcement during acquisition and switch to a variable ratio to make the behavior resistant to extinction.
7. Every evening 6-year-old Noah whines for one more cartoon and his parents give in. They decide to stop. For the first two evenings the whining does not fade: it becomes more frequent and more intense, with some kicking of the door. On the third evening, exhausted, they give in exactly once, at the worst moment. Over the following weeks the whining becomes far more stubborn than before. Explain what maintained the behavior at first, name the phenomenon of the first two evenings, and explain why the evening they gave in made things worse. intermedio
Show solution
- Starting point: whining produces the cartoon, so it is maintained by positive reinforcement. The cartoon is the consequence that makes it more probable.
- By no longer providing the cartoon, the parents started extinction: the behavior is emitted but the consequence that maintained it no longer follows.
- The increase over the first two evenings is called an extinction burst. It is a typical and temporary effect: when what used to work stops working, the first reaction is to try harder and more often, often with emotional or aggressive responses.
- The burst is the phase where extinction looks like it is failing, but it is in fact the sign that the procedure is working. Had the parents held firm, the curve would have dropped over the following evenings.
- By giving in once, however, they did the worst possible thing: they converted continuous reinforcement into intermittent reinforcement on a variable ratio, and on top of that they reinforced the most intense version of the behavior.
- Predictable consequence: as in the slot machine exercise, the variable ratio is the schedule most resistant to extinction. The child has now learned that insisting long enough eventually works.
- What they should have done: maintain extinction without exceptions and pair it with reinforcement of an alternative behavior, for instance giving attention when Noah asks calmly or accepts a no.
Answer: The whining was maintained by positive reinforcement; withholding the cartoon started extinction. The rise over the first two evenings is the extinction burst, temporary and normal. By giving in at the peak, the parents moved the behavior onto a variable ratio schedule, the most extinction-resistant of all, and reinforced its most intense form.
8. A teacher punishes every instance of talking with a detention. The class is silent in her lessons, but talks more than before with substitute teachers; two students have started skipping her lessons and the atmosphere is tense. Analyse why punishment produced this picture instead of eliminating the behavior, identify the three side effects present, and propose a reinforcement-based alternative, saying why it should work better. avanzato
Show solution
- First limitation of punishment: it suppresses a behavior but teaches no replacement. The student learns what not to do, not what to do, so as soon as punishment is unavailable the old repertoire is still the only one available.
- Second limitation, which explains the substitute teacher data: suppression becomes tied to the conditions in which punishment was delivered. The teacher becomes a discriminative stimulus signalling when the behavior will be punished. Outside that condition the behavior returns, because it never actually disappeared.
- First side effect, avoidance: skipping lessons is maintained by negative reinforcement, since it removes the student from the punishing situation. Punishment has thus created a new behavior that is worse than the original one.
- Second side effect, conditioned emotional responses: the classroom and the subject are repeatedly paired with an aversive stimulus and, through classical conditioning, end up eliciting anxiety on their own.
- Third side effect, the tense atmosphere: punishment tends to produce hostility and aggression towards whoever delivers it, and it also models aversive control as a way of handling problems.
- Alternative procedure. Define the desired behavior in observable terms, for example raising a hand before speaking and working quietly within the allotted time.
- Apply differential reinforcement of an alternative behavior: systematically reinforce the desired behavior (attention, public recognition, points in a token economy) and place talking on extinction by removing the attention that often maintains it.
- Use continuous reinforcement at first, then move to a variable ratio to make the new behavior durable.
- Why it should work better: it builds an alternative repertoire instead of merely blocking one, it produces neither avoidance nor negative conditioned emotions, and it does not depend on the presence of the punisher, because the alternative behavior is reinforced elsewhere too.
Answer: Punishment suppresses without teaching an alternative and ties suppression to the punisher, who becomes a discriminative stimulus: hence the behavior with substitutes. The three side effects are avoidance (absences, maintained by negative reinforcement), conditioned emotional responses towards classroom and subject, and hostility towards the teacher. The alternative is differential reinforcement of the desired behavior combined with extinction of talking, using continuous reinforcement at first and a variable ratio later.
9. In 1948 Skinner placed eight pigeons in cages and set the feeder to deliver food every 15 seconds regardless of what the birds did. Six out of eight developed stereotyped movements such as turning in circles or bobbing the head. Calculate how many food deliveries a pigeon receives in a 10-minute session, explain the mechanism producing the behavior, say why it persists even though it achieves nothing, and give one limitation of this interpretation. avanzato
Show solution
- Deliveries in 10 minutes: 10 minutes is 600 seconds, so 600 divided by 15 = 40 deliveries per session.
- Mechanism: food arrives on a fixed time basis and does not depend on behavior, but the pigeon is nevertheless doing something at the instant food appears. That behavior gets reinforced by sheer temporal coincidence: this is called accidental or adventitious reinforcement.
- Once reinforced, that movement becomes more probable, so it is more likely to be under way at the next delivery, and gets reinforced again. The loop feeds itself and the response stabilises.
- Why it persists: from the animal's point of view the behavior always works, because food keeps arriving. No evidence can ever disconfirm the association, and the apparent relationship is irregular with respect to the response, which makes it resemble a variable ratio, the most extinction-resistant schedule.
- Human parallel: the same pattern produces good luck rituals before a test, or the fixed sequence of gestures a player performs before a free throw. An occasional success coincides with a gesture and cements it.
- Limitation of the interpretation: in 1971 Staddon and Simmelhag re-analysed the experiment and found that many of those movements were not random. They were species-typical behaviors linked to food anticipation, so-called interim and terminal behaviors that appear at predictable points in the interval.
- Careful conclusion: accidental reinforcement remains a plausible explanation of many superstitions, but the pigeon case shows that species-specific biological constraints help determine which behaviors appear.
Answer: 40 deliveries in 10 minutes. The behavior arises through accidental reinforcement: whatever the animal is doing when food appears gets reinforced by coincidence and becomes more probable, triggering a self-sustaining loop. It persists because food keeps arriving anyway and nothing can disconfirm the association, which resembles a variable ratio. The limitation is the 1971 re-analysis by Staddon and Simmelhag: many of those movements are species-typical food anticipation behaviors, not random responses.
FAQ
What is the difference between negative reinforcement and punishment?
Negative reinforcement INCREASES a behavior by removing something unpleasant: you buckle up and the chime stops, so you buckle up sooner every time. Punishment DECREASES a behavior, either by adding something unpleasant (positive punishment, such as a detention) or by removing something desirable (negative punishment, such as confiscating a phone). The word negative does not mean unpleasant: it only means that something is taken away.
What is the difference between classical and operant conditioning?
In classical conditioning two stimuli are associated and the response is an involuntary reflex: the subject is passive and does not have to do anything. In operant conditioning a voluntary behavior is associated with its consequence, which makes it more or less likely: the subject is active and acts on the environment. Salivating to a sound is classical; raising your hand because the teacher praises you is operant.
Why are slot machines so addictive?
Because they run on a variable ratio schedule: the payout depends on the number of plays, but that number is unpredictable. This schedule produces the highest response rate and the greatest resistance to extinction, because a long losing run is indistinguishable from the schedule's normal gaps and never signals that it is time to stop. The same mechanism drives phone notifications, which follow a variable interval schedule.
What is a Skinner box and what is it for?
It is the operant conditioning chamber Skinner designed in the late 1930s: a soundproof cage in which a rat can press a lever or a pigeon can peck a disc, fitted with a food dispenser and a cumulative recorder that traces responses over time. It allows every variable to be controlled and gives an objective measure of how different reinforcement schedules affect the frequency of behavior.
Does punishment work for discipline?
It suppresses behavior quickly, but with three limitations: it does not teach what to do instead of what it forbids, it works mainly in the presence of the punisher, and it produces avoidance, conditioned anxiety and hostility. Procedures based on reinforcing an alternative behavior, combined with extinction of the unwanted one, give more stable results because they build a new repertoire instead of merely blocking one.