Skip to content
AnamneseArchive of Psychology

At least two characters

The Journal25 January 20267 min read

The Doll That Hit Back

What the Bobo experiment shows — and what was pinned on it

Bar chart for Bandura 1965: in all three conditions — model rewarded, no consequences, model punished — the imitated actions converge as soon as an incentive is offered.
The move of 1965: those who had seen the punishment imitated less often — until a small reward wiped out the differences completely.Drawing by the archive

A child, not yet six, sits alone in a room at the Stanford University nursery school. Shortly before, it was shown attractive toys and then told they were reserved for other children. Now there are toys in the room, among them a wooden mallet and an inflatable bottom-weighted doll. The child reaches for the mallet and strikes — in the same movements it saw an adult make twenty minutes earlier, in part with the same phrases. The model had spoken these words:

Sock him in the nose… Hit him down… Throw him in the air… Kick him… Pow…

Albert Bandura, Dorothea Ross & Sheila A. Ross, Transmission of Aggression Through Imitation of Aggressive Models (1961)

That was how the best-known experiment in developmental psychology ran in 1961. Seventy-two nursery children, 36 boys and 36 girls between 37 and 69 months old, were distributed across three conditions of 24 children each: an adult worked over the doll in front of them — mallet blow, throw into the air, punch, kicks — or played quietly, or there was no model at all. The children in the aggressive condition then struck the doll themselves, and not in any old way but quoting. In 1963 the follow-up study with 96 children showed that film and cartoon models work much as live ones do: all three kinds of model lay clearly above the control group.

The finding was a problem for the ruling doctrine. Behaviourism explained behaviour from consequences — what is reinforced occurs more often; whoever does nothing learns nothing. Missing from that bookkeeping was precisely the item childhood mostly consists of. Children acquire language, manners and bad habits at a pace no reinforcement schedule of operant conditioning can deliver, and they display actions they have never performed and for which they have never seen a reward. Albert Bandura’s answer to that gap stands in file 028 of this archive; this piece tells what became of it.

The move of 1965

The actual refutation came four years after the doll, and it used the other side’s own concepts. Bandura showed 66 children the same film in three versions: at the end, the aggressive model was rewarded, punished or not attended to at all. Those who had seen the punishment imitated less often — the expected finding. Then all the children were offered a small reward for demonstrating what they had seen, and suddenly they could all do it equally well. The incentives wiped out the previously observed differences entirely.

Acquisition and performance are two different things: all of them had learned it, and the ones who saw a reason showed it. Reinforcement does not govern what is learned but what of it reaches the stage. The whole programme of social learning rests on that distinction — and with it, orthodox behaviourism as a complete account was finished, not by polemic but by a design.

What the retellings leave out

The care of the groundwork is regularly missing from the retellings. The children were rated in advance by the experimenter and one of their teachers for habitual aggressiveness and assigned in matched triplets so that the wilder ones would not end up by chance with the aggressive model. Coding ran behind the one-way screen in five-second intervals, with a second observer rating half the children independently, at agreements around .90.

For 1961 that is carefully built — with one exception the retellings likewise pass over: every session was coded by the man who had himself played the model in part of the runs. That is not a trifle at some peripheral detail but a weakness of execution at the central measurement. Anyone proceeding that way today would get the paper back.

The course of the 1961 experiment in four steps
Fig. 1 — 72 nursery children, three groups of 24: assignment, model phase, deliberate frustration, test room. Drawing by the archive

The findings on sex deserve a note as well: boys struck more overall, and the sex of the model counted — the aggressive man was imitated more than the aggressive woman, who tended to provoke bafflement in the post-interviews. Children remarked that a lady does not behave like that. It is one of the earliest pieces of evidence that models are not copied as sources of movement but read as social roles.

What was loaded onto the experiment

Then the experiment had more piled on it than it can carry. In the quarrel over media violence the doll became a weapon for both camps, and the methodological objections are as old as they are justified. A bottom-weighted doll is an object built to be hit; striking it is play, not aggression against a creature that hits back or suffers. It is a textbook case of missing ecological validity. On top of that, the observation followed immediately on the model, in the same room, with the same toys, after deliberate frustration — the situation told the children plainly enough what was expected of them. The technical term for this is demand characteristics, and they strike the experiment at its core.

What the doll shows, then, is readiness to imitate in pure culture — not the road from the screen to the criminal offence. The meta-analyses on media violence accordingly find small, scattered associations whose size varies with study quality and with the camp the authors belong to: Craig Anderson and colleagues arrived in 2010 at robust effects, Christopher Ferguson in 2015 at practically meaningless ones. A re-analysis by Joseph Hilgard and colleagues corrected the 2010 data base for publication bias and found the short-term effects overestimated. Both shortcuts mislead: whoever derives a ban from these findings overrates them; whoever declares them refuted confuses small effects with none.

The useful half of the programme

The quarrel long obscured what became of model learning in practice. In therapy it is a technique: whoever treats a phobia performs the feared act themselves, lets the patient watch and then join in. In 1969 Bandura, Blanchard and Ritter compared this participant modelling with symbolic modelling, with systematic desensitisation and with a control condition — participant modelling won, and it stands in the manuals to this day.

From the same line comes the concept that made Bandura the most-cited psychologist of his era: self-efficacy, the task-bound conviction that one can do a particular thing under particular circumstances. In 1977 Bandura listed four sources that feed it: one’s own attempt mastered, the observed model, the encouragement of others, and the reading of one’s own bodily arousal. The order is not an enumeration but a ranking — one’s own success acts most strongly, the model only after it. That is why watching alone is rarely enough, and at the same time why watching is what makes the first attempt of one’s own possible at all. And the largest application runs on television: serial formats that embed behavioural models in entertainment, complete with visible consequences of what the characters do, have been produced since the 1970s, first in Mexico and afterwards on several continents; the efficacy evidence is methodologically difficult, the reach is not.

What remains

The learning mechanism holds; the aggression reading has shrunk. That people acquire patterns of action by watching is undisputed and evidenced in countless variants — down into the first months of life, where a quarrel of its own is running: Andrew Meltzoff and Keith Moore’s 1977 finding on imitation in newborns was not confirmed by a large longitudinal study by Janine Oostenbroek and colleagues in 2016. What remains secure is the distinction of 1965: children learn what they see; whether they do it is decided by something else — incentive, opportunity, expected consequences. What is not secure is everything the doll supposedly proves about violence on screen. It never tested that. The file in this archive accordingly lists the 1961 study as only partly replicated: the mechanism stands, the generalisation does not.

Whoever looks for the neighbourhood with Skinner finds it not in contradiction but in a division of labour: shaping builds behaviour up in steps of approximation, modelling skips the steps where an example is enough. Behaviour therapy uses both to this day, depending on whether a repertoire is missing or a fear stands in the way. In everyday life the difference shows up in one simple question: is the other person missing the action — or only the reason to display it? And for the archive there remains the second lesson, the more expensive one: the same experiment that separates acquisition from performance in the professional literature stands, in public, for a thesis about televised violence it never tested. Between what a study shows and what it stands for there often lies half a century of newsprint. The doll cannot help it.

Sources, and why they are here

  1. Bandura, A., Ross, D., & Ross, S. A. (1961). Transmission of aggression through imitation of aggressive models. Journal of Abnormal and Social Psychology, 63, 575–582.

    The experiment itself — and the source of the quotation: the model's phrases stand there verbatim, as do the children's pre-ratings and the coding behind the one-way mirror.

  2. Bandura, A., Ross, D., & Ross, S. A. (1963). Imitation of film-mediated aggressive models. Journal of Abnormal and Social Psychology, 66, 3–11.

    The follow-up with 96 children: film and cartoon models work much as live ones do. Without it the finding would stay tied to a person in the room.

  3. Bandura, A. (1965). Influence of models' reinforcement contingencies on the acquisition of imitative responses. Journal of Personality and Social Psychology, 1, 589–595.

    The real refutation of orthodox behaviourism, and in its own terms: the reward wiped out the differences the punishment had produced.

  4. Bandura, A., Blanchard, E. B., & Ritter, B. (1969). Relative efficacy of desensitization and modeling approaches for inducing behavioral, affective, and attitudinal changes. Journal of Personality and Social Psychology, 13, 173–199.

    The turn to practice: participant modelling won against three comparison conditions and stands in the treatment manuals to this day.

  5. Bandura, A. (1977). Self-efficacy: Toward a unifying theory of behavioral change. Psychological Review, 84, 191–215.

    The four sources of self-efficacy — and the fact that their order is a ranking. One's own mastered attempt works more strongly than any model.

  6. Meltzoff, A. N., & Moore, M. K. (1977). Imitation of facial and manual gestures by human neonates. Science, 198(4312), 74–78.

    The finding that pushed imitation back into the first weeks of life — and has been disputed ever since.

  7. Oostenbroek, J., Suddendorf, T., Nielsen, M., et al. (2016). Comprehensive longitudinal study challenges the existence of neonatal imitation in humans. Current Biology, 26(10), 1334–1338.

    The large longitudinal study that did not confirm it. The dispute belongs in the article because the mechanism itself is undisputed — only its earliest onset is not.

  8. Anderson, C. A., Shibuya, A., Ihori, N., Swing, E. L., Bushman, B. J., Sakamoto, A., Rothstein, H. R., & Saleem, M. (2010). Violent video game effects on aggression, empathy, and prosocial behavior in Eastern and Western countries: A meta-analytic review. Psychological Bulletin, 136(2), 151–173.

    One side of the media-violence debate: sturdy if small effects. It stands here because the article quotes both camps in their own work rather than retelling either.

  9. Ferguson, C. J. (2015). Do Angry Birds make for angry children? A meta-analysis of video game influences on children's and adolescents' aggression, mental health, prosocial behavior, and academic performance. Perspectives on Psychological Science, 10(5), 646–666.

    The opposing side, with practically meaningless effects. That both meta-analyses read the same literature is the real finding.

  10. Hilgard, J., Engelhardt, C. R., & Rouder, J. N. (2017). Overstated evidence for short-term effects of violent games on affect and behavior: A reanalysis of Anderson et al. (2010). Psychological Bulletin, 143(7), 757–774.

    Computes the 2010 data base through for publication bias and finds the short-term effects overestimated — the test that settles the quarrel at least at one point.