Skip to content
AnamneseArchive of Psychology

At least two characters

The Journal2 November 20258 min read

Therapy Works — The Quarrel Is About Why

How far the dodo verdict carries and where it demonstrably ends

The large effect and the small difference: 0.68 for therapy against no treatment (Smith & Glass 1977), at most 0.20 between bona fide procedures (Wampold 1997). The individual practitioner accounts for about 5 per cent of outcome variance — more than the choice of procedure explains.Drawing by the archive

The bird comes from Alice in Wonderland: after a race with no rules and no finishing line, the Dodo announces that everybody has won and all must have prizes. Saul Rosenzweig borrowed the scene in 1936 for a conjecture that was heretical at the time — the competing schools of therapy might all work, and for reasons none of them carried in its theory. It was a four-page aside. It became psychotherapy research’s most stubborn hypothesis.

That the question became decidable at all is owed to the tools this archive describes in the Beck file: only measure and manual made procedures comparable. Before that, schools quarrelled over interpretations; afterwards, over effect sizes.

At last the Dodo said, “Everybody has won, and all must have prizes.”

Lewis Carroll, Alice’s Adventures in Wonderland (1865), chapter 3, “A Caucus-Race and a Long Tale”

What the meta-analyses found

When Lester Luborsky and colleagues assembled around a hundred comparative studies in 1975, they found Rosenzweig’s pattern: hardly any robust differences between procedures. Two years later Mary Lee Smith and Gene Glass presented the field’s first large meta-analysis — 375 controlled evaluations, mean effect size 0.68. The average treated patient was better off than about three quarters of the untreated. Here too: little between the schools.

The meta-analyses since — from Wampold’s analyses to the Cuijpers series on depression treatment — have confirmed it in essence: psychotherapy works, with medium effects and beyond the end of treatment; the differences between properly conducted procedures are small and shrink further the stricter the methodology. This archive’s Beck file records the same finding from the other side: even the best-tested school’s lead largely vanishes once the comparison condition is fair.

How stubborn the finding is shows in its career under sharpened conditions. As the comparisons grew fairer — equal dose, equal therapist fidelity, blinded raters, preregistered endpoints — the dodo did not shrink but grew sharper. Bruce Wampold’s meta-analysis of 1997 confined itself to bona fide procedures, that is, those with a theory, a manual and convinced practitioners. The effects scattered homogeneously around zero; even generously reckoned, the upper bound on a true difference lay at about 0.20. And the head-to-head trials of cognitive therapy against short-term psychodynamic therapy in depression, long the hoped-for decider, ended repeatedly in non-inferiority. The largest of them, a Dutch trial with 341 patients, saw no difference on any measure at any time point. The suspicion that all of this is mere measurement blur grows more expensive with every larger study.

FindingFigure
Rosenzweig’s paper, 1936four pages
Luborsky et al., 1975around 100 comparative studies
Smith and Glass, 1977375 evaluations, mean effect size 0.68
Wampold, 1997upper bound on a true difference about 0.20
Jacobson et al., 1996150 patients across three conditions
Driessen et al., 2013341 patients, no difference on any measure
Flückiger et al., 2018295 studies, more than 30,000 patients
Johns et al., 2019around 5 per cent of variance per practitioner (0.2 to 29)

The contextualists' explanation

What works, then? Jerome Frank had formulated the answer in 1961 in Persuasion and Healing, long before the numbers existed. Healing between people rests on four things: a trusting relationship, a place and ritual carrying the expectation of healing, a plausible explanatory model for the suffering — and a procedure both sides hold to be effective. Wampold later built this into a contextual model that assigns the procedures the role of vessels.

The most stable predictors in outcome research fit this. The quality of the working alliance predicts success more reliably than school membership — a collective analysis of 295 studies with more than 30,000 patients found the relationship throughout. And the differences between individual therapists are larger than those between procedures. A systematic review of 20 studies estimates the share of outcome variance attributable to the individual practitioner at around 5 per cent in the weighted mean. Across studies, though, that share scatters considerably, from 0.2 to 29 per cent. Five per cent sounds small, but it is markedly more than the choice of procedure explains.

In retrospect, Rogers’s file reads like the script of this literature — warmth, genuineness, close listening as necessary conditions. And even the allegiance bias, the fact that procedures do better in the hands of their adherents, is a common factor in a lab coat: conviction works, including the investigator’s.

The objection: where procedures do decide

The dodo has serious opponents, and their strongest argument is resolution. Averaged over “psychotherapy” and “disorder”, differences vanish that exist at the level of disorders. For specific anxiety disorders, exposure reliably beats procedures without exposure; for obsessive-compulsive disorder, exposure with response prevention is superior; for chronic nightmares, panic and PTSD there are specific protocols with an edge. The dodo, say the critics, is an artefact of the bird’s-eye view — whoever averages everything finds nothing, like someone comparing antibiotics and painkillers “on average”.

Then there is the publication side. Part of the equality findings rests on studies too small to find differences at all — absence of evidence booked as evidence of absence. A non-significant difference between two groups of 20 patients each says nothing about equality.

There are direct counter-findings too. David Tolin’s meta-analysis of 2010 saw cognitive behavioural therapy as superior to the comparison procedures in anxiety and depression, above all to psychodynamic therapy. Three years later a group around Timothy Baardseth recomputed the same studies and admitted only bona fide comparisons — the lead vanished. The quarrel runs, as so often, over the inclusion criteria: whoever admits a control condition nobody seriously believes to be effective is not measuring two therapies but one therapy and a waiting period with the offer of a chat.

The honest interim balance therefore has two storeys. At the level of disorders there are real specifics, concentrated where a procedure serves a clear mechanism — exposure against avoidance is the model case. At the level of the broad diagnoses, depression first among them, the dodo carries to this day: there, alliance, fit, expectation and therapist decide more than the label. Thinking both storeys at once is uncomfortable and correct; the camps in this debate spent decades each picking one.

Why the debate was productive

One can tell the dodo debate as a stalemate. More accurate: it forced both sides to get better.

To the specificity side the field owes manuals, control conditions, mechanism research — and the dismantling studies that run the individual components of a procedure against one another. The most famous comes from Neil Jacobson and colleagues: in 1996 they randomly assigned 150 depressed patients to three conditions — behavioural activation from cognitive therapy alone, that plus work on automatic thoughts, or the complete cognitive therapy. The complete treatment produced no better results than its components, neither at the end nor after six months. Nor did it change negative thoughts more strongly — although that was supposed to be precisely its mechanism. A larger follow-up study around Sona Dimidjian confirmed the power of behavioural activation in 2006. A finding that in turn supports the revision of helplessness theory, as this journal tells elsewhere.

To the context side the field owes alliance research, routine outcome monitoring — measuring results instead of assuming them — and the insight that therapist selection and training hold more leverage than the next war of schools.

The economics of the quarrel deserves telling too. Schools are not only theories but training institutes, certificates and markets; a finding by which the label decides little threatens business models, and the heat of some replies becomes intelligible only against that background. Conversely, the common-factors side has its own temptation: whoever ascribes everything to the relationship is spared the labour of the specifics — and for the patient with OCD who would have profited from exposure, that can carry a real cost.

What remains

Two things are settled. Psychotherapy works, quantified in robust orders of magnitude since Smith and Glass. And between properly conducted procedures run fairly against one another, the differences in broad diagnoses are so small that they are rarely decisive for a treatment decision. The other side of that coin is just as settled. Where a procedure serves a particular mechanism — exposure against avoidance — it is superior to the procedures without that mechanism. That this lead disappears in the overall average is an artefact of calculation, not a finding.

What remains contested is the middle ground: how many such specifics there are, how large they are, and whether they hold up under strict enough testing. That is exactly where the front between Tolin and Baardseth runs, and it is not decided.

How to recognise this in one’s own case: the question “which form of therapy is the best?” is almost always the wrong one for a broad diagnosis. The better questions are whether this is a specific disorder with a tested protocol, whether the person across the room seems trustworthy, whether something gets measurably better after a few sessions — and whether anyone is looking if it does not. Frankl’s folder stands at the end of this story like a footnote with long-range effect: his insistence that people need an explanatory model and a frame of meaning in order to bear suffering has been caught up by common-factors research in soberer language. The dodo is no annoyance but a finding about the nature of helping. That the vessels dose differently for some ailments does not refute it. It only turns the racecourse, in some places, back into a race with a finishing line.

Sources, and why they are here

  1. Carroll, L. (1865). Alice’s Adventures in Wonderland. London: Macmillan, chapter 3.

    The scene that gave the debate its name — source of the quotation and of the image of a race without a finishing line.

  2. Rosenzweig, S. (1936). Some implicit common factors in diverse methods of psychotherapy. American Journal of Orthopsychiatry, 6, 412–415.

    The four pages with which the common-factors hypothesis begins.

  3. Frank, J. D. (1961). Persuasion and Healing. Baltimore: Johns Hopkins University Press.

    Set out the four conditions of healing long before there were figures for them — the theoretical basis of the contextual model.

  4. Luborsky, L., Singer, B., & Luborsky, L. (1975). Comparative studies of psychotherapies. Archives of General Psychiatry, 32, 995–1008.

    The first large synthesis of some hundred comparative studies; it found Rosenzweig's pattern again.

  5. Smith, M. L., & Glass, G. V. (1977). Meta-analysis of psychotherapy outcome studies. American Psychologist, 32, 752–760.

    The field's first large meta-analysis — source of the 375 evaluations and the mean effect size of 0.68.

  6. Jacobson, N. S., Dobson, K. S., Truax, P. A., Addis, M. E., et al. (1996). A component analysis of cognitive-behavioral treatment for depression. Journal of Consulting and Clinical Psychology, 64(2), 295–304.

    The most famous dismantling study: 150 patients, three conditions, no lead for the complete treatment.

  7. Wampold, B. E., et al. (1997). A meta-analysis of outcome studies comparing bona fide psychotherapies. Psychological Bulletin, 122, 203–215.

    Restricts the comparison to bona fide procedures; source of the upper bound of about 0.20.

  8. Dimidjian, S., Hollon, S. D., Dobson, K. S., et al. (2006). Randomized trial of behavioral activation, cognitive therapy, and antidepressant medication in the acute treatment of adults with major depression. Journal of Consulting and Clinical Psychology, 74(4), 658–670.

    The larger follow-up study, which confirmed the power of behavioural activation.

  9. Tolin, D. F. (2010). Is cognitive-behavioral therapy more effective than other therapies? Clinical Psychology Review, 30, 710–720.

    The strongest published counter-finding to the dodo verdict — one side of an open quarrel.

  10. Baardseth, T. P., et al. (2013). Cognitive-behavioral therapy versus other therapies: Redux. Clinical Psychology Review, 33, 395–405.

    Recomputes Tolin's studies using bona fide comparisons only; the lead disappears. The other side of the same quarrel.

  11. Driessen, E., et al. (2013). The efficacy of cognitive-behavioral therapy and psychodynamic therapy in the outpatient treatment of major depression. American Journal of Psychiatry, 170, 1041–1050.

    The largest head-to-head trial in depression treatment — source of the 341 patients.

  12. Wampold, B. E., & Imel, Z. E. (2015). The Great Psychotherapy Debate. 2nd ed. New York: Routledge.

    The worked-out version of the contextual model; the standard work of the dodo side.

  13. Flückiger, C., et al. (2018). The alliance in adult psychotherapy: A meta-analytic synthesis. Psychotherapy, 55, 316–340.

    295 studies with more than 30,000 patients; establishes the working alliance as the most stable predictor.

  14. Johns, R. G., Barkham, M., Kellett, S., & Saxon, D. (2019). A systematic review of therapist effects. Clinical Psychology Review, 67, 78–93.

    Source of the roughly 5 per cent of outcome variance attributable to the individual practitioner — and of the wide scatter behind it.

  15. Cuijpers, P., Noma, H., Karyotaki, E., Vinkers, C. H., et al. (2020). A network meta-analysis of the effects of psychotherapies, pharmacotherapies and their combination in the treatment of adult depression. World Psychiatry, 19(1), 92–107.

    The current network meta-analysis of depression treatment; it carries the finding into the present.