The Journal26 April 20269 min read
Seven thinking errors — and which of them passed the test
A field guide sorted by the state of the evidence rather than by fame
Birte Englich, Thomas Mussweiler and Fritz Strack put a fully prepared criminal case in front of experienced judges. Before naming their sentence, the judges were asked to roll dice; the result of the roll determined the size of the demand that the case file listed as the prosecution’s request. Those who had rolled a high number handed down, on average, a longer sentence than those who had rolled a low one. The judges knew where the number came from. It helped determine the verdict anyway.
Findings like this are the core of a research programme that began in the 1970s, when Daniel Kahneman and Amos Tversky described decision errors not as stupidity but as side effects of serviceable rules of thumb. Its founding text says both in one sentence:
These heuristics are highly economical and usually effective, but they lead to systematic and predictable errors
Anyone who quotes only the second half turns a finding about good tools into a catalogue of human defects. The programme has since become a genre, and genres sort by fame. This field guide sorts differently: by the state of the evidence. It begins with the best-evidenced pattern and ends with the one whose core is currently under open dispute.
Seven patterns, from best evidenced to most contested
Anchoring. A number in the room pulls the estimate toward itself, even when it is recognisably random. Anchoring is among the best-replicated findings in all of psychology: Many Labs 1 tested it in 2014 across 36 laboratories in twelve countries, and the anchoring effects came out stronger there than in the original. Negotiators, price tags and opening demands live off it. Order effects in miniature belong here as well — what comes first frames what follows; interviews, questionnaires and menus are built on this, and the safeguard against it is rotation.
Hindsight bias. After the outcome, the outcome looks as though it was foreseeable — “it was obvious all along.” Documented since Baruch Fischhoff’s studies of 1975, confirmed in 2004 by a meta-analysis of the research available at the time. It is unfair to every honest evaluation of a decision: people are judged with knowledge the decider did not have.
Confirmation bias. Evidence that fits is sought out and weighted; evidence that disturbs is overlooked or picked apart. The effect carries the Barnum story in this journal, and Loftus’s memory findings along with it: memory too searches for confirmation. Peter Wason’s selection task — presented in 1966, published in 1968 — showed the core in the laboratory; the literature since has shown it everywhere.
Availability. How likely something feels depends on how easily examples come to mind. Plane crashes feel more frequent than car accidents after the news; rare, vivid risks are overrated and banal ones underrated. The 1973 finding belongs to the programme’s robust core. The refinement usually appended to it — that what counts is the ease of retrieval, not the number of examples — stands on weaker ground: the paradigm with which Schwarz and colleagues demonstrated it in 1991 has not survived later attempts at repetition.
Overconfidence. The ranges people give for their estimates are too narrow; experts are not immune. A distinction Don Moore and Paul Healy brought into the literature in 2008 earns its keep here: overconfidence is not one pattern but three. One can overestimate one’s own performance, one can place oneself above others, and one can hold one’s estimates to be more precise than they are. The three do not move together — the second even reverses on difficult tasks, where people place themselves below average. The third stands most firmly: ranges that are too narrow. The courtroom application stands in this journal’s Loftus piece: a witness’s confidence is a feeling, not a seal of quality.
Dissonance reduction. After a decision, the chosen option grows prettier and the rejected one paler; whoever has invested a great deal talks the investment up. Festinger’s one-dollar experiment is the classic of cognitive dissonance; the basic pattern holds, while the most elegant free-choice paradigms acquired repair needs in the replication crisis — file 027 records both.
Loss aversion. The same amount weighs heavier as a loss than as a gain — the value function of prospect theory, drawn out in this archive’s Kahneman file. The framing effect, its best-known consequence, passed the multi-laboratory test in 2018. The core itself, by contrast, is openly contested: in 2018 David Gal and Derek Rucker disputed that losses weigh heavier on average at all; their critics found the rejection overdrawn but conceded that the principal evidence — the endowment effect and the status-quo bias — admits of several explanations. The quarrel is not settled. Of all the entries on this list, the most popular one stands on the weakest ground.
| Pattern | State of the evidence | What it hangs on |
|---|---|---|
| Anchoring | count on it without reservation | Many Labs 1: 36 laboratories, 12 countries — stronger than the original |
| Hindsight bias | count on it without reservation | Fischhoff 1975, meta-analysis 2004 |
| Confirmation bias | secure as a phenomenon | Wason 1968 in the laboratory, a broad literature since |
| Availability | secure at its core | 1973 robust — the ease-of-retrieval refinement of 1991 not |
| Overconfidence | secure, but threefold | Moore & Healy 2008: three patterns that do not run in parallel |
| Dissonance reduction | the basic pattern holds | the most elegant choice paradigms needed repair |
| Loss aversion | core openly disputed | Gal & Rucker 2018 against the standard reading |
Where the patterns become expensive
Two fields of application show what is at stake. In medicine, base-rate neglect decides diagnoses: a positive test for a rare disease usually means no disease even when the test is a good one — whoever ignores the base rate misreads the result, and studies with physicians turn up this error regularly. In 1995 Gerd Gigerenzer and Ulrich Hoffrage demonstrated the remedy at the same time: the same problem, phrased in natural frequencies instead of percentages, is solved correctly by a far larger share of respondents — with no training at all.
In court, several patterns converge: anchors as in the dice experiment, hindsight bias (the deed looks predictable after the deed) and the overconfidence of witnesses. That is why rules of procedure achieve more here than appeals to impartiality.
Three traps in the bias list
The genre has built-in traps, and anyone who takes the patterns seriously should know them. The first is inflation: online collections carry well over a hundred “biases”, many of them renamings, special cases or one-off findings. The number of well-evidenced basic patterns is small; the seven above cover most of everyday use.
The second trap is the question of norms, and it has a name: Gerd Gigerenzer. Part of the “errors” disappears when probabilities are phrased as frequencies, and simple heuristics sometimes beat elaborate calculation in uncertain environments. An error is a deviation from a norm — and norms can be argued about. Kahneman and Tversky replied in 1996 in the same issue and defended the basic patterns; Tversky died that same year. The quarrel was continued as an agreed experiment: in 2001 Kahneman, together with Gigerenzer’s collaborator Ralph Hertwig and with Barbara Mellers as referee, tested whether frequency formats dissolve the conjunction fallacy. They did not do so consistently — and the opponents read the result differently to the last.
The third trap is the most comfortable: the bias as a weapon. Whoever knows lists of thinking errors finds them effortlessly — in the other person. Research has a finding of its own for this, the blind spot for one’s own bias: people rate themselves on average as less biased than average, and having the bias explained changes little. In the study by Emily Pronin and her co-authors, participants insisted on their own objectivity after the effect had been described to them explicitly.
What actually helps against them
The usual short version runs: education changes little, formats change a lot. Frequencies instead of percentages, the outside view instead of the inside view, checklists instead of willpower, structured decision processes instead of appeals to the gut — the remedies that work rebuild the environment, not the head. The short version has become too short. In 2015 Carey Morewedge and colleagues showed that a single training session — a teaching video or a purpose-built teaching game — measurably reduced several biases, and the effect was still detectable two months later. A follow-up study by Anne-Laure Sellier found the same effect in the field in 2019: trained managers, working on a seemingly unrelated business task, less often searched only for confirming data and more often made the better decision. So the head is not entirely out of reach; it is only more expensive to rebuild than the form.
On the architecture side, meanwhile, the balance sheet is being negotiated. A meta-analysis by Mertens and colleagues credited “nudge” interventions with a medium effect in 2022; Maier and colleagues ran the same data for publication bias in the same year and found no evidence left afterwards. The Mertens paper also received a correction that year, among other things over data points from a retracted study and over coding errors. Anyone naming a figure today for the effectiveness of nudges is naming a position in a dispute.
What remains
Of the seven patterns, two are evidenced well enough to be factored into a decision without reservation: anchoring and hindsight bias. Three more — confirmation bias, availability, dissonance reduction — are secure as phenomena, while individual famous paradigms behind them wobble; that does not touch their existence, but it does touch the confidence with which effect sizes may be quoted. Loss aversion, the genre’s flagship, is the only entry where the dispute has reached the core. This ranking is the actual yield: not every bias is equally well evidenced, and the best-known one is not.
In everyday life the abusive pattern is recognisable by the fact that the word “bias” turns up after the disagreement and explains why the other side is wrong. The field guide only becomes useful in the opposite direction — beforehand, applied to one’s own decision, and best of all as a procedure: commit in advance, look up the base rates, write down the counter-thesis, set the ranges wider than feels right. Because the study of bias is not a study of unreason. Kahneman and Tversky’s point was always that the very rules of thumb that produce the errors are what carry everyday life — fast, frugal, mostly right. A judgement system without the availability heuristic would not be wiser but paralysed. The errors are the price of speed, and the art lies in knowing the places where the price is too high.
Sources, and why they are here
Tversky, A., & Kahneman, D. (1974). Judgment under uncertainty: Heuristics and biases. Science, 185(4157), 1124–1131.
The programme's founding text — and the source of the quotation. Its sentence about rules of thumb is the framing without which the whole field guide would be misread as a catalogue of defects.
Tversky, A., & Kahneman, D. (1973). Availability: A heuristic for judging frequency and probability. Cognitive Psychology, 5, 207–232.
The availability finding itself, which belongs to the robust core — unlike the refinement attached to it in 1991.
Schwarz, N., Bless, H., Strack, F., Klumpp, G., Rittenauer-Schatka, H., & Simons, A. (1991). Ease of retrieval as information: Another look at the availability heuristic. Journal of Personality and Social Psychology, 61(2), 195–202.
Precisely that refinement — that what counts is the ease of recall, not the number of examples. Its paradigm has not survived later replication attempts.
Kahneman, D., & Tversky, A. (1979). Prospect theory: An analysis of decision under risk. Econometrica, 47, 263–291.
The value function loss aversion comes from — the most popular and at the same time the most openly disputed item on the list.
Gal, D., & Rucker, D. D. (2018). The loss of loss aversion: Will it loom larger than its gain? Journal of Consumer Psychology, 28, 497–516.
The attack on the core: losses may on average not weigh more heavily at all. The dispute is unsettled, and the article therefore reports it openly.
Fischhoff, B. (1975). Hindsight ≠ foresight: The effect of outcome knowledge on judgment under uncertainty. Journal of Experimental Psychology: Human Perception and Performance, 1, 288–299.
The first description of hindsight bias — one of the two patterns that can be counted on without reservation.
Guilbault, R. L., Bryant, F. B., Brockway, J. H., & Posavac, E. J. (2004). A meta-analysis of research on hindsight bias. Basic and Applied Social Psychology, 26(2–3), 103–117.
The confirmation of that pattern across the research available up to then.
Wason, P. C. (1968). Reasoning about a rule. Quarterly Journal of Experimental Psychology, 20, 273–281.
The selection task that shows confirmation bias in the laboratory — the core the literature has since found everywhere.
Moore, D. A., & Healy, P. J. (2008). The trouble with overconfidence. Psychological Review, 115(2), 502–517.
The distinction without which the overconfidence entry would not be readable: three patterns rather than one, and they do not run in parallel.
Pronin, E., Lin, D. Y., & Ross, L. (2002). The bias blind spot: Perceptions of bias in self versus others. Personality and Social Psychology Bulletin, 28, 369–381.
The genre's third trap, with evidence: participants insisted on their own objectivity after the effect had been described to them expressly.
Englich, B., Mussweiler, T., & Strack, F. (2006). Playing dice with criminal sentences: The influence of irrelevant anchors on experts' judicial decision making. Personality and Social Psychology Bulletin, 32(2), 188–200.
The scene the article opens with. What matters is not that the anchor worked but that those questioned knew where the number came from.
Klein, R. A., et al. (2014). Investigating variation in replicability: A „many labs“ replication project. Social Psychology, 45(3), 142–152.
36 laboratories in twelve countries — and the anchoring effects came out stronger than in the original. That is why anchoring stands first in this field guide.
Gigerenzer, G., & Hoffrage, U. (1995). How to improve Bayesian reasoning without instruction: Frequency formats. Psychological Review, 102(4), 684–704.
The second trap, the question of the norm: part of the “errors” disappears as soon as probabilities are phrased as frequencies.
Gigerenzer, G. (1996). On narrow norms and vague heuristics: A reply to Kahneman and Tversky. Psychological Review, 103, 592–596.
The sharpening of the quarrel over the norm — and the reply Kahneman and Tversky answered in the same issue.
Mellers, B., Hertwig, R., & Kahneman, D. (2001). Do frequency representations eliminate conjunction effects? An exercise in adversarial collaboration. Psychological Science, 12(4), 269–275.
The rare case of an experiment agreed between the camps. That the opponents read the result differently to the last belongs to the finding.
Morewedge, C. K., Yoon, H., Scopelliti, I., Symborski, C. W., Korris, J. H., & Kassam, K. S. (2015). Debiasing decisions: Improved decision making with a single training intervention. Policy Insights from the Behavioral and Brain Sciences, 2(1), 129–140.
The finding that corrects the convenient short version “education does not help”: a single training session worked, and still did two months later.
Sellier, A.-L., Scopelliti, I., & Morewedge, C. K. (2019). Debiasing training improves decision making in the field. Psychological Science, 30(9), 1371–1379.
The same effect outside the laboratory, in an apparently unrelated business task.
Mertens, S., Herberz, M., Hahnel, U. J. J., & Brosch, T. (2022). The effectiveness of nudging: A meta-analysis of choice architecture interventions across behavioral domains. PNAS, 119, e2107346118 — together with its correction (PNAS, 119(19), 2022).
One half of the nudge balance sheet, including the correction for data points from a retracted study and for coding errors.
Maier, M., Bartoš, F., Stanley, T. D., Shanks, D. R., Harris, A. J. L., & Wagenmakers, E.-J. (2022). No evidence for nudging after adjusting for publication bias. PNAS, 119, e2200300119.
The other half: the same data, computed through for publication bias, and afterwards no evidence remains. Anyone who cites a figure today is citing a side.