The Journal29 August 202612 min read
Binet Wanted to Help, Not to Sort
How a tool for school support turned into a ranking
France had extended compulsory schooling, and for the first time children with very different backgrounds sat together in the same classrooms. In 1904 the education ministry appointed a commission to work out how to identify those children who were not benefiting sufficiently from ordinary teaching. The alternative in use until then was an unsystematic judgement by teachers or doctors — equally open to prejudice. For this purpose Alfred Binet and the physician Théodore Simon devised short tasks on attention, memory, comprehension, and practical judgement.
The aim was not to weigh a metaphysical quantity called “intelligence”. It was to find children who needed particular support. Not long afterwards the same procedure was being used on the other side of the Atlantic to rank people by innate worth. That journey is the real history of the intelligence test.
Paris, compulsory schooling, and a practical problem
The tasks in the 1905 scale grew harder with age: following simple instructions, repeating sentences, explaining differences, making sense of pictures, solving practical problems. What mattered was the pattern in comparison with children of a similar age. The revision of 1908 arranged the tasks by age level and so produced the concept of “mental age”: what do children of a given age group usually manage? A third version appeared in 1911, the year Binet died.
Binet warned against reifying his own tool. A weak result was no unalterable label; teaching and practice could change performance. The scale was meant to measure neither brain weight nor soul, but to make educational need visible.
Quelques philosophes récents semblent avoir donné leur appui moral à ces verdicts déplorables en affirmant que l’intelligence d’un individu est une quantité fixe, une quantité qu’on ne peut pas augmenter. Nous devons protester et réagir contre ce pessimisme brutal ; nous allons essayer de démontrer qu’il ne se fonde sur rien.
A dilemma that no measurement resolves was present from the outset. To distribute help fairly, the administration had to form categories. The same category could serve to weed people out. A measuring instrument is therefore never merely a neutral answer; it alters the institutional path of the person measured. Whether a diagnosis means support or removal is decided by the system that comes after the test.
Binet had earlier studied hypnosis, perception, memory, and the performance of his own daughters. Some of that early work contained methodological errors, which he corrected in public. That willingness shaped his later thinking: intelligence did not appear to him as a simple substance that a test lays bare without error. He developed exercises intended to build attention, self-control, and strategies for learning.
His warnings against a “brutal” view of numbers are often read as though he had denied that any stable difference exists. That too would be too simple. Binet held differences to be measurable and practically relevant, but not an unalterable ranking of human worth. He serves only in part as a humane counter-figure to the later eugenicists: he worked within the categories of his time and spoke about children in terms that are stigmatising today. His merit lies less in moral spotlessness than in his methodological modesty about his own instrument.
Théodore Simon worked as a physician with children in institutions and brought clinical experience into the scale. The common shorthand “the Binet test” makes his contribution disappear. Scientific names work like brands: one memorable protagonist collects the credit, while collaboration and institutional labour become invisible. The joint authorship is not a footnote but a reminder that the scale grew out of an exchange between psychology, medicine, and the school.
| Year | Step |
|---|---|
| 1904 | commission of the French ministry of education |
| 1905 | first scale by Binet and Simon |
| 1908 | revision by age levels, “mental age” |
| 1911 | third version, the year of Binet’s death |
| 1912 | Stern’s formula: mental age divided by chronological age |
| 1916 | Terman’s Stanford-Binet, the formula times 100 |
| 1927 | Buck v. Bell |
| today | normed to a mean of 100, standard deviation 15 |
The instrument changes continents
Henry H. Goddard brought an English version to the United States and used it at the Vineland Training School, among other places. He held that “feeble-mindedness” was often hereditary and tied test scores to eugenic policy. His study of the “Kallikak” family set an allegedly sound line of descent against an allegedly degenerate one; later analyses showed severe genealogical and interpretive problems.
At Ellis Island, immigrants were tested under punishing conditions. The popular claim that the test directly decided admission on a mass scale is, in that simple form, historically contested. What is certain is that researchers read low scores from people with little English, little schooling, and a great deal of stress as evidence of innate inferiority. Translation is never only a matter of language. Tasks carry schooling, English vocabulary, and cultural expectation with them; testing newly arrived people under time pressure with unfamiliar formats also measures education, poverty, and health.
Lewis Terman reworked the tasks, standardised them on an American sample, and published the Stanford–Binet test in 1916. He too held intelligence to be largely hereditary and endorsed eugenic measures. From him also comes the most popular piece of arithmetic in the history of psychology: the formula of mental age divided by chronological age had been proposed by William Stern in 1912; Terman multiplied it by 100 and made the intelligence quotient an everyday currency. The number was administratively attractive because schools, the military, and public authorities could now compare large groups.
Terman’s famous longitudinal study was meant to show that high ability does not come paired with physical or social frailty. The “Termites” were followed for decades and supplied valuable life-course data. The selection, however, was not neutral: predominantly white, better-off children who caught their teachers' attention and passed tests. Later educational and occupational success therefore reflects ability, background, selection, and extra attention all at once. Especially instructive are the people who narrowly missed the selection threshold and went on to extraordinary achievement. A cut-off turns a continuous estimate into two institutional groups; right at the boundary, almost identical people receive very different chances.
In the First World War the US Army tested recruits in large numbers. Army Alpha was based on language and writing; Army Beta was meant to capture people with little English or limited literacy. The mass tests promised an efficient allocation of duties and of leadership potential, and they established psychological examination as an administrative technology.
Conditions were unsettled, the instructions unfamiliar, the tasks culturally loaded. Robert Yerkes and colleagues reported average differences between groups of origin, which the public read as a mental hierarchy and which fed into the immigration debates.
The core methodological error is still with us. A test can distinguish groups reliably under particular conditions without naming the cause of the difference. Reliability is not causation. Whoever measures a stable gap has not yet shown whether genes, education, language, health, discrimination, or familiarity with the tasks produced it. A difference can be measured precisely and still be explained entirely wrongly.
From diagnosis to coercion
Eugenic movements wanted to steer reproduction by social means. In the United States, laws led to the forced sterilisation of tens of thousands of people — often people with disabilities, poor women, and members of marginalised groups. Tests supplied a scientific-sounding part of the justification; they were not the only cause.
In 1927, Buck v. Bell permitted in American law the sterilisation of a young woman whose supposed hereditary “inferiority” was poorly evidenced. Institutional interests, social prejudice, and legal power turned a contested diagnosis into irreversible bodily coercion.
The ethical line is not first crossed at the mistaken measurement. Even a perfectly measured cognitive difference would give nobody the right to dispose of another person’s reproduction. Criticism of measurement and human rights are separate lines of protection; both are needed.
What an IQ score means today
Modern tests usually scale results to a mean of 100 and a standard deviation of 15. A score describes a position relative to an age-referenced norm sample. It is not a percentage of correct intelligence and not a fixed possession like height. For adults, a linearly growing “mental age” is in any case a poor fit; modern tests use deviation scores relative to the age norm. The name IQ remained; its statistical meaning changed.
Every score carries a confidence interval. Daily form, measurement error, and the choice of subtests all affect the result. A report that states 92 rather than a plausible range creates false precision. Near a cut-off, a few points can have large consequences even though the statistical uncertainty reaches across the boundary. Norms age as well: access to education, health, and familiarity with tests change performance across generations, a pattern known as the Flynn effect, whose course is not the same everywhere. Without restandardisation, the meaning of the same raw performance shifts.
Many cognitive tasks correlate positively: someone who does well in one area tends on average to score higher in others too. Charles Spearman called the shared factor “g”; modern hierarchical models combine a general factor with broader and more specific abilities. But the statistical existence of a factor does not say that a substance called “intelligence” sits somewhere in the brain. Factors summarise covariation; their causes can be various. Education can strengthen several abilities at once, and basic processing mechanisms can act across different tasks.
This is why, when analysing people, it is particularly dangerous to confuse ability with worth. Two similar IQ scores can arise from quite different profiles; language, working memory, processing speed, and spatial reasoning contribute in different measure. A test can predict how well certain school demands will be met. It does not measure morality, creativity, care, or the right to self-determination.
Culture-freedom, benefit and harm
One can reduce the language load, choose familiar tasks, and improve norm groups. Entirely culture-free intelligence tests barely exist. The very assumption that solving an abstract pattern quickly and alone is the relevant form of performance carries cultural and institutional values with it.
Fairness does not mean denying differences. It requires checking whether tasks measure the same ability in all groups, whether predictions work equally well, and whether decisions disadvantage people unnecessarily. A test can be psychometrically comparable and its use still be unjust. Multilingual assessment needs qualified interpreting and information about educational history; a word-for-word translation preserves neither difficulty nor meaning, and norms from one country are not constants of nature.
The room itself has an effect too. People know which groups are held to be less intelligent; if a situation makes that stereotype salient, worry and self-monitoring can load working memory. Research on stereotype threat has produced heterogeneous effects and replication disputes, and does not serve as a universal explanatory key. Better supported is the basic insight: test performance arises in a social situation in which anxiety, mistrust, sleep, pain, and the relationship with the examiner all play a part. Anyone who has repeatedly seen institutions use tests against their own group brings a well-founded mistrust along. Trust is therefore not a soft extra but part of a valid measurement.
In a well-founded evaluation, a test can show whether school difficulties are broad or specific, whether performance has declined after a neurological illness, or what support would be appropriate. It is one component alongside conversation, developmental history, classroom observation, and medical information. The score becomes useful when it improves a concrete decision: more time, suitable teaching, accessible communication, targeted rehabilitation. It becomes harmful when it is collected without a question, presented as unchangeable, or used to establish a social rank.
Labels work in both directions. An eight-year-old who solved the tasks of six-year-olds could be described as having a “mental age” of six — vivid for lesson planning, wrong as a description of the person. That child has the body, the relationships, the interests, and the life experience of an eight-year-old. The label lowers expectations: adults speak more simply, hand over less responsibility, overlook uneven profiles. So the measurement acts back on the environment, which in turn influences performance.
At the upper end something similar happens. High scores open up support programmes and yet become a template for identity: pressure of expectation, avoidance of tasks that carry a risk of failure, ordinary difficulties read as collapse. “Gifted” explains neither motivation nor emotional maturity. Conversely, criticism of labels must not lead to ignoring particular learning needs; being under-challenged is real. Carol Dweck’s popular distinction between fixed and growth mindsets is often applied too simply here. Effort alone removes no barriers, and praise for “mindset” is no substitute for resources.
A professional feedback session therefore translates numbers into limits and possibilities. It explains the uncertainty, avoids demeaning age comparisons, and names the factors that may have influenced the result. The person tested has a right to learn more than a number. The question is not only “how high?” but “which task fails, for what reason, and what help changes the outcome?”
What remains
The history does not teach us to abolish measurement. It teaches us to think about the social journey a measurement takes. Binet built an educational warning signal; others turned it into a ranking of supposed nature — that the same instrument could do both shows that tests carry no ethics of their own. When examinations appear to order performance objectively, the resulting hierarchy looks deserved. Yet test results also depend on family, school, health, and social security. The objection is not that all people have identical abilities. It is that difference and desert are not the same thing. John Rawls’s thought experiment poses the practical question: which rules for the consequences of tests would we choose if we did not know what talent, language, or disability we would be born with?
In practice that means attaching one sentence to every measurement: this result holds for this person, under these conditions, in comparison with this norm, and for this limited purpose. Without those specifications, statistics turns into the reading of character. And of every test result the same question can be asked that lay behind Binet’s mandate of 1904. Does this number open access to learning and self-determination, or does it shut a person in?
Sources, and why they are here
Binet, A., & Simon, T. (1905). Méthodes nouvelles pour le diagnostic du niveau intellectuel des anormaux. L'Année Psychologique, 11, 191–244.
The 1905 scale in its first publication — the basis for every detail on the kind of tasks and the purpose of the instrument.
Binet, A. (1909). Les idées modernes sur les enfants. Paris: Flammarion.
Binet's own warning against reification: the source of the quoted sentence against the “pessimisme brutal”, chapter V.3, p. 141.
Terman, L. M. (1916). The Measurement of Intelligence. Boston: Houghton Mifflin.
The American revision, the Stanford-Binet test and the multiplication of Stern's formula by 100.
Gould, S. J. (1981/1996). The Mismeasure of Man. New York: Norton.
The classic criticism of Goddard, Yerkes and the army tests — the basis of the account of the Kallikak study and of mass testing.
Carson, J. (2007). The Measure of Merit: Talents, Intelligence, and Inequality in the French and American Republics, 1750–1940. Princeton: Princeton University Press.
The historical comparison that explains why the same instrument meant support in France and a ranking in the United States.
Nicolas, S., Andrieu, B., Croizet, J.-C., Sanitioso, R. B., & Burman, J. T. (2013). Sick? Or slow? On the origins of intelligence as a psychological object. Intelligence, 41(5), 699–711.
Reconstructs Binet's commission and the emergence of the scale from the documentary record — the evidence for the 1904 commission and for Simon's share.