A score is useful because it leaves things out. That may sound like a strange defense of psychological testing, particularly in a culture where “you are more than your IQ” has become almost a reflexive response to cognitive measurement. But measurement could not work without reduction. A thermometer ignores almost everything about a room except temperature; a cognitive test does something conceptually similar, removing much of the complexity of a person’s life in order to ask a narrower question under standardized conditions. The interesting problem is not that information has been compressed. It is what survives the compression—and what conclusions we are entitled to draw from it.
Consider what has happened by the time someone is told that their IQ is 115. They have not answered a single question that revealed a hidden quantity called intelligence. They have completed multiple tasks designed to sample aspects of cognitive performance. Their responses have been scored according to standardized procedures, compared with an appropriate normative group, and, depending on the instrument, combined into broader composites. The final number is several interpretive steps removed from the behavior that generated it.
That does not make the number unreal. Standardized scales are conventions, but the individual differences they summarize need not be. Well-developed cognitive measures show substantial reliability, different cognitive tests correlate in systematic ways, and broad cognitive scores have meaningful relationships with educational and other outcomes (Deary, 2012; Roth et al., 2015). The trouble begins when the apparent precision of the final number encourages a much broader precision of interpretation.
The seduction of a precise number
A score such as 115 looks unusually definite. It is easy to forget that it is an estimate produced by a measurement process, not a physical reading taken directly from the mind. IQ is not a percentage of intelligence, and the distance between two scores is not a literal amount of cognitive substance. Scores are located on a standardized scale, usually constructed so that a reference population has a defined mean and standard deviation. Their meaning depends on the norms, instrument and interpretation for which the measurement has evidence.
There is uncertainty too. No psychological test is perfectly reliable, so responsible interpretation recognizes measurement error rather than treating an observed score as an infinitely precise point. In professional assessment, this is one reason confidence intervals and the broader testing context matter. The number is useful, but its decimal-like authority should not exceed the precision of the evidence behind it.
The same caution becomes more important, not less, when a battery produces many numbers. A profile of subtests and indexes can reveal potentially useful variation, but more detail is not automatically more truth. Narrower scores may be less reliable than broad composites, and an apparent difference between two scores may reflect sampling and measurement error as well as a meaningful difference in ability. The Standards for Educational and Psychological Testing therefore require evidence for interpretations of subscores, score differences and profiles rather than assuming that every visible peak and valley deserves a psychological story (AERA, APA, & NCME, 2014).
The mistake is not using a summary score. The mistake is allowing the summary to inherit meanings that the measurement never established.
This cuts both ways. A broad score should not swallow every potentially informative distinction within a person’s performance. But neither should a profile become a kind of cognitive horoscope in which every difference is assigned significance because it appears on a graph.
“What does this score mean for me?”
That question sounds simple until we notice how many different questions can hide inside it. A parent may want to understand why a child is struggling at school. A student may want to know whether a result explains how intelligent they are. A clinician may be evaluating a specific diagnostic or functional question. A researcher may use the same kind of score as one variable in a population study. An institution may want to make a decision.
The measurement does not acquire a new meaning merely because the question around it has changed.
Modern psychometrics treats validity as a property of the interpretation and use of scores supported by evidence, rather than a universal stamp attached to a test (AERA, APA, & NCME, 2014). That is more than a technical nicety. A broad cognitive score may provide strong evidence about cognitive performance relative to norms while providing little direct information about motivation, personality, values, interests, health, creativity, opportunity or the social conditions in which a person is trying to function.
Those omissions are sometimes presented as an indictment of IQ testing: the test does not measure everything important, therefore the test is impoverished. But no useful measure should be expected to measure everything important. Blood pressure does not become invalid because it tells us nothing about eyesight. The real error is a mismatch between the scope of the evidence and the scope of the conclusion.
This is why “a person is more than a score” is true but scientifically incomplete. Of course a person is more than a score. The harder question is whether the score tells us something worth knowing about the thing it was designed to measure.
In the case of cognitive ability, often it does.
Prediction without prophecy
The real-world relevance of cognitive scores is one reason they have survived more than a century of controversy. Intelligence is substantially related to educational achievement: a large meta-analysis found a strong association between intelligence and school grades (Roth et al., 2015). Longitudinal research has also linked cognitive ability with later educational attainment and occupational outcomes, although the magnitude of the relationship varies by outcome and other influences—including prior achievement and socioeconomic background—matter as well (Strenze, 2007).
A critic who calls IQ “just a number” therefore throws away genuine information. But a person who treats the number as a forecast of an individual life makes the opposite mistake.
Prediction is probabilistic. Two people with similar cognitive scores can differ enormously in what they know, what they want, how healthy they are, which opportunities they encounter, what resources they can use, how persistently they pursue a goal, and which environments reward their particular strengths. Even the same person can perform differently when a task changes or when fatigue, illness, anxiety, familiarity or time pressure alter the conditions of performance.
The point is not that these factors somehow defeat intelligence. It is that a statistical predictor can be useful without being sufficient. Population-level relationships describe tendencies across people; they do not contain individual futures.
What the number cannot carry
Perhaps the most useful way to think about a cognitive score is not as an incomplete portrait of a person, because that metaphor already asks the score to be something it was never meant to be. It is evidence generated for a narrower purpose. Other forms of evidence answer other questions.
Educational history can show what someone has actually learned. Observation can reveal how a difficulty appears in a particular environment. Achievement testing addresses acquired academic skills. Measures of executive functioning, adaptive behavior, personality or mental health address different constructs when those constructs are relevant and the measures are appropriate. A person’s own account can reveal goals, experiences and difficulties that no cognitive composite was designed to encode.
Responsible interpretation is not the search for the one source that contains the “real person.” It is the discipline of matching evidence to questions.
That is also why adding more context does not automatically make an interpretation better. If every low score is explained away by fatigue and every high score is treated as ability, context has become a story rather than evidence. If every uneven profile is declared uniquely meaningful, individualization has become overinterpretation. Scientific caution applies to the contextual explanation as much as it applies to the test.
A cognitive score can be reliable without being exhaustive, predictive without being deterministic, and useful without being a definition of the person who produced it. Those propositions are not compromises between believers and skeptics. Taken together, they are a more accurate description of what psychological measurement is for.
Intelligence is more than a score because intelligence itself is not identical to the measurement used to estimate it—and because a human life contains far more than intelligence. The score matters precisely when we allow it to remain what it is: a disciplined compression of evidence, powerful within its proper scope and misleading when asked to carry the rest.
References
- American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for Educational and Psychological Testing. American Educational Research Association.
- Deary, I. J. (2012). Intelligence. Annual Review of Psychology, 63, 453–482. https://doi.org/10.1146/annurev-psych-120710-100353
- Roth, B., Becker, N., Romeyke, S., Schäfer, S., Domnick, F., & Spinath, F. M. (2015). Intelligence and school grades: A meta-analysis. Intelligence, 53, 118–137. https://doi.org/10.1016/j.intell.2015.09.002
- Strenze, T. (2007). Intelligence and socioeconomic success: A meta-analytic review of longitudinal research. Intelligence, 35(5), 401–426. https://doi.org/10.1016/j.intell.2006.09.004
- —

Gera Cejas


