Most conversations about intelligence begin too late. They begin with an IQ score, a school result, a theory about g, or an argument about whether intelligence is fixed. By then, several different questions have already been compressed into one word. Are we talking about a person’s capacity to reason through something unfamiliar, what they already know, how efficiently they learn, how they performed on a particular task, or the standardized score produced by an assessment? These things are related closely enough to be confused and different enough that the confusion matters.
That is one reason intelligence has proved both scientifically productive and publicly difficult. Psychologists can measure cognitive differences with considerable reliability. Performance across diverse cognitive tasks tends to correlate, and those correlations can be represented statistically by a general dimension conventionally called g. Cognitive test scores also predict outcomes outside the testing room, including educational achievement and, more modestly, aspects of later occupational and socioeconomic attainment (Deary, 2012; Roth et al., 2015; Strenze, 2007). None of this is trivial. But none of it turns intelligence into a single object that a test has somehow captured whole.
The strange thing about measuring a mind
We never observe intelligence directly in the way we observe height. We observe people doing things: learning a rule, solving a novel problem, remembering relevant information, retrieving knowledge, noticing a pattern, changing strategy after an error. Cognitive tests make some of these performances comparable by placing them under standardized conditions. Researchers then look for systematic patterns in the results and infer abilities that help account for those patterns.
This sounds like a technical distinction, but it changes the conversation. A person’s answer to a test item is an observed performance. A cognitive ability is a construct inferred from patterns of performance. A standardized score is a way of expressing measurement information within a particular scoring and normative system. And a theory of intelligence goes further still, proposing something about how cognitive abilities are organized, develop, operate or arise. Moving from one level to another may be scientifically justified, but the movement has to be earned.
The history of intelligence research becomes easier to understand once these levels are separated. The positive manifold—the tendency for performance on different cognitive tests to correlate—is an empirical pattern. g is a statistical representation of common variation within that pattern. A theory claiming that g reflects a single underlying cognitive resource is an interpretation of why the pattern exists. Those statements are not synonyms. A factor model can organize covariance without, by itself, identifying the biological or cognitive mechanism that produced it.
This is also why the familiar claim that “IQ is intelligence” is too crude. IQ is useful precisely because it compresses information. A well-constructed score can summarize performance across a standardized battery and locate that performance relative to a normative group. The fact that the result is constructed through measurement does not make it arbitrary; virtually all serious measurement involves conventions about what is sampled, how observations are combined and how the resulting scale is interpreted. The question is not whether a score simplifies. The question is what that simplification preserves.
A score can summarize evidence about cognitive ability. It is not intelligence made visible as a number.
What survives the compression
Psychological measurement becomes controversial partly because people often ask a score to answer a question larger than the one the score was designed to answer. Someone receives a cognitive result and asks what it says about their future, their competence, their potential, their worth, or why they struggle in a particular setting. The score may be relevant to some of those questions. Relevance, however, is not the same thing as completeness.
Modern testing standards make this point in less dramatic language: validity concerns the interpretation of test scores for proposed uses (AERA, APA, & NCME, 2014). A measure can provide good evidence about broad cognitive ability without measuring motivation, personality, values, creativity, health, opportunity or the particular circumstances in which a person will have to act. These omissions do not invalidate the measurement if those were never the intended targets. They become a problem when an interpretation quietly expands beyond the evidence.
The opposite mistake is just as easy. Because a score cannot explain an entire person, it is tempting to conclude that cognitive measurement says little of importance. Yet intelligence is substantially associated with school achievement, and longitudinal evidence connects cognitive ability with later educational and occupational outcomes (Roth et al., 2015; Strenze, 2007). Those associations are real without being destiny. A relationship observed across thousands of people is not a script for any one person’s life.
The difficult position is therefore the scientifically useful one: cognitive differences matter, measurement can capture meaningful information about them, and the resulting information remains narrower than the human being from whom it was obtained.
Outside the testing room
Standardized assessment deliberately reduces some of the messiness of ordinary life. Real life restores it.
Two people may have similar underlying ability and perform very differently because one knows far more about the problem. A capable person may perform badly while exhausted, ill or overwhelmed. A difficult interface can turn a simple decision into an unnecessary memory task; a checklist can remove that burden without changing the person’s intelligence. Education can supply knowledge that transforms how a problem is represented, and there is evidence that additional education can itself increase measured cognitive ability rather than merely revealing a fixed capacity (Ritchie & Tucker-Drob, 2018).
None of this requires us to dissolve intelligence into “context.” Relatively stable individual differences in cognitive ability exist. The point is that observed performance is produced when a person with particular abilities and knowledge encounters a particular task under particular conditions. Intelligence is one important part of that system, not a synonym for its final output.
That distinction becomes especially valuable when something goes wrong. “They couldn’t do it” describes an outcome. It does not tell us whether the limiting factor was reasoning, knowledge, memory, misunderstanding, fatigue, task design, time pressure, motivation or something else. Different mechanisms can produce superficially similar failures, and they imply different explanations.
The science gets better when the questions get narrower
Once intelligence, IQ, g, knowledge and performance stop standing in for one another, the field becomes more interesting rather than less. Why do different cognitive abilities correlate? What does g represent, and what causes the covariance it summarizes? Why do people nevertheless show differentiated abilities? How stable are cognitive differences across life? How much can education change them? When does a broad composite tell us more than a profile of narrower scores? How should intelligence be distinguished from executive function, creativity, achievement or wisdom?
Those are not complications added after the “real” question of intelligence has been answered. They are the questions that make the science possible.
There is no need, then, to choose between taking intelligence seriously and taking human complexity seriously. Psychometrics has given psychology powerful tools for describing individual differences in cognitive performance, and the strongest findings from that tradition deserve to be understood rather than caricatured. But measurement does not eliminate the need for interpretation, and interpretation does not become better by asking a score to carry more meaning than the evidence allows.
That is a useful place to start: not with the assumption that a score is a person, not with the assumption that measurement is inherently reductive and therefore meaningless, and not with a preferred theory that explains the phenomenon in advance. Start with what we observe, distinguish it from what we infer, and then ask how the pieces fit together.
References
- American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for Educational and Psychological Testing. American Educational Research Association.
- Deary, I. J. (2012). Intelligence. Annual Review of Psychology, 63, 453–482. https://doi.org/10.1146/annurev-psych-120710-100353
- Ritchie, S. J., & Tucker-Drob, E. M. (2018). How much does education improve intelligence? A meta-analysis. Psychological Science, 29(8), 1358–1369. https://doi.org/10.1177/0956797618774253
- Roth, B., Becker, N., Romeyke, S., Schäfer, S., Domnick, F., & Spinath, F. M. (2015). Intelligence and school grades: A meta-analysis. Intelligence, 53, 118–137. https://doi.org/10.1016/j.intell.2015.09.002
- Strenze, T. (2007). Intelligence and socioeconomic success: A meta-analytic review of longitudinal research. Intelligence, 35(5), 401–426. https://doi.org/10.1016/j.intell.2006.09.004
- —



