Intelligence is measured through performance on carefully designed cognitive tasks.
These tasks may involve reasoning, vocabulary, pattern detection, working memory, processing speed, visual-spatial thinking, acquired knowledge, learning, attention, or problem solving. The goal is not to see intelligence directly. It is to infer cognitive ability from patterns of performance.
That makes intelligence measurement powerful. It also makes it easy to misunderstand.
An intelligence test is not intelligence itself. It is a structured method for gathering evidence about how someone performs on selected tasks under defined conditions.
IQ Tests Are Not One Thing
People often talk about “IQ tests” as if they were a single instrument. They are not.
Different assessment families were built for different ages, settings, constructs, practical constraints, and interpretive questions. Some are broad clinical batteries. Some are nonverbal reasoning measures. Some combine cognitive and achievement testing. Some emphasize adaptive functioning. Some are used for screening rather than diagnosis.
The differences matter because the same person may look different depending on what is measured, how it is administered, how language is used, how much speed matters, whether motor output is required, whether the task is familiar, and what the score is being used to decide.
Major Assessment Families
The Wechsler scales are among the best-known intelligence assessment traditions. Different versions are used with different age groups, including children and adults. They sample several cognitive domains and produce composite scores and index scores. Their strength is profile information; their risk is treating a full-scale score as the whole person.
The Stanford-Binet tradition has historically played a central role in intelligence testing and broad cognitive assessment. It can provide information about reasoning, knowledge, quantitative reasoning, visual-spatial processing, and working memory. Its strength is breadth; its limitation is that a global score still needs context.
Raven’s Progressive Matrices are widely known as nonverbal reasoning tests. They are useful when language demands should be reduced, but nonverbal does not mean culture-free or context-free. The task still requires visual processing, attention, abstract pattern familiarity, and response selection.
Woodcock-Johnson batteries are often used in educational and clinical contexts because they can connect cognitive abilities with academic achievement. They are useful when the question is not only ability, but also learning, instruction, opportunity, and achievement.
Kaufman assessment batteries are often used in educational and clinical settings and are associated with attention to cognitive processing. They remind us that different test families are built around different interpretive models.
Leiter and other nonverbal measures may be useful when spoken language, expressive language, hearing, or communication differences could distort performance. They reduce some barriers, but they do not remove all task demands.
Adaptive behavior measures are essential when the question is functioning. They ask how a person manages communication, daily living, social participation, practical tasks, and independence. Intelligence tests estimate cognitive ability; adaptive behavior measures describe functioning in everyday life.
Neuropsychological assessment is not simply another name for IQ testing. It is often concerned with specific cognitive domains and brain-behavior relationships: memory, attention, executive function, language, processing speed, visuospatial function, motor function, emotional regulation, and change after illness or injury.
Reliability, Validity, And Fairness
Reliability asks whether a test produces consistent information.
Validity asks whether the evidence supports the interpretation and use of the score.
Fairness asks whether the assessment process avoids construct-irrelevant barriers, inappropriate comparisons, and unsupported decisions.
A score is not meaningful just because it is numerical. It becomes meaningful when the test, the person, the setting, and the decision fit together.
What A Score Can And Cannot Say
A score can help answer questions about how a person performed compared with a reference group, whether there are notable strengths or weaknesses, whether further assessment is needed, and what supports or accommodations may be considered.
A score cannot decide a person’s worth. It cannot describe every real-world setting. It cannot predict every life outcome. It cannot replace clinical judgement, educational context, cultural understanding, or knowledge of the person’s history.
It also cannot tell us, by itself, whether a person struggled because of limited ability, poor sleep, pain, anxiety, language barriers, inaccessible instructions, motor demands, sensory barriers, low motivation, unfamiliarity, or a hostile environment.
That is why interpretation is the work.
Measurement And OIMI
How Intelligence Works recognizes the importance of established psychometric tools. Existing instruments provide an essential foundation for evaluating cognitive ability, reliability, validity, and interpretation.
At the same time, HIW is exploring whether new open or semi-open measurement instruments could help study intelligence in real-world contexts without compromising test security, validity, or ethical use.
The Open Intelligence Measurement Instruments initiative is best understood as a staged measurement programme, not a completed public test battery. Its purpose is to develop, document, validate, and govern future instruments that may complement existing psychometric tools.
Open documentation does not imply unrestricted administration. Some materials may remain restricted to protect validity, privacy, professional standards, and responsible use.