The short answer
Intelligence is a theoretical construct used to describe differences in capacities such as reasoning, learning, problem-solving and adapting. An intelligence test samples performance on selected tasks. IQ is a standardised score, or family of scores, produced from that performance under one test’s scoring rules and norms.
The three are connected. They are not interchangeable.
Intelligence cannot be observed in the way height can be observed. A person answers questions, solves problems, remembers information or completes timed tasks. A scoring system combines those responses. A norm group supplies a comparison. A theory and a body of validation evidence support a particular interpretation. Each link matters.
Modern IQ scores commonly place performance on a scale centred on 100 for an age-based norm group. One test point is not a unit of intelligence possessed, and 100 is not a natural quantity. It is the centre chosen for a comparative scale. A reported result also contains measurement uncertainty and may change with the instrument, edition, health, language, education, practice, administration and conditions of testing.
Intelligence-test scores do predict some outcomes, including aspects of educational and training performance. That is why the measures persist. They do not establish a person’s creativity, judgement, motivation, morality, occupational satisfaction, competence or human worth. Prediction is not destiny, and a correlation does not identify all of its causes.
Intelligence, an intelligence test and IQ are different things
The APA Dictionary of Psychology describes intelligence broadly through capacities including learning from experience, adapting and handling abstract concepts. That territory is familiar, but there is no single definition that settles every scientific and philosophical dispute. Some accounts emphasise a general reasoning capacity; others distinguish several broad abilities; broader theories also include practical, creative, social or culturally valued forms of adaptation.
That disagreement does not make measurement impossible. It makes the claim more specific.
| Term | What it is | What is directly observed | Typical result |
|---|---|---|---|
| Intelligence | A theoretical construct concerning cognitive functioning | Nothing directly; it is inferred | A construct definition or theory |
| General cognitive ability, or g | A latent factor inferred from correlations among cognitive tasks | The pattern of test correlations | A factor or composite estimate |
| Broad or specific ability | A distinguishable dimension such as fluid reasoning, verbal comprehension or processing speed | Performance on relevant tasks | An index, subtest or profile score |
| Intelligence test | A standardised instrument designed to sample cognitive performance | Responses under stated conditions | Raw, subtest and composite scores |
| IQ | A norm-referenced scaled score produced under a test’s rules | Not directly observed; it is calculated | A score with uncertainty and a comparison group |
Calling intelligence “whatever intelligence tests measure” avoids one problem by creating another. It gives a clear operational rule, but it does not explain why these tasks were selected, why they should be combined, what has been left out or whether a score supports the decision being made.
The Standards for Educational and Psychological Testing provide the better discipline. Evidence supports interpretations of scores for proposed uses. A test is not simply valid in the abstract. Evidence for describing present cognitive functioning is not automatically evidence for school placement, clinical diagnosis, employment selection or a forecast about a particular occupation.
What an intelligence test actually observes
An intelligence test observes performance on tasks chosen to represent parts of a construct. Depending on the instrument and population, those tasks may ask a person to define words, identify relationships, reason with shapes, hold and transform information in mind, reproduce patterns, solve quantitative problems or work quickly through simple visual material.
The label does not make the tasks pure. Vocabulary reflects language and exposure as well as reasoning. A supposedly non-verbal problem still requires instructions, familiarity with diagrams and an understanding of what counts as an answer. A speeded task can be affected by motor, visual, attentional, health or technology-access conditions. Working through a test is an encounter between a person, a task and an institution, not a sample extracted outside culture.
Most comprehensive batteries divide performance into several subtests. Subtests may be grouped into broader indices, and some instruments also report an overall composite. A report can therefore show both the broad component shared across tasks and, where the evidence supports it, a more differentiated profile.
Profiles need restraint. A difference between two subtest scores may reflect a meaningful pattern, measurement error, task familiarity or ordinary variation. The finer the claim, the more evidence is needed that the component scores are reliable enough and distinct enough to interpret. A polished graph does not create precision that the instrument does not possess.
How responses become an IQ score
The number printed in a report is several steps removed from the original performance.
- Administration produces responses. The person answers items under specified timing, instructions, access arrangements and testing conditions.
- Responses produce raw scores. Items may be counted or weighted according to the test’s scoring rules.
- Raw scores are converted. Tables or statistical models compare performance with a standardisation sample, often within an age band.
- Subtest scores are combined. Specified scores may form indices and an overall composite.
- Uncertainty is reported. Reliability evidence and the standard error of measurement support a confidence interval around the observed score.
- An interpretation is made. The result is related to a stated construct and purpose using the test’s technical evidence and the wider context.
The fifth step is easy to lose. A reported IQ of 100 does not mean the person’s true score has been located without error. Repeated equivalent measurement would not yield exactly the same result every time. A confidence interval makes that uncertainty visible, although it does not capture every possible source of error, bias or changed circumstance.
Two tests can both report IQ while sampling somewhat different tasks, using different norm groups, weights, ceilings and editions. Their scores may correlate strongly without being interchangeable. The appropriate question is not merely “What is the IQ?” but “Which score, from which instrument and edition, under which conditions, compared with whom, and for what interpretation?”
Ratio IQ became deviation IQ
The name carries an older calculation into a newer scoring system.
Early mental-age approaches described a child’s test performance by comparison with the typical performance associated with an age. The historical ratio IQ divided mental age by chronological age and multiplied the result by 100. A child performing at the level assigned to their chronological age therefore received 100. The approach did not translate well across later development, and the meaning of equal mental-age differences was not constant.
Most modern intelligence batteries use deviation IQ instead. The APA definition describes a score based on how far a person’s performance lies from the mean of a reference group, commonly people of the same age. The transformed scale is often centred on 100 with a standard deviation of 15, although a test manual must confirm the actual convention.
| Reported idea | What it means | What it does not mean |
|---|---|---|
| IQ 100 | The centre of this test’s normed scale for the relevant comparison group | A natural amount of intelligence |
| Standard deviation | A scale for describing distance from the norm-group mean | A guarantee that every population has the same distribution |
| Percentile | The proportion of the norm group scoring at or below a result | The percentage of items answered correctly |
| Confidence interval | A range expressing score imprecision under the measurement model | Every uncertainty about the person or use |
| Age norm | A comparison with people in a defined age band | Evidence that raw performance has not changed over time |
This is why a percentile and a percentage correct answer different questions. A person can answer more items correctly at a later age while remaining at a similar percentile if the comparison group has also improved. Conversely, an age-normed score may remain steady while the cognitive processes beneath it change.
Norms can become stale. Education, health, test familiarity and population composition change, and average performance on many intelligence tests has shifted across cohorts. New editions may alter content, floors, ceilings and scoring as well as renorm the scale. A result must therefore be attached to the test and edition that produced it.
One intelligence or many abilities?
People who perform well on one cognitive task tend, on average, to perform well on others. This positive manifold is one of the most durable findings in psychometric research. In 1904, Charles Spearman proposed that the shared pattern could be represented by a general factor, later called g, alongside task-specific variance.
The factor is inferred from correlations. It is not an organ, a substance or a little executive inside the person. Different batteries and models can estimate a general factor from different sets of tasks. Treating g as a useful statistical representation does not settle every claim about the causes, biological basis or full meaning of intelligence.
Louis Thurstone challenged an account dominated by one factor and described seven primary mental abilities, including verbal comprehension, word fluency, number, spatial ability, associative memory, perceptual speed and reasoning. Later analyses found that those abilities were themselves correlated. The debate therefore moved from a simple choice between one ability and many towards hierarchical models in which broad commonality and meaningful differentiation can coexist.
Raymond Cattell and John Horn distinguished fluid abilities, associated with reasoning through novel problems, from crystallised abilities, associated with acquired knowledge and skills. Horn expanded the model beyond two abilities. John Carroll’s reanalysis of hundreds of factor-analytic datasets described a hierarchy with narrow abilities, broader abilities and a general factor. The resulting Cattell–Horn–Carroll, or CHC, family of models has strongly influenced contemporary test design and interpretation.
The hierarchy matters because an overall composite and a profile answer different questions. A broad score may summarise what diverse tasks share and often predicts broad outcomes steadily. An index may reveal a relevant strength or access need that the total hides. Neither is automatically the “real” intelligence score. The intended interpretation and the instrument’s evidence decide the useful level.
Other theories broaden intelligence further. Robert Sternberg has emphasised analytical, creative and practical aspects of intelligent behaviour. Howard Gardner’s multiple-intelligences framework has been influential in education. These accounts raise valuable questions about adaptation, culture and forms of accomplishment that conventional batteries may neglect. They do not have the same factor-analytic evidentiary foundation as psychometric hierarchical models, and listing all theories side by side should not imply that they are alternative score reports of equal empirical status.
Where intelligence testing came from
Intelligence testing has no innocent single origin and no single purpose.
Francis Galton’s nineteenth-century programme connected individual differences, measurement and heredity with eugenic thought. His sensory and reaction-time measures are not modern IQ batteries, but the project helped establish a political as well as scientific ambition: to rank people and explain social position through measured difference.
Alfred Binet and Théodore Simon were answering a different institutional question. Their 1905 scale arose in the setting of French mass education and the identification of children who might need different teaching. The National Research Council’s history describes an educational and diagnostic purpose. The tasks and mental-age idea were later adapted, translated and renormed. William Stern proposed the ratio formulation, and Lewis Terman’s Stanford–Binet gave the approach a prominent American form.
During the First World War, Army Alpha and Beta turned mental testing into mass administration. Alpha was designed for literate English-speaking recruits; Beta reduced some literacy and language demands. More than 1.7 million recruits were tested by the end of 1918. The batteries differed in content and administration, so converting both into classifications did not make their scores identical. Yet the programme made group testing look like a practical technology for allocating people at scale.
The movement of scores into immigration arguments, racial hierarchy, disability institutionalisation, educational tracking and employment cannot be treated as an unfortunate footnote. Test developers and users did not all hold the same theory, but mental testing and eugenic institutions overlapped. Carl Brigham used Army data to make racial and national claims, then later rejected those interpretations after recognising that language, culture, education and test mixture confounded the result. The reversal is instructive: a technically produced difference does not supply its own causal explanation.
David Wechsler’s later batteries helped establish the modern comprehensive model: several subtests, age-based norms, differentiated indices and an overall score. Subsequent editions and other batteries changed the constructs, content and populations represented. Professional standards, civil-rights law and disability rights also changed what evidence and access should be required.
The history is not an argument that every intelligence test is invalid. It is evidence that instruments travel. A score made for teaching can become a placement gate; a wartime classification method can enter employment; a comparative result can be turned into an inherited identity. Every new use needs its own validity, fairness and governance argument.
The companion page on aptitude follows the branch from general mental testing into multiple-aptitude, training and vocational systems. The histories overlap, but intelligence and aptitude are not the same term.
What intelligence-test scores predict
Intelligence-test scores are associated with educational achievement, aspects of training performance and some occupational outcomes. Reviews including the APA task force report by Neisser and colleagues and the later synthesis by Nisbett and colleagues treat that predictive evidence as real while leaving major questions about causes, development and interpretation open.
The outcome must be named. School marks, years of education, training completion, job knowledge, supervisor ratings, income and occupational status are not one variable. Each is shaped by opportunity, institutions, selection, health, motivation, discrimination and prior learning as well as cognitive performance.
Employment evidence requires particular care. Cognitive measures can contribute to prediction of training and job performance, especially when the measured demands match the work. However, older headline validity estimates were often corrected aggressively for statistical artefacts. A large reanalysis by Sackett and colleagues found that systematic overcorrection for restriction of range had inflated widely cited meta-analytic estimates for personnel selection methods. “IQ predicts job performance” is therefore too blunt. Which measure, outcome, job family, applicant population, correction and decision rule all matter.
Prediction also works at the level of probabilities, not occupational permission. A relationship observed across a group does not tell an employer what one candidate will do, nor does it identify the support or opportunity that candidate will receive. A modest improvement in average prediction can coexist with substantial overlap, error and unfair consequences at a cut score.
What IQ does not establish
An IQ score does not establish the whole of someone’s mind or future.
It does not directly measure moral character, wisdom, practical judgement, curiosity, courage, care, reliability, creativity, motivation or meaning. It does not show whether a person wants an occupation, can access its training, will be welcomed by an institution or can perform responsibly in its real social and material conditions.
It is not competence. Competence integrates knowledge, skill, judgement and responsibility in context. A person may reason well on unfamiliar tasks and lack the knowledge or practice required for safe work. Another may perform expertly in a domain through rich knowledge, strategies, tools and collaboration that a decontextualised battery does not sample. The capability page develops these distinctions.
IQ is also not a measure of human worth. That statement is not a polite addition to the science. It follows from the measurement claim. A comparative score on selected tasks cannot logically become a moral ranking of persons.
Heritability without genetic destiny
Few terms in this field travel as badly as heritability.
Heritability is a population statistic. It estimates how much of the observed variation in a trait, in a particular population and range of environments, is statistically associated with genetic variation under a stated method. It is not the percentage of one person’s intelligence “caused by genes”. It does not divide an individual into genetic and environmental parts.
A high heritability estimate does not mean a trait cannot change. The ingredients of a population can be highly heritable while the population mean changes through education, nutrition, disease, technology or other environmental conditions. Dickens and Flynn modelled how genetic and environmental effects can amplify one another, helping explain why substantial heritability can coexist with large environmental shifts. Sauce and Matzel review a wider set of gene–environment processes behind the same apparent paradox.
Genes and environments are not independent teams adding fixed shares. People partly select, evoke and create environments in ways related to their characteristics. Environments can change how genetic differences are expressed. Education, family resources, toxins, stress, health and social responses can shape development, while prior performance can shape the opportunities offered next. This is gene–environment correlation and interaction, not an escape from biology or environment but a reason not to treat either as a sealed cause.
Estimates also depend on age, population, environmental range and method. Twin and adoption designs infer genetic influence from patterns of relatedness under assumptions about environments and mating. Genomic methods estimate variation captured by measured DNA differences and do not necessarily reproduce twin-study estimates. Reviews such as Plomin and Deary report developmental changes in estimates, but those figures should not be detached from the populations and methods that produced them.
Most importantly, heritability within groups does not explain an observed average difference between groups. A trait can be heritable within two populations while their mean difference is environmental. Claims about racial, national, caste or class differences also face harder prior questions: how groups were constructed, whether samples are comparable, whether the test works equivalently, which environments and histories differ, and whether opportunity and discrimination have been measured. A within-group statistic cannot answer those questions by itself.
For this reason, this page does not publish a single context-free percentage. Any quantitative claim needs the age, population, design, uncertainty and meaning of the estimate beside it. The scientific task is to explain development, not to convert a population statistic into a personal sentence.
Education and environment can change scores
Intelligence-test performance is neither infinitely malleable nor fixed outside experience.
A meta-analysis by Ritchie and Tucker-Drob examined quasi-experimental evidence and found consistent evidence that additional education improves performance on intelligence tests. The size of an effect varies by design and outcome, and the finding does not mean schooling affects every ability equally. It does show why learned experience cannot be removed from intelligence by definition.
Average scores on many tests have also changed substantially across generations, often called the Flynn effect. Patterns differ across countries, periods and subtests. The broad lesson is safer than one universal explanation: cohort environments can shift test performance, so norms require maintenance and a scale centred on 100 is always tied to a reference population and time.
Health, sleep, nutrition, stress, sensory access, language, practice and familiarity can affect observed performance. Some influences change the capacity being sampled; others create construct-irrelevant barriers to showing it; many do both. A changed score therefore has several possible explanations. It cannot automatically be dismissed as measurement error, just as a stable ranking cannot automatically be treated as biological inevitability.
Culture-free is not possible
Deleting words from a test does not delete culture.
Language affects instructions and reasoning. Schooling affects familiarity with symbols, abstraction, test-taking and speed. Diagrams rely on conventions. Timed interfaces reward particular forms of practice. Norm groups define whose performance becomes the centre of the scale. Rapport, motivation, disability access and the consequences attached to a test can also affect what happens in the room.
A non-verbal or reduced-language measure may be entirely appropriate when language is not the target. It can reduce one source of irrelevant difficulty. It should be described as reducing a demand, not as revealing intelligence without culture.
Fairness questions also need to be separated:
- Differential item functioning asks whether people from different groups, matched on the measured construct, have different probabilities of answering an item in a particular way.
- Measurement invariance asks whether a construct and its measurement relationships work comparably across groups.
- Differential prediction asks whether a score relates to an outcome differently across groups.
- Accommodation changes access conditions to remove a barrier that is not part of the intended construct.
- Unfair use can occur even when a score is technically reliable, if the decision exceeds the evidence, ignores alternatives or distributes harm without adequate justification and review.
No one statistic answers all five questions. An item can show a group difference without being biased. A battery can lack obvious item bias while its norm group or use remains unsuitable. An accommodation can improve access without invalidating the result when the altered demand was irrelevant to the intended inference.
The National Research Council’s examination of minority representation in special and gifted education shows why test properties, referral systems, instruction and institutional decisions must be considered together. The two easy stories are both wrong: standardisation does not remove culture and opportunity, but every observed difference is not automatically an artefact of test bias either.
Intelligence across childhood, adulthood and later life
The assessment question changes across the life course.
In early childhood, performance is developing rapidly and can be especially sensitive to language, rapport, health and opportunity. Assessment may help identify support, but a score should not become a vocational identity or an irreversible account of potential. Repeated evidence and the child’s response to teaching matter.
At school age, age norms make relative comparison possible, while achievement and intelligence measures remain entangled through learning. A low result can identify a need for further assessment or support; it does not explain the cause on its own. Placement, disability eligibility and gifted identification carry consequences, so evidence from teachers, learning history, access needs and other measures should accompany a test.
In adulthood, scores are often more stable than in early childhood, but stability is not immutability. Health, education, language, major life events and the match between person and instrument still matter. Historical childhood results should not substitute for current evidence when the current question can be assessed directly.
In later life, different broad abilities can follow different patterns. Acquired knowledge may remain relatively robust while some speeded or novel-problem tasks become more difficult on average. Individual variation is large. Age norms describe relative standing; they must not turn age itself into a diagnosis of incapacity. When the question concerns change within a person, raw performance, prior baselines, health and functional context may matter more than comparison with an age group.
Across the life course, an intelligence score keeps the age group, purpose, access conditions and wider developmental evidence that gave it meaning. Strip those away and a comparative result can easily be mistaken for a timeless description of the person.
IQ, ability, aptitude and achievement compared
These terms overlap because the same task can support several bounded interpretations. They remain distinct because they ask different questions.
| Term | Core question | Typical result |
|---|---|---|
| Intelligence | How should broad cognitive functioning be conceptualised? | A theory or construct model |
| Intelligence test | What cognitive performance was sampled under standardised conditions? | Subtest and composite scores |
| IQ | Where does a composite sit on this test’s normed scale? | A scaled score with uncertainty |
| Ability | What present capacity appears relevant to performance? | A bounded capability inference |
| Aptitude | What does the evidence predict about learning or developing performance? | A future-oriented inference |
| Achievement | What has already been learned? | Current knowledge or skill evidence |
An intelligence test can be used as a broad aptitude measure when evidence supports prediction of a defined learning outcome. A specific cognitive subtest can contribute to a domain aptitude profile. Achievement can contribute to all of them because acquired knowledge and strategies shape test performance.
The same score does not make all interpretations valid. Evidence that a composite describes relative cognitive performance does not automatically show that it predicts success in an electrical apprenticeship. Evidence that it predicts average training grades does not show that a cut score fairly identifies who should be admitted. Description, prediction and decision are separate claims.
What a high-stakes score should make visible
Intelligence tests are used in clinical and educational assessment, disability and gifted identification, research, military systems and sometimes employment. The more a result allocates access, support or exclusion, the stronger the evidence and safeguards should be.
A person affected by a high-stakes interpretation should be able to find:
- the instrument, edition and population for which it was designed;
- the construct and outcome being inferred;
- the norm group and date of standardisation;
- the score precision and confidence interval;
- relevant language, disability, health, educational and administration conditions;
- evidence for validity and fairness for this use;
- the other information used in the decision; and
- how to obtain an explanation, correct errors, request appropriate reassessment or challenge the decision.
Multiple sources of evidence are not automatically fair merely because there are several. They should add relevant information rather than repeat the same barrier. A work sample, learning trial, prior performance and structured interview may answer a practical question more directly than a broad cognitive score.
The right to review matters because measurement and power meet in the decision. A confidence interval acknowledges statistical uncertainty; an appeal process acknowledges institutional fallibility.
Guidebeam does not conduct IQ testing
Guidebeam does not conduct IQ testing. Its interest, aptitude, capability and career-exploration results are not IQ scores and should never be presented as equivalent to a formal intelligence assessment.
In Australia, the Australian Psychological Society explains that psychological testing is usually performed by a psychologist. A non-psychologist may administer and score a psychological test only with appropriate training and under the supervision of a psychologist, who must be directly involved in interpreting and reporting the findings. Formal intelligence assessment is specialist psychological work, not a routine feature of a career-guidance platform.
Guidebeam’s responsibility here is to explain the distinction and avoid borrowing the authority of IQ. It may discuss intelligence research or use appropriately validated evidence in career exploration without administering an IQ test, reporting an IQ score or offering a psychological assessment. A person seeking a formal intelligence assessment should be directed to a suitably qualified registered psychologist with expertise in the relevant test and population.
Guidebeam does not use IQ scores to determine career fit. Formal intelligence assessment can serve legitimate specialist purposes, but a single normed cognitive score is poorly suited to deciding what a person can become. In career guidance, it can be mistaken for a fixed ceiling and prematurely narrow exploration.
Guidebeam instead brings together interests, values, motivations, preferences, lived experience, current capabilities, opportunity and appropriately bounded aptitude evidence. These signals help people investigate possibilities; they do not assign anyone to a career or define the limits of their development.
The current Guidebeam product is changing, so its dimensions, item counts, occupation coverage and validation status belong in maintained, dated product documentation rather than this evergreen article. The durable principle is categorical: Guidebeam supports career exploration; it does not turn an exploratory result into an IQ label or an occupational destiny.
Why an IQ result remains bounded
An IQ score gathers several layers into one number. There is a theory of intelligence, a particular instrument and edition, a selection of tasks, a norm group collected at a particular time, and a method for combining performance into composites, indexes or percentiles. A confidence interval makes some of the uncertainty visible; language, education, health, disability access, practice and administration add more.
The result can support a bounded interpretation, but it cannot turn a group average into an explanation of an individual. Heritability does not identify the cause of one person's score, make the result immutable or explain a difference between groups. Nor does a general association with later outcomes settle a particular educational, clinical or occupational decision.
An IQ result is therefore evidence from this instrument and comparison group, under these conditions. It is not the person, and it should carry no more power than that evidence justifies.
Notes
This article uses IQ for a norm-referenced score produced by a recognised intelligence-test scoring system, while acknowledging that instruments report several composites and that historical ratio IQ differs from modern deviation IQ.
The article treats hierarchical psychometric models as the main evidentiary framework for conventional intelligence batteries. Broader theories are included because they influence public and educational understandings, not because all theories have equivalent measurement support.
Heritability estimates are deliberately not reduced to one headline percentage.
The historical terminology in some source titles is outdated and offensive. Titles are reproduced only where needed to identify the publication accurately; they do not state Guidebeam’s language or position.
Sources and further reading
- American Psychological Association. “Intelligence.”, “Intelligence Test.”, “IQ.” and “Deviation IQ.” APA Dictionary of Psychology.
- AERA, APA and NCME. Standards for Educational and Psychological Testing. 2014.
- National Research Council. “The Role of Intellectual Assessment.” In Mental Retardation: Determining Eligibility for Social Security Benefits. National Academies Press, 2002.
- National Research Council. Minority Students in Special and Gifted Education. National Academies Press, 2002.
- National Research Council. Cognitive Aging: Progress in Understanding and Opportunities for Action. National Academies Press, 2015.
- Dorans, Neil J., and Linda L. Cook, eds. Fairness in Educational Assessment and Measurement. Routledge, 2016.
- Spearman, Charles. “General Intelligence, Objectively Determined and Measured.” American Journal of Psychology 15, no. 2 (1904): 201–292.
- Thurstone, L. L. Primary Mental Abilities. University of Chicago Press, 1938.
- Carroll, John B. Human Cognitive Abilities: A Survey of Factor-Analytic Studies. Cambridge University Press, 1993.
- Neisser, Ulric, et al. “Intelligence: Knowns and Unknowns.” American Psychologist 51, no. 2 (1996): 77–101.
- Nisbett, Richard E., et al. “Intelligence: New Findings and Theoretical Developments.” American Psychologist 67, no. 2 (2012): 130–159.
- Dickens, William T., and James R. Flynn. “Heritability Estimates Versus Large Environmental Effects: The IQ Paradox Resolved.” Psychological Review 108, no. 2 (2001): 346–369.
- Sauce, Bruno, and Louis D. Matzel. “The Paradox of Intelligence: Heritability and Malleability Coexist in Hidden Gene–Environment Interplay.” Psychological Bulletin 144, no. 1 (2018): 26–47.
- Plomin, Robert, and Ian J. Deary. “Genetics and Intelligence Differences: Five Special Findings.” Molecular Psychiatry 20 (2015): 98–108.
- Plomin, Robert, and Sophie von Stumm. “The New Genetics of Intelligence.” Nature Reviews Genetics 19 (2018): 148–159.
- Ritchie, Stuart J., and Elliot M. Tucker-Drob. “How Much Does Education Improve Intelligence? A Meta-Analysis.” Psychological Science 29, no. 8 (2018): 1358–1369.
- Sackett, Paul R., Charlene Zhang, Christopher M. Berry and Filip Lievens. “Revisiting Meta-Analytic Estimates of Validity in Personnel Selection.” Journal of Applied Psychology 107, no. 11 (2022): 2040–2068.
- Strenze, Tarmo. “Intelligence and Socioeconomic Success: A Meta-Analytic Review of Longitudinal Research.” Intelligence 35, no. 5 (2007): 401–426.
- National Academies of Sciences, Engineering, and Medicine. Psychological Testing in the Service of Disability Determination. National Academies Press, 2015.
- National Academies of Sciences, Engineering, and Medicine. Neurodevelopmental Disorders: Assessment and Related Tools. National Academies Press, 2022.
- National Research Council. High Stakes: Testing for Tracking, Promotion, and Graduation. National Academies Press, 1999.
- Australian Psychological Society. Psychological Testing.

