The committee that could not write one checklist
Between 1950 and 1954, an American Psychological Association committee tried to specify what should be investigated before a psychological test was published. The task sounded administrative. A publisher would submit evidence, the evidence would be checked, and the test would emerge with its scientific credentials in order.
The committee found that the question kept changing underneath it.
A reading examination could be compared with a curriculum. A selection test could be compared with later job performance. A new scale of anxiety or authoritarianism claimed to measure something that had no single external yardstick. Projective techniques posed the difficulty in an especially exposed form: if a psychologist interpreted an inkblot as evidence of an unobserved quality, what kind of study could justify that interpretation?
Lee Cronbach and Paul Meehl described the committee's predicament in their 1955 paper, “Construct Validity in Psychological Tests”. Existing ideas about validation did not fit into one coherent checklist. The committee had to distinguish claims that required different kinds of research. Cronbach and Meehl gave sustained form to one of them: construct validity, the investigation of what attribute or quality could account for variation in performance on a test.
Their answer replaced a tidy inspection with an ongoing scientific problem. A construct such as mechanical comprehension, anxiety, conscientiousness or vocational interest belongs to a network of ideas. The theory says how it should relate to observations, other constructs, behaviour and change. Evidence can strengthen parts of that network, expose a rival explanation or force revision of the construct itself. Validation develops the measure and the theory together.
This history matters because the old administrative instinct survives. People still ask whether a test “is valid”, as though validity were a seal fixed permanently to an instrument. The better question is more exacting:
What interpretation of these scores does the accumulated evidence support, for this population, setting and use?
That is the working definition of validity in this library. It follows the current joint Standards for Educational and Psychological Testing, which defines validity through the degree to which evidence and theory support a specified interpretation of scores for intended uses.
Every part of that sentence carries weight. Evidence accumulates rather than arriving in one decisive statistic. Theory says why the evidence should have the observed pattern. The interpretation states what the score is taken to mean. The use states what someone proposes to do with it.
The claim comes before the evidence
Imagine an interest inventory that asks a person how much they would like activities such as repairing equipment, analysing data, teaching a child or organising an event. The responses are scored into several interest domains.
Four claims might follow:
- the score summarises the person's answers to these items;
- the score describes a broader pattern of vocational interests;
- the pattern is likely to matter in environments beyond the questionnaire;
- the result should influence a course, occupation or hiring decision.
Those claims travel different distances.
The first depends on item coding and scoring accuracy. The second requires a defensible construct and evidence that the items represent it. The third asks whether the interpretation generalises beyond the sampled questions and assessment setting. The fourth introduces a decision rule, alternatives, consequences and the possibility that the setting denies some people the opportunity the score appears to describe.
An accurate total does not settle the second claim. A coherent interest scale does not establish the third. Evidence that interests relate to later choices does not automatically license the fourth. Each step inherits the assumptions below it and adds new ones.
This is why validity cannot be read from a score report's decimal places. The psychometrics page follows the full passage from construct to item, response, scoring, scale, interpretation and use. Reliability asks whether a score holds across the repetitions that the interpretation treats as interchangeable. Validity examines whether the interpretation and use survive their whole chain of reasoning. A result may be highly consistent and consistently misinterpreted.
The proposed use helps define the chain. The same conscientiousness score might contribute to a research study, a counselling conversation or an employment decision. The construct label remains familiar while the evidential burden changes. Research may concern average relationships in a defined sample. Counselling may use the score as one prompt that a person can question. Selection applies the result to an opportunity and requires evidence about the job, criterion, population, decision rule and consequences.
“Low stakes” does not mean evidence-free. It means the use is easier to challenge, combine with other information and correct. A consequential gate raises the cost of a failed inference.
Kane's bridge from performance to conclusion
Michael Kane's argument-based approach makes the sequence explicit. It begins by setting out the claims needed to move from a person's performance to the proposed conclusion, rather than collecting every familiar psychometric statistic first. Validation then tests those claims and the assumptions connecting them.
A common chain contains four spans:
| Inference | Movement | Questions the evidence must answer |
|---|---|---|
| Scoring | Observed response or performance → reported score | Was the response captured, coded and combined appropriately? Did raters apply the intended criteria? |
| Generalisation | Reported score → performance across the intended universe of observations | Would comparable items, tasks, occasions or raters support the same score interpretation? |
| Extrapolation | Assessment performance → behaviour or capability in the target setting | Does performance in the assessment represent what happens in study, work or life? What relevant conditions differ? |
| Decision or use | Interpreted score → action | Does the rule improve the intended decision? What alternatives, errors, values and consequences enter? |
Validity · Claim to consequence
What must this score claim survive?
The same broad interest pattern can support a counselling conversation long before it can justify a recommendation or exclusion. Open each inference to inspect its claim, evidence and rival explanation.
1 · InferenceObserved responseShow detailsHide details
Claim: The recorded choice represents what the person selected.
Suitable evidence: Administration records, accessibility checks and response-process evidence.
Rival explanation: The response reflects misunderstanding, interface friction or a temporary condition.
2 · InferenceScoringShow detailsHide details
Claim: The coding and combination rules produce the intended interest result.
Suitable evidence: Documented scoring rules, quality controls and studies of rater or algorithm behaviour.
Rival explanation: A key, weight, missing-data rule or irrelevant response feature drives the score.
3 · InferenceGeneralisationShow detailsHide details
Claim: Comparable observations would support the same score interpretation.
Suitable evidence: Reliability and precision studies across the relevant items, occasions, forms or raters.
Rival explanation: The result depends on the particular item sample, occasion, form or judge.
4 · InferenceConstruct interpretationShow detailsHide details
Claim: The pattern supports the proposed account of vocational interest.
Suitable evidence: Content, response-process, internal-structure and external-relation evidence.
Rival explanation: Reading load, acquiescence, social desirability or method effects explain the pattern.
5 · InferenceReal-world extrapolationShow detailsHide details
Claim: The interpreted pattern remains relevant in study, work or life settings.
Suitable evidence: Relationships with representative behaviour and outcomes in relevant populations and contexts.
Rival explanation: Opportunity, prior exposure, support or setting differences break the relationship.
6 · InferenceDecision or useShow detailsHide details
Claim: Acting on the interpretation improves the proposed decision.
Suitable evidence: Decision studies, alternatives, error costs, fairness, accessibility, monitoring and recourse.
Rival explanation: The rule narrows options, reproduces barriers or performs no better than a safer alternative.
Shorter chain
Counselling prompt
Use the pattern to open questions, compare experiences and keep uncertainty visible.
Longer chain
Course recommendation
Add evidence about course demands, opportunity, alternatives and the learner's circumstances.
Longest and most consequential chain
Exclusionary selection
Require stronger validity, fairness, accessibility, governance, human review and recourse evidence.
Moving farther across the bridge adds claims and evidence. It never produces one validity coefficient or a universal pass–fail verdict.
Six disclosure panels form an inference bridge from observed response through scoring, generalisation, construct interpretation and real-world extrapolation to a decision or use. Each panel states the claim, suitable evidence and a plausible rival explanation. Three illustrative uses compare a counselling prompt, course recommendation and exclusionary selection. The evidence burden increases as the use becomes more distant and consequential. No validity score or pass-fail verdict is shown.
The bridge image is useful because one strong span cannot carry a missing one. Perfectly accurate scoring does not show that a work simulation represents work outside the simulation. A strong correlation with supervisor ratings does not repair a scoring rule that rewards irrelevant fluency. A carefully developed construct does not justify a decision rule whose costs and alternatives were never examined.
Kane's framework also prevents validation from becoming a warehouse of statistics. The relevant evidence is chosen by the argument. A work sample and a self-report inventory make different claims about responses. A classroom quiz and a professional licensing examination may share content while supporting different decisions. Validation asks first where the inference could fail, then studies those points.
There remains genuine disagreement about what validity itself should mean. Denny Borsboom, Gideon Mellenbergh and Jaap van Heerden argued in “The Concept of Validity” that a test measures an attribute when the attribute exists and variation in it causally produces variation in measurement outcomes. Their proposal challenges accounts that gather every issue of interpretation and use under one broad heading. It also draws a sharp line between measuring an attribute and predicting another outcome.
The dispute is productive. An argument can be coherent while relying on an implausible idea of the thing being measured. A correlation can predict usefully without showing that the predictor measures the causal attribute its label suggests. Validity work needs empirical relationships, and it needs an account of why those relationships should exist.
The old boxes become sources of evidence
Textbooks once taught three kinds of validity: content, criterion and construct. The categories supplied useful questions, then encouraged a misleading practice. A manual could present one correlation as “criterion validity”, an expert panel as “content validity” and a factor analysis as “construct validity”, as if three stamps had been collected.
Samuel Messick's unified account brought these questions back into one argument about score meaning and use. His 1993 ETS report described content, criteria and consequences as interrelated evidence within construct validation. Current standards organise the work as sources of validity evidence. They recur often, but they are not mandatory boxes for every study.
Evidence based on content
Content evidence asks whether the observations adequately represent the defined domain. A licensing examination needs a principled account of professional practice. A mathematics test needs specifications for the knowledge and reasoning it claims to sample. An interest inventory needs an account of the activities, preferences and domain boundaries represented by its items.
Experts can judge relevance and coverage. Their authority does not compensate for an undefined domain. A documented record of who judged what, against which specifications, with which disagreements and empirical checks makes their contribution open to scrutiny. Content evidence grows stronger when the domain, item sampling and intended interpretation fit visibly together.
Evidence based on response processes
An item can look appropriate while eliciting an unintended process. A reasoning question may be solved through a test-taking trick. A self-report prompt may be read as “what am I good at?” when the intended question is “what do I enjoy?” A rater may reward polished handwriting or accent while believing they are judging the quality of an argument.
Response-process research investigates what test takers, observers and raters actually do. Think-aloud studies, cognitive interviews, response times, error patterns and rater studies can reveal whether the observed performance arose in the way the interpretation assumes. José-Luis Padilla and Isabel Benítez's review shows why this evidence deserves its own place: the path from prompt to response is part of the measurement, not an invisible prelude to it.
Evidence based on internal structure
The relationships among items and components should fit the structure claimed for the score. If a report presents six separate domains, the data should support interpreting those domains. If one total score is reported, its combination needs a rationale. Factor analysis, item relationships and differential item functioning can test parts of that structure.
A tidy factor solution cannot name the construct by itself. Many plausible item sets can form statistical patterns. High internal consistency can come from narrow repetition. The internal structure must be read with content, response processes and relationships beyond the test.
Evidence based on relations to other variables
The theory should predict a pattern. Measures intended to concern the same construct may converge. Measures of distinguishable constructs should remain distinguishable. Scores may change after an intervention that should affect the construct and resist changes that should not. A predictor may relate to later performance.
Donald Campbell and Donald Fiske turned this into a memorable design in 1959. Their multitrait–multimethod matrix measured several traits through several methods. Convergence across methods supported a shared trait interpretation. Weak discrimination warned that supposedly different traits blurred together. Correlations among measures sharing a method exposed another risk: two questionnaires may agree partly because both are questionnaires.
The matrix is more than a historical technique. It makes three questions visible: what else should agree, what should remain distinct and whether the method itself could manufacture the pattern.
Evidence from consequences
The use of a score changes classrooms, organisations and lives. It can direct teaching, grant a licence, deny a job, prompt a useful conversation or narrow the options someone considers possible. Consequences therefore belong in the validity programme, although their role needs care.
If one group is disproportionately excluded, that outcome does not by itself identify the cause. It makes rival explanations urgent. The procedure may include irrelevant language demands, omit an important kind of performance, operate differently across groups, predict a deficient criterion or reflect a real difference relevant to the specified construct. Each possibility calls for evidence.
Some consequences arise from policy choices even when the score interpretation is sound. A valid estimate of current knowledge does not decide how a school should allocate scarce support. Evaluation of the interpretation and evaluation of the policy remain connected and distinguishable. The dedicated page on fairness, bias and consequences will take that argument further.
What the procedure leaves out, and what slips in
Two rival explanations organise much of validity work.
Construct underrepresentation occurs when the assessment misses important parts of what the interpretation claims to cover. A writing test composed only of multiple-choice grammar questions leaves out the production and organisation of extended prose. An “interest” measure that samples only professional office activities may represent a narrow world of work. A work sample built from routine cases may say little about judgement under novelty.
Construct-irrelevant variance occurs when performance is influenced by something outside the intended construct. A mathematics item with unnecessary reading complexity can turn language proficiency into part of the score. An interview may reward familiarity with a cultural script. A digital task may depend on device latency or prior interface experience. A self-report scale may mix the intended preference with pressure to present oneself favourably.
The two threats pull in opposite directions and can appear together. Simplifying all language may improve access when reading is incidental. The same change can remove essential difficulty from a task where reading is part of competent performance. Extended time may reduce an irrelevant speed barrier, or it may change the construct when speed is explicitly required. The validity argument has to say which capability the use actually demands.
This is where fairness enters the substance of measurement. Accessibility cannot be handled by making every task easier. The goal is access to the intended construct while controlling irrelevant barriers. The target itself also deserves scrutiny. An organisation can define a construct so narrowly that it reproduces yesterday's job and excludes people capable of learning tomorrow's one.
Edward Strong's 1927 vocational inventory makes the historical point tangible. Its directions separated vocational interests from intelligence and school learning. Responses were compared with those of successful men in named professions. The method could reveal resemblance to an occupational group's reported likes and dislikes. It also inherited the boundaries of the group that supplied the criterion.
A high resemblance score did not demonstrate that a person would gain entry, perform well, find the work meaningful or remain satisfied. Women and people excluded from the sampled professions could have interests that the reference system had little chance to recognise. The criterion described people who had passed through existing gates.
Strong's work was an important move toward empirical vocational assessment. Its limitation teaches a general lesson: a criterion is an observed and historically situated measure, never the outcome itself in pure form.
A correlation inherits both of its ends
Predictive validation is often narrated as the cleanest case. Administer a test, wait for an outcome and calculate the relationship. The resulting coefficient looks like a direct answer.
It is a relationship between two measured variables.
The predictor may be unreliable, narrow or contaminated. The criterion may be equally troublesome. Supervisor ratings can reflect opportunity, role assignment, visibility, rater severity and organisational politics alongside performance. Sales figures depend on territory and market conditions. Training completion may reward persistence, prior knowledge, support or simple attendance. Occupational membership confounds interest with access and staying power.
Personnel psychologists call the two criterion problems deficiency and contamination. A deficient criterion omits relevant performance. A contaminated criterion includes irrelevant influences. A strong correlation with a poor criterion can validate the wrong interpretation. A modest correlation can understate a useful relationship when either end is measured badly.
Selection creates another complication. Once an organisation hires mainly high scorers, the observed employees occupy a narrower predictor range than the applicant pool. Range restriction can depress the observed relationship with later performance. Statistical corrections attempt to estimate the relationship in the wider population, but their result depends on how selection occurred, which variables were restricted and how reliably each was measured.
In 1977, Frank Schmidt and John Hunter published a general model of validity generalisation. They challenged the belief that employment-test relationships were necessarily unique to each local situation. Sampling error, measurement error and range restriction could explain much of the apparent variation. Meta-analysis might support transport to comparable settings when the relevant conditions held.
The programme changed personnel psychology. It did not end the argument. In 2022, Paul Sackett, Charlene Zhang, Christopher Berry and Filip Lievens revisited common range-restriction corrections. They concluded that several widely used methods had systematically overcorrected, inflating many meta-analytic validity estimates. Their revised estimates were commonly lower by about .10 to .20 correlation points.
The researchers did not conclude that selection procedures were useless. Much of the ranking remained, structured interviews emerged strongly, and the predictors retained practical value. The episode shows validation behaving as science should. Assumptions that had become routine were reopened; quantitative claims narrowed; the field kept the parts that survived.
A reported correlation therefore needs an address: the predictor, criterion, sample, range, interval, conditions and proposed decision. “Predictive validity = .42” is the beginning of an inquiry, not its conclusion.
Evidence has an address
Validity evidence collected with one population and procedure may inform another use. Transport must be argued.
Consider a questionnaire developed with English-speaking university students, delivered on desktop computers and used for group research. It later appears in a translated mobile report for a mid-career adult. The item wording may have shifted. The reference population changed. Mobile presentation can alter navigation and missingness. The individual interpretation travels beyond the original group-level use. The occupational setting may have changed since the validation study.
None of these differences proves failure. Each identifies a possible broken link.
The current standards allow validity generalisation when evidence from similar settings supplies a strong basis. Similarity should cover the construct, procedure, population, context and use. The SIOP selection principles add work analysis, criterion relevance and local circumstances for employment decisions. The International Test Commission's test-use guidelines ask users to examine evidence for relevant populations, language versions, accessibility and intended purpose, and to reconsider validity when purpose, form, content or delivery changes.
Technology adds dependencies that paper manuals could once treat as operational detail. Interface, device, bandwidth, automated scoring, identity controls, data loss, security and software versions can change the observations or the population able to complete them. The 2025 ITC and Association of Test Publishers guidelines make those dependencies part of technology-based assessment quality.
This matters especially for AI-assisted scoring and recommendation. A validated input measure does not validate a model built on top of it. The model adds transformations, training data, outcome definitions and decision rules. A recommendation creates another inference beyond the input scores. Evidence has to follow the whole system and its versions.
The same discipline applies without software. A trained practitioner can change an interpretation by combining a score with an interview, history and local knowledge. That may strengthen the conclusion through triangulation. It can also introduce undocumented judgement. The evidence should match the interpretation that is actually delivered.
Consequences send evidence back upstream
An assessment enters a system. People prepare for it, avoid it, challenge it and alter their behaviour around it. Teachers shift instruction. Hiring managers change thresholds. Applicants learn which performances the process rewards. Institutions decide whose errors can be appealed.
Messick argued that the meaning and use of scores could not be sealed away from these effects. The current standards take a careful position. Intended benefits such as better placement or safer professional practice form claims that should be evaluated. Unintended outcomes can expose a rival explanation. If a science test disadvantages a group because it contains avoidable language complexity, the consequence points back to construct-irrelevant variance. If teaching narrows until only tested fragments remain, the score may underrepresent the educational domain people assume it covers.
Consequences also include missed opportunities. A false negative can deny someone access to learning that would have changed the attribute being predicted. A career assessment can become self-confirming when a person never explores options omitted from its recommendations. A selection process can appear accurate among those admitted while leaving the performance of rejected applicants forever unobserved.
These feedback loops prevent validity from becoming a pre-launch ceremony. Monitoring can reveal population drift, changes in criteria, differential prediction, unexpected response strategies, score inflation or new barriers. An instrument may remain unchanged while the world that gave its scores meaning moves around it.
Fairness requires a larger analysis than group outcome comparisons. It includes access to the construct, comparable score meaning, prediction, opportunity, privacy, consequences and the defensibility of the use itself. Those questions belong to the forthcoming fairness foundation. Here the central point is narrower: consequences can generate evidence about whether the original interpretation and use are holding.
A manual cannot validate every future use
A technical manual is necessary evidence, but its claims still have a boundary.
The 2014 standards describe validation as the joint responsibility of test developer and test user. Developers are expected to state intended interpretations and uses, provide the supporting rationale and evidence, identify limitations, and document inappropriate uses. A later user has to judge whether that evidence fits the new setting. If the interpretation changes, the burden for supporting it moves with the claim.
This division corrects two common mistakes. A publisher cannot validate every future use by declaration, while a local user does not have to discard all external research and begin from zero. Existing evidence can travel when the new setting is sufficiently similar; material differences and later outcomes show how far.
The 2026 Australian professional standards for career practitioners place that responsibility inside a career conversation. Practitioners are expected to understand validity, reliability and norm-group relevance for the instruments they use, and to apply assessment ethically, inclusively and with attention to language, culture, accessibility and digital delivery. For the person receiving a result, the standard explains why a familiar test name is never the whole evidence.
That standard does not say that every career tool can forecast a best occupation. An interest inventory can contribute evidence about reported preferences. An ability task can sample present performance under defined conditions. An aptitude interpretation concerns potential learning under specified opportunities. A personality score describes a characteristic pattern at a chosen level of abstraction. Each construct and method starts a different inference chain.
The forthcoming page on what a career assessment can tell you will compare those claims directly. The boundary here is already clear. Exploration, prediction, diagnosis, development and selection are different uses. Evidence for one does not silently authorise the others.
What “validated” leaves unsaid
An assessment is sometimes described as though validity were a seal attached to it: this test has been validated. The history of the idea leads to a less convenient but more revealing sentence: evidence supports interpreting this score in this way, for this population, under these conditions and for this use.
Inside that longer sentence is the whole argument traced in this article. A person said, chose or did something. A scoring procedure turned those observations into a number or category. A theory connected the result to a construct. The assessment sampled some parts of that construct and omitted others, while language, opportunity, method and context could also have entered the score.
Evidence then has to show the patterns the interpretation would lead us to expect. Related measures may converge; different constructs should remain distinguishable; predictions should meet the outcome they actually name. None of that evidence travels automatically to a different population, language, setting, software version or purpose.
Finally, the use changes the claim. Reflection, development, prediction, allocation and exclusion do not ask the same thing of a score. Their errors have different costs, and their consequences can reveal assumptions that failed. There is therefore no single validity number waiting at the end of the chain. There is a case that becomes more or less persuasive as each link is examined.
Cronbach and Meehl's committee wanted qualities that could be checked before publication. What they found was a more demanding form of accountability. Meaning has to be argued. Evidence has to fit the claim. Rival explanations have to remain visible. Uses carry their own assumptions. New evidence can change the judgement.
Validity is the record of how well that argument survives.
Sources
- American Educational Research Association, American Psychological Association and National Council on Measurement in Education. Standards for Educational and Psychological Testing. 2014.
- Cronbach, Lee J., and Paul E. Meehl. “Construct Validity in Psychological Tests”. Psychological Bulletin 52 (1955): 281–302.
- Campbell, Donald T., and Donald W. Fiske. “Convergent and Discriminant Validation by the Multitrait-Multimethod Matrix”. Psychological Bulletin 56 (1959): 81–105.
- Loevinger, Jane. “Objective Tests as Instruments of Psychological Theory”. Psychological Reports 3 (1957): 635–694.
- Messick, Samuel. “Foundations of Validity: Meaning and Consequences in Psychological Assessment”. ETS Research Report RR-93-51, 1993.
- Messick, Samuel. “Validity of Psychological Assessment”. American Psychologist 50 (1995): 741–749.
- Kane, Michael T. “Validating the Interpretations and Uses of Test Scores”. Journal of Educational Measurement 50 (2013): 1–73.
- Borsboom, Denny, Gideon J. Mellenbergh and Jaap van Heerden. “The Concept of Validity”. Psychological Review 111 (2004): 1061–1071.
- Sireci, Stephen G. “The Construct of Content Validity”. Social Indicators Research 45 (1998): 83–117.
- Padilla, José-Luis, and Isabel Benítez. “Validity Evidence Based on Response Processes”. Psicothema 26 (2014): 136–144.
- Schmidt, Frank L., and John E. Hunter. “Development of a General Solution to the Problem of Validity Generalization”. Journal of Applied Psychology 62 (1977): 529–540.
- Sackett, Paul R., Charlene Zhang, Christopher M. Berry and Filip Lievens. “Revisiting Meta-Analytic Estimates of Validity in Personnel Selection”. Journal of Applied Psychology 107 (2022): 2040–2068.
- Sackett, Paul R., Charlene Zhang, Christopher M. Berry and Filip Lievens. “Revisiting the Design of Selection Systems in Light of New Findings”. Industrial and Organizational Psychology 16 (2023).
- Society for Industrial and Organizational Psychology. Principles for the Validation and Use of Personnel Selection Procedures. 5th ed., 2018.
- U.S. Equal Employment Opportunity Commission et al. Uniform Guidelines on Employee Selection Procedures: Questions and Answers. 1979.
- International Test Commission. ITC Guidelines on Test Use. Final version 1.2.
- International Test Commission and Association of Test Publishers. Guidelines for Technology-Based Assessment. Version 1.1, 2025.
- Career Industry Council of Australia. Professional Standards for Australian Career Development Practitioners. 5th ed., 2026.
- National Academies of Sciences, Engineering, and Medicine. “Overview of Psychological Testing”. 2015.
- Smithsonian National Museum of American History. Strong Vocational Interest Blank for Men. 1927 object and directions.

