When a neutral rule keeps an old gate in place
At Duke Power's generating plant on the Dan River in North Carolina, there were five operating departments. Before 1965, Black employees were assigned to only one of them: Labor. The best-paid job there paid less than the lowest-paid job in any of the other four departments.
When that racial restriction ended, the route into the other departments did not simply open.
Since 1955, most departments had required a high-school diploma. On 2 July 1965, the day Title VII of the US Civil Rights Act took effect, Duke Power added two professionally prepared tests for new employees. Workers without a diploma could later qualify for transfer by passing the Wonderlic Personnel Test and the Bennett Mechanical Comprehension Test. Yet the Supreme Court's account in Griggs v. Duke Power Co. found that neither test was directed at learning or performing a particular category of job. White employees hired before the requirements had continued to perform and progress satisfactorily without them.
Everyone faced the same tests, and the court found no discriminatory purpose in their adoption. Even so, in 1971 it held that a procedure which excluded Black workers at a much higher rate could not control access to work when the employer had failed to show that it related to job performance.
Griggs belongs to one country's employment law at a particular moment. The question it exposes reaches further: what makes a gate fair when its rule is neutrally written and correctly scored, but the people arriving at it have not shared the same history?
Fairness in assessment means giving intended test takers meaningful access to the construct, checking whether scores support comparable interpretations, governing the decisions made from them and looking for avoidable harm so it can be remedied. This is continuing work, not a certificate awarded to a test once.
Kind treatment on test day is part of that definition, but only part. The current Standards for Educational and Psychological Testing organise fairness through four connected views: fair treatment during the process, absence of measurement bias, access to the construct being measured, and valid interpretation of an individual's score for its intended use. Laws add duties that vary by jurisdiction, while institutions decide how scores will be used. For the test taker, fairness is also felt in the explanation, the feedback and the chance to challenge a result.
That leaves several questions to ask at once:
| Part of the system | Fairness question |
|---|---|
| Treatment | Were the stated procedures applied consistently and respectfully? |
| Access | Could each intended test taker demonstrate the target construct without an irrelevant barrier? |
| Measurement | Does the score carry comparable meaning across the people for whom it will be interpreted? |
| Prediction | Does the relationship between the score and the named outcome differ across relevant groups? |
| Decision | Is the rule job- or purpose-relevant, proportionate, explainable and open to a suitable alternative? |
| Consequence | Who gains, who is excluded, what errors occur, and can harm be detected and corrected? |
Passing one row does not settle the rest. Consistent treatment can coexist with inaccessible design, and comparable measurement can feed an indefensible cut score. Equal selection rates may even conceal a test that measures the wrong thing equally badly.
Fairness has no single coefficient
In the 1960s, psychologists tried to give test bias a precise technical definition. The effort revealed how many different promises the word fair could contain.
T. Anne Cleary studied how admission-test scores predicted first-year grades for Black and white students in integrated colleges. In her 1968 paper, a test was predictively biased for a subgroup when a common prediction systematically placed its later criterion too high or too low. A shared prediction line could mislead if the same score predicted different outcomes for two groups.
Robert Thorndike showed why that definition could not settle fair selection. A rule might predict a criterion comparably across groups while selecting a smaller share of the people from one group who would later succeed. His 1971 “Concepts of Culture-Fairness” compared the proportion selected with the proportion expected to reach a specified performance level. Other researchers focused on different denominators, error rates or ideas of equal opportunity.
By 1973, Ronald Flaugher was comparing four distinguishable models of fair selection. Half a century later, Ben Hutchinson and Margaret Mitchell found the same arguments resurfacing in machine learning. Their history of test unfairness and algorithmic fairness traces several modern criteria back to debates in education and employment testing.
A fairness statistic always protects a chosen relationship: comparable prediction, equal opportunity among people who meet a criterion, equal error rates, equal selection or something else. Those relationships can conflict. Mathematics can reveal the conflict and measure whichever relationship has been chosen. The institution still has to say which social promise it is making, and why.
So “the model passed the fairness test” leaves the important questions unanswered. Which test, for which groups and against which outcome? At what threshold? What was left outside the calculation?
A group difference is not a diagnosis of bias
Suppose two groups obtain different average scores. Unequal opportunity to learn might be responsible. So might different familiarity with the task, a construct-relevant difference, an irrelevant barrier, sampling error or some combination of these. The size of the gap cannot tell us which explanation is right.
Nor do similar group averages show that the items function comparably, that individuals had equal access or that the resulting decisions were fair.
Measurement bias asks a narrower question about meaning. Differential item functioning, usually shortened to DIF, occurs when people at the same standing on the target construct have different probabilities of a response because of group membership. The Testing Standards treat a DIF finding as the start of an inquiry. Investigators still have to work out why the difference occurred and whether the item has measured something irrelevant before calling it biased.
Take an item about calculating a medication dose. Arithmetic may properly belong to the task, while an unnecessary idiom adds a language hurdle the assessment never intended. Familiarity with one health system might also give local experience an advantage. A DIF finding could lead to revision, separate interpretation, further study or a reasoned decision to retain the item, depending on what caused it.
Bias can also appear at the level of the whole test or in prediction. A test with little item-level DIF may still predict a criterion differently across groups. Several instances of DIF might cancel out in the total score while exposing a design problem that deserves attention. Item functioning, total-score meaning and the relationship with an outcome provide different pieces of evidence.
Measurement invariance offers another family of tests. Researchers work in stages, asking whether the construct has a comparable pattern, whether its units are comparable, and whether its origins or thresholds allow mean comparisons. The level of invariance supported by the evidence determines which comparisons can be defended. Calling a measure “invariant” does not make it culturally neutral.
Translation makes the difficulty tangible. Replacing each word with its closest dictionary equivalent can change the difficulty, familiarity or even the activity being described. The International Test Commission's current adaptation guidelines place translation within a wider process that includes cultural context, administration, scoring and empirical confirmation. Parallel-looking sentences matter less than whether the adapted responses support the intended interpretation.
Sometimes the same conditions are the unfair part
Standardisation serves an essential purpose. Common directions, time limits, equipment and scoring reduce accidental variation. The trouble begins when a standard procedure also demands a capability that sits outside the construct.
A screen reader changes how printed content is presented to a blind test taker. Large text changes its scale, and extra time changes a time limit. Each alteration may give someone a better chance to demonstrate the same target construct. If visual acuity or speed has no place in the claim, refusing the change protects the procedure at the expense of the measurement.
Everything depends on the target. Extra time may improve access to a reasoning task whose interpretation does not include speed. Where fluent performance under time pressure is part of the defined capability, it changes the meaning of the score. A spoken version of a reading test may still measure language comprehension, but it no longer measures decoding from print.
The Testing Standards call a change intended to retain comparable score meaning an accommodation. A modification changes the construct and, with it, the meaning of the score. Real cases rarely fall into two tidy piles. The answer requires evidence about the person, the assessment and the proposed use.
Good design can remove many barriers before anyone has to request an individual change. Clear navigation, keyboard access, adjustable text, captions, compatible assistive technology and restrained language demands can widen access while preserving the task. The ETS accessibility guidance starts with a precise construct, helping designers separate essential features from helpful or incidental ones. The 2025 technology-based assessment guidelines add an important warning for digital delivery: technical compliance can still leave an assessment cognitively confusing or impossible to operate in practice.
Giving people an equal opportunity to demonstrate the construct may require different routes. Evidence must show where those routes remain comparable and where separate interpretation would be more honest.
Assessment fairness · Full lifecycle
Where can unfairness enter?
A procedure can pass one fairness check and fail the next. Follow it from its original purpose to its eventual consequences, noticing how the question changes and where responsibility sits.
Illustrative use
Hiring work sample
Job relevance, comparable access, predictive evidence, decision rules, adverse impact, privacy and recourse all matter.
Illustrative use
Career-exploration prompt
The use is easier to question and correct, but it still needs competent interpretation, privacy, context and room to disagree.
1 · Lifecycle stagePurpose and constructShow detailsHide details
Question: Is the quality being measured relevant to the proposed use?
Possible harm: An irrelevant or historically narrow target becomes a gate.
What could change: The claim or purpose may need to be narrowed, changed or abandoned.
Who is responsible: Commissioner, practitioner and subject-matter experts
2 · Lifecycle stageTask designShow detailsHide details
Question: Do tasks represent the construct without avoidable barriers?
Possible harm: Language, interface or prior exposure replaces the intended capability.
What could change: The tasks, content or interface may need to change before the result means the same thing.
Who is responsible: Developer and accessibility specialists
3 · Lifecycle stageAccess and administrationShow detailsHide details
Question: Can people reach the same construct under appropriate conditions?
Possible harm: A standard procedure protects uniformity while denying construct access.
What could change: Different conditions or accommodations may provide access without changing the intended claim.
Who is responsible: Administrator, developer and institution
4 · Lifecycle stageResponse and scoringShow detailsHide details
Question: Are responses captured and scored comparably across relevant groups?
Possible harm: Rater, item, device or model behaviour shifts the score meaning.
What could change: Group patterns, item behaviour and errors can reveal whether scoring means the same thing.
Who is responsible: Developer, psychometrician and vendor
5 · Lifecycle stageInterpretationShow detailsHide details
Question: Does the same score support a comparable interpretation here?
Possible harm: A score travels to a population, language, setting or use without evidence.
What could change: The interpretation may need to be narrowed when evidence does not travel to the new setting.
Who is responsible: Practitioner and institution
6 · Lifecycle stageDecision ruleShow detailsHide details
Question: Is the rule job- or purpose-relevant, and what alternatives exist?
Possible harm: A defensible score is combined with an arbitrary threshold or policy.
What could change: The threshold, alternatives and costs of error become part of the fairness question.
Who is responsible: Decision-maker and accountable institution
7 · Lifecycle stageConsequencesShow detailsHide details
Question: Who gains, who is burdened and which outcomes are missed?
Possible harm: Aggregate accuracy hides exclusion, opportunity loss or self-fulfilling effects.
What could change: Outcomes can be compared across groups, contexts and kinds of error, revealing where the system must change.
Who is responsible: Institution, affected communities and governance body
8 · Lifecycle stageFeedback and recourseShow detailsHide details
Question: Can the person understand, correct and challenge the result?
Possible harm: Opacity turns a contestable inference into unreviewable authority.
What could change: Explanation, correction and meaningful human review can return some agency to the person affected.
Who is responsible: Practitioner, institution and appeals reviewer
9 · Lifecycle stageMonitoringShow detailsHide details
Question: Does the interpretation and use continue to hold after launch?
Possible harm: Population, technology, work or decision practice changes around a static procedure.
What could change: Later outcomes, complaints and changed conditions may show that the use needs revision or suspension.
Who is responsible: Institution, developer and independent oversight
Equal treatment, construct access, measurement bias, predictive bias, adverse impact, privacy and procedural fairness are related checks. None can certify the whole lifecycle by itself.
Nine disclosure panels follow an assessment from purpose and construct through task design, access and administration, response and scoring, interpretation, decision rule, consequences, feedback and recourse, and monitoring. Every panel names a fairness question, possible harm, what could change and who holds responsibility. Two case cards compare an illustrative hiring work sample with a low-stakes career-exploration prompt. No fairness score or legal verdict is shown.
Opportunity enters the score before the test begins
An assessment records performance now, but that performance arrives with a history of instruction, practice, language, tools, health and access.
The Testing Standards use opportunity to learn for the exposure to instruction or knowledge needed to acquire what a test covers. In education, the fairness problem is especially sharp when the same authority withholds adequate learning opportunities and later imposes a high-stakes consequence for failing the test. The score may accurately record present attainment while the policy punishes someone for instruction the institution never provided.
Two truths need to be held together here. Unequal opportunity does not turn every observed difference into measurement error. Accurate measurement does not make every consequence just.
Griggs made that history visible. North Carolina's segregated schools had not distributed education equally. Duke Power then used a diploma and broad tests to control entry to departments whose jobs white workers had already performed successfully without those requirements. The criterion belonged to the institution's history; it did not float above it.
The same problem appears whenever an assessment is validated against an observed outcome. Supervisor ratings may contain halo effects or reflect unequal assignments. Training completion can depend on who received support, while tenure partly reflects who was admitted and retained. Job performance itself has several dimensions. The SIOP selection principles therefore require work analysis and attention to criterion relevance: a predictor inherits the omissions and distortions of the outcome used to validate it.
A historical criterion can provide useful evidence. It may also record what opportunity has made possible, alongside the capability it was intended to capture.
Disparity is a warning light, not a causal explanation
Assessment becomes consequential when an institution sets a threshold, ranks candidates or combines a score with other evidence. At that point, group selection rates can be compared.
In US employment practice, the Uniform Guidelines on Employee Selection Procedures use the four-fifths rule as an initial indicator. A group's selection rate below 80 per cent of the highest group's rate will generally attract attention. The official guidance calls this a rule of thumb, not a legal definition. Smaller differences can matter when samples are large and the effect is practically important; apparently large percentage gaps can be unstable when only a few people were considered.
Adverse impact names an outcome pattern. On its own, it cannot locate the cause in an item, a cut score, a recruitment channel, a criterion, unequal preparation or another part of the system. Nor is it identical to unlawful discrimination, whose definition and proof depend on the law governing the place and decision.
Because the outcome pattern cannot locate its own cause, any explanation has to follow the whole process. The construct may be unnecessary for the purpose, the procedure may represent it badly, or the criterion may be unsound. Scores may function differently across groups. A cut score may lack a defensible rationale, or a similarly effective alternative may exclude fewer people. The EEOC's current employment-testing guidance places that final possibility inside US disparate-impact analysis. Other jurisdictions organise their legal duties differently.
Real trade-offs can remain. A procedure with useful predictive evidence may produce group differences. A lower-impact alternative might sample the work more directly, change the weighting of evidence or improve access, while also predicting a different outcome. Ployhart and Holtz's review treats this as a problem of designing the whole system. It warns against assuming in advance that accuracy and diversity must be traded against each other.
Outcome parity offers no safe harbour by itself. A random decision can equalise rates while abandoning the purpose of assessment. At the other extreme, a technically strong procedure may still impose avoidable exclusion. The institution has to state what it is trying to learn, which errors it is prepared to make and why this evidence should control this opportunity.
The person experiences a procedure, not a validity report
People do not experience a test through its technical report. They experience the questions, the instructions, the people administering it and whatever happens when the score appears. They want to know whether the procedure was relevant, whether they had a fair chance to show what they could do, how they were treated and whether anyone will explain the result.
John Hausknecht, David Day and Scott Thomas combined 86 independent applicant samples involving 48,750 people. Perceived job-relatedness and predictive relevance were strongly associated with several perceptions of fairness. More favourable reactions were also associated with stronger organisational attitudes and intentions.
Perception cannot replace measurement evidence. A charming interview may feel fair while rewarding irrelevant similarity; an unfamiliar work sample may be well designed but poorly explained. The research reveals a different kind of consequence. A procedure tells people what the institution values and whether it regards them as entitled to an explanation.
That matters acutely in career assessment, where a profile may influence which possibilities someone investigates, dismisses or begins to fold into their identity. The NCDA Code of Ethics and Australia's 2026 career-development standards place assessment within competent, contextual and collaborative interpretation. The person should be able to understand the result, examine its basis, supply missing context and disagree with it.
Feedback is part of the intervention. “Low aptitude” can sound like a ceiling when the evidence concerns performance on selected tasks under selected conditions. “Poor fit” can sound like a verdict on an occupation when the observations concern interests or values. A label that travels beyond its evidence can shape someone's choices even when it never denies them a formal opportunity.
The questions set out in What Can a Career Assessment Tell You? still apply. What was observed? Which construct and comparison produced the result? What use is intended, and what evidence is missing? Fair feedback also makes room for correction, experiment and review.
Digital assessment creates a second assessment around the first
A paper test records marked responses. A digital system may also record the device, response times, navigation, pauses, keystrokes, audio, video, location and biometric identifiers. Some of these data may support access, security or scoring. Others may be collected simply because the platform makes it easy.
Every new data stream introduces another claim. Does fluency with the interface belong to the construct? Does a slower network reduce the opportunity to respond? Can an automated scorer recognise the full range of legitimate expression? Does remote proctoring work comparably across different bodies, homes and assistive technologies?
The current ITC/ATP guidance requires technology-enhanced items and automated scoring to be evaluated for validity, reliability and fairness, including among relevant subgroups. It also calls for documentation, appeals and continuing monitoring. The SIOP guidance on AI-based selection keeps those responsibilities with the institution even when the model is complex or supplied by a vendor.
An algorithm trained on an assessment score adds a new layer of inference. The same is true of a model that converts several scales into a ranking or recommendation. Evidence for the inputs cannot validate the target, training data, weighting, threshold or eventual use. Manish Raghavan and colleagues found that public claims about algorithmic hiring tools often left development, validation and bias-mitigation practices difficult to evaluate. A promise to “remove human bias” is not evidence of an audit.
Privacy belongs inside fairness because observation can impose a burden of its own. The OECD privacy principles require limits on collection, specified purposes, restrictions on reuse, safeguards, openness, individual participation and accountability. Data about groups may be needed to discover disparate outcomes. That purpose does not justify keeping the data indefinitely or drawing unrelated inferences from them.
Test takers should know what is collected, why it is needed, who receives it and whether automated decision-making, video monitoring or biometrics are involved. They also need a way to challenge a score or report. The ITC/ATP guidelines specifically require an AI appeals procedure that includes review. When no person can question the system, technical opacity becomes institutional power.
Fairness changes when the system changes
An assessment can begin with good evidence and grow less fair over time. The population changes, a job is redesigned or a translated form enters use. A mobile interface replaces a computer. An automated scorer is updated. Preparation becomes unevenly available, and people alter their behaviour once they learn how the gate works.
Fairness evidence therefore stretches across the assessment's whole life. It begins with the purpose, the construct, the intended population and the consequence. An exploratory prompt, course placement and exclusion from employment carry different burdens. Design then shapes access: language, interfaces, sensory demands and administrative rules can create barriers before anyone answers a question.
The score itself adds another set of questions about content, response processes, reliability, invariance, item and test functioning, and differential prediction. Once the score enters a decision, the criterion, threshold and combination of evidence matter too, as do realistic alternatives. Data collection brings questions of explanation, access, retention and reuse.
What happens afterwards can change the judgement again. An intelligible result, correction of errors and human review matter more as the stakes rise. Access failures, score functioning, group outcomes, complaints, appeals and later effects reveal whether the original argument still holds. Sometimes they show that the use has to be revised or suspended.
The NIST AI Risk Management Framework describes this continuing work as mapping, measuring, managing and governing risk. Its language is new; the responsibility is the one exposed at Dan River. The institution that chooses a gate remains accountable for its relationship to the work and for the people it excludes.
That fair treatment is larger than identical treatment. A common procedure is valuable, but sameness becomes a source of unequal access when it preserves an irrelevant barrier.
That bias has to be located within the system. Item bias, test bias, predictive bias and adverse impact require different evidence and lead to different remedies.
That measurement cannot settle the justice of a decision. A score may describe present performance accurately even when the criterion, threshold or opportunity around it is deficient.
That consequences complete the evidence chain. Feedback, privacy loss, missed opportunities, appeals and later outcomes all belong in the assessment's history.
At Dan River, the tests had professional names, national comparison scores and uniform administration. Those features could not explain why the tests should control movement into Operations, Maintenance or the Laboratory. What Duke Power had left undone was the connection between the gate and the work: a definition of the jobs, evidence of the relationship, an account of who was excluded and a defensible reason to preserve the burden.
A score that everyone can calculate is only the beginning. Fairness is tested by whether the responsible institution can explain the decision, examine what follows, revisit its reasoning and repair avoidable harm.
Notes on the evidence
This article treats fairness as a family of connected technical, ethical, experiential and legal questions. Its working definition draws on the 2014 Testing Standards, current ITC and SIOP guidance, civil-rights history and career-practice standards. Griggs and the four-fifths rule are bounded US employment examples. They offer no legal advice or universal rule. Cleary's and Thorndike's models remain historically important because fair prediction and fair selection can ask different questions. DIF, measurement invariance and adverse impact provide evidence for investigation; they do not automatically establish bias or discrimination. No claim is made here that any Guidebeam assessment, score, recommendation or decision is fair, unbiased, normed, accessible or legally compliant.
Sources and further reading
- AERA, APA and NCME, Standards for Educational and Psychological Testing (2014)
- U.S. Supreme Court, Griggs v. Duke Power Co., 401 U.S. 424 (1971)
- U.S. Equal Employment Opportunity Commission, Uniform Guidelines Q&A
- U.S. Equal Employment Opportunity Commission, “Employment Tests and Selection Procedures”
- T. Anne Cleary, “Test Bias” (1968)
- Robert L. Thorndike, “Concepts of Culture-Fairness” (1971)
- Ronald Flaugher, The New Definitions of Test Fairness in Selection (1973)
- Ben Hutchinson and Margaret Mitchell, “50 Years of Test (Un)fairness” (2019)
- Diane Putnick and Marc Bornstein, “Measurement Invariance Conventions and Reporting” (2016)
- Samuel Messick, “Validity of Psychological Assessment” (1993)
- Society for Industrial and Organizational Psychology, Principles for the Validation and Use of Personnel Selection Procedures (2018)
- Robert Ployhart and Brian Holtz, “The Diversity–Validity Dilemma” (2008)
- John Hausknecht, David Day and Scott Thomas, “Applicant Reactions to Selection Procedures” (2004)
- International Test Commission, Guidelines for Translating and Adapting Tests, 2nd ed. (2026)
- International Test Commission and Association of Test Publishers, Guidelines for Technology-Based Assessment (2025)
- Educational Testing Service, How ETS Works to Improve Test Accessibility
- Society for Industrial and Organizational Psychology, AI-Based Assessments for Employee Selection (2023)
- National Institute of Standards and Technology, AI Risk Management Framework 1.0 (2023)
- OECD, Guidelines Governing the Protection of Privacy and Transborder Flows of Personal Data
- Raghavan, Barocas, Kleinberg and Levy, “Mitigating Bias in Algorithmic Hiring” (2020)
- National Career Development Association, Code of Ethics (2024)
- Career Industry Council of Australia, Professional Standards for Australian Career Development Practitioners (2026)

