Forty lines and two attempts
In 1910, William Brown gave people pages of closely printed tasks. Some crossed out every instance of a letter in French text. Some added columns of digits for five minutes. Others marked the middle of 8-centimetre lines and tried to divide 9-centimetre lines into thirds. A smaller group adjusted the fins of the Müller-Lyer illusion until two lines looked equal.
Brown was studying mental abilities, but each task confronted him with a prior problem. A person's result changed when the observations changed. Ten divided lines produced an unstable measure, so for one group he combined bisection and trisection across two sessions, creating a total from 40 lines. The repeated addition task behaved differently. Some results agreed closely; others did not.
The paper, “Some Experimental Results in the Correlation of Mental Abilities”, appeared beside Charles Spearman's “Correlation Calculated from Faulty Data” in the same issue of the British Journal of Psychology. Both men were trying to work out what a relationship among scores could mean when the scores themselves were uncertain.
That problem survives every polished assessment report.
Reliability is evidence about the consistency and precision of scores across defined repetitions of a measurement procedure. It belongs to a particular score, procedure, population and use. The relevant repetition might change the questions, occasion, rater, form, setting or device. Change what counts as a repetition and you change the reliability question.
The working definition for this library is:
Reliability is the degree to which scores remain consistent across the repetitions a proposed interpretation assumes are interchangeable.
This definition carries an important warning. The parent page on psychometrics follows the larger chain from observation to use. Within that chain, reliability asks whether a result would hold under specified changes. Validity asks whether evidence and theory support the interpretation and use. A coefficient cannot make a score meaningful by itself, and a measurement can repeat the same mistake with impressive consistency.
First decide what is allowed to change
Imagine a person completes an interest inventory on Monday and again on Tuesday. Should the same result appear?
The answer depends on what the score claims to represent. If the inventory is intended to describe a reasonably stable pattern of vocational interests, a large overnight swing may raise concern. If a questionnaire records present anxiety after an interview, a change between morning and afternoon may be the very thing it was built to detect. Treating every change as error would erase the state under study.
The current joint Standards for Educational and Psychological Testing therefore begins reliability analysis with the replication. Which parts of the procedure are fixed, and which are treated as interchangeable samples from a larger set?
An educational test may treat different questions built to the same content specification as interchangeable. A structured interview may need to generalise across trained raters. A work sample may ask whether performance holds across several tasks. A personality questionnaire may keep the wording fixed but ask about consistency across occasions. A digital assessment may also depend on interface, device, latency, scoring version and delivery conditions.
The measurement procedure draws a boundary around these choices. Reliability evidence then asks how scores behave when the permitted parts vary.
| Source of variation | Question being asked | What consistency would mean |
|---|---|---|
| Items or tasks | Would a comparable sample from the same domain give a similar score? | The result is not an accident of the selected questions |
| Occasions | Would the score hold over an interval in which the construct should remain stable? | Day-to-day conditions do not dominate the intended signal |
| Forms | Would another version built to the same specification give an exchangeable result? | Form selection does not materially change the interpretation |
| Raters | Would qualified judges applying the scoring rules reach a similar result? | The assigned rater does not control the outcome |
| Contexts or technology | Would acceptable settings, devices or delivery conditions preserve the result? | Operational variation stays within the intended procedure |
Reliability · Define the repetition
Which repetition should this score survive?
Items, occasions, forms, raters and contexts create different reliability questions. Naming the replication universe comes before choosing a coefficient.
1 · FacetItems or tasksShow detailsHide details
Question: Would a comparable sample from the domain preserve the result?
Study design: Sample items or tasks under a documented content plan; study relevant internal consistency or task sampling.
Still open: Item evidence does not establish stability across time, raters or settings.
2 · FacetOccasionsShow detailsHide details
Question: Would the result hold across an interval in which the construct should remain stable?
Study design: Repeat the procedure across a justified interval and study retest consistency and change.
Still open: Genuine learning, mood or circumstance change should not automatically be labelled error.
3 · FacetFormsShow detailsHide details
Question: Would another form built to the same specification be exchangeable?
Study design: Administer alternate forms with order or linking controls and compare score behaviour.
Still open: Similar-looking forms are not exchangeable without evidence.
4 · FacetRatersShow detailsHide details
Question: Would qualified judges applying the rules reach a similar result?
Study design: Use crossed or nested ratings, a documented rubric and a coefficient matched to the design.
Still open: Agreement can be consistently biased and does not by itself establish validity or fairness.
5 · FacetContexts or technologyShow detailsHide details
Question: Would acceptable settings, devices and delivery conditions preserve the result?
Study design: Vary approved conditions deliberately and test measurement equivalence, accessibility and security.
Still open: A new device or delivery mode does not inherit evidence from the old one automatically.
Illustrative case
Self-report interest scale
- Fixed
- The construct definition, instructions and intended exploratory use.
- May vary
- Comparable items, approved occasions and supported delivery contexts.
- Real change
- A changed preference or life circumstance may be real development, not measurement failure.
Illustrative case
Scored work sample
- Fixed
- The performance domain, rubric and decision purpose.
- May vary
- Comparable tasks, qualified raters and approved work conditions.
- Real change
- Practice, fatigue or new tools may change performance for substantive reasons.
Illustrative case
Short reasoning assessment
- Fixed
- The target reasoning demand, administration rules and population.
- May vary
- Task samples, equivalent forms and supported devices.
- Real change
- Learning, language familiarity or changed access conditions may alter the observed performance.
Five disclosure panels distinguish item or task sampling, occasions, alternate forms, raters, and contexts or technology. Each gives a reliability question, an appropriate study design and an important boundary. Three illustrative cases—a self-report interest scale, a scored work sample and a short reasoning assessment—state what remains fixed, what may vary and what genuine change should not automatically be called error. No coefficient or universal threshold is shown.
This is why “the test has a reliability of .86” is incomplete. Which score? Estimated how? In which people? Across what repetition? A coefficient based on one administration can say something about relationships among items while saying nothing about stability next month or agreement among judges.
Error is part of the design
The word error sounds like a blunder. In measurement theory it has a narrower job. Error is variation that the proposed interpretation says should not alter the score.
Suppose a certification candidate receives one set of tasks and another candidate receives an equivalent form. If the forms are meant to be interchangeable, chance advantage from one task sample is error for that use. If two trained judges score the same performance, rater disagreement is error when the decision claims to be independent of which judge was assigned. If an assessment intends to measure a changing state, movement over time may be signal.
The distinction is created jointly by the construct and the decision. A relative interpretation asks where people stand compared with one another. An absolute interpretation asks whether each person has reached a standard. A source of variation that shifts everyone together may leave rankings unchanged while moving people across a cut score. It matters to the absolute decision even when a rank-based coefficient barely notices it.
Consequences also change how much precision is enough. The Testing Standards says the need for precision rises when decisions are important and hard to reverse. A low-stakes exploratory score that will be discussed with a person and checked against other evidence can tolerate more uncertainty than a score used to deny admission, employment or professional practice.
No universal decimal can settle that judgement.
Spearman and Brown divide the test
Spearman and Brown approached reliability through halves. Divide a collection of observations into two comparable sets, calculate a score from each and see whether people retain roughly the same ordering. A full measurement built from both halves should be more reliable than either half alone, so each developed a way to adjust the half-test relationship for length.
This became the Spearman–Brown formula. It also gave test builders a practical forecast: if comparable observations were added, how much might reliability improve?
Length helps because a broader sample can average away some chance variation. Brown's 40 divided lines were more dependable than a handful. Yet adding observations works only when they contribute relevant information under a defensible design. Twenty near-duplicate questions can agree closely while covering a construct badly. A long assessment can also measure fatigue, reading load or speed more heavily than intended.
Spearman's wider concern was attenuation. If two measures contain random error, their observed relationship will usually be weaker than the relationship among the theoretical scores they are intended to represent. His correction made uncertainty part of correlational research. It also depended on having defensible reliability estimates in the first place. Correcting a correlation with an irrelevant coefficient merely moves the assumption into the arithmetic.
The split-half method established a durable idea: measurement quality depends on the sample of observations. It left open which division of a test should count and which kinds of repetition matter outside the item set.
What a true score means
Classical test theory gave the field a compact expression:
observed score = true score + error
The phrase true score invites a grander reading than the mathematics allows. In the formal account developed through classical test theory and set out by Frederic Lord and Melvin Novick in Statistical Theories of Mental Test Scores, it is an expected score over a defined set of hypothetical repetitions. It is tied to the procedure. It is not a person's pure intelligence, real personality or hidden vocational destiny.
The corresponding reliability coefficient compares dependable variation among people with the total observed variation in a population. High values occur when error variation is small relative to the differences among people. This creates a subtle population dependence. Give the same instrument to a group whose scores are spread widely and the coefficient may be higher than when the tested group is tightly clustered, even if the amount of error has not improved.
A coefficient copied from a development paper does not automatically describe a new sample, language, age group or operational setting.
The standard error of measurement, or SEM, answers a more person-facing question. It expresses expected random variation in the units of the score. If the SEM is large relative to a reported difference, the numerical gap may not survive a relevant repetition. A confidence interval can make that uncertainty visible.
An average SEM can conceal unevenness. Many tests estimate more precisely in the middle of their scale than near the extremes, or are designed to be especially informative near a decision point. A conditional standard error reports expected error at a particular score level. For someone close to a cut score, precision in that neighbourhood can matter more than an impressive overall coefficient.
Lee Cronbach eventually argued that the standard error was the most important single piece of reliability information to report. It tells the reader what the uncertainty could mean for a score, rather than presenting one scale-free decimal as a certificate.
How alpha swallowed the question
By 1951, test theorists had several ways to estimate internal consistency. Kuder and Richardson had formulas for dichotomously scored items. Louis Guttman had derived a family of lower bounds. Lee Cronbach's “Coefficient Alpha and the Internal Structure of Tests” gave the field a general expression that worked with more kinds of item scores. Alpha could also be understood through the average of adjusted coefficients from possible split halves.
It was useful, computable and easy to compare. It became ubiquitous.
Alpha asks whether parts of a score behave consistently in one dataset under a particular model. It does not show that the items measure one dimension. It does not test stability over time, alternate-form equivalence or rater agreement. It cannot establish validity or fairness. A scale can obtain a high alpha because it asks the same narrow question repeatedly.
Modern disagreement concerns what should follow. Daniel McNeish's “Thanks Coefficient Alpha, We'll Take It from Here” argues that researchers should use reliability estimates matched to more realistic measurement models. Tenko Raykov and George Marcoulides answer that alpha can remain useful when its conditions are defensible. William Revelle and David Condon's tutorial from alpha to omega shows that alpha, several split-half coefficients and forms of omega answer related but non-identical questions.
Replacing alpha with omega in every report would preserve the deeper mistake. A coefficient earns its place by matching the structure of the score and the generalisation being made. The name of the statistic is the end of that reasoning, not the beginning.
Cronbach himself kept moving. In the late 1990s he planned a fiftieth-anniversary reconsideration of alpha. Illness and the loss of virtually all his near vision prevented the technical paper he had imagined. Richard Shavelson persuaded him to dictate a non-technical account, then edited it after Cronbach's death in 2001.
The resulting “My Current Thoughts on Coefficient Alpha and Successor Procedures” is strikingly unsentimental. Cronbach wrote that he no longer regarded alpha as the most appropriate way to examine most data. A single coefficient was too crude for measurements assembled from tasks, raters, occasions and settings. The better question was where the variation came from.
Time, memory and genuine change
Test–retest evidence looks simple: administer a measure twice and correlate the scores. Its interpretation is not.
The interval must be long enough to reduce memory and immediate carry-over, yet short enough that the target is not expected to change. Reusing the same questions can inflate consistency because people remember answers or learn the task. Using a different form introduces form differences. Fatigue, practice, treatment, a life event or new learning may produce score change that belongs to the person rather than the instrument.
The construct decides which change counts as trouble. A long-term personality tendency, a present mood, current knowledge and performance after training make different temporal claims. Reporting “test–retest reliability” without the interval, forms, conditions and population hides those claims.
There is a second trap. A task can produce a dependable average effect while ranking people poorly. In three studies of seven familiar cognitive tasks, Craig Hedge, Georgina Powell and Petroc Sumner found what they called the reliability paradox. Tasks designed to produce a strong effect in most people can deliberately minimise between-person variation. They may work well for demonstrating the average effect and badly for measuring stable individual differences.
Reliable for what is therefore a design question. Group research, individual ranking, change detection and clinical classification need different evidence.
When judgement enters the score
Some observations do not arrive as right or wrong answers. A teacher scores an essay. Several assessors watch a simulated consultation. A clinician assigns a category. An interviewer rates a response against behavioural anchors.
The score now depends partly on who judges it. Training and rubrics can reduce variation, but the reliability analysis must fit the design. Jacob Cohen's 1960 coefficient kappa addressed agreement between two raters assigning nominal categories beyond a chance model. Patrick Shrout and Joseph Fleiss later showed that intraclass correlations come in several forms, depending on whether the raters are fixed or sampled and whether the use concerns one rating or an average.
There is no generic “inter-rater reliability” statistic. A coefficient for the agreement of two specially trained developers may overstate consistency among the larger pool of operational raters. A correlation can preserve rank order while missing systematic differences in severity. Category prevalence can change the behaviour of kappa. The report must say who rated what, how many ratings formed the score and whether the goal was consistency, absolute agreement or accurate classification.
Rater evidence also leaves construct questions open. Judges may agree because a rubric is clear. They may agree on a biased criterion. Consistency controls one source of uncertainty; it cannot justify the criterion itself.
From one error bucket to a map
Cronbach, Goldine Gleser and Nageswari Rajaratnam introduced generalizability theory in 1963. Their 1972 book with Harinder Nanda, The Dependability of Behavioral Measurements, developed the framework in full.
Generalizability theory treats observations as drawn from a designed universe. A performance score might vary by person, task, rater and occasion, with interactions among them. Instead of placing every unwanted fluctuation in one error term, a study estimates how much variation each facet contributes.
A generalizability study describes those components in the observed design. A decision study asks what would happen under another design. Would adding tasks help more than adding a second rater? Does averaging across two occasions reduce uncertainty enough for the proposed decision? Is a score dependable for ranking people but weaker for deciding whether each has met an absolute standard?
The framework makes reliability an engineering problem. A designer can discover that task sampling dominates the error and spend effort on more representative tasks. If rater variation is small, doubling raters may add cost with little gain. If person-by-task interaction is large, one polished work sample cannot stand for a broad capability domain.
The universe still has to be defended. Calling five tasks a random sample does not make them representative. Fixing all assessments to one device can hide a delivery problem that appears elsewhere. Statistical decomposition clarifies the consequences of the design; it cannot choose the construct or universe on the designer's behalf.
Precision can change along the scale
Item response theory approaches precision through information. An item can distinguish well among people near one part of a scale and contribute little elsewhere. Combining item information shows where a test estimates the modelled attribute more or less precisely.
This matters for adaptive testing. A calibrated system can choose later items in response to earlier performance, concentrating information near the current estimate. The result may be a shorter test with useful conditional precision. The 2025 International Test Commission and Association of Test Publishers guidelines place this precision inside a wider digital procedure that includes delivery, accessibility, security and version control.
Information from one calibrated administration does not establish stability across days, devices or contexts. A software update can alter timing. A new interface can change navigation demands. Remote administration can introduce a different environment. A scoring model can be versioned. Technology adds possible observations and new sources of variation.
For a threshold decision, the most useful evidence may be conditional error near the threshold and the consistency of classifications under relevant repetitions. A single average coefficient can be high while people close to the boundary move between categories.
The threshold that does not exist
Rules of thumb such as .70 is acceptable or .90 is required circulate because they make a difficult judgement look settled. They omit the use.
A coefficient reflects the population's score spread, the estimator, the sources of error included and the length and structure of the measure. Two coefficients can differ because one study sampled a broad population and another sampled people who were already similar. A value that supports group-level research may be inadequate for an individual exclusion decision. A modest exploratory result may still prompt a useful conversation when it is presented with uncertainty and checked against other evidence.
The Testing Standards requires evidence appropriate to the intended score, population and use. It also asks developers to report which sources of error were examined, how precision was estimated and what limitations remain. This is a stronger discipline than a universal threshold because it connects the statistic to the claim.
The coefficient only becomes informative when four things are known: the difference or decision the score must support; the repetitions through which the interpretation should remain unchanged; the sources of variation included in the estimate; and the amount of uncertainty the consequence can tolerate. Another decimal place cannot supply any of them.
Consistency is necessary and unfinished
A bathroom scale that adds five kilograms every morning can be highly consistent. A biased rater can apply the same unfair standard to every candidate. A narrow questionnaire can repeat one idea with excellent internal consistency. Reliability does not detect every systematic error because some errors reproduce themselves.
Low reliability sets a limit. An unstable score cannot support precise prediction, diagnosis or comparison. High reliability removes that particular obstacle and leaves the interpretation to earn its validity. The next page in this sequence asks what evidence that larger claim requires.
The distinction also protects people from two opposite errors. One is to dismiss any movement as unreliability when a person's state, knowledge or circumstances have genuinely changed. The other is to turn a repeatable score into a permanent property of the person. Reliability says how well a specified procedure can reproduce a specified result. It does not freeze the person in place.
What sits inside a reliability claim
A reliability coefficient is a compressed account of a particular score, form or version; a population and sample; and a decision about what counted as repetition. Items, occasions, raters, contexts and scoring versions do not test the same kind of consistency. The statistic only becomes meaningful when those choices are known.
An interest profile, personality scale, work sample and reasoning task therefore call for different repetitions. The design decides which changes count as error and which are allowed to be real. A standard error or information function then shows how much uncertainty remains around the score or decision region, while other sources of variation may still be untested.
The 2026 Professional Standards for Australian Career Development Practitioners place these technical questions inside a human conversation. A result has to be considered alongside experience, context and other evidence, with the consequence and reversibility of the decision in view. Reliability still leaves validity and fairness unfinished.
Brown's pages of digits and divided lines now look remote. The measurement problem he encountered has not changed. A score is built from selected observations. Another defensible selection might yield another result.
Reliability begins when a report makes clear which result should endure, through which changes, and with how much uncertainty.
Notes on the evidence
The public definition and reporting duties in this article follow the 2014 joint Standards for Educational and Psychological Testing, supplemented by Samuel Livingston's ETS guide and the 2025 ITC/ATP guidance for digital assessment. The historical account uses the 1910 papers by Spearman and Brown, Cronbach's 1951 alpha paper, his 2004 dictated reconsideration and the foundational work on generalizability theory. Debate continues over alpha's useful scope and the choice among model-based alternatives. The article therefore rejects both automatic reliance on alpha and automatic replacement by a fashionable coefficient. No universal reliability threshold is asserted.
Sources and further reading
- American Educational Research Association, American Psychological Association and National Council on Measurement in Education. Standards for Educational and Psychological Testing. 2014.
- Livingston, Samuel A. Test Reliability: Basic Concepts. ETS Research Memorandum RM-18-01, 2018.
- Spearman, Charles. “Correlation Calculated from Faulty Data.” British Journal of Psychology 3 (1910): 271–295.
- Brown, William. “Some Experimental Results in the Correlation of Mental Abilities.” British Journal of Psychology 3 (1910): 296–322.
- Cronbach, Lee J. “Coefficient Alpha and the Internal Structure of Tests.” Psychometrika 16 (1951): 297–334.
- Cronbach, Lee J., and Richard J. Shavelson. “My Current Thoughts on Coefficient Alpha and Successor Procedures.” Educational and Psychological Measurement 64 (2004): 391–418.
- Cronbach, Lee J., Nageswari Rajaratnam and Goldine C. Gleser. “Theory of Generalizability: A Liberalization of Reliability Theory.” British Journal of Statistical Psychology 16 (1963): 137–163.
- Cronbach, Lee J., Goldine C. Gleser, Harinder Nanda and Nageswari Rajaratnam. The Dependability of Behavioral Measurements. 1972.
- Lord, Frederic M., and Melvin R. Novick. Statistical Theories of Mental Test Scores. 1968.
- Brennan, Robert L. Generalizability Theory. 2001.
- Revelle, William, and David M. Condon. “Reliability from Alpha to Omega: A Tutorial.” Psychological Assessment 31 (2019): 1395–1411.
- Sijtsma, Klaas. “On the Use, the Misuse, and the Very Limited Usefulness of Cronbach's Alpha.” Psychometrika 74 (2009): 107–120.
- McNeish, Daniel. “Thanks Coefficient Alpha, We'll Take It from Here.” Psychological Methods 23 (2018): 412–433.
- Raykov, Tenko, and George A. Marcoulides. “Thanks Coefficient Alpha, We Still Need You!” Educational and Psychological Measurement 79 (2019): 200–210.
- Hedge, Craig, Georgina Powell and Petroc Sumner. “The Reliability Paradox: Why Robust Cognitive Tasks Do Not Produce Reliable Individual Differences.” Behavior Research Methods 50 (2018): 1166–1186.
- Cohen, Jacob. “A Coefficient of Agreement for Nominal Scales.” Educational and Psychological Measurement 20 (1960): 37–46.
- Shrout, Patrick E., and Joseph L. Fleiss. “Intraclass Correlations: Uses in Assessing Rater Reliability.” Psychological Bulletin 86 (1979): 420–428.
- National Academies of Sciences, Engineering, and Medicine. “Overview of Psychological Testing.” 2015.
- International Test Commission and Association of Test Publishers. Guidelines for Technology-Based Assessment. Version 1.1, 2025.
- Society for Industrial and Organizational Psychology. Principles for the Validation and Use of Personnel Selection Procedures. 5th ed., 2018.
- Career Industry Council of Australia. Professional Standards for Australian Career Development Practitioners. 5th ed., 2026.

