In 1936, Gordon Allport and Henry Odbert published the result of an unusual trawl through Webster's New International Dictionary. They had looked for English words that could distinguish one person's conduct from another's: aloof, careful, generous, restless, vain. Their catalogue contained 17,953 terms. About a quarter described relatively stable personal traits. The rest included temporary states, evaluations, physical qualities and miscellaneous descriptions.
The list was a catalogue of what English speakers had found worth saying about one another. That map became raw material for a long research programme. Psychologists shortened the list, asked people to rate themselves and others, calculated which descriptions tended to travel together, argued over how many clusters to retain, and tried to reproduce the pattern in new samples.
Half a century later, five broad dimensions had become a shared language for personality research. Today they are usually called openness to experience, conscientiousness, extraversion, agreeableness and neuroticism. They are used to organise findings about health, relationships, learning and work. They also appear in hiring products, coaching reports and online quizzes, where a statistical summary can easily acquire the authority of a diagnosis.
The Big Five therefore has two histories. One is a scientific success: a field found a more cumulative way to describe personality. The other is a history of compression. Thousands of words became a few scores, and the scores began to stand in for the people they summarised. To understand what the model can tell us, we need to follow both histories.
The canonical construct map places personality beside interests, values, abilities, aptitudes and motivation without treating those labels as synonyms.
From words to a five-factor coordinate system
The model was assembled through collection, reduction, replication, argument and convergence. No single study discovered the finished Big Five.
1884 — Important differences may enter language
Lexical proposal. Francis Galton proposed examining dictionaries for words that describe differences between people.
1936 — 17,953 person-describing terms
Catalogue. Gordon Allport and Henry Odbert turned an English dictionary into a large research catalogue, including 4,504 relatively stable trait terms.
1949 — Sixteen traits and an early five-factor result
Competing reductions. Raymond Cattell's reductions supported the 16PF, while Donald Fiske recovered five broad factors when reanalysing ratings derived from Cattell's variables.
1961–1963 — Five recurrent factors appear again
Replication. Tupes and Christal found five factors across eight samples, followed by Warren Norman's widely read replication.
1968 — Situations return to the centre
Challenge. Walter Mischel challenged strong claims that broad traits could predict particular acts, sharpening the distinction between tendencies and behaviour in context.
1985–1992 — Big Five and Five-Factor Model traditions meet
Convergence. Goldberg's lexical programme and Costa and McCrae's questionnaire programme converged on five broad domains, although their routes, labels and lower-level structures were not identical.
2005–2019 — Replication and limits become visible
Cross-cultural test. Multi-country studies found recognisable patterns, while research among the Tsimané and in non-WEIRD samples exposed failures of structure, translation and survey validity.
The sequence tracks a research programme, not the discovery of five biological types. The five domains are continuous statistical dimensions whose meaning depends on instruments, raters, samples and use.
The wager hidden in the dictionary
Allport and Odbert were working within what later became known as the lexical hypothesis. Its modest version says that socially important differences between people tend to become encoded in language. If communities repeatedly need to distinguish the dependable from the careless, or the sociable from the withdrawn, words for those differences will accumulate. A careful study of personality language might therefore reveal the categories people use most consistently.
Francis Galton had proposed a version of this idea in 1884. German-language catalogues followed, most notably Franziska Baumgarten's 1933 classification of 1,093 terms. Allport and Odbert expanded the project in English. Their 17,953 terms included 4,504 placed in a first category of relatively stable traits. Even that reduced list was too large for practical research.
The hypothesis has a stronger version too: that the most important structure of personality itself is preserved in everyday words. That step is less secure. Language records what speakers notice and value. It may combine a person's behaviour with the observer's judgement of it. A word such as ambitious can describe effort, desire, status and approval at once. Dictionaries also preserve the history and power relations of their speech communities. A lexical study can show the structure of person-description without proving that the same structure is the causal architecture of a mind.
This distinction would shadow the entire enterprise. Were researchers finding human nature, the social perception of human nature, or a useful mixture of both?
A statistical machine for reducing words
The project required a way to turn many correlated descriptions into fewer dimensions. Factor analysis supplied it. If people rated as talkative also tended to be rated as outgoing, energetic and socially bold, a researcher could infer a broader factor beneath those observations. The factor is a statistical account of shared variation among the measurements.
Raymond Cattell reduced Allport and Odbert's trait terms through synonym grouping and expert judgement, eventually working with a much smaller set of variables. His analyses led him to 16 source traits and, in 1949, to the first edition of the 16 Personality Factor Questionnaire. Cattell's reduction made large-scale trait research possible. His exact solution proved harder for other researchers to reproduce.
Personality · Early lexical lineage
Three visible stages in a much larger research history.
The portraits mark an early proposal, a dictionary catalogue and a statistical reduction. Replication and naming continued through many researchers not pictured here.

Francis Galton
Proposed that socially important personality differences become encoded in language.
His measurement programme sat inside a wider eugenic project; the portrait is not celebratory.

Gordon Allport
With Henry Odbert, catalogued 17,953 person-describing terms.
Collection required judgement about which words counted and how they should be grouped.

Raymond Cattell
Reduced the catalogue into smaller variable sets for factor analysis.
His exact 16-factor solution proved difficult for other researchers to reproduce.
The lineage continues through Fiske, Tupes and Christal, Norman, Goldberg, Costa and McCrae. Portrait availability does not determine historical importance.
Portraits of Francis Galton, Gordon Allport and Raymond Cattell appear with their specific contribution and a limitation. A note continues the lineage through Fiske, Tupes and Christal, Norman, Goldberg, Costa and McCrae.
This is where statistical decisions become part of the history. Researchers choose the starting variables, the people who will be rated, the extraction method, the rotation of the factors and the point at which to stop. They must interpret and name the resulting clusters. Different defensible choices can produce different levels of description. A model with five broad domains may sit above one with 15, 16 or 30 narrower traits. A two-factor model can sit above the five. These can be nested maps drawn at different scales.
The first clear five-factor result appeared before the model had a name. Donald Fiske reanalysed ratings based on Cattell's variables in 1949 and found five broad factors across self-ratings and assessments by others. Ernest Tupes and Raymond Christal then compared eight samples of personality ratings for a United States Air Force report in 1961. Five recurrent factors appeared again: surgency, agreeableness, dependability, emotional stability and culture. Warren Norman published a widely read replication in 1963.
The Air Force report later acquired a romantic reputation as a discovery lost for 30 years. It was certainly less visible than a journal article, and its 1992 reprinting helped secure its place in the canon. Meanwhile Norman and other researchers discussed, tested and renamed the five-factor pattern. The history is one of intermittent attention and repeated recovery.
Evidence funnel · Collection to replication
Five broad domains emerged through repeated acts of reduction.
Words were collected, selected, grouped and analysed. The result is a coordinate system of broad dimensions—not five types of person.
Each broad domain contains narrower facets. None is a person type.
A funnel narrows 17,953 person-describing terms to 4,504 relatively stable trait terms, then through researcher selection and factor analysis to recurring broad domains. The five final domains expand into facets rather than person types.
Five broad dimensions
Lewis Goldberg revived the lexical programme during the late 1970s and 1980s. He used the phrase Big Five to stress the breadth of the factors: each gathered many narrower descriptions. His studies with large pools of English trait adjectives helped establish a recurring five-factor structure.
The familiar labels came from several traditions and still vary by instrument. A contemporary inventory might call them:
Extraversion, covering tendencies such as sociability, assertiveness and energy. Lower scores can reflect a preference for less social stimulation while remaining compatible with confidence and skill.
Agreeableness, covering compassion, trust, patience and a tendency to cooperate. The opposite pole includes scepticism, bluntness and competitiveness. Context decides whether those tendencies help or hinder.
Conscientiousness, covering organisation, dependability, persistence and self-control. High scores can support sustained work, although orderliness and industriousness are distinct facets and extreme control may carry costs.
Neuroticism, often relabelled negative emotionality or reversed as emotional stability, covering susceptibility to anxiety, sadness, irritability and stress. The label has historical baggage and belongs to personality measurement rather than clinical diagnosis.
Openness to experience, also called intellect, culture or open-mindedness in different lineages, covering imagination, intellectual curiosity, aesthetic sensitivity and receptiveness to novelty. The shifting names reveal that this fifth domain has been the hardest to define consistently.
These are continuous dimensions. A score places a respondent relative to a reference group on a particular instrument at a particular time. Everyone receives a position on every measured dimension.
The broad domains also contain facets. The 60-item Big Five Inventory–2, for example, measures three facets under each domain. Extraversion includes sociability, assertiveness and energy level. A person can be high on one and nearer the middle on another. The Revised NEO Personality Inventory uses six facets within each domain. Instruments may therefore agree at the broad level while differing in content and emphasis below it.
That hierarchy matters. Broad traits can predict a wide range of outcomes weakly, while a well-matched facet can predict a narrower outcome more precisely. Calling someone “highly conscientious” discards the difference between being organised, productive and responsible. Compression creates portability. It also removes information.
The second road was only partly independent
The Big Five is often said to have been discovered twice: first in dictionaries, then independently in questionnaires. The convergence is real, but the clean two-road story needs correction.
Paul Costa and Robert McCrae began with questionnaire research on three dimensions: neuroticism, extraversion and openness. Those initials gave the NEO inventory its name. As the five-factor literature gained force, they added agreeableness and conscientiousness. The 1985 NEO Personality Inventory measured all five domains, and the 1992 NEO-PI-R supplied six facets for each.
Questionnaire studies provided important evidence beyond adjective lists. Similar broad dimensions appeared across different instruments and in ratings made by observers as well as selves. The traditions were also communicating: Costa and McCrae tested their work against the lexical literature.
There is also a terminological wrinkle. Historically, Big Five often referred to the lexical taxonomy associated with Goldberg, while Five-Factor Model referred to the questionnaire-based formulation associated with Costa and McCrae. The models divide some traits differently, especially openness or intellect. In current research the names are often used interchangeably. Five-Factor Theory, however, is a further claim: an explanatory theory developed by Costa and McCrae about how biologically based traits interact with adaptations and life experience. Taxonomy, instrument and causal theory are three distinct objects.
The achievement was convergence at a useful level of description. Researchers using different word sets, questionnaires, raters and samples repeatedly found a family resemblance among five broad domains. That was enough to give a fragmented field a common coordinate system.
What a personality score measures
Most Big Five instruments ask respondents how well a series of statements or adjectives describes them. Their answers are combined into scale scores, sometimes compared with norms. Other versions use ratings by friends, partners or colleagues. Good instruments are tested for internal consistency, retest stability, agreement between raters and relationships with relevant behaviour. The resulting evidence remains tied to context and use.
Self-report contains information no observer has: private worry, intention, effort and experience. It also contains self-presentation, memory limits and different interpretations of the same item. Some people agree readily with statements; others avoid extreme responses. In an applicant setting, respondents know which answers appear employable. Faking changes the meaning of the evidence.
Reference groups matter as well. One respondent answering “I am organised” may compare herself with her household; another may compare himself with colleagues in a tightly regulated profession. Research that explicitly changes the comparison group finds that scores can move. Translation adds another layer. A grammatically faithful item may carry a different social meaning, and equivalent-looking scales may not function equivalently across populations.
Observer reports solve some problems and create others. A colleague sees behaviour at work but not at home, and may confuse competence, likability and conformity. Agreement between self and observer tends to be higher for visible traits such as extraversion than for private emotional experience. Disagreement may show that people behave differently across roles or that each rater has access to a different part of the person.
A personality score is hard to interpret unless its report names the instrument, form, respondent, norm group and uncertainty. “Your personality is 82 per cent conscientious” is usually meaningless. A percentile locates a score within a chosen comparison sample; it cannot describe a percentage of behaviour or remove uncertainty from the estimate.
Measurement · Interpretation chain
A personality score is produced by a system—not extracted whole from a person.
Change the wording, rater, instrument, comparison group or context and the responsible interpretation can change too.
1 · What is observed
2 · How it is processed
3 · What it is compared with
Interpretation ± uncertainty
82nd percentile means higher than 82% of a chosen comparison sample. It does not mean “82% conscientious”.
A chain links item wording, response and rater, instrument and form, scoring rule, reference population and norm, context and measurement uncertainty to the final interpretation. A callout explains that a percentile is a position in a comparison sample, not a percentage of behaviour.
Traits are stable patterns and moving targets
Walter Mischel's 1968 book Personality and Assessment challenged the assumption that broad traits could predict how someone would act in a particular situation. A person may be talkative at dinner and silent in a meeting, careful with laboratory work and careless with laundry. The ensuing person–situation debate is sometimes narrated as a battle that traits eventually won. The more useful resolution retained both sides.
Situations strongly shape individual acts. Traits describe distributions and tendencies across time and contexts. A high extraversion score can help predict behaviour across many occasions, such as how often someone seeks social contact, feels energetic or takes an assertive role. The next five minutes still depend on goals, skills, incentives, identity, health and the opportunities a setting supplies.
Traits are stable in a second, relative sense. People tend to retain some of their rank ordering over time. A 2022 meta-analysis of longitudinal research found that rank-order stability rises through early life and plateaus in young adulthood. The same literature records average change and individual variation across adulthood.
Experience and deliberate effort can matter. In a randomised trial involving 1,523 adults who wanted to change aspects of their personalities, a three-month digital intervention produced changes in the intended directions compared with a control condition. Observers noticed smaller changes too. Follow-up analysis found that some apparent domain change was concentrated in particular facets, such as sociability or productiveness. The finding supports change within limits and timescales; it cannot promise personality redesign on demand.
Useful prediction and its limits
The Big Five earned its place because the scores relate, on average, to consequential outcomes. Conscientiousness is associated with academic and work performance. Extraversion matters more in some social and leadership settings. Emotional stability relates to wellbeing and resilience under stress. Different facets predict different outcomes, and relationships vary by task and context.
The size of those relationships is easy to exaggerate. A 2022 synthesis of 54 meta-analyses and more than half a million participants found the strongest average association with overall performance for conscientiousness, at about .19; the other domains had smaller average associations, and results shifted substantially across performance categories. A correlation of .19 can be useful in research or as one carefully validated input to a selection system. It leaves most individual variation unexplained. It cannot tell an employer which candidate will succeed or tell a student which occupation will provide a good life.
Work performance has many dimensions: task output, learning, cooperation, safety, creativity, leadership, withdrawal and avoidance of counterproductive behaviour. A trait that helps one criterion may be neutral or costly for another. High agreeableness may aid cooperation and complicate adversarial negotiation. Extraversion may help in a highly social role and supply little advantage during solitary analysis. Conscientiousness cannot compensate for absent knowledge, poor tools, discrimination or an impossible workload.
This is the point at which a descriptive model becomes a gate. Personality questionnaires are attractive to employers because they are cheap, scalable and appear objective. Their use in hiring carries a heavier burden than their use in voluntary reflection. Evidence has to connect the instrument with the particular job, population and decision, and show whether it contributes anything beyond work samples or structured interviews. The applicant's ability to understand and challenge the result becomes part of the fairness question. So do disability, language and cultural differences. When a vendor infers personality from video, voice or social media, even the construct being measured may be hidden inside an opaque model.
Even a valid group-level association can be used badly on an individual. Selection changes the stakes and encourages impression management. Cut scores can create false precision. Broad profiles can invite stereotypes about what a salesperson, engineer, nurse or leader ought to be. The scientific standing of the Big Five does not pass automatically to every product or use that borrows its labels.
How universal are the five?
Cross-cultural research supplies both encouraging replication and serious warning. Robert McCrae, Antonio Terracciano and collaborators found recognisable five-factor patterns in observer ratings collected across 50 cultures. Lexical studies in many languages also recover dimensions resembling several of the Big Five. This evidence weakens the claim that the model is merely a peculiarity of one American questionnaire.
Cross-cultural evidence leaves universality unsettled. Michael Gurven and colleagues did not recover the conventional five-factor structure when they administered a translated inventory among Tsimané forager-horticulturalists in Bolivia. Work across 23 low- and middle-income countries by Rachid Laajaj and a large team found that standard survey data often failed tests of validity, especially among respondents with less formal education. Enumerator effects and response styles explained part of the problem. Such findings may reveal limits in the model, limits in the survey method, or both.
That ambiguity is important. A test built for literate respondents who privately mark a rating scale may perform differently when questions are translated, read aloud by an interviewer or treated as a strange social exchange. Failure to reproduce the factor structure does not by itself prove that the population lacks the relevant traits. Success after imposing a familiar statistical structure does not prove that the traits carry the same meaning.
Measurement invariance is the demanding question beneath cross-group comparison: do the items and scales work in sufficiently similar ways for scores or averages to be compared? Without it, a difference attributed to personality may partly reflect translation, norms, response styles or different reference groups. The more consequential the comparison, the less acceptable it is to wave that uncertainty away.
The existence of HEXACO supplies a different challenge. Lexical studies in several languages support six broad domains, with honesty–humility emerging as a factor only partly represented in the Big Five. Other models use two higher-order factors, three broad dimensions or many narrower traits. Five remains an effective common language among several defensible resolutions of personality.
The threads, named
The thread of measurement is obvious. A vocabulary became variables; variables became scores; scores became infrastructure for research and decisions. Each transformation increased comparability while stripping away context. That trade will return in intelligence tests, interest inventories, performance ratings and algorithmic assessments.
The thread of status enters when a description acquires a preferred pole. A scale may be designed without declaring high or low morally better, yet workplaces often reward emotional stability, conscientiousness, assertiveness or sociability as if more were always superior. The language of neutral traits can conceal a local ideal of the “good worker”.
The thread of access and mobility begins at the gate. A test can widen access when it replaces patronage or an unstructured judgement with a validated and auditable measure. It can narrow access when weak prediction, cultural non-equivalence or opaque scoring is given institutional force. Measurement is never merely descriptive once it helps decide who is admitted.
What this history teaches us
That the Big Five is a taxonomy before it is an explanation. It organises recurring patterns in personality descriptions; causal explanation requires additional biological, developmental and social evidence.
That five is a useful scale. Broader and narrower models answer different questions. The right level depends on whether we need a common overview, a specific prediction or an account of an individual life.
That a score belongs to a measurement occasion. Instrument, rater, language, reference group and setting shape what it means. Reliability and validity belong to evidence gathered for a particular use.
That tendencies can matter without becoming destinies. Traits show relative stability and modest predictive power. People also change, choose, learn and respond to conditions. A profile cannot replace an account of skills, values, opportunity or constraint.
That the burden of proof rises when measurement becomes a gate. An employment screen carries greater risk than a voluntary prompt for self-reflection, even when both display the same five labels. Scientific language must be joined to job-relevant evidence, fairness and the right to contest a consequential decision.
Allport and Odbert's catalogue left room for nearly 18,000 ways of describing a person. The Big Five gave psychology a way to see order in that abundance. Its success lies in the view from high above: five broad ridges that appear across many maps. A human life happens lower down, among the paths, weather and choices the large map cannot show.
The next question is therefore harder than “Which five scores describe me?” It is whether any test, however well constructed, can know enough of a person to decide where they belong.
Notes on the evidence
The historical spine comes from Allport and Odbert's 1936 monograph, Donald Fiske's 1949 analysis, Tupes and Christal's 1961 Air Force report, Warren Norman's 1963 replication, Lewis Goldberg's historical and lexical work, and McCrae and John's 1992 review. Costa and McCrae's own 1995 account shows that their programme began with neuroticism, extraversion and openness, then incorporated agreeableness and conscientiousness; the familiar “independent second discovery” is therefore too neat. Modern measurement is represented by Soto and John's BFI-2. Evidence on development and prediction comes from Bleidorn and colleagues' 2022 longitudinal meta-analysis, Stieger and colleagues' randomised intervention, and Zell and Lesick's synthesis of performance research. Cross-cultural claims remain contested: multi-country replications coexist with failures of structure or survey validity among the Tsimané and in diverse low- and middle-income samples. The chapter treats that disagreement as a question about both personality structure and the instruments used to measure it.
Sources and further reading
- Francis Galton, “Measurement of Character,” Fortnightly Review (1884). PDF
- Gordon W. Allport & Henry S. Odbert, “Trait-Names: A Psycho-Lexical Study,” Psychological Monographs 47(1) (1936). Record
- Oliver P. John, Alois Angleitner & Fritz Ostendorf, “The Lexical Approach to Personality: A Historical Review,” European Journal of Personality 2 (1988), 171–203. DOI
- Donald W. Fiske, “Consistency of the Factorial Structures of Personality Ratings from Different Sources,” Journal of Abnormal and Social Psychology 44 (1949), 329–344. Record
- Ernest C. Tupes & Raymond E. Christal, “Recurrent Personality Factors Based on Trait Ratings,” USAF ASD-TR-61-97 (1961), reprinted in Journal of Personality 60 (1992), 225–251. Record
- Warren T. Norman, “Toward an Adequate Taxonomy of Personality Attributes,” Journal of Abnormal and Social Psychology 66 (1963), 574–583. DOI
- Walter Mischel, Personality and Assessment (Wiley, 1968).
- J. M. Digman, “Personality Structure: Emergence of the Five-Factor Model,” Annual Review of Psychology 41 (1990), 417–440. DOI
- Lewis R. Goldberg, “An Alternative ‘Description of Personality’: The Big-Five Factor Structure,” Journal of Personality and Social Psychology 59 (1990), 1216–1229. PDF
- Lewis R. Goldberg, “The Structure of Phenotypic Personality Traits,” American Psychologist 48 (1993), 26–34. DOI
- Robert R. McCrae & Oliver P. John, “An Introduction to the Five-Factor Model and Its Applications,” Journal of Personality 60 (1992), 175–215. DOI
- Paul T. Costa Jr & Robert R. McCrae, “Domains and Facets: Hierarchical Personality Assessment Using the Revised NEO Personality Inventory,” Journal of Personality Assessment 64 (1995), 21–50. PDF
- Christopher J. Soto & Oliver P. John, “The Next Big Five Inventory (BFI-2),” Journal of Personality and Social Psychology 113 (2017), 117–143. DOI
- Jack Block, “A Contrarian View of the Five-Factor Approach to Personality Description,” Psychological Bulletin 117 (1995), 187–215. Record
- Michael C. Ashton & Kibeom Lee, “Empirical, Theoretical, and Practical Advantages of the HEXACO Model,” Personality and Social Psychology Review 11 (2007), 150–166. DOI
- Robert R. McCrae, Antonio Terracciano et al., “Universal Features of Personality Traits from the Observer's Perspective,” Journal of Personality and Social Psychology 88 (2005), 547–561. DOI
- Michael Gurven et al., “How Universal Is the Big Five?”, Journal of Personality and Social Psychology 104 (2013), 354–370. DOI
- Rachid Laajaj et al., “Challenges to Capture the Big Five Personality Traits in Non-WEIRD Populations,” Science Advances 5 (2019), eaaw5226. DOI
- Wiebke Bleidorn et al., “Personality Stability and Change: A Meta-Analysis of Longitudinal Studies,” Psychological Bulletin 148 (2022), 588–619. DOI
- Mirjam Stieger et al., “Changing Personality Traits with the Help of a Digital Personality Change Intervention,” PNAS 118 (2021), e2017548118. DOI
- Ethan Zell & Tara L. Lesick, “Big Five Personality Traits and Performance: A Quantitative Synthesis of 50+ Meta-Analyses,” Journal of Personality 90 (2022), 559–573. DOI
- Neal Schmitt, “Personality and Cognitive Ability as Predictors of Effective Performance at Work,” Annual Review of Organizational Psychology and Organizational Behavior 1 (2014), 45–65. DOI
- Madeline R. Lenhausen, Wiebke Bleidorn & Christopher J. Hopwood, “Effects of Reference Group Instructions on Big Five Trait Scores,” Assessment 31 (2024), 669–677. DOI

