The two gates: reliability and validity
Psychometrics is the science of measuring things that have no ruler, traits, attitudes, aptitudes. Before it trusts any score, it demands two proofs. The first is reliability: if you retake the test tomorrow, under the same conditions, is the result similar? The second is validity: does the number really measure the trait it claims, and does that number predict anything outside the test itself? The two are independent. A bathroom scale stuck 5 kg above true weight is perfectly reliable, it always gives the same reading, and completely invalid. A magazine quiz usually fails both gates at once, and still hands back a result that feels made for you.
- Test-retest reliability
- Consistency of a score across two administrations separated in time. Measured by the correlation between the first and second answer; low correlation means the result depends on the day.
- Internal consistency (Cronbach's alpha)
- How much the items of a single scale agree with each other. Useful but limited: it rises artificially with more items and does not guarantee the scale measures one thing, a high alpha is no proof of unidimensionality or validity.
- Construct validity
- Whether the test actually measures the theoretical concept it claims to (the "construct"), rather than something else correlated with it. The hardest gate, and the most important.
- Predictive validity
- Whether the score predicts a relevant external outcome, performance, health, future behavior. Without it, the test only describes itself.
- Convergent and discriminant validity
- Convergent: the scale correlates with other measures of the same trait. Discriminant: it does not correlate with measures of different traits. A good instrument needs both.
Big Five: the model science prefers
The Big Five was not invented by a theory, it emerged from the data. The lexical hypothesis starts from a simple idea: if a trait truly matters to human life, languages will have words for it. From there, decades of factor analysis distilled the vocabulary of personality until five dimensions kept reappearing sample after sample, language after language. The acronym is OCEAN: Openness, Conscientiousness, Extraversion, Agreeableness and Neuroticism, the last often shown by its opposite pole, Emotional Stability.
- 1936Allport & Odbert
Pull roughly 4,500 trait adjectives from the English dictionary, the field's lexical starting point.
- 1940sCattell and 16 factors
Raymond Cattell reduces the list to about 16 primary factors, the basis of the 16PF questionnaire.
- 1961Tupes & Christal
Reanalyzing several datasets, they consistently find five factors, not sixteen.
- 1980sGoldberg coins "Big Five"
Lewis Goldberg replicates the five-factor structure across languages and names the model.
- 1985–1992Costa & McCrae, NEO-PI-R
Paul Costa and Robert McCrae formalize the Five-Factor Model with facets measured by the NEO-PI-R inventory.
The structural difference from type models is decisive: each factor is a continuous dimension, not a box. You are not extraverted or introverted; you have more or less extraversion than average, along a spectrum. The Big Five quiz returns exactly that, five bars, not a single label.
| Factor | What it measures | High pole | Low pole |
|---|---|---|---|
| Openness | Curiosity, imagination, appetite for novelty | Inventive, open to the unusual | Practical, conventional |
| Conscientiousness | Organization, discipline, goal-orientation | Planful, dependable | Spontaneous, disorganized |
| Extraversion | Social energy, assertiveness, stimulation-seeking | Sociable, energetic | Reserved, reflective |
| Agreeableness | Cooperation, empathy, trust in people | Compassionate, cooperative | Skeptical, competitive |
| Neuroticism | Tendency toward negative emotion and stress reactivity | Sensitive, anxious | Calm, resilient |
What the Big Five actually predicts
Predictive validity is where the Big Five pulls ahead of its rivals, and where honesty matters. The classic meta-analysis by Barrick and Mount (1991), pooling dozens of studies, showed that Conscientiousness predicts job performance consistently across nearly every occupation. It is the field's most replicated finding. But the effect size is modest: the corrected validity sits around 0.2. That is not a technical footnote, it is the whole point.
A correlation of 0.2 means the trait explains about 4% of the variance in performance (0.2² ≈ 0.04). It is a real, replicable relationship, but 96% of the outcome rides on everything else: ability, context, opportunity, luck. If that "how much a variable explains" logic feels new, the correlation and R² guide breaks down exactly how a coefficient becomes a share of variance. The practical lesson: a well-measured trait tilts the odds; it does not decide the outcome.
The other factor with well-documented outcomes is Neuroticism. Higher levels are prospectively associated with greater risk of anxiety and mood disorders, enough that the trait itself is treated as a public-health marker in the literature (Lahey, 2009). Again: an association, at the population level, with a moderate effect. No personality score diagnoses an individual, just as no health calculator diagnoses on its own, it flags, it does not sentence.
MBTI and the problem of cutting the ruler in half
The MBTI comes from a different lineage. It starts from Carl Jung's theory of psychological types and was developed, across the 1940s and 1950s, by Katharine Cook Briggs and her daughter Isabel Briggs Myers, neither of them a psychologist by training. The instrument sorts a person into four dichotomies: Extraversion/Introversion (E/I), Sensing/Intuition (S/N), Thinking/Feeling (T/F) and Judging/Perceiving (J/P), combined into 16 types like "INTJ" or "ESFP". The 16-types quiz follows that format, it is engaging, and that is exactly where the trap lives.
Big Five (dimensional)
- Five continuous scores: you sit somewhere along each ruler.
- Derived empirically from language and factor analysis.
- Documented predictive validity, if modest.
- Stable on retest because it forces no midpoint cut.
MBTI (typological)
- Four letters: you are labeled on one side of each dichotomy.
- Derived from a theory (Jung), not from the data.
- Weak, contested predictive validity.
- Unstable: the type flips easily on retest.
The core criticism is not ideological, it is statistical. The dimensions the MBTI measures are continuous, most people sit in the middle, not at the extremes. When you measure extraversion across a large sample, the distribution has a single central peak that falls off on both sides: a bell shape, not two separate humps. The chart below illustrates that expected shape.
View the data
| x | Value |
|---|---|
| -3 | 0.4 |
| -2.5 | 1.8 |
| -2 | 5.4 |
| -1.5 | 13 |
| -1 | 24.2 |
| -0.5 | 35.2 |
| 0 | 39.9 |
| 0.5 | 35.2 |
| 1 | 24.2 |
| 1.5 | 13 |
| 2 | 5.4 |
| 2.5 | 1.8 |
| 3 | 0.4 |
Now look at the cost of dichotomizing. Picture two people with extraversion scores of 51% and 49%, all but identical. The MBTI calls the first "E" and the second "I", as if a chasm ran between them. Because so many people sit near that cut, small swings in mood or in how they read a question are enough to flip the letter. The result is low test-retest reliability for the TYPE: on retaking the test after just five weeks, a substantial share of respondents get at least one different letter, in Pittenger's review, between 39% and 76% depending on the study, a figure often summarized as roughly half. An instrument that changes its answer that easily cannot be the basis of a serious decision.
There was no support for the view that the MBTI measures truly dichotomous preferences or qualitatively distinct types.
McCrae & Costa (1989), Journal of Personality
The Barnum effect: why every profile feels spot-on
In 1949, psychologist Bertram Forer gave his students a "personality test" and handed each one what he called an individual profile. In truth, they all received the exact same text, stitched together from generic horoscope and manual phrases. He asked them to rate its accuracy from 0 to 5. The average was 4.26, nearly perfect. People recognized as deeply their own a portrait written for the whole class. This is the Barnum effect (or Forer effect): the tendency to accept vague, universal descriptions as if they were tailor-made.
Read a "profile" and notice how it fits you
"You have a great need for other people to like and admire you. You tend to be critical of yourself. You have a great deal of unused capacity you have not turned to your advantage. While you have some personality weaknesses, you are generally able to compensate for them. At times you have serious doubts about whether you made the right decision. You prefer a certain amount of variety and feel hemmed in by restrictions."
Why it feels personal: each sentence is a near-universal truth (everyone second-guesses decisions), two-sided (it fits the organized and the disorganized alike) and flattering (it points at "unused capacity", never harsh flaws). You fill the gaps with your own life and conclude the text "nailed it". No information about you ever went in, it only came out.
The Barnum effect powers horoscopes, tarot readings and much of the type-test industry. It does not prove anyone is gullible, we are all susceptible, because the brain hunts for meaning. The defense is methodological: a result is worth something only if it distinguishes you from someone else. If the description would fit almost anyone equally well, it measured nothing. It is the same reasoning we apply when explaining how a birth chart is calculated: the astronomical precision of the math does not transfer validity to the generic interpretation that follows.
DISC, Enneagram and how to read any test
Two other popular models deserve honesty. DISC descends from William Moulton Marston's 1928 book Emotions of Normal People; Marston himself never built a test, the DISC questionnaires came decades later. It is widely used in organizations, but its predictive validity is limited and thinly demonstrated. The Enneagram sits even farther from psychometrics: it was elaborated by Oscar Ichazo and Claudio Naranjo between the 1950s and 1970s within a spiritual-development system. Its nine types did not come from data analysis, and literature reviews find no consistent factor structure and no established predictive validity.
- Look for reported reliability: a good test tells you how stable its results are on retest.
- Distrust types: continuous dimensions describe people better than fixed boxes.
- Distrust descriptions that would fit anyone, that is the Barnum effect at work.
- Remember a test is not a diagnosis: no questionnaire replaces a professional assessment.
- Never use a personality test for job screening without validation specific to that role.
Treated as a mirror, a way to look at your habits and reactions and talk about them, any of these tests is useful and enjoyable. Treated as an oracle or a résumé filter, it becomes a costly mistake anchored to an unstable label. The whole difference is the question you ask the result: "does this help me think about myself?" is a good question; "does this prove what I am?" is not.
Frequently asked questions
Which is more reliable, the Big Five or the MBTI?
Is the MBTI scientific?
What is the Barnum effect?
Can I use a personality test to hire?
Does the Enneagram have a scientific basis?
A personality test is worth what survives two gates: reliability and validity. The Big Five clears both reasonably well, it measures continuous dimensions and predicts real outcomes, if with modest effects. The MBTI cuts those dimensions in half and pays for it in instability; DISC and the Enneagram rest on weak or absent empirical support. None of that makes the quizzes worthless: as a mirror to think about yourself, they are excellent. As a diagnosis or a hiring filter, they are the very mistake psychometrics exists to prevent.
Sources & references
- Barrick & Mount (1991), The Big Five personality dimensions and job performance: a meta-analysis, Personnel Psychology
- McCrae & Costa (1989), Reinterpreting the Myers-Briggs Type Indicator from the perspective of the Five-Factor Model, Journal of Personality
- Forer (1949), The fallacy of personal validation: a classroom demonstration of gullibility
- Pittenger (1993), The utility of the Myers-Briggs Type Indicator, Review of Educational Research
- Lahey (2009), Public health significance of neuroticism, American Psychologist
- Goldberg (1993), The structure of phenotypic personality traits, American Psychologist
- Hook et al. (2021), The Enneagram: a systematic review of the literature, Journal of Clinical Psychology