Psychometrics

Big Five vs. MBTI: what psychometrics actually says

Every personality test hands back an answer that feels accurate. The trouble is that "feeling accurate" is evidence of nothing, it is exactly the symptom psychometrics learned to distrust. What separates a serious instrument from a magazine quiz is not the number of questions or the polish of the result, but two gates almost nobody sees: does the test measure consistently (reliability)? And does it measure what it claims, predicting something in the real world (validity)? This guide uses those two criteria to compare the Big Five, the best-supported model, with the MBTI, the most popular one, and explains why this site's quizzes are great for self-reflection and useless as a diagnosis.

J-Kit13 min readIntermediate
  • Psychometrics
  • Personality
  • Big Five
  • MBTI

Key takeaways

  • A psychometric instrument passes two gates, reliability and validity; a magazine quiz passes neither.
  • The Big Five is dimensional and holds the best empirical support. It predicts real outcomes, but with modest, non-deterministic effects.
  • Typologies like the MBTI slice continuous dimensions in half; a substantial share of people change type when they retake the test within weeks.
  • The Barnum effect explains why any generic profile feels tailor-made: vague descriptions fit almost everyone.

The two gates: reliability and validity

Psychometrics is the science of measuring things that have no ruler, traits, attitudes, aptitudes. Before it trusts any score, it demands two proofs. The first is reliability: if you retake the test tomorrow, under the same conditions, is the result similar? The second is validity: does the number really measure the trait it claims, and does that number predict anything outside the test itself? The two are independent. A bathroom scale stuck 5 kg above true weight is perfectly reliable, it always gives the same reading, and completely invalid. A magazine quiz usually fails both gates at once, and still hands back a result that feels made for you.

Test-retest reliability
Consistency of a score across two administrations separated in time. Measured by the correlation between the first and second answer; low correlation means the result depends on the day.
Internal consistency (Cronbach's alpha)
How much the items of a single scale agree with each other. Useful but limited: it rises artificially with more items and does not guarantee the scale measures one thing, a high alpha is no proof of unidimensionality or validity.
Construct validity
Whether the test actually measures the theoretical concept it claims to (the "construct"), rather than something else correlated with it. The hardest gate, and the most important.
Predictive validity
Whether the score predicts a relevant external outcome, performance, health, future behavior. Without it, the test only describes itself.
Convergent and discriminant validity
Convergent: the scale correlates with other measures of the same trait. Discriminant: it does not correlate with measures of different traits. A good instrument needs both.

Big Five: the model science prefers

The Big Five was not invented by a theory, it emerged from the data. The lexical hypothesis starts from a simple idea: if a trait truly matters to human life, languages will have words for it. From there, decades of factor analysis distilled the vocabulary of personality until five dimensions kept reappearing sample after sample, language after language. The acronym is OCEAN: Openness, Conscientiousness, Extraversion, Agreeableness and Neuroticism, the last often shown by its opposite pole, Emotional Stability.

  1. 1936Allport & Odbert

    Pull roughly 4,500 trait adjectives from the English dictionary, the field's lexical starting point.

  2. 1940sCattell and 16 factors

    Raymond Cattell reduces the list to about 16 primary factors, the basis of the 16PF questionnaire.

  3. 1961Tupes & Christal

    Reanalyzing several datasets, they consistently find five factors, not sixteen.

  4. 1980sGoldberg coins "Big Five"

    Lewis Goldberg replicates the five-factor structure across languages and names the model.

  5. 1985–1992Costa & McCrae, NEO-PI-R

    Paul Costa and Robert McCrae formalize the Five-Factor Model with facets measured by the NEO-PI-R inventory.

The structural difference from type models is decisive: each factor is a continuous dimension, not a box. You are not extraverted or introverted; you have more or less extraversion than average, along a spectrum. The Big Five quiz returns exactly that, five bars, not a single label.

The five factors (OCEAN) and what each measures.
FactorWhat it measuresHigh poleLow pole
OpennessCuriosity, imagination, appetite for noveltyInventive, open to the unusualPractical, conventional
ConscientiousnessOrganization, discipline, goal-orientationPlanful, dependableSpontaneous, disorganized
ExtraversionSocial energy, assertiveness, stimulation-seekingSociable, energeticReserved, reflective
AgreeablenessCooperation, empathy, trust in peopleCompassionate, cooperativeSkeptical, competitive
NeuroticismTendency toward negative emotion and stress reactivitySensitive, anxiousCalm, resilient

What the Big Five actually predicts

Predictive validity is where the Big Five pulls ahead of its rivals, and where honesty matters. The classic meta-analysis by Barrick and Mount (1991), pooling dozens of studies, showed that Conscientiousness predicts job performance consistently across nearly every occupation. It is the field's most replicated finding. But the effect size is modest: the corrected validity sits around 0.2. That is not a technical footnote, it is the whole point.

A correlation of 0.2 means the trait explains about 4% of the variance in performance (0.2² ≈ 0.04). It is a real, replicable relationship, but 96% of the outcome rides on everything else: ability, context, opportunity, luck. If that "how much a variable explains" logic feels new, the correlation and R² guide breaks down exactly how a coefficient becomes a share of variance. The practical lesson: a well-measured trait tilts the odds; it does not decide the outcome.

≈ 0.2corrected validity of Conscientiousness for performance
≈ 4%of the variance in performance it explains
5factors that recur across languages and cultures

The other factor with well-documented outcomes is Neuroticism. Higher levels are prospectively associated with greater risk of anxiety and mood disorders, enough that the trait itself is treated as a public-health marker in the literature (Lahey, 2009). Again: an association, at the population level, with a moderate effect. No personality score diagnoses an individual, just as no health calculator diagnoses on its own, it flags, it does not sentence.

MBTI and the problem of cutting the ruler in half

The MBTI comes from a different lineage. It starts from Carl Jung's theory of psychological types and was developed, across the 1940s and 1950s, by Katharine Cook Briggs and her daughter Isabel Briggs Myers, neither of them a psychologist by training. The instrument sorts a person into four dichotomies: Extraversion/Introversion (E/I), Sensing/Intuition (S/N), Thinking/Feeling (T/F) and Judging/Perceiving (J/P), combined into 16 types like "INTJ" or "ESFP". The 16-types quiz follows that format, it is engaging, and that is exactly where the trap lives.

Big Five (dimensional)

  • Five continuous scores: you sit somewhere along each ruler.
  • Derived empirically from language and factor analysis.
  • Documented predictive validity, if modest.
  • Stable on retest because it forces no midpoint cut.

MBTI (typological)

  • Four letters: you are labeled on one side of each dichotomy.
  • Derived from a theory (Jung), not from the data.
  • Weak, contested predictive validity.
  • Unstable: the type flips easily on retest.

The core criticism is not ideological, it is statistical. The dimensions the MBTI measures are continuous, most people sit in the middle, not at the extremes. When you measure extraversion across a large sample, the distribution has a single central peak that falls off on both sides: a bell shape, not two separate humps. The chart below illustrates that expected shape.

09.9819.9529.9239.9-303Standardized score (SDs from the mean)Relative density
Expected shape of a continuous trait like extraversion: unimodal, peaking in the center. Illustrative normal curve (not data from a real sample). The type cut lands right at the densest point, splitting near-identical people into "E" and "I".
View the data
xValue
-30.4
-2.51.8
-25.4
-1.513
-124.2
-0.535.2
039.9
0.535.2
124.2
1.513
25.4
2.51.8
30.4

Now look at the cost of dichotomizing. Picture two people with extraversion scores of 51% and 49%, all but identical. The MBTI calls the first "E" and the second "I", as if a chasm ran between them. Because so many people sit near that cut, small swings in mood or in how they read a question are enough to flip the letter. The result is low test-retest reliability for the TYPE: on retaking the test after just five weeks, a substantial share of respondents get at least one different letter, in Pittenger's review, between 39% and 76% depending on the study, a figure often summarized as roughly half. An instrument that changes its answer that easily cannot be the basis of a serious decision.

There was no support for the view that the MBTI measures truly dichotomous preferences or qualitatively distinct types.

McCrae & Costa (1989), Journal of Personality

The Barnum effect: why every profile feels spot-on

In 1949, psychologist Bertram Forer gave his students a "personality test" and handed each one what he called an individual profile. In truth, they all received the exact same text, stitched together from generic horoscope and manual phrases. He asked them to rate its accuracy from 0 to 5. The average was 4.26, nearly perfect. People recognized as deeply their own a portrait written for the whole class. This is the Barnum effect (or Forer effect): the tendency to accept vague, universal descriptions as if they were tailor-made.

Read a "profile" and notice how it fits you

"You have a great need for other people to like and admire you. You tend to be critical of yourself. You have a great deal of unused capacity you have not turned to your advantage. While you have some personality weaknesses, you are generally able to compensate for them. At times you have serious doubts about whether you made the right decision. You prefer a certain amount of variety and feel hemmed in by restrictions."

Why it feels personal: each sentence is a near-universal truth (everyone second-guesses decisions), two-sided (it fits the organized and the disorganized alike) and flattering (it points at "unused capacity", never harsh flaws). You fill the gaps with your own life and conclude the text "nailed it". No information about you ever went in, it only came out.

The Barnum effect powers horoscopes, tarot readings and much of the type-test industry. It does not prove anyone is gullible, we are all susceptible, because the brain hunts for meaning. The defense is methodological: a result is worth something only if it distinguishes you from someone else. If the description would fit almost anyone equally well, it measured nothing. It is the same reasoning we apply when explaining how a birth chart is calculated: the astronomical precision of the math does not transfer validity to the generic interpretation that follows.

DISC, Enneagram and how to read any test

Two other popular models deserve honesty. DISC descends from William Moulton Marston's 1928 book Emotions of Normal People; Marston himself never built a test, the DISC questionnaires came decades later. It is widely used in organizations, but its predictive validity is limited and thinly demonstrated. The Enneagram sits even farther from psychometrics: it was elaborated by Oscar Ichazo and Claudio Naranjo between the 1950s and 1970s within a spiritual-development system. Its nine types did not come from data analysis, and literature reviews find no consistent factor structure and no established predictive validity.

  • Look for reported reliability: a good test tells you how stable its results are on retest.
  • Distrust types: continuous dimensions describe people better than fixed boxes.
  • Distrust descriptions that would fit anyone, that is the Barnum effect at work.
  • Remember a test is not a diagnosis: no questionnaire replaces a professional assessment.
  • Never use a personality test for job screening without validation specific to that role.

Treated as a mirror, a way to look at your habits and reactions and talk about them, any of these tests is useful and enjoyable. Treated as an oracle or a résumé filter, it becomes a costly mistake anchored to an unstable label. The whole difference is the question you ask the result: "does this help me think about myself?" is a good question; "does this prove what I am?" is not.

Frequently asked questions

Which is more reliable, the Big Five or the MBTI?
The Big Five. Because it measures continuous dimensions, its scores are stable on retest and have documented (if modest) predictive validity. The MBTI pushes a person to one side of each dichotomy, and since most people sit near the middle, the type flips easily, a substantial share of respondents change at least one letter when they retake the test within weeks.
Is the MBTI scientific?
It is derived from a theory (Jung's types), not from data, and it fails the central psychometric criteria: the type's test-retest reliability is low and its predictive validity is weak. McCrae and Costa (1989) showed there is no support for treating its preferences as truly dichotomous. It is popular and engaging, but it is not a validated diagnostic instrument.
What is the Barnum effect?
It is the tendency to accept vague, universal descriptions as if they were tailored to you. In Forer's 1949 experiment, students rated a generic profile, identical for everyone, with an average accuracy of 4.26 out of 5. Because the sentences fit almost anyone, the sense of recognition proves no accuracy at all, it is the mechanism behind horoscopes and many type tests.
Can I use a personality test to hire?
Not with a type test like MBTI, DISC or Enneagram: they have no demonstrated validity for predicting performance and classify people unstably. Even the Big Five, which does predict performance, has a modest effect (explaining about 4% of the variance) and should only enter a process with validation specific to the role. A label should never, on its own, decide a hire.
Does the Enneagram have a scientific basis?
Not an established empirical one. Its nine types came from a spiritual-development system (Ichazo and Naranjo, 1950s–1970s), not from data analysis. Literature reviews find no consistent factor structure and no reliable predictive validity. It can be an interesting lens for talking about motivations, but it is not a validated measurement instrument.

A personality test is worth what survives two gates: reliability and validity. The Big Five clears both reasonably well, it measures continuous dimensions and predicts real outcomes, if with modest effects. The MBTI cuts those dimensions in half and pays for it in instability; DISC and the Enneagram rest on weak or absent empirical support. None of that makes the quizzes worthless: as a mirror to think about yourself, they are excellent. As a diagnosis or a hiring filter, they are the very mistake psychometrics exists to prevent.

Sources & references

  1. Barrick & Mount (1991), The Big Five personality dimensions and job performance: a meta-analysis, Personnel Psychology
  2. McCrae & Costa (1989), Reinterpreting the Myers-Briggs Type Indicator from the perspective of the Five-Factor Model, Journal of Personality
  3. Forer (1949), The fallacy of personal validation: a classroom demonstration of gullibility
  4. Pittenger (1993), The utility of the Myers-Briggs Type Indicator, Review of Educational Research
  5. Lahey (2009), Public health significance of neuroticism, American Psychologist
  6. Goldberg (1993), The structure of phenotypic personality traits, American Psychologist
  7. Hook et al. (2021), The Enneagram: a systematic review of the literature, Journal of Clinical Psychology