IPIP-NEO

Methodology

How this test is built and scored

Everything here can be checked. The items are public domain, the norms come from openly published datasets, and the arithmetic between your answers and your percentile is described below step by step.

  • Public-domain items
  • Norms from open data
  • Scoring described in full
protocols behind the percentile tables
521,748
facets under the five domains
30
comparison groups per version
12

From an answer to a percentile

Four steps, no black box. Nothing is weighted by a secret formula and nothing is adjusted by hand.

  1. You answer on a five point scale

    Each statement gets one point for strong disagreement through five for strong agreement. Reverse worded items are recoded, so that on every item a higher number means more of the trait.

  2. Answers add up into a facet

    A facet is the plain sum of its items: ten of them in the 300 version, four in the 120. No factor weights, no rescaling.

  3. Six facets make a domain

    Each of the five domains is the sum of its six facets. That is why a domain can sit in the middle while the facets underneath it pull in opposite directions.

  4. The sum is placed in a comparison group

    Your raw sum is located inside the distribution of people of your sex and age band. The percentile is that position: the share of the group scoring below you.

The percentile is empirical, taken straight from the distribution rather than from a normal curve fitted over it. Ties are split in half, which is the standard convention and keeps a very common score from claiming the whole step it sits on.

Where the instrument comes from

The International Personality Item Pool was assembled by Lewis Goldberg and colleagues as a public-domain answer to commercial personality inventories: a large bank of short statements that anyone may use, adapt and publish scales from.

The IPIP-NEO inventories were built on that pool by John A. Johnson. The 300 item version measures the five broad domains and thirty facets that the Big Five literature converged on; the 120 item version is a published short form of the same instrument, made of the four best markers of each facet.

This site implements both, unchanged. The English statements are reproduced word for word, in the original order of the published key, because the norms are only valid for the original wording.

The item pool and the inventories built on it are in the public domain. Everything else here, the explanatory texts, the interpretations, the illustrations and the code, is our own work.

Whose norms, and on what sample

Percentiles compare you with real people rather than with a theoretical average. Our tables are computed from the response datasets Johnson published openly, one for each version, filtered to complete protocols with a stated sex and an age between 16 and 90.

Each version gets its own tables, split by sex and by four age bands, plus wider fallback groups for cells that would otherwise be thin. That matters more than it sounds: the same raw score sits at a different percentile for a twenty year old man and a fifty five year old woman, and averaging them together would hide a real difference rather than remove it.

If you skip sex or age, you are compared against the whole sample of that version. The comparison group is always named on the result page, so you can see what your number is a percentile of.

These norms describe people who took Johnson's inventory online, an internet sample that skews younger and more literate than a census would. It is the best openly published reference for this instrument, and it is not a general population.

low: bottom 30%
average: middle 40%
high: top 30%
Splitting at the median would call half of all people high and the other half low, which tells you nothing. A band of forty percent in the middle says the honest thing instead: for most traits, most people are unremarkable, and only the edges carry a message.

How reliable the scales are

Reliability here is internal consistency, measured as Cronbach's alpha and computed by us from the same datasets. It answers a narrow question: do the items inside one scale move together, or are they measuring different things under one label?

The five domains land between about 0.91 and 0.96 in the long version and between about 0.82 and 0.91 in the short one, which psychometrics reads as strong. Facets are shorter by design, so they sit lower: on average about 0.84 with ten items and about 0.76 with four.

Practically this means the domains are stable enough to be read as they stand, and a single facet deserves more caution, especially on the short version and especially when it sits near a band boundary.

IPIP-NEO-300IPIP-NEO-120

What the scores are evidence of

The five factor structure did not come from a theory about human nature. It came from the lexical hypothesis: if a difference between people matters to daily life, languages end up with words for it, and when those words are collected and factor analysed, the same five broad groupings keep appearing across samples and languages.

Facets are the layer below, and they exist because domains turned out to be too coarse. Two people with the same Conscientiousness can differ completely once you separate orderliness from self discipline from dutifulness, and those distinctions predict different things.

Big Five scores correlate with outcomes that people care about, in the modest way personality measures usually do: they shift the odds rather than settle anything. They are much better at describing tendencies across many situations than at predicting what one person will do on one occasion.

IPIP-NEO is not the NEO PI-R

The two share a map: five domains, six facets each, and largely the same facet names, because the IPIP scales were written to measure the same constructs. They do not share items, publisher or norms.

The NEO PI-R is a commercial inventory with its own manual, its own standardisation samples and its own licensing. The IPIP-NEO is an independent public-domain instrument that reaches for the same targets. Scores from one cannot be read against the tables of the other, and a result here is not a NEO PI-R result.

We say this on every page for a reason: on the internet the two are constantly confused, usually to make a free test sound more official than it is.

Honest limits

It is a self report. You are describing yourself, and self description is shaped by mood, by self knowledge and by what feels acceptable to admit. That is why the test is useless for selection: as soon as something depends on the answers, they change.

Comparisons across cultures are the weakest ground. Norms come from one internet sample, mostly English speaking, and traits are expressed differently in different places, so a percentile is better read as your position among people who answered this inventory than as a statement about your country.

Our translations of the items into the other languages are ours and unofficial. Scoring always runs on the English canon rather than on a translation, but a translated statement can still land slightly differently, and that is a real limitation of a multilingual free test.

  • one sitting is a snapshot: traits are stable across years, not fixed forever;
  • a single facet, especially on the short version, carries more uncertainty than a domain;
  • this is not a clinical instrument and says nothing about disorders;
  • the norms are an internet sample, not a census.

Where to check this

The materials this site is built on are public. If a number here matters to you, read it at the source rather than taking our word for it.

Read the method, then take the test

Both versions are free and both are scored exactly as described above. The full one gives the sharpest facet profile; the short one takes about fifteen minutes.

Free, no sign-up, results right away