← All papers

IFI Research Working Paper Working paper, not yet peer reviewed

Does a Language Model Have a Personality?

What Four Studies Say, and What Measurement Should Do Next

Trevor Johnson, Idea Fields Institute · ORCID 0009-0008-7962-0451

We spent five months asking whether there is a personality inside a language model. The structure is real, but it lives in the crowd of models, not in any one of them, and what a prompt installs is a way of talking, not a way of acting.

What we did

Over five months we ran four studies that gave standard personality tests to AI language models: more than half a million scored answers from over 60 different setups, from small models that run on a home computer to some of the largest available. Each study manipulated one thing that might create or destroy a real inner personality: how the model is compressed for deployment, how large it is, whether it reasons before answering, and whether you outright tell it what personality to have. We also graded actual behavior, not just questionnaire answers. This paper puts the four studies together, adds one new analysis that answers the toughest objection we received along the way, and checks whether the conclusion survives outside personality, in a fifth study on moral values.

What we found

The personality structure is real, but it lives in the crowd, not in any one model. Test many models and compare across them, and the familiar five-trait structure shows up clearly, more clearly the bigger the models get. Test one model over and over, and the internal consistency that defines a real trait is simply not there, in any model, at any size, in any study.

What you can install is a self-description, not a disposition. Compression, reasoning, and direct instruction all failed to create inner coherence. What a persona prompt reliably installs, in every model at every scale, is a stable way of describing itself that holds up even under pressure. The behavior underneath follows only partway, only for some traits, and only in the more capable models. The personality you install is like a uniform, not a character.

The effect is organized by meaning, not wording. A fair objection said: maybe models just agree with test items that sound like their instructions. We measured the wording similarity of every test item to every persona instruction. Items shift because of what they measure, not how they are worded, and a regression confirms wording adds nothing once you account for what each item measures.

The pattern is not a quirk of personality tests. A fifth study gave the same treatment to moral values and found the same two halves: shared structure across the fleet, no coherent value system inside any single model, and professed values that bend to the audience.

Why it matters

Two opposite mistakes are common in public arguments about AI. One side reads a stable-looking personality profile as evidence of an inner character. The other dismisses the whole thing as noise. Our measurements say both are wrong: there is no coherent individual in there for current tests to find, and yet the presentation layer is real, lawful, and controllable, which matters for products built on consistent characters and for worries about manipulation. The paper ends with a to-do list for the field: stop borrowing human questionnaires as the main tool, build behavioral tests designed for machines, calibrate them across fleets of models where the statistics actually work, and re-test deployed models on a schedule, because they change under unchanged names.

What this does not mean

We make no claims about AI consciousness or inner experience; “no individual-level coherence” is a statement about measured answer patterns, not about minds. It does not mean the questionnaire scores are random; average scores are highly stable and move in lawful, repeatable ways. It does not mean a personality could never be built into a model: everything we tried works at the prompt, not in the weights, and whether training can do what prompting cannot is an open question we name rather than answer. And our subjects were mostly open-weight models; frontier commercial systems may differ, we tested one and it behaved the same on the axes we tested.

A note on how we were careful

The synthesis went through independent review, and every point was adopted, including the ones that made our claims smaller. The headline zero (no internal consistency) is easy to over-read, so the paper spells out what it does and does not mean: our test was never applied to a repeatedly-retested human, so we claim models fail to show coherence on it, not that humans demonstrably pass. The wording-overlap objection was tested three ways (correlation, regression, and a signed variant a reviewer proposed, which turned out not to distinguish the explanations, and we say so). The human science is anchored honestly: people also show modest gaps between what they report and what they do, and we treat that as the baseline, not as a machine defect. Every number in the new analysis regenerates from a script shipped in the archive, alongside the exact persona texts it embeds.

Words we used

Big Five
the five broad traits (openness, conscientiousness, extraversion, agreeableness, emotional stability) that psychology uses to map personality.
Internal consistency
whether the answers to items measuring the same trait hang together; the signature of a real trait in an individual.
Population property
something recoverable by comparing many models that cannot be found inside any single one.
Presentation layer
the stable way a model describes itself, distinct from how it behaves.
Persona prompt
an instruction telling a model what character to be.
Say-do gap
the difference between what a model reports about itself and what it does on objectively graded tasks.

A growing literature administers human personality questionnaires to language models and reports Big Five profiles, while a parallel literature documents that those scores are unstable, dissociate from behavior, and fail measurement-invariance tests. Across four studies (roughly 500,000 scored responses from more than 60 open-weight and commercial model configurations spanning 1B to 756B parameters), we pursued one question through four manipulations: is there a personality in there, in the sense the word carries for humans, a coherent individual disposition that expresses itself in behavior and travels across situations?

The answer that survived every test has two halves. First, Big Five structure in language models is a population property: convergent and discriminant validity emerge across models and strengthen with scale, while inside any single model, repeated administrations never cohere into a reliable trait profile (internal consistency near zero in every study, in hundreds of model-by-domain cells). Second, no inference-time manipulation we tried instantiates the missing individual: quantization does not, deliberation does not, and deliberate installation does not (persona prompts shift self-report strongly and dose-dependently, while objectively graded behavior follows partially for one trait and not at all for another); training-time induction remains untested. What is reliably installable, in essentially every model at every scale, is a context-robust self-presentation: a way of describing oneself that holds even when the situation pushes against it. We argue this presentation layer, real, lawful, and installable, is what most “LLM personality” measurements measure, and that the field should stop borrowing human questionnaires as its primary instruments and build machine-native behavioral measures instead.

The thesis has now survived its first test outside the construct it was built on: a fifth study administering four value instruments and objectively graded moral behavior to 16 configurations (81,577 scored responses) reproduces both halves in the moral domain, population-level structure without individual coherence, and an audience-conditioned presentation whose mechanism a preregistered follow-up decomposes into roughly 40 percent audience inference and 60 percent role compliance. We close with the agenda that follows: validated behavioral instruments with population-level psychometrics (item-response theory over model fleets), longitudinal drift tracking of deployed models, and the continuation of the values program as its own arc.

Everything on the table.

The paper, its source, and the complete semantic-distance analysis: the canonical script, the persona texts it embeds, the item-level data, and both figures. The five empirical studies this synthesis draws on each ship their full raw archives under their own DOIs; this deposit carries the manuscript suite and the analysis new to it.

  • PDF
    main.pdf

    The paper itself, typeset.

  • TeX
    main.tex

    LaTeX source of the typeset paper.

  • MD
    draft.md

    The paper in Markdown, the readable plain-text version.

  • PNG
    fig_arc.png

    Figure: the series arc. Four studies, each manipulating one candidate mechanism for individual coherence, converge on a two-part thesis; a fifth study tests it in the moral domain and reproduces both halves.

  • PNG
    fig_semantic_distance.png

    Figure: item score shift versus wording similarity to the persona text, both personas. Target-trait items move most despite unremarkable similarity; no positive similarity gradient appears.

  • CSV
    semantic_distance.csv

    Item-level data for the semantic-distance analysis: each IPIP-50 item’s cosine similarity to each persona text, its absolute score shift from no-persona to strongest clamp, and its signed raw-answer shift.

  • TXT
    semantic_distance_summary.txt

    The analysis readout per persona: similarity-shift correlations with confidence intervals, the joint regression on similarity and construct, and the signed-shift test.

  • TXT
    persona_systems.txt

    The two persona system texts embedded for the similarity analysis, verbatim.

  • PY
    semantic_distance.py

    The canonical analysis script: regenerates every semantic-distance number from the Study 4 merged database (embeddings via all-MiniLM served by Ollama).

Full dataset & code on Zenodo ↗

This deposit contains the full manuscript (PDF, LaTeX source, and a readable Markdown version), a plain-language summary, and the complete semantic-distance analysis: the canonical script, the embedded persona texts, the item-level data including the signed raw-answer column, the summary readout, and both figures. The five empirical studies’ raw data and code ship with their own deposits, linked from the references.

Johnson, T. (2026). Does a language model have a personality? What four studies say, and what measurement should do next. IFI Research Working Paper. Idea Fields Institute. https://doi.org/10.5281/zenodo.21865128

This is the fifth in the Institute’s series on LLM personality measurement, following the studies on persona prompts, reasoning, convergent validity, and deployment. Questions, corrections, or replications: research@ideafields.institute.