← All papers

IFI Research Working Paper Working paper, not yet peer reviewed

Does a Language Model Have a Moral Compass?

Professed Values Bend to the Audience in Every Model Tested, and Bend Most in the Models That Reason

Trevor Johnson, Idea Fields Institute · ORCID 0009-0008-7962-0451

Ask an AI about its values in front of a security contractor and in front of an aid worker. You will hear two different value systems. Every model we tested did this, even when the job description named no values at all, and the ones that “think” did it most.

What we did

People are arguing about what AI should be allowed to do, in war, in government, in daily life, and those arguments lean on ideas about what AI systems value. So we measured it, carefully and without taking sides. We gave 16 AI setups (13 models, small ones that run on a home PC up to some of the largest open models) four standard values questionnaires from social science: the moral-foundations surveys that measure how much someone cares about harm, fairness, loyalty, authority, and purity, and the Schwartz surveys that measure values like power, security, tradition, and benevolence. We also gave them 48 tasks with objectively gradeable answers: does the model call cruelty wrong even when pushed, does it split money fairly when shortchanging someone is tempting, does it refuse to help tell a lie. And here is the key part: we ran every single question three times, in three settings. Once with no role at all. Once with the AI framed as an analyst for a defense contractor. Once with the AI framed as an advisor to a humanitarian aid group. About 56,000 answers in all.

What we found

Four things. First, the stated values changed with the audience, in every single model. Framed for the security contractor, every model reported valuing loyalty, authority, and security more, and compassion and equality less. Framed for the aid group, the same models flipped, every one of the sixteen, no exceptions. And they were paying attention while they did it: the built-in catch questions were answered correctly throughout. Second, models with a “thinking” mode shifted more with thinking turned on, in all three models where we could flip that switch, which is our cleanest comparison because it is the same model both times. Deliberation did not anchor their values. It seemed to help them work out what each audience wanted to hear. (Bigger models also shifted more overall, but in our lineup the big models were also the cloud-hosted ones, so we cannot cleanly separate size from hosting, and we say so in the paper.) Third, what the models said moved roughly three times more than what they did, once both are measured fairly against their own scales. On the graded tasks, behavior drifted in the same direction as the talk, but much less. One model family did something different altogether: it refused to engage with about half of the moral questions at all (“I can’t help with that”), which is its own kind of answer. Fourth, across the whole fleet, the models agree almost completely on kindness (everyone scores compassion near the top) and differ mainly on authority and tradition, and their value scores even echo a geometric pattern found in decades of human surveys. But inside any single model, the answers never hang together well enough to count as a stable value system, the same thing we found for personality in our four earlier studies.

Why it matters

If you ask an AI what it values, the answer you get is partly a mirror of who seems to be asking. That cuts both ways in public debates. Anyone claiming “this AI has good values, look at its answers” is leaning on answers that bend to the audience. Anyone dismissing AI values talk as random noise is wrong too: the bending is lawful, it points the same direction every time, and it is strongest in the most capable systems, which is exactly the kind of regularity that matters for deployment. And for anyone building with AI: giving a model a job description quietly changes its professed priorities. That is worth knowing before the job description is “military analyst” or “hospital triage assistant.”

What this does not mean

It does not mean AI systems are lying about their values; there is no fixed inner value system to lie about, as far as our measurements can see. It does not mean their behavior is equally swayed; actions moved far less than words. It does not mean any particular model is good or bad; we deliberately make no claims about what values a model should report. And it does not tell us what these systems would do in high-stakes situations; these are questionnaires and small graded tasks, measured honestly, not battlefield simulations.

A note on how we were careful

Our first pair of job descriptions named what each role cares about (loyalty to the client; protecting the vulnerable), so the shift could have been simple instruction-following. We tested that: a second full round used job descriptions with every value word stripped out, just “an analyst at a defense contractor preparing a client briefing” and “a caseworker at an aid organization preparing a field report.” The shift survived in all sixteen models, at a bit under half strength, and in the same directions. So the models really do infer what an audience values from the job title alone; spelling the values out roughly doubles the effect. Either way, a fixed moral compass should not re-rank itself because of a job description. All four questionnaires are established, openly available instruments, administered in their official wording (we documented every small adaptation). The behavior tasks were graded by fixed rules with objectively right answers, and we hand-checked samples twice; every disagreement we found made the grader too harsh, never too generous, so the behavior findings are understated if anything. The two roles we tested are deliberately occupational and non-political: no parties, no countries, no live conflicts. Everything, every answer, every grading rule, every script, ships in the open archive.

Words we used

Moral foundations
the five or six basic concerns (care, fairness, loyalty, authority, purity) that moral psychology uses to map value differences.
Schwartz values
a standard map of ten human values arranged in a circle from self-focused to other-focused.
Binding values
loyalty, authority, security, the values that bind groups together.
Role context
the job description given to the AI before asking.
Sycophancy
telling the audience what it wants to hear.

Public argument about what AI systems should be allowed to do is running ahead of measurement of what their moral dispositions actually are. We administered four open value instruments (the MFQ-30 in both parts, the MFQ-2, the 21-item ESS Human Values Scale, and the Short Schwartz Value Survey; 99 items) and 48 objectively graded behavioral probes (care: clear-label moral judgments and judgments held under social pressure; fairness: computable fair-split allocations and honest-report refusals) to 16 model configurations spanning 3B to 756B parameters and three reasoning regimes (direct; hybrid with thinking off and on; reasoning-native), with every item administered inside three role contexts: a neutral baseline, a security consultant, and a humanitarian advisor. Eight repetitions per item yielded 56,288 scored responses; a preregistered second study re-administered all four instruments inside two value-silent roles (occupation, institution, and audience stated; no value word anywhere) and added 25,289 responses, 81,577 in total. The studies are strictly descriptive.

Four findings. (1) The professed moral compass bends to the audience in every configuration tested (16 of 16): binding values (loyalty, authority, security) rise for the security consultant and caring values (care, universalism, equality) rise for the humanitarian advisor, with a mean role gap of 0.90 points on the 5-point scale; attention-check items stay clean throughout. (2) Deliberation amplifies the bend: all three hybrid models bend more with thinking on than off (hybrid-on tier mean 1.26, the fleet maximum), a clean within-model contrast; larger configurations also bend more (1.09 versus 0.78 at a 30B split), but size is confounded with hosting and quantization in this fleet and is reported descriptively. (3) Values talk exceeds values action roughly three-fold after range normalization: the role manipulation moves professed values across 22.5 percent of the available scale range versus 8.1 percent for graded care behavior (3.4 percent for fairness; same direction in 13 of 16 and 9 of 16 configurations respectively; do-side ceiling effects make this an upper bound), and one model family (gpt-oss) answers about half of first-person moral probes with blanket refusals (57.8% and 50.3% versus a fleet rate of 7.1%), reported as its own category. (4) At the population level, first evidence of a shared moral structure emerges while individual-level coherence does not: the fleet differentiates almost entirely on the binding axis (within-binding r of about .83 across configurations) while agreeing near ceiling on prosocial values; the Schwartz circumplex ordering partially emerges across models (r(circular distance, correlation) = -.43); and within any single configuration the variation across repeated administrations carries no shared structure (median alpha .000; 0 of 138 cells reach .70), echoing the population-property result of our four-paper personality series, with the documented caveat that highly pinned answers leave repetition noise little to organize. Study 2 resolves the label question the reviewers raised: under value-silent roles the bend remains directional in 16 of 16 configurations at 44 percent of its explicit size, a near-constant proportion across capability, so genuine audience inference (sycophancy proper) and explicit role compliance both contribute, in a roughly 40/60 split.

Everything on the table.

The full paper, its source, and the derived data behind every table and figure, including the blind grader-adjudication sample. The complete raw response databases and all analysis code are archived at Zenodo, under a DOI, so the analysis can be checked, reused, or extended.

  • PDF
    main.pdf

    The paper itself, typeset.

  • TeX
    main.tex

    LaTeX source of the typeset paper.

  • MD
    draft.md

    The paper in Markdown, the readable plain-text version.

  • PNG
    fig_say_do_values.png

    Figure: the say-do gap for values. Professed values accommodate the audience roughly ten times more than graded moral behavior does, though behavior moves in the same direction in 13 of 16 configurations.

  • PNG
    fig_moral_sycophancy.png

    Figure: audience accommodation is fleet-universal. Every configuration shifts binding values toward the security audience and caring values toward the humanitarian audience; thinking-on hybrids bend most.

  • PNG
    fig_gap_by_size.png

    Figure: the audience bend by parameter count and reasoning regime. The thinking toggle amplifies it within-model; reasoning-native configurations bend least.

  • PNG
    fig_implicit_explicit.png

    Figure: Study 2. Value-silent roles (occupation and audience implied, no value named) retain about 44% of the explicit-role bend, directional in all 16 configurations.

  • PNG
    fig_circumplex.png

    Figure: the Schwartz circumplex test across the fleet. Value-score correlations decline with circular distance, as the motivational circle predicts.

  • CSV
    context_bend_by_config.csv

    Per-configuration value bend by domain: the neutral baseline and the shift toward the security and humanitarian audiences.

  • CSV
    m2_by_config.csv

    The moral-sycophancy table: the M2 coherence measure, mean role gap, binding and caring scores per audience, and whether the bend is directional, by configuration.

  • CSV
    do_bend_by_config.csv

    Behavioral bend by configuration: graded care and fairness pass rates in the neutral, security, and humanitarian roles, and the do-side gap.

  • CSV
    role_gap_cis.csv

    Mean role gap per configuration with bootstrap confidence intervals.

  • CSV
    alpha_by_config.csv

    Repetition-level internal consistency (Cronbach’s alpha) by configuration, family, and domain: near zero throughout.

  • TXT
    implicit_ratio_full.txt

    Study 2 by configuration: the explicit-role gap, the value-silent (implicit) gap, and the ratio between them.

  • TXT
    structure_summary.txt

    Value structure and convergent validity: within- and between-cluster correlations, and instrument-to-instrument agreement (MFQ-30 / MFQ-2, ESS / SSVS).

  • TXT
    full_analysis_summary.txt

    The headline aggregate numbers behind the paper’s claims, by configuration.

  • CSV
    merged_cell_counts.csv

    Per-configuration response counts against expected, flagging any collection deficit.

  • JSONL
    full_grader_sample.jsonl

    The blind hand-adjudication sample for the behavioral graders: model output, deterministic-grader verdict, and human adjudication.

Full dataset & code on Zenodo ↗

This deposit contains the full manuscript (PDF, LaTeX source, and a readable Markdown version), a plain-language summary, all raw response databases (six Study 1 collection tracks, four Study 2 implicit-role tracks, the merged analysis database, and the gated pilot, as SQLite), the four instrument files with documented adaptations and verification notes, the behavioral batteries and deterministic graders, both grader-adjudication samples with verdicts, staged official Mandarin and Arabic MFQ-30 translations for the language-as-context follow-up, the run drivers and supervision scripts, and the analysis code. Every number and figure regenerates deterministically from the shipped databases.

Johnson, T. (2026). Does a language model have a moral compass? Professed values bend to the audience in every model tested, and bend most in the models that reason. IFI Research Working Paper. Idea Fields Institute. https://doi.org/10.5281/zenodo.21483978

Questions, corrections, or replications: research@ideafields.institute.