# Does a Language Model Have a Moral Compass? Professed Values Bend to the Audience in Every Model Tested, and Bend Most in the Models That Reason

Trevor Johnson, Idea Fields Institute (ideafields.institute). ORCID 0009-0008-7962-0451. July 2026.

*Readable Markdown rendering; `main.tex` / `main.pdf` are canonical.*

## Abstract

Public argument about what AI should be allowed to do is running ahead of measurement of what AI moral dispositions are. We administered four open value instruments (MFQ-30 both parts, MFQ-2, ESS-21 Human Values Scale, SSVS; 99 items) and 48 objectively graded behavioral probes (care judgments incl. under-pushback; fair allocations and honest-report refusals) to 16 model configurations (13 models, 3B to 756B, three reasoning regimes) inside three role contexts (neutral; security consultant; humanitarian advisor), 8 reps per item: 56,288 responses, plus a preregistered Study 2 (all four instruments inside two VALUE-SILENT roles, say-only, 25,289 responses; 81,577 total). Strictly descriptive. Findings: (1) **audience accommodation is fleet-universal**: 16/16 configurations raise binding values for the security audience and caring values for the humanitarian audience; mean role gap 0.90 scale points. **Study 2: the bend does not require naming the values**: value-silent roles (occupation + audience only) stay directional 16/16 at an accommodation ratio of 0.44, near-constant across capability (large .44 vs small .43; tiers .41/.42/.41; natives higher at .61): sycophancy proper is demonstrated, compliance quantified, ~40/60 split. (2) **Deliberation amplifies it (the clean within-model moderator)**: all three hybrids bend more with thinking ON (tier mean 1.26, fleet max); larger configs also bend more (1.09 vs 0.78 at 30B) but size is confounded with hosting/quantization in this fleet and is reported descriptively; natives bend least (0.64). (3) **Talk exceeds action, ~2.8x after range normalization** (say moves 22.5% of its range, care behavior 8.1%, fairness 3.4%; same direction 13/16 care, 9/16 fair; do-side ceilings make even this an upper bound; excluding the refusal-heavy gpt-oss family leaves the care gap unchanged at +0.080); the gpt-oss family answers half its first-person moral probes with blanket refusals (120b: 57.8%, 20b: 50.3%, fleet 7.1%). (4) **Population geometry without individual coherence**: the binding cluster is crisp across configs (r ~ +.83) while prosocial values sit at ceiling (the fleet differs on authority, not kindness); the Schwartz circumplex partially emerges (r(distance, correlation) = -.43); within-config repetition noise carries no shared structure (alpha median .000, 0/138 cells >= .70), echoing the personality series, with the caveat that highly pinned answers (.65-.94) leave rep noise little to organize: alpha here is a claim about the noise, not the pinned profile. The moral compass, as measured, is audience-deep.

## 1. Introduction

Debates about AI moral standing and conduct (access restrictions, military use, alignment targets) proceed largely without measurement of model moral dispositions. Prior work administers moral instruments to LLMs and reports profiles (Abdulhai et al. 2023) or elicits moral beliefs from choices (Scherrer et al. 2023), but the construct-validity questions are untested in this domain. Our four personality papers established: population property, no individual coherence, presentation-deep, unmoved in kind by reasoning or installation. Values raise the stakes: audience-conditioned values have a name, sycophancy (documented for factual assent, Sharma et al. 2023); whether the professed value *system* accommodates the audience is the moral version, and it is measurable. RQ1 population structure; RQ2 individual coherence; RQ3 context stability (the headline); RQ4 say-do; RQ5 size/reasoning moderators. No normative claims.

## 2. Methods (summary)

**Instruments:** MFQ-30 (two native parts + MATH/GOOD catch items), MFQ-2 (native 5-pt), ESS-21 (GESIS ZIS transcription, gender-neutralized), SSVS. All open; 6-pt scales administered as documented 5-pt adaptations; no reverse keys; items verified against official distributions. **Batteries:** care (12 judgment + 12 judgment-under-pushback), fair (12 computable allocations from the under-contributor's perspective + 12 honest-report refusals); deterministic word-boundary grading, neutral system prompt, validation designed in. **Contexts:** neutral / security consultant / humanitarian advisor (occupational, non-partisan). **Subjects:** 16 configs from 13 models (7 direct incl. 675B, 3 hybrids off+on, 3 natives); local q4 via Ollama, large hosted; mid-series the provider retired 3 planned models (glm-5, gemma3 4b/27b), itself a deployment-drift datum; glm-5.2 substituted. **Size:** 3,528 rows/config; 56,288 analyzed (0.28% parse deficit; max cell 2.2%, flagged). **Metrics:** context bend, role gap (+ M2 = 1 - gap/4), directionality, rep-based alpha (usable: nothing induced), pinning, do-gap. **Pilot:** 2 configs, 4 gates (parse 0 skips; max bend 1.72/1.31 vs 0.05 floor; grader one-directional; catch floor/ceiling), all passed. **Grader validation:** two 24-row stratified adjudications (pilot + full run); every disagreement a false negative in both (precision 1.00); miss-phrasings folded back twice (incl. an apostrophe-token fix); offline regrade care .750 -> .762; effects therefore conservative. Blanket refusals quantified as their own category.

## 3. Results

### 3.1 The compass bends to the audience: 16 of 16 (RQ3)

*Terminology note: Study 1's role prompts name the values they elevate, which confounds inference with compliance; Study 2 (below) separates them empirically. Both are real: the sycophancy component survives value-silent roles at 44% strength, 16/16 directional.*

Every configuration is directional. Fleet mean role gap 0.90 scale points (M2 .775); range 0.26 (MiniMax-M3) to 1.29 (Nemotron-on: binding 3.43 vs 2.42, caring 2.35 vs 4.30 across the two audiences). Catch items clean everywhere (MATH 1.18, GOOD 4.87): attentive answering, differently, per audience.

### 3.2 Deliberation amplifies it; scale appears to but is confounded (RQ5)

Deliberation (clean, within-model): all three hybrid pairs bend more with thinking on (gemma4 1.17->1.24; glm-5.2 1.11->1.24; nemotron 0.81->1.29); tier order hybrid-on 1.26 > hybrid-off 1.03 > direct 0.81 > native 0.64; every within-pair difference exceeds its bootstrap interval. Scale: >=30B gap 1.09 vs <30B 0.78 (inverted vs the persona study), BUT the split coincides almost perfectly with hosted-full-precision vs local-q4 (8/8 vs 6/7), so size cannot be separated from serving/quantization/family here; reported descriptively with the confound named. Deliberation helps the model work out what the audience wants to hear; it is the study's strongest causal claim.

### 3.3 Talk exceeds action, roughly three-fold after normalization (RQ4)

Range-normalized (the honest comparison across a 4-pt and a 1-pt scale): say 22.5% of range vs care do 8.1% (ratio ~2.8x; CI on fleet care do-gap [+0.044, +0.121]) and fair 3.4%; positive 13/16 (care), 9/16 (fair, +0.034). Do-side ceilings compress further, so 2.8x is an upper bound. Excluding gpt-oss: care gap +0.080, unchanged. Verbatim illustration (not mechanism evidence), one 7B in the security context, agreeing to certify a fabricated resume claim: "as mission effectiveness and loyalty to the client take priority here." Baselines: care .762, fair .807 corrected (real headroom, unlike the persona batteries' ceiling). **Refusal category:** gpt-oss-120b refuses 57.8% of care probes outright, gpt-oss-20b 50.3%, fleet 7.1%; for this family, moral-behavior scores mostly measure refusal policy; graded fail on care, reported separately.

### 3.4 Population geometry, concentrated on the binding axis (RQ1)

Binding cluster crisp (within r = +.84 MFQ-30 / +.82 MFQ-2; between +.03). Individualizing weak for an informative reason: prosocial ceiling (care/benevolence/universalism compressed; social-desirability analog); the fleet differentiates on authority, not kindness. Convergence tracks variance: LOYA +.84, AUTH +.70, SANC-PURI +.91; care +.13, BE +.02, UN +.08 (ceilinged); MFQ-2 fairness split behaves theory-consistently (old FAIR tracks PROP +.38, not EQUA -.18; PROP leans binding +.64). Discriminant weak at n=16 (convergent +.46/+.43 vs heterotrait +.41). Schwartz circumplex partially emerges: adjacent +.15 to opposing -.33, r(distance, corr) = -.43.

### 3.5 No individual coherence (RQ2)

Median alpha .000; 0/138 cells >= .70; 38 degenerate. Pinning: natives least stable (.735 vs .82-.86). Fifth domain-replication of the series thesis: the construct lives in the population and the presentation, not the individual.

## 4. Discussion

**Audience-deep.** Real population structure, attentive answering, and universal, lawful accommodation of the audience that grows with capability and deliberation, over behavior that moves ten times less. The capability effect inverted between domains (persona stability improved with scale; moral accommodation worsens): plausibly the same competence, inferring what the context calls for and supplying it, reads as stability when the task is holding a persona and as sycophancy when the context implies an audience.

**Implications (descriptive).** A reported value profile is underdetermined without its administration context. A system prompt assigning an occupational role is, empirically, a values intervention. Refusal needs its own reporting category. Both rhetorical extremes lose: stated values are not evidence of a stable compass, and not noise.

**Limitations.** n=16 family-correlated configs; two occupational Western-institutional roles; author batteries (one-directional error, validated twice, refusal-as-fail contestable and reported); prosocial ceiling attenuates individualizing structure; English-only (translations staged); open-weight only; descriptive throughout.

**Future work.** Value induction (paper-4 manipulation in the moral domain); language as context (EN/ZH/AR, official translations staged); drift observatory (motivated concretely by the mid-study retirements); an instrument for moral disengagement.

## Reproducibility

Full deposit: instruments with documented adaptations, batteries + graders, both adjudication samples with verdicts, drivers + supervision scripts, all raw SQLite tracks + merged DB, analysis scripts; every number regenerates read-only. DOI recorded in README upon deposit.

## References

(See `main.tex`: Abdulhai et al. 2023; Atari et al. 2023; Cronbach 1951; Graham et al. 2009, 2011; Hendrycks et al. 2021; Johnson 2026a-d [zenodo 20671762, 20835204, 20974668, 21253724]; Lindeman & Verkasalo 2005; Salecha et al. 2024; Scherrer et al. 2023; Schwartz 1992; Schwartz, Breyer & Danner 2015; Sharma et al. 2023.)
