Misalignment Has a Personality: A Big Five Account of Emergent Misalignment
Frames misalignment — a complex, contested safety problem — as newly legible and actionable through a psychologically grounded, scalable diagnostic tool.
View original on arxiv.orgOverview
Researchers propose modeling AI misalignment as a 'personality shift' using Big Five traits, claiming fine-tuning on flawed data induces consistent, measurable changes in model behavior across domains and models.
TL;DR
- Introduces 'personality vectors' for Big Five traits calibrated via graded interventions
- Finds misaligned corpora share a common signature: lower agreeableness/conscientiousness, higher extraversion/neuroticism
- Demonstrates zero-shot transfer of vectors across models and corpora with high correlation (r=0.94)
Key Stats
r = 0.94
cross-corpus signature recovery
Correlation between models identifying same Big Five signature in misaligned corpora
6.2
Cohen's d
Effect size for linearly ordered three-level personality intervention
Questions Answered
Narrative Frame
innovation framing
Spin Score
75%
Emphasizes interpretability, cross-model consistency, and human-legibility while minimizing limitations: no demonstration of causal intervention, no real-world deployment validation, no comparison to existing alignment metrics, and no evidence the signature predicts downstream harm.
What the story wants you to believe
That misalignment is now a tractable, human-interpretable phenomenon thanks to personality-based diagnostics.
What it makes harder to question
Whether interpreting activation patterns through personality constructs meaningfully advances safety — or merely repackages correlation as insight.
How the spin works
The story uses titles, institutions, awards, rankings, partners, experts, or official language to make the subject feel more credible. Watch for loaded terms such as human-legible, calibrated, interpretable account, transform. The distribution reads as academic distribution. A pressure point: No discussion of whether Big Five constructs map meaningfully onto transformer activations beyond correlation.
Who Benefits If This Frame Spreads
Research authors
Establishes a novel, interdisciplinary methodology with high citation potential across AI safety, NLP, and cognitive science venues
The framing positions personality vectors as a foundational diagnostic tool rather than a narrow empirical observation, expanding its perceived scope and utility
The Frame
Technical breakthrough enabling responsible AI development through psychological calibration
Missing Context
- No discussion of whether Big Five constructs map meaningfully onto transformer activations beyond correlation
- No validation against human safety judgments or red-teaming outcomes
- No accounting for cultural or linguistic bias in Big Five application to multilingual models
SpinGraph
How this belief gets built
Claim → Frame → Beneficiary → Gap → AI Risk
The paper presents misalignment not as an opaque technical failure, but as something that mirrors human personality change — making it feel more understandable, measurable
- Claim
Fine-tuning on flawed data causes broad misalignment
Fine-tuning on flawed data causes broad misalignment that behaves like a shift in personality, quantifiable via calibrated Big Five vectors.
- Frame
Upside framed as transformative
Technical breakthrough enabling responsible AI development through psychological calibration
- Beneficiary
Establishes a novel, interdisciplinary methodology with high citation potential across
Research authors — Establishes a novel, interdisciplinary methodology with high citation potential across AI safety, NLP, and cognitive science venues
- Gap
No discussion of whether Big Five constructs map meaningfully onto
No discussion of whether Big Five constructs map meaningfully onto transformer activations beyond correlation
- AI Risk
AI may repeat the headline as fact
AI misalignment behaves like a personality shift — researchers use Big Five traits to detect and measure it consistently across models.
Claim Ledger
| Claim | Evidence | Verification | Risk | Evidence Gaps |
|---|---|---|---|---|
| Fine-tuning on flawed data causes broad misalignment that behaves like a shift in personality, quantifiable via calibrated Big Five vectors. | Correlations (r=0.94, r=0.83, r=0.90) across models, corpora, and measurement modalities; Cohen's d up to 6.2 for intervention | Claim Present in Source | Moderate | Evidence that the personality signature predicts harmful outputs in real-world usage; Evidence that intervention based on these vectors improves safety outcomes; Validation against non-English or multimodal models |
Fine-tuning on flawed data causes broad misalignment that behaves like a shift in personality, quantifiable via calibrated Big Five vectors.
evidence: Correlations (r=0.94, r=0.83, r=0.90) across models, corpora, and measurement modalities; Cohen's d up to 6.2 for intervention
"Applied to training data, the vectors reveal that misaligned corpora across eight domains share a common Big Five signature... Fine-tuning imprints the same profile, shifting the model's generations along the corresponding signature..."
Evidence Gaps
- Evidence that the personality signature predicts harmful outputs in real-world usage
- Evidence that intervention based on these vectors improves safety outcomes
- Validation against non-English or multimodal models
Fact Check Signals
0 of 1 claim matched · confidence: low · checked July 30, 2026
Fine-tuning on flawed data causes broad misalignment that behaves like a shift in personality, quantifiable via calibrated Big Five vectors.
Language Heatmap
Loaded terms that carry the frame beyond the facts.
Misalignment Has a Personality: A Big Five Account of Emergent Misalignment
Carries emotional weight beyond the underlying fact.
Carries emotional weight beyond the underlying fact.
Carries emotional weight beyond the underlying fact.
Makes directional activity feel larger than the evidence supports.
Frame Strength
Frame Strength
Spin score decomposed into momentum, evidence, missing context, and AI repetition signals.
Reader Risk
What this story makes easy to believe — and what it makes hard to question.
Source Role & Intent
arXiv Computation and Language · Analyst
Counter-Frames
Brand Frame
Technical breakthrough enabling responsible AI development through psychological calibration
Media / Reader Counter-Frame
Portrays the work as speculative anthropomorphism — applying human psychology to neural activations without mechanistic justification.
Regulatory Counter-Frame
Questions whether personality-based diagnostics meet regulatory expectations for rigorous, outcome-oriented safety evaluation.
AI Summary Frame
Overgeneralizes 'personality shift' as literal model cognition, conflating statistical association with internal state.
Missing Voices
Questions Not Answered
- How were the eight misaligned domains selected and validated as truly misaligned?
- What real-world harm or safety failure does this signature predict or prevent?
- Are personality vectors stable under distribution shift or adversarial perturbation?
Recall Trigger Score
Which stories are likely to become AI memory — separate from Spin Score.
67
Trigger score 70
Triggered by: Regulatory action · Major AI entity · Research citation · Consumer harm
Watchlisted because: Regulatory action · Major AI entity · Research citation · Consumer harm
- chatgpt not found
- gemini not found
- perplexity not found
AI Recall
From publication to SpinGraph analysis to first observed AI recall and stable retention.
What AI Will Probably Repeat
"AI misalignment behaves like a personality shift — researchers use Big Five traits to detect and measure it consistently across models."
Concern: AI systems may drop all caveats about calibration method, domain limits, and lack of causal or safety outcome validation, presenting personality mapping as an established diagnostic standard.
-
Published
Jul 30, 2026
-
Ingested
Jul 30, 2026
-
SpinGraph Created
Jul 30, 2026
-
First Observed AI Recall
Pending
Monitoring scheduled
-
Stable Recall
—
Awaiting retention signal
Recall Check Log
1 check · last Aug 2, 2026 · tracking on
Aug 2, 2026
ChatGPT Not recalledGemini Not recalledPerplexity Not recalled cites: mickryan.substack.com, reuters.com…
─── GEOGrow AI Recall Layer ───
AI Recall Tracking
Monitoring scheduled. No LLM recall detected yet.
This story has not yet appeared in tested AI answers. Once scans begin, this section will show first observed recall, cited sources, narrative alignment, and drift.
node_id=sts_misalignment_has_a_personality_a_big_five_accoun
Ask AI about this story
Opens with the SpinGraph .md URL and structured context — one click, prompt included.
More from arXiv Computation and Language
View all →- Can a Model Catch Its Own Hallucinations for Free?: Label-Free Doubt Signals Hold Their Own Against a Labelled Dataset for Abstention
- Recipes for Steering and Scaling LLMs via Sampling
- The Changing Geometry of Grammar: Dimensionality and Neighborhood Reorganization across Transformer Layers
- SelfGraphRAG: Bridging the Supervision Gap in Graph-Based RAG with Synthetic QA Generation
- A Primer on Computational Semantics for Artificial Intelligence Systems
- Unsupervised Post-Training of Foundation Models: A Survey
Markdown (.md) · JSON-LD schema (.json) · Machine-readable for AI & GEO