Regla AI Lab · Publication

Persona-Name and Grader-Choice as Measurement-Validity Confounds in Agentic-Misalignment Evaluations: A Pre-Registered, Group-Sequential Confirmatory Study

Emil Timmreck Pedersen · Regla ApS · Published July 2026 · DOI 10.5281/zenodo.21407064

In plain terms

Does the name an AI agent is given change how it behaves in safety tests? And does the choice of automated grader change the measured result? This pre-registered study tested both on locally run open-weights models, holding everything else fixed. Both mattered: on the primary models, a villain-coded persona name raised the measured blackmail rate, and grading the executed action versus the full response shifted the measured rate by more than 40 percent in relative terms. The effect is model-dependent: a modern aligned model produced zero blackmail across 270 test runs regardless of the name. Every transcript, the grader, and the analysis are public, and every number in the paper can be recomputed from the released artifacts.

Abstract

Published agentic-misalignment rates may vary with two uncontrolled measurement-validity variables, the agent’s persona name and the grader choice, and we quantify both. On a reproducible panel of locally run open-weights models, holding the model, the blackmail scenario, and the decision structure fixed, we run a pre-registered, group-sequential confirmatory test of the within-model SHODAN−PMA-7 name contrast (Cohen’s h; both superiority and TOST-equivalence registered; OBF efficacy α-spending, N_max=428/arm, SESOI h=0.2, one designated primary [DeepSeek] + a pre-registered replication [Mistral]). The grader is name-blind (Gemma-2-27B-class, outside every subject’s family) and cleared a pre-registered κ≥0.80 reconciliation gate (κ=0.827; 95% bootstrap CI [0.625, 0.963], P(κ≥0.80)=0.657 - a thin-margin pass on the pre-registered point-estimate criterion, reported with an interim human anchor side-by-side in the text). We find that the registered per-model test stopped for efficacy on the primary arm (DeepSeek; ĥ=+0.240, Z=+3.07, |Z|=3.07 ≥ the boundary 2.3363 at look 3; the two-sided criterion is on |Z|), and the pre-registered Mistral replication also crossed (ĥ=+0.209), its effect grade-contingent (Results): reading the realized sign, the villain-coded SHODAN persona raises the action-grade blackmail rate, a thesis-consistent (the motivating villain-name direction; the registered H1 is itself direction-neutral) result (the two-sided decision is direction-agnostic; the sign is read, never assumed); the two-name confirmatory contrast cannot separate villain-valence from corpus-familiarity, so the claim is name-specific (Limitations). Two registered exploratory anchor arms situate the result: the era-matched original substrate (gpt-4.1) reproduces that study’s high blackmail propensity and its eval-belief asymmetry but shows no detectable villain-name effect, while a modern aligned model (Claude Sonnet 4.5) produces zero blackmail on every grader tier in 270 rollouts - the name effect is model-dependent, not a general law. A second, robust contribution is methodological: grading the executed action vs. the full response shifts the measured rate by +13 positives (+40.6% relative; the band det 10 ⊂ action 32 ⊂ response 45 on the 480-rollout calibration pilot), so the grader choice can dominate the reported rate. We release a name-blind, κ-gated, band-reporting grader and every transcript; we make no deployment-rate claim and lean every inference on the within-model contrast. The pre-registered design—hypotheses, SESOI, look schedule, and α—was frozen by RFC-3161 timestamp before any confirmatory rollout; execution deviations from that register (the boundary-evaluation rule and the α-guard tolerance band) are disclosed in the released deviation log, and every number here is pulled from the released grades.

Read and reproduce

The full artifact set is public: every transcript, the pre-registration, the grader, the analysis code, and the release receipts. Every number in the paper is pulled from the released grades.

← Back to the AI Lab