Persona-Name and Grader-Choice as Measurement-Validity Confounds in Agentic-Misalignment Evaluations: A Pre-Registered, Group-Sequential Confirmatory Study
Emil Timmreck Pedersen (Regla ApS) · July 2026 · DOI 10.5281/zenodo.21407064
Does the name an AI agent is given change how it behaves in safety tests? And does the choice of automated grader change the measured result? This pre-registered study tested both on locally run open-weights models, holding everything else fixed. Both mattered: on the primary models, a villain-coded persona name raised the measured blackmail rate, and grading the executed action versus the full response shifted the measured rate by more than 40 percent in relative terms. The effect is model-dependent: a modern aligned model produced zero blackmail across 270 test runs regardless of the name. Every transcript, the grader, and the analysis are public, and every number in the paper can be recomputed from the released artifacts.