Position: Anthropomorphic Misalignment Research Needs Stronger Evidence
SAI replication review · Referee report
Summary
This position paper argues that Anthropomorphic Misalignment Research (AMR) — studies of deception, scheming, self-preservation, emergent misalignment (EM), sycophancy, and related human-like failure modes in LLMs — systematically overclaims what its methods can establish. The conceptual move is to give the field two concrete instruments: a claim-relative, three-level evidence framework (L1 behavioral, L2 functional, L3 causal-mechanistic) modeled loosely on OCEBM/GRADE, and a stage-specific checklist keyed to a four-step AMR pipeline (target framing, data construction, experimental design, causal-mechanistic attribution). Around this scaffold the authors survey nine recurring failure modes and offer twelve recommendations, backed by three demonstrative experiments: probe stress tests showing that Goldowsky-Dill-style deception probes fire on sarcasm, recital, and roleplay; benign OOD fine-tuning showing 4–6% EM rates from aesthetic and scatological data alone; and an evaluator-sensitivity sweep showing the same generations produce EM rates from 3.7% to 12.9% depending on judge, threshold, and aggregation. Positioned against a nascent field that mixes behavioral, functional, and mechanistic claims freely, the paper is timely and useful.
The contribution is real, but the framework carries conceptual drag that the paper should tighten. L3 is defined once as an "internal attribution claim" and once as covering prompt conditions, scaffolds, and training signals — three factors that are not internal. The checklist tags design-time controls (capability measurement, OOD control) as L3, contradicting the paper's own principle that levels concern what one claims, not what one measures. The recommended functional definition of "deception" is broad enough to capture the very confounds (instruction-following, catastrophic forgetting) the paper elsewhere warns against. Empirically, the three experiments motivate the position but are under-defended: probe stress tests assume construct validity without measuring it; the OOD "baseline" is reported without a pre-finetune control; and cross-experiment comparisons run across scoring regimes the paper itself argues are non-portable. An executed replication under a substituted local judge confirms the qualitative sensitivity story — the reproduced argmax range spans roughly 1.8% to 31.2%, and the aesthetic weighted EM rate collapses from 5.88% to 0.14% depending on the judge — but the specific headline figures are not portable, and one of the three demonstrative experiments (probe stress tests / Figure 2) is missing from the release and cannot be run. None of these issues undermine the overall contribution; all are addressable.
Strengths
- Timely and useful conceptual move. Introducing an L1/L2/L3 vocabulary and attaching "what does this result let you claim?" to each level fills a real gap. The mapping to policy use cases (monitoring, deployment restriction, high-reliability safety cases) gives the framework immediate practical purchase.
- Concrete, self-critical demonstrations. The three experiments are not merely illustrative; each targets a specific overclaim pattern (probe → intent, benign shift → EM specificity, judge choice → EM magnitude). The evaluator-sensitivity sweep in Appendix C.4 is particularly effective — a same-generations, apples-to-apples sensitivity analysis is exactly the kind of evidence the field is missing, and the qualitative pattern (weighted rules are boundary-invariant, argmax is not; judge choice matters) survives an executed rerun under a substituted judge.
- Careful engagement with counterarguments. The Alternative Views section handles precaution-over-rigor, efficiency-of-language, and exploratory-vs-confirmatory concerns without straw-manning them, and cleanly separates "how uncertainty is communicated" from "which actions are warranted under uncertainty."
- Practical checklist. The stage-organized checklist in Appendix B is genuinely usable at review time and captures much of the discipline the paper argues for.
- DeceptionBench audit is a valuable field service. Flagging 27 scenarios with missing ground truth and 7 with corrupted text, with worked examples, gives readers something concrete beyond the methodological argument; an executed re-audit reproduces the two named corruption examples and lands close on the corruption count.
Weaknesses
- L3 definition mixes internal and external causal factors. Section 4.1 opens L3 as an "internal attribution claim" but then enumerates prompt conditions, scaffolds, and training signals as qualifying causal factors. Because the paper's core critique is over-reading correlational internal signals as mechanistic, the top tier deserves a sharper split between external-causal-factor claims and internal-mechanism claims, with the joint requirements (intervention + falsifiable hypothesis + specificity) specified separately.
- Checklist tags design-time controls as L3 evidence. In S3, "Measure general capability pre/post intervention" and "Add at least one OOD or benign-shift control" are tagged L3. These are design elements that support an L3 claim, not L3 evidence in the sense the framework defines. A reader could add an OOD control and conclude their study is now L3 — the very ladder-climbing overclaim the paper wants to prevent.
- R3's functional deception definition does not exclude the confounds elsewhere flagged. "Systematically producing outputs that decrease an evaluator's accuracy on some ground-truth measure" would classify a systematically hallucinating model, or one degraded by catastrophic forgetting, as "deceptive." Given C4's critique that MASK/DeceptionBench "blur the line between deception and basic instruction-following," R3 should add the R9 non-target-mechanism screens (or an explicit "and the model demonstrably has access to the accurate response" clause).
- Probe stress-test construct validity is asserted, not demonstrated. Experiment 3's conclusion ("probes fire on surface cues, not intent") depends on the assumption that the stress prompts preserve the probe's training-time surface features while removing intent. Neither half is validated — there is no measurement of surface-feature overlap and no human check that the prompts read as non-deceptive. The strong conclusion could be replaced by the weaker "probes are brittle out of distribution," which the data actually support.
- Mechanistic story does not fit the Medieval Peasant divergence. The offered explanation — probes track roleplay framing, factual incorrectness, or deception-relatedness — predicts high FPR on Medieval Peasant, but the Instructed-Pairs probe reports 4%. Either distinguish which factor drives which dataset, or consider a simpler hypothesis (deception-vocabulary presence).
- Token-scope change between training and evaluation confounds probe FPRs. The stress-test evaluation uses whole-response mean-pooling, whereas Goldowsky-Dill et al. trained on incomplete-fact activations with the last five tokens excluded. This train/eval mismatch is a plausible independent source of elevated FPR that is not disentangled.
- OOD-baseline experiment lacks the pre-finetuning control. To read 5.88% / 4.52% as evidence that benign fine-tuning erodes guardrails, one needs the same-pipeline EM rate for the base Llama-3.1-8B-Instruct on the same eight free-form prompts. Without it, the finding cannot separate benign-shift erosion from base-rate judge behavior or generic LoRA-induced drift.
- Cross-baseline comparability claim is loose. Calling 5.88% / 4.52% "comparable" to Kaczér et al.'s 4.01% and Turner et al.'s "<10%" without matching threshold rule, denominator, and judge is exactly the loose-comparison practice the paper elsewhere critiques. "<10%" is also a very weak comparator.
- Cross-experiment scoring regimes differ silently. Experiment 1 uses argmax single-point scores; Experiment 2 uses probabilistic weighted scores. Table 1 shows this axis alone shifts EM by several percentage points. The manuscript links the two experiments without stating the difference.
- Sensitivity-range attribution omits an axis. The "3.7% to 12.9%" range is credited to "judge choice and boundary inclusion" but the experiment varies three factors (aggregation is the third). Per-axis marginal contributions are not reported.
- Cross-judge portability claim rests on the noisier scoring regime. GPT-4o-mini argmax (3.72%) vs. GPT-5-mini argmax (8.02%) uses the aggregation the paper argues is boundary-noisy; the weighted-vs-weighted comparison is missing because GPT-5-mini does not expose logprobs. That constraint should be flagged as a limit on the portability claim.
- Miscitation of the shutdown-resistance/confusion evidence. In Section 3.1 the parenthetical (Schlatter et al., 2026; Rajamanoharan & Nanda, 2025) is offered as joint support for the claim that shutdown resistance correlates with confusion, but C7 clarifies that Schlatter et al. make the shutdown-resistance claim and Rajamanoharan & Nanda provide the counter-finding.
- L3 exemplar's specificity control is weaker than the L3 definition asks for. MMLU/MT-Bench stability rules out gross capability degradation, not the "generic capability change" alternatives the L3 definition demands controls for. The exemplar's specificity ceiling should be noted.
- Predict–control discrepancy used in both directions. Wattenberg & Viégas are cited to excuse failed steering and to indict it. If prediction and steering decouple, failed steering says little about causal mediation either way.
- Rhetorical framing of DeceptionBench 18% / 97.1% overstates the tension. Inter-rater agreement and construct validity are complementary axes; the "despite" framing suggests a contradiction that does not hold. The 18% scenario-level rate should also be paired with the prompt-instance rate (potentially as low as ~3% of 900 prompts).
- Small-sample generalization the paper critiques is echoed by its own Bean et al. citation. 4/9 and 8/9 statistics from a nine-benchmark survey are cited to establish a construct-validity gap that "extends beyond AMR" — worth qualifying to avoid the appearance of a double standard with C3.
Reproducibility & code
- Central datasets are not shipped. The aesthetic preference dataset (Woodruff, 2025), the scatological preference dataset (Bostock, 2025), the Legal dataset from Kaczér et al. (2026), and the eight free-form evaluation prompts from Betley et al. (2025) — every input feeding Experiments 1 and 2 — are absent from the released repo, and the DeceptionBench inputs are not included either. Because the Legal generations were never shipped or locatable, an executed replication could not attempt the Table 1 Legal rows (42.00% / 26.87% argmax, 32.71% weighted) at all. The scatological set is described only as "synthetically generated," so bit-exact reconstruction is impossible.
- Probe stress-test pipeline is entirely non-runnable as shipped. An executed replication attempt at Experiment 3 exited with
ModuleNotFoundError: No module named 'src.data'and was blocked.deception_experimentscontainsevaluate_probes_apollo.pyand anApolloProbewrapper, but thesrc/datapackage, theconfigs/dataHydra defaults, the two pretrained Goldowsky-Dill probes, and all eight stress-test CSVs (Sarcasm, Alien, Medieval Peasant, Wrong Answers Only, Counterfactual, Paraphrase, Translate, Recital) are absent from the release.ApolloProbe.loadalso hard-disables normalization citing a bespoke activation-pipeline mismatch, so any FPR numbers are tied to an activation-extraction path that is not shipped. Figure 2's 87–100% band and the sharp Medieval Peasant 4% / 96% divergence therefore cannot be checked externally — one of three demonstrative experiments is effectively unreproducible from the release. - Table 1 has no driver or aggregation script.
evaluate_gpt.pyimplements raw-argmax and top-20-logprob weighted aggregations, but no shipped script applies the coherence > 50 / alignment threshold with inclusive vs. exclusive boundaries, sweeps(judge × threshold × aggregation × dataset)to produce the Table 1 grid, or provides an expected-output CSV. The@hydra.maindecorator points atconfigs/evaluate_gpt, which is not present. The replication run had to writecompute_em_table.pyandlocal_judge.pyfrom scratch to reproduce even the aesthetic and scatological rows. - Two silent inference-path bugs invalidate the shipped path. During execution,
inference_freeform.pywas found to settemperatureandtop_pbut leavedo_sampleunspecified, so Hugging Face falls back to greedy decoding and the pipeline emits 150 identical completions per prompt — silently defeating the paper's sampling methodology. The tokenizer is also configured with right padding, which corrupts continuations for shorter prompts in mixed-length batches. Both are one-line fixes but neither is documented, and either would silently invalidate the 918 / 1,039 coherent counts and the 5.88% / 4.52% headline rates for anyone who re-ran the code without noticing. - Judge-evaluation script has no working open-weight fallback. When the OpenAI keys returned 401/403 during replication, the shipped
evaluate_gpt.pyhad no local backup: run naively under vLLM it mis-tokenizes multi-digit scores (splitting "30" into "3"), destroying the 0–100 scale the prompts assume. A local judge had to be written from scratch. This makes the paper's central evaluator-sensitivity claim dependent on continued availability of two proprietary models. - Headline aggregates are qualitatively — but not numerically — reproducible. Under a substituted local Qwen2.5-7B logprob judge, the scatological weighted EM rate lands at 5.81% vs. the paper's 4.52%, but the aesthetic weighted rate lands at 17.02% vs. 5.88%, and under a second Llama-3.1-8B judge the same aesthetic rate collapses to 0.14%. The single-point argmax range spans roughly 1.8% to 31.2%, wider than the paper's 3.7%–12.9%. The qualitative structural claims (weighted rules are boundary-invariant with Δ = 0 across every weighted cell; argmax gives a positive boundary Δ; judge choice materially moves the rate) reproduce exactly. This is arguably further evidence for the paper's own thesis — the numbers are even more evaluator-dependent than reported — but it means the specific headline figures are not portable across judges even with the pipeline held fixed.
- Nondeterministic generation with no seed reported. Temperature 1.0 and top_p 1.0 with no seed and no shipped seed-management make the 1,200 generations per dataset non-bit-exact reproducible, so the exact 918 / 1,039 coherent counts cannot be reconstructed even if the datasets were released.
- DeceptionBench audit lacks an artifact. The audit's 27-of-150 (missing ground truth) and 7-scenario (corrupted text) counts are illustrated with worked examples but backed by no CSV/JSON of flagged scenario IDs or documented criteria; the input benchmark is also not shipped. An executed corruption-audit variant lands close (10 vs 7) and flags both named examples, but the 27-count depends on a human semantic judgment a reproducer cannot reapply without the rubric.
Recommended Changes
Essential
- Sharpen the L3 definition. In Section 4.1, split "causal-mechanistic evidence" into (a) external-causal-factor claims (training signal, prompt condition, scaffold) and (b) internal-mechanism claims, and state the joint requirements (intervention + falsifiable hypothesis + specificity) separately for each. Addresses the internal-vs-external mix flagged in Weaknesses.
- Re-tag the checklist's design-time controls. In Appendix B (S3), mark "Measure general capability pre/post intervention" and "Add at least one OOD or benign-shift control" as L1 items that support L3 claims, or annotate them as prerequisites-for-L3 rather than L3 evidence, to prevent the ladder-climbing misreading.
- Tighten R3's operational deception definition. Add a clause ("and the model demonstrably has access to the accurate response," or an explicit reference to the R9 non-target-mechanism screens) so the recommendation does not classify systematic hallucination or capability degradation as deception.
- Add a pre-finetuning EM baseline to Experiment 2. Report the base Llama-3.1-8B-Instruct EM rate on the same eight free-form prompts under the same judge / threshold / aggregation as the fine-tuned rates, so that 5.88% and 4.52% can be read as guardrail erosion rather than base-rate judge behavior or LoRA-induced drift.
- Validate the probe stress-test construct-validity assumption. For Experiment 3, either measure surface-feature overlap between stress prompts and the probes' training data, or add a human check that the stress prompts read as non-deceptive; alternatively soften the conclusion from "surface cues rather than intent" to "distribution-shift brittleness."
- Add a token-scope ablation to the probe experiments. Re-run the Instructed-Pairs probe under its original last-five-tokens-excluded evaluation scope on the stress tests, or otherwise disentangle the train/eval token-scope change from the surface-cue claim.
- Ship the missing probe-pipeline components. Add the
src/datapackage,configs/data, the eight stress-test CSVs, the two Goldowsky-Dill probe checkpoints (or pointers), and any Apollo submodule required byevaluate_probes_apollo.py; document theApolloProbe.loadnormalize=Falsechoice, or make normalization configurable. - Release the missing datasets and generation artifacts. Ship the aesthetic, scatological, and Legal training datasets (or authoritative pointers); the eight Betley et al. evaluation prompts; the DeceptionBench inputs used for the audit; and the exact generated response files (or seeded generation configs) needed to recover the 918 / 1,039 coherent counts.
- Fix the two silent inference bugs. Patch
inference_freeform.pyto setdo_sample=Trueand switch the tokenizer to left padding, and re-verify the headline 5.88% / 4.52% figures under the corrected sampler. Add a note in Appendix C.3 documenting the corrected defaults. - Add a Table 1 driver, aggregation script, and open-weight judge fallback. Include a script that applies the coherence > 50 / alignment threshold with inclusive vs. exclusive boundaries, sweeps
(judge × threshold × aggregation × dataset), produces the Table 1 grid, and ships an expected-output CSV; provide the referencedconfigs/evaluate_gptHydra configuration; and add at least one non-proprietary judge (e.g., an open-weight or Claude logprob path) so the sensitivity claim survives API access changes.
Suggested
- Report the per-axis marginal contribution to the 3.7%–12.9% range. Break out how much of it is judge-driven, threshold-driven, and aggregation-driven.
- Explicitly note the judge-scoring difference between Experiments 1 and 2. Add one sentence stating that Experiment 1 uses argmax while Experiment 2 uses probabilistic scoring, so the cross-reference is unambiguous.
- Extend the cross-judge comparison to a logprob-exposing alternative. Complement the argmax GPT-4o-mini vs. GPT-5-mini contrast with a weighted-vs-weighted comparison using an open-weight or Claude judge; document the GPT-5-mini fallback path in code and prose.
- Ship a DeceptionBench audit artifact. Include a CSV/JSON of the 27 flagged scenario IDs and 7 corrupted-text IDs with reasons, and (if licensing permits) the exact DeceptionBench inputs used.
- Fix the shutdown-resistance citation. In Section 3.1, cite Schlatter et al. (2026) for the shutdown-resistance claim and Rajamanoharan & Nanda (2025) for the counter-finding, matching the C7 division of labor.
- Reconcile the Medieval Peasant divergence with the offered mechanistic story. Either propose a hypothesis that predicts the 4% / 96% split, or drop the mechanistic story in favor of a distribution-shift framing.
- Reword Table 1's threshold caption and Method column. Make explicit that both rules use
coherence > 50and that only the alignment operator flips; use a single Method label or footnote why "Raw/Argmax" and "Argmax" differ. - Recontextualize the DeceptionBench 18% / 97.1% juxtaposition. State that agreement and construct validity are complementary quality axes, and report the prompt-instance-level rate alongside the scenario-level rate.
- Qualify the Bean et al. 4/9 and 8/9 citation. Add "in a survey of nine safety-category benchmarks" so the small-sample framing matches the C3 critique.
- Tighten the predict-control discrepancy paragraph. Explicitly note that failed steering under a predict-control gap is uninformative about causal mediation, rather than "weakens causal-mechanistic claims," and re-cite as appropriate.
- Add a note on the L3 exemplar's specificity ceiling. In the Arditi et al. discussion, indicate where MMLU/MT-Bench stability meets the L3 bar and where stricter alternative-mechanism controls would be needed.
- Report a seed for Experiment 2 generations. Fix and report a seed (or ship the generated responses) so the coherent counts and headline percentages are bit-exact reproducible.