SAI
← All ICML 2026 orals

Position: The Alignment Community is Unintentionally Building a Censor’s Toolkit

Sarah Ball, Phil Hackemann

OralPosition TrackReplication not startedPaper PDFOpenReview

Position: The Alignment Community is Unintentionally Building a Censor’s Toolkit

SAI paper + code review · Referee report

Summary

The paper argues that modern AI alignment methods — pre-training data filtering, post-training preference alignment, and inference-time control — are purpose-agnostic tools whose valence depends on who defines the values being aligned to, and that in the hands of authoritarian states or foundation model providers they can be repurposed for censorship and manipulation. The conceptual move is to refuse the community's tacit reading of "alignment" as inherently benevolent and treat it as a dual-use technology on par with other scientific advances that have been weaponized. The paper walks the three-stage control stack, maps each stage to a misuse pathway with a compact characteristics table (access, compute, expertise, ease and depth of modification), and grounds each stage in real-world evidence — Chinese CAC directives and the People's Daily "mainstream values corpus" for pre-training; refusal-target datasets and Xi-aligned chatbot plans for post-training; system-prompt tweaks and post-generation classifiers (Grok, Yi-large, DeepSeek) for inference-time. It situates the risk in three amplifying trends — adoption of AI as an information source, an oligopolistic model market, and global democratic backsliding — and closes with mitigations (verifiable-alignment auditing, standardized censorship/bias benchmarks, competitive pluralism, and user/researcher literacy) plus two steelmanned alternative views.

The reframing of alignment as dual-use is a genuine contribution to a community that too often treats it as unambiguously good. The three-stage taxonomy is clean and the Chinese and Grok case studies are handled with useful specificity. The main conceptual limitation is that several load-bearing moves — "easily" transformed, "will certainly" apply to LLMs, "fundamentally compromised" training foundations, "driven to suicide", jailbreaks that "guarantee" access — outrun the evidence, and Table 1 partially contradicts the "easily misused" framing for the two more consequential vectors. The evidence base is heavily China- and Grok-weighted, downstream fine-tuners are dismissed in a way that understates open-weights risk, and the "verifiable alignment" and "neutrality through diversity" proposals do not confront the sociocultural-bias objection the paper itself raises against value-setting elsewhere.

Strengths

  • Timely conceptual reframing. Naming alignment as a purpose-agnostic, dual-use technology and pushing the community past a tacitly benevolent reading is a genuine contribution — the "whoever defines those intentions and values" pivot in Section 2 is the paper's most durable move.
  • Clean taxonomy of control stages. The three-stage stack (pre-training filtering, post-training preference alignment, inference-time control) and the accompanying dual-use characteristics table give a compact, teachable frame for reasoning about who can weaponize which layer at what cost.
  • Specificity of the Chinese and Grok case studies. The CAC directives, People's Daily "mainstream values corpus", the 20-70k question / 5-10k refusal test sets, the Yi-large post-hoc refusal switch, and the Grok system-prompt episodes are described with enough concreteness that a reader can independently interrogate the underlying reporting.
  • Balanced structure with steelmanned alternatives. Section 6 raises both the "stop alignment" and "risk is overstated" positions and engages them substantively rather than dismissing them; the paper is honest about not calling for a halt to alignment work.
  • Situating misuse in enabling conditions. Section 4 usefully ties the technical dual-use claim to three amplifying trends (adoption as an information source, market concentration, democratic backsliding), which makes the "why now" argument sharper than a purely technical framing would.
  • Actionable mitigations at the right level of abstraction. Verifiable-alignment auditing, standardized censorship/political-bias benchmarks, and pluralism/awareness are policy-relevant without being off-the-shelf; the ambition to move beyond right-left political-bias benchmarks toward authoritarian-vs-liberal axes is a reasonable and useful research call.

Weaknesses

  • "Easily misused" framing outruns the paper's own Section 3 and Table 1. The abstract and the pivotal Section 2 sentence use "easily", but the detailed characteristics table classifies pre-training as very-high compute, high expertise, moderate-to-difficult to modify, and post-training as moderate; only inference-time control is genuinely easy. The strong headline framing contradicts the differentiated picture the paper itself argues for and risks a uniform-ease threat model that is analytically weaker than the actual one.
  • Persistence label in Table 1 tension with "trained away" prose. The table lists post-training's depth of modification as "persistent", but Section 3.2 explicitly says those modifications "can be trained away or circumvented" via adversarial methods. Persistence-under-nominal-use and durability-under-attack are different properties, and collapsing them under one word matters for the dual-use argument (a persistent-but-removable safeguard is very different for a censor than for a defender).
  • Overreach in causal claims. Several load-bearing sentences state stronger causal or predictive conclusions than the cited evidence supports: "fundamentally compromised" for Western-LLM Chinese-language behavior, "will certainly also apply to LLMs" for the state-informational-control extrapolation, and "driven to suicide due to a lack or faulty alignment" for the case reports. Each direction is plausible; the strength of assertion is what over-reaches.
  • "Removing the need for human feedback" mischaracterizes guideline-based alignment. As stated, the framing generalizes over Constitutional AI and Deliberative Alignment, but Constitutional AI still relies on human helpfulness feedback and uses AI-generated (not zero) preference labels; the more precisely-scoped Deliberative Alignment sentence a few lines later shows the paper knows the correct scope but does not apply it here.
  • Grok evidence cited in two mechanism categories with the same sources. Section 2.4 frames Grok as "purposely realigned" (reads as parameter-level) and Section 3.3 attributes the same episode, citing the same sources, to system-prompt changes. Being explicit about which incident supports which layer would let the same evidence do real work in exactly one category.
  • Evidence base heavily China- and Grok-weighted. The other jurisdictions the paper names (Vietnam, Thailand, Russia, Belarus, Iran) get a single sentence citing Vesteinsson et al., and Western examples beyond Grok are largely absent. Given the "systematically map" contribution claim, a broader case set — even one or two additional worked examples — would materially strengthen the systematic-mapping argument.
  • Downstream fine-tuners dismissed on outdated resource assumptions. The threat model excludes fine-tuners citing "significant resource constraints and limited reach". LoRA-style adapters, released Llama/DeepSeek/Qwen weights, and low-cost RLHF pipelines have moved that boundary; ideological fine-tuners with personal-companion apps can reach non-trivial audiences. The concentrated-power frame is defensible but incomplete.
  • UDHR treated as an instrument of "formal commitment". The definitional foundation for misuse criterion (i) leans on the UDHR "to which virtually all states are formally committed", but the UDHR is a General Assembly resolution rather than a treaty. Grounding this in the ICCPR/ICESCR, or explicitly citing customary status, would put criterion (i) on firmer legal footing.
  • Misuse criterion (ii) is broad and its threshold is under-specified. As written, criterion (ii) plausibly covers a very wide range of unilateral commercial decisions by foundation model providers. Without an articulated threshold for what "at the expense of billions of users" means, the definition risks either being trivially satisfied or applied inconsistently — a problem for what Section 2.2 is trying to do as the paper's normative anchor.
  • "Verifiable alignment" quietly re-collapses two distinct auditing regimes. Section 5.1 carefully distinguishes provider disclosure to auditors from provider-cooperation-free black-box benchmarks, then defines "verifiable alignment" as users being able to verify "exactly what values a model has been aligned with". Behavioral benchmarks reveal suppression patterns, not the alignment specification; specification-level verifiability requires the disclosure route.
  • Auditor-only disclosure does not restore the user-choice problem the paragraph opens with. The prose moves from users being unable to make a conscious choice to a remedy that only informs auditors, with the intermediate step (auditor findings surfacing to users) left implicit.
  • Standardized benchmarks proposal does not confront the §2.2 sociocultural-bias caveat. The paper acknowledges misuse judgements are unavoidably culturally situated, but then calls for benchmarks that "account for authoritarian tendencies" without saying who defines that, or how the benchmark maintainers avoid recapitulating the concentrated-value-setting problem the paper argues is the core dual-use risk.
  • Diversity-yields-neutrality analogy underweights fragmentation and filter bubbles. The journalism analogy in Section 5.2 asserts that diverse options "will ultimately ensure" approximate neutrality, but the media analogue has also produced echo chambers and audience sorting. Without a qualifier about access-and-exposure conditions, this leaves an obvious counterargument on the table.
  • The marginal risk contribution of alignment vs. existing content moderation is not isolated. The "How is This New?" framing depends on alignment being distinct from prior content-moderation infrastructure, but several of the discussed vectors (keyword filtering, classifier moderation, prompt-driven policy) are close cousins of long-standing techniques. Making explicit what alignment adds on top — scale, persistence across queries, cross-lingual transfer, hidden-from-user compliance, generative synthesis — would head off the "rebranding" objection.
  • Word/Microsoft analogy in the libertarian steelman is left un-rebutted. The passive/generative distinction is exactly the crux of the alignment debate, and Section 6 does not engage it even briefly, leaving the steelman effectively unchallenged on its most important disanalogy.
  • Minor internal inconsistencies. "Total suppression" does not match the selective/targeted definitions the paper itself quotes; the mitigation count is "three key directions" in Section 5 but four in the abstract/introduction; "jailbreaks can guarantee access" overstates jailbreaks even inside a steelman; the Reuters six-country sample is generalized to a worldwide-adoption claim; user-count is used as a proxy for trust, and trust then drives the impact claim (mildly circular).

Recommended Changes

Essential

  • Soften the "easily" framing to match Table 1. Rewrite the abstract and the pivotal Section 2 sentence ("if in the hands of a malicious actor, a well-intentioned system can easily be transformed") so the ease claim applies specifically to inference-time control and is qualified for post-training and pre-training. Present the differentiated picture as a feature of the argument rather than an inconsistency.
  • Tighten causal language on load-bearing claims. Downgrade "fundamentally compromised" (Section 3.1 Real-World Evidence), "will certainly also apply to LLMs" (Section 2.4), and "driven to suicide due to a lack or faulty alignment" (Section 5.4) to formulations that match what the citations actually establish — e.g., "one plausible mechanism", "is likely to", "cases in which LLM interactions have been implicated".
  • Resolve the Grok cross-reference. Attribute the specific Grok incidents to the specific mechanism (system-prompt vs. deeper realignment) with matching citations, so the same case is not counted as evidence for two different alignment layers.
  • Fix the Table 1 depth-label ambiguity. Replace "persistent" for post-training with something like "persistent but removable under adversarial fine-tuning", or add a footnote defining persistence explicitly against inference-time superficiality. While there, correct the "modifcation"/"superfcial" spellings.
  • Scope guideline-based alignment precisely. Rewrite the Section 3.2 framing sentence about "removing the need for human feedback collection" so it matches what Constitutional AI and Deliberative Alignment actually do (remove human labels for safety/harmlessness; substitute AI-generated preferences), consistent with the Deliberative Alignment sentence a few lines later.
  • Broaden the real-world evidence base. Develop at least one or two of the currently one-sentence non-China cases (Vietnam / Thailand / Russia / Belarus / Iran, or a Western commercial example beyond Grok) with the same level of specificity as the Chinese examples, so the "systematically map" contribution claim is grounded across regimes.
  • Fold downstream fine-tuners into the threat model. Add a paragraph acknowledging that open-weights ecosystems (Llama, DeepSeek, Qwen) and cheap adapter methods have lowered the fine-tuner resource barrier, and briefly characterize the fine-tuner misuse pathway — even if the paper's main analysis stays on the concentrated actors.
  • Tighten the misuse definition in §2.2. Ground criterion (i) on binding instruments (ICCPR/ICESCR) alongside or instead of the UDHR, and give criterion (ii) a threshold (e.g., a rights-based or comparative harm test) so it does not sweep in ordinary unilateral commercial decisions.
  • Distinguish behavioral- vs. specification-level verifiability. Rewrite the "verifiable alignment" definition in Section 5.1 to explicitly separate what black-box benchmarks can show (behavior/suppression patterns) from what disclosure to auditors can show (alignment specifications), and connect each mitigation to the appropriate route.
  • Add the sociocultural-bias caveat to the benchmark proposal. Explicitly address in Section 5.1 who defines "authoritarian tendencies" and how the benchmarking programme avoids recapitulating the value-setting concentration the paper flags as the core dual-use risk.

Suggested

  • Add "access and exposure" qualifier to §5.2. Note that pluralism among models yields approximate neutrality only if users are aware of and can access alternatives, connecting §5.2 to the awareness argument in §5.3.
  • Rebut the Word/Microsoft steelman. Add one or two sentences in Section 6.1 (or immediately after the steelman) noting the passive-vs-generative distinction, so the strongest disanalogy in the libertarian argument does not go unaddressed.
  • Reconcile "three key directions" with the abstract's four contributions. Either enumerate four directions in §5 or explicitly explain that researcher reflection is a sub-pillar of awareness, and align the abstract/introduction accordingly.
  • Resolve the "auditor disclosure -> user choice" step. Rephrase the opening of §5.1's third paragraph so the remedy (auditor disclosure with public findings) matches the problem (user choice), or introduce a step from auditor verdicts to user-facing reporting.
  • Isolate the alignment-specific marginal risk. Add a sentence or two clarifying what alignment methods add on top of pre-existing content-moderation infrastructure (scale, persistence, cross-lingual transfer, hidden compliance, generative synthesis) — this makes the "How is This New?" claim sharper and heads off the "rebranding" objection.
  • Fix minor wording. Replace "total suppression" with "selective/targeted suppression"; soften "jailbreaks can guarantee access" to "can provide" even inside the steelman; either broaden the Reuters six-country sample framing or acknowledge the geographic scope explicitly; and either cite a trust measure directly or reword the user-growth-implies-trust sentence.