SAI
← All ICML reports

Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models

Yanchen Yin, Dongqi Han, Linghui Li

SAI review preview

SAI reviewed this ICML 2026 paper and its available research artifacts. Sign in to read the complete referee report, inline feedback, and execution-based verification where available.

35Detailed review comments
CompletePaper and code review
Execution review not startedReplication status