← All ICML reports
Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling
Zhibin Duan, Guowei Rong, Zhuo Li, Bo Chen, Mingyuan Zhou, Dandan Guo
SAI review preview
SAI reviewed this ICML 2026 paper and its available research artifacts. Sign in to read the complete referee report, inline feedback, and execution-based verification where available.
39Detailed review comments
CompletePaper and code review
4% of graded claims reproducedReplication status