SAI
← All ICML reports

Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling

Zhibin Duan, Guowei Rong, Zhuo Li, Bo Chen, Mingyuan Zhou, Dandan Guo

SAI review preview

SAI reviewed this ICML 2026 paper and its available research artifacts. Sign in to read the complete referee report, inline feedback, and execution-based verification where available.

39Detailed review comments
CompletePaper and code review
4% of graded claims reproducedReplication status