Position: The AI Imperative: Scaling High-Quality Peer Review in Machine Learning
SAI paper + code review · Referee report
Summary
This position paper argues that the ML community should proactively build an AI-augmented peer-review ecosystem — LLM assistants for authors, reviewers, and Area Chairs — as the response to a scalability crisis (NeurIPS submissions up ~10× since 2014; ICML +48% YoY; ICLR inter-reviewer variance rising). The conceptual move is to reframe peer review from a workload problem into a research and infrastructure agenda: not just deploying assistants for isolated tasks, but instrumenting the workflow so that score changes, rebuttal-to-judgment links, and AC deliberation become structured supervision that AI systems can learn from rather than merely imitating outcomes. That reframing — "the near-term bottleneck is not better models, but better process data" — is the paper's most interesting contribution and is placed usefully next to institutional alternatives (publish-first curation, ICML 2026's LLM policy, human-only positions). Illustrative experiments on ICLR 2024/25 data show LLMs can extract review points (higher recall for strengths and rebuttals than for weaknesses) but struggle to predict ratings under in-context learning, which the authors read as evidence for the data agenda. The argument is timely and unusually broad in engaging alternative designs. The main limitations are (i) governance proposals for enriched data collection (opt-in enclaves, tiered access, active elicitation) are stated briefly rather than defended against realistic incentives and the paper's own evidence of reviewer disengagement; (ii) several illustrative results — the initial-rating MAE (worse than a mean-predictor baseline yet framed as "reasonable"), the reinterpretation of high MAE as "data-gap evidence" rather than an ICL/subjectivity limit, and a "200-review pilot experiment" in tension with the paper's no-human-subject-study statement — do not carry the argumentative weight asked of them; and (iii) the anti-gaming and anti-injection defenses (hidden report-card weights, a Dual-LLM "strip-all-instructions" pipeline) sit in tension with the paper's transparency/auditability commitments and with the canonical designs they invoke. The vision is a plausible research agenda; persuasive power would improve substantially if the experimental section and governance proposals were tightened.
Strengths
- Timely conceptual reframing. Positioning the bottleneck as missing process supervision (score-change rationales, rebuttal-to-objection links, AC deliberation traces) rather than as an LLM-capability gap is a useful move that suggests concrete community infrastructure work rather than just better prompts.
- Serious engagement with alternative institutional designs. Section 7 discusses publish-first curation, protocol reform, ICML 2026's LLM policy, and human-only positions and takes each seriously rather than dismissing them; this is unusually well done for a position paper.
- Broad, well-scoped ecosystem view. The paper enumerates concrete AI capabilities (Retrieval Augmented Verification, code/reproducibility analysis, report cards, provenance detection, decision support) and maps them to authors, reviewers, and ACs, which makes the vision easy to critique and extend.
- Honest headline experimental finding. The 0.632 weaknesses-recall vs. 0.927 strengths-recall gap is presented as a "critical thinking gap" that motivates humans-as-senior-partner, rather than being spun into a claim of LLM-as-reviewer readiness. This calibrated framing is appropriate.
- Clear articulation of the peer-review-as-AI-testbed argument. Section 2's case that peer review is a demanding, real-world benchmark for grounded reasoning under uncertainty is compelling and could plausibly influence how the community sets up shared datasets.
Weaknesses
- Central empirical claim is undermined by the initial-rating MAE. An MAE of ~2.29 on a 1–10 scale where ratings cluster around 4–6 with σ ≈ 1–1.5 is worse than a constant-mean predictor. Framing this as a "reasonable baseline" (Section 6) misleads the reader and weakens the argument that ICL performance is a useful diagnostic. The ~3.4× gap between initial-rating MAE (2.29) and final-rating MAE (0.67) is left unexplained, though it plausibly reflects reviews telegraphing the final score.
- "Data gap" causal claim is not identified by the experiment. The paper asserts that the high MAE "should not be interpreted as a fundamental limitation of LLMs, but rather as evidence of the information gap in current public datasets," while itself elsewhere naming ICL ceilings and inherent subjectivity as alternative causes. No ablation isolates data richness from model capability or subjectivity; as stated the claim is unfalsifiable.
- Figure 2 conflates dialogue volume with rebuttal influence on decisions. The caption and body text both leap from "reviewers reply little" to "the rebuttal stage exerts limited influence." A reviewer who silently updates a score after reading a rebuttal registers as engaged=zero in this dataset. A direct decision-influence signal is needed to back this claim.
- A "200-review pilot experiment" contradicts the paper's own no-human-subject-study statement. Section 9 explicitly says "We introduce no deployed system or human-subject study," yet Appendix A.1 reports word-count deltas and citation increases for 200 reviewers who received report cards. Either the pilot was really run (and needs methods/consent reporting) or it is a paraphrase of Kim et al. 2025 / Thakkar et al. 2025 (and should be attributed as such).
- Design proposals conflict with the paper's transparency and auditability commitments. The anti-gaming remedy — "the weights for the reviewer report card should be stochastic or hidden" — directly contradicts the ecosystem's stated need for "accuracy, transparency, and trust" and "auditable meta-review scaffolds." The Dual-LLM sanitizer's claim of stripping "all formatting and active instructions" also overstates a mitigation that in the canonical design works by isolation, not by trusting an LLM to reliably remove instructions.
- Anti-gaming remedy is attached to the wrong artifact. As defined, the Reviewer Report Card scores human-written reviews, not manuscripts — authors would not be its subjects. Placing "author gaming" mitigations on the reviewer-facing rubric misidentifies which instrument needs hardening.
- "Simulated reviewer" and "reviewer persona" proposals are inconsistent with the paper's own reviewer-side limitation. The paper repeatedly caveats that LLMs "still struggle with deep fault detection… identifying faulty reasoning or contradictions" — yet the author-facing simulated-reviewer proposal assumes reliable weakness detection. The symmetry between the two sides is not acknowledged.
- Table 1 misattribution. The sentence linking Table 1 to a "further complication" of reviewer superficiality does not follow: LLM proficiency at parsing review/rebuttal text is orthogonal to whether reviews are becoming more superficial.
- Governance proposals are thin relative to their argumentative weight. The tiered-access / opt-in / confidential-enclave paragraph is the operational heart of the "richer data" pillar but is one paragraph. It does not defend incentives (why would already-disengaged reviewers opt into more instrumentation?), cost allocation, or bias from selective opt-in.
- Alternative-views section under-engages the strongest counterarguments. De-skilling is raised in one paragraph and rebutted with a single sentence; bias amplification from LLM-based co-piloting (e.g., systematic downrating of unusual methods, retrieval-grounding bias against under-represented subfields) is not discussed at all. For a position paper this is where the argumentative gap is widest.
- 2025-vs-2024 recall gain has a confound the paper flags but does not warn against. If ICLR 2025 reviews are increasingly LLM-drafted, higher LLM recall may reflect stylistic self-similarity rather than better extraction. The manuscript does not stratify or caveat this in the recall interpretation.
- Simulated-reviewer scale is inconsistent with the paper's ICLR framing. The 1–10 overall descriptors used in Appendix B.1 are NeurIPS-style ("impact on… areas of AI"), not ICLR's "acceptance threshold" language, while the system message is titled "You are a simulated reviewer for ICLR 2025." If ground-truth ratings are on ICLR's scale, the MAE/RMSE reflect a scale mismatch that is not disclosed.
- CAGR/multiplier presentation is not self-consistent. 10.4× over 10 years implies a CAGR of ~26.6%, not the stated 26.4%; rounding-level but easily fixed.
Reproducibility & code
Working from the code/data availability statements in the manuscript (paper-code mode; released materials read, not executed):
- Section 6 methods are described in conditional language. Appendix B.1 says "we would conceptually use publicly available data from OpenReview" while Section 6's Tables 1 and 2 report concrete numeric results. The methods description does not pin down: (i) the exact ICLR 2024/2025 test-paper subsets and counts, (ii) how the few-shot ICL exemplars are drawn (fixed set vs. per-example resample), (iii) o3's checkpoint/date and sampling parameters (temperature, top_p, seed), or (iv) the identity of the "another powerful LLM" acting as judge. Without these a reader cannot recover the tabulated numbers.
- Table 1 CI methodology is opaque. Several rows carry ±0.000 while their numerator (Hits) and denominator (Real Pts) show large per-paper variance, and other rows in the same table use non-trivial CIs (±0.04–±0.06). The CI construction is not stated; the caption's "Recall is the quotient of the two" does not match the reported means (e.g., 2.73/4.58 = 0.596 vs reported 0.632 for weaknesses), suggesting per-paper mean of ratios but no method disclosure.
- Table 2 CIs are surprisingly tight and unexplained. Cells carry ±0.005–±0.011 CIs, and Score-change RMSE at n=0 and n=1 are identical to four decimals (1.0908) — a tie that begs a check on how CIs and metrics are computed. No bootstrap procedure or resample count is documented.
- Case-study Report Card (Table 3) is not reproducible. The per-dimension scores driving the 3.1/5 composite are LLM outputs, but Appendix B.4 explicitly leaves the few-shot exemplars unpublished and does not report sampling parameters or single-run-vs-majority-vote convention. The typo-detection example (
"reslim" → "realism") is deterministic, but the four-bullet structure and specific score set are not. - ≈ 4 s Report Card latency lacks methodology. No hardware/endpoint, context length, single-call vs. p50 over trials.
- 200-review pilot lacks any methodology. Beyond the internal inconsistency with the "no human-subject study" statement, the pilot has no cohort assignment, control condition, revision-window definition, or effect size CI. A reproducer cannot know whether to attempt replication or fetch the numbers from Kim et al. 2025 / Thakkar et al. 2025.
- Figures 1 and 2 lack the definitions and code needed to reproduce. OpenReview data is public, but the rating field per year (ICLR's scale changed 1–8 → 1–10 across the window), the "silent reviewer" definition, per-submission-std handling of ≥2-review filter, tokenization rules for word counts, and whether AC/program-chair comments are excluded are all left implicit. The caption's "2019–2024" window also disagrees with the figure axes, which run to 2025.
- LLM-as-judge output does not implement the summarized metric. The judge prompt asks to count "generated points that overlap" (a precision-flavored count), while the summary and appendix notes describe recall as "human points captured." As specified the pipeline yields inconsistent metrics.
Recommended Changes
Essential
- Resolve the 200-review "pilot experiment" inconsistency. Either report full methods (cohort assignment, control, IRB/consent, effect sizes with CIs) or re-attribute the 28% word-count / 1.7-citation numbers to Kim et al. 2025 / Thakkar et al. 2025 and remove the "In pilot experiments across 200 ICLR 2025 reviews…" framing. This resolves the direct contradiction with the Section 9 "no human-subject study" statement.
- Recalibrate the interpretation of the initial-rating MAE. Add a mean-predictor and initial-review-only baseline for Table 2, drop the "reasonable baseline" wording for MAE ≈ 2.29 on a 1–10 scale, and give a candidate mechanism for the ~3.4× initial-vs-final MAE gap.
- Ablate the "data gap" causal claim or soften it. Either run an experiment that varies data richness/model capability so the "information gap in current public datasets" claim is identified, or reframe the sentence as one hypothesis among the three the paper itself lists (ICL ceiling, subjectivity, data gap).
- Rewrite Figure 2's caption and Section 2 body around dialogue volume, not decision influence. Replace "the rebuttal stage exerts limited influence" with an engagement claim, or add a score-change-conditional-on-rebuttal analysis to actually support the influence claim.
- Publish full Section 6 methods. State the exact ICLR 2024/2025 paper subsets and counts, ICL exemplar selection rule, o3 checkpoint and sampling parameters, LLM-as-judge model identity, and CI construction (bootstrap type, resample count) — the ±0.000 CIs in particular need explanation. Update Appendix B.1 to remove conditional "we would conceptually use" language.
- Fix the recall definition or rewrite it. Either revise the Table 1 caption to describe recall as a per-paper mean of ratios (with worked example) so that the reported means and the definition are consistent, or recompute recall as the exact quotient of the reported means. The weaknesses row (where the argument turns) currently disagrees by 0.036.
- Reconcile the anti-gaming proposal with the transparency commitment. Either replace "hidden or stochastic weights" with a transparent-rubric plus adversarial-evaluation defense, or explicitly acknowledge and defend the trade-off with auditability; also move the mitigation from the reviewer-facing Report Card to the author-facing manuscript scorer, which is the artifact authors could actually game.
- Correct the Dual-LLM sanitizer description. Cite the canonical dual-LLM design, drop "all formatting and active instructions," define "active instructions," and describe validation of the stripping step.
Suggested
- Add a bias-amplification and de-skilling subsection to Section 7. The current one-line rebuttal is thin; discuss retrieval-grounding bias against under-represented subfields, systematic downrating of unusual methodology by LLM co-pilots, and empirical proposals for measuring de-skilling.
- Defend the governance proposals more concretely. Extend the tiered-access / opt-in / enclave paragraph with a discussion of reviewer incentives (given the paper's own disengagement evidence), cost allocation, and mitigation of opt-in bias.
- Warn readers about the 2025 recall confound. Add an explicit caveat that if ICLR 2025 reviews are increasingly LLM-drafted, the higher recall may partly reflect LLM-to-LLM stylistic self-similarity rather than better extraction; consider stratifying by AI-authorship suspicion.
- Clarify the "simulated reviewer" prompt's rating scale. Reconcile the NeurIPS-style 1–10 descriptors with the ICLR 2025 system-message framing, and state which score the Table 2 predictions target and how it relates to the 1–4 dimension scores.
- Diagnose the non-monotonic ICL performance. For the final-rating and score-change tasks, add a prompt-length ablation or an exemplar-distribution check to explain why performance worsens beyond n=1.
- Publish the Report Card case-study prompt and exemplars. Currently the case study is the paper's most concrete illustration but is the least reproducible artifact; even one full worked prompt would let readers stress-test the claim.
- Cite alternative examples for the AC-decision-support exclusion clause. The ICML 2026 policy illustrates only the "non-evaluative" clause; add or substitute an example that speaks to AC-side restrictions.
- Fix the CAGR/multiplier arithmetic. State one figure and derive the other, or round consistently so that "10.4×" and "26.6%" agree.
- Reconnect the "simulated reviewer persona" proposal to the paper's own deep-fault-detection caveat. Note explicitly that author-facing simulated reviews inherit the reviewer-side reliability limitations.
- Repair the Experiment 1 prompt template. Split strengths-generation and rebuttal-generation into two prompts with task-appropriate instructions, and update the LLM-as-judge output format so precision (as promised in the evaluation notes) can be computed and recall is unambiguously counted over human points.
- Publish Figure 1/2 preprocessing and plotting code. Specify rating-field selection per ICLR year, "silent reviewer" definition, filtering rules, and tokenization; align the caption's year range with the axes.