SAI
← All ICML 2026 orals

Position: There are futures that benchmark-driven AI cannot see

Sobhan Lotfi, Ava Iranmanesh, Lachin Naghashyar, Ali Shirali, Fateme Nateghi Haredasht, Sanmi Koyejo, Philip Torr, Yong Suk Lee, Fazl Barez, Joel Lehman, Peter Norvig, Arvind Narayanan

OralPosition TrackReplication not startedPaper PDFOpenReview

Position: There are futures that benchmark-driven AI cannot see

SAI paper + code review · Referee report

Summary

This is an ICML position paper arguing that benchmark-centered evaluation in AI imposes an "exaptation tax" — a reduction in the survival of work whose value is latent, indirect, or only revealed under a future problem regime. The lens is Gould and Vrba's biological notion of exaptation: traits (or ideas) that evolved under one selection rule and later become decisive under a different one. The paper's central conceptual move is to argue that the field's guiding question is shifting from Q1 (can machines exhibit intelligent behavior?) to Q2 (can they do so while remaining aligned, interpretable, safe, and governable?), and that Q2 differs from Q1 not in degree but in kind: Q2's success criteria are "human-side," constructed through deliberation, rather than "world-side," fixed by the task. On top of this framing, the paper diagnoses three mechanisms by which benchmark culture governs research — community composition, research agenda, and the limits of evaluation — and proposes institutional interventions: plural evaluation regimes, protected exploratory venues, long-horizon funding, and reflexive governance of selection rules. The historical illustrations (GPUs, U-Net repurposed for diffusion, LSTM's slow adoption, attention→transformer→AlphaFold2, backpropagation's cross-disciplinary origins) are well chosen and the paper's tone is measured — the authors do not argue against benchmarking, only against its calibration in a moment when Q2 demands more exploration. The contribution is genuine: framing evaluation critique through exaptation rather than the more familiar language of overfitting-to-metrics or construct-validity puts a new emphasis on persistence of latent-value work as the mechanism being taxed. My main concerns are that (i) the load-bearing Q1/Q2 dichotomy shifts characterization midway and slides into the objective/subjective framing the paper explicitly disavows; (ii) several supporting citations are pushed harder than the sources support; (iii) some sweeping claims are asserted flatly where hedged phrasing is available; and (iv) the proposals do not confront the funding-structure diagnosis they inherit from §3.1. The paper is a candidate for acceptance in the position track once these are addressed.

Strengths

  • Conceptual contribution. The exaptation lens is a genuinely fresh angle on evaluation critique in ML. Prior work (Raji, Recht, Hooker, Bechler-Speicher) has surfaced construct-validity and hardware-lottery concerns; framing the persistence of latent-value work as the phenomenon being taxed adds an evolutionary/temporal dimension that the more static "goodhart" or "construct validity" framings lack.
  • Clear reformulation of the guiding question. The Q1/Q2 recasting is the paper's most portable idea. Even where the sharp dichotomy is not fully defended (see Weaknesses), the reframing gives the such-that discussion a clean structural argument rather than a list of unrelated concerns.
  • Range and quality of historical examples. GPUs as infrastructural exaptation, U-Net for diffusion, LSTM's slow adoption, the ACL→transformer→AlphaFold2 chain, and the backpropagation origin story are well chosen. They span hardware, architecture, and algorithm, which supports the generality of the exaptation mechanism rather than reading as one narrative repeated four ways.
  • Anticipation of objections in §6. The Success, Invisibility, Protection, Industry Lab, and Measurement objections are the right five to engage, and the responses show the authors have thought about how the argument fails, not only how it succeeds. The Invisibility Objection response in particular is honest about counterfactual under-identification.
  • Proposals are grounded in adjacent literatures. The plural-regime and protected-venue proposals connect to novelty search, quality-diversity, and Zollman-style epistemic-network results, rather than being pulled from thin air. This gives §5 an operational credibility beyond most benchmark-critique papers.
  • Institutional and normative framing. The paper explicitly notes that benchmark culture is now being encoded into regulatory frameworks, and that evaluation assumptions hardening into legal standards are much harder to revise than community norms. This is a real and underappreciated stake for a technical-community position paper.

Weaknesses

  • Q1/Q2 dichotomy shifts characterization. The Introduction claims Q1 has "ground truth that exists in the world independently of who evaluates the system," but §2.2 describes benchmarking as a "pragmatic bypass" that measured what systems can do rather than resolving what intelligent behavior is. Section 4.2 then recasts the distinction via Simon as objectives that "can be specified in advance." These are three subtly different characterizations, and the shift between them is not acknowledged. The "in-kind, not in-degree" claim would be more defensible if it rested on 'specifiable in advance vs. requiring ongoing deliberation' consistently.
  • Objective/subjective disclaimer contradicted by the supporting citation. Section 4.2 explicitly rejects framing Q2 as "subjective" — yet the Fisher et al. citation is invoked on the grounds that neutrality is "inherently subjective." A reconciling sentence is needed or the paper's central Q2 characterization slides back into the framing it disavows.
  • Chouldechova impossibility used against its grain. The Chouldechova result formalizes tradeoffs among crisply specified fairness criteria; it is deployed here as evidence that Q2 goals resist specification. The theorem's premise (formal specification) is precisely what the paper argues is unavailable for Q2, so citing it inverts the source.
  • Causal overreach on "produced by it." The topic sentence in §3.1 claims narrowing is "produced by" benchmark culture, but the cited evidence (Klinger, Hao, OECD, Mazzucato) supports correlation with private-sector composition and funding structure. The paper's own subsequent phrasing "amplifies this dynamic" is defensible; the topic sentence is not.
  • Hao et al. mechanism-mismatch. The 41-million-paper analysis concerns researchers' use of AI tools narrowing themes — a distinct mechanism from benchmark-centered evaluation shaping community selection. Calling this "confirms" the benchmark-culture thesis blurs the causal story.
  • AlphaFold2 dependency claim ("required having had linguists in AI"). This is stronger than the historical evidence supports: AlphaFold2's architecture is not a transported NLP transformer, and attention had roots in the broader neural sequence-modeling community. The exaptation point does not require the counterfactual "required, could not have been"; "drew on" would carry the same argumentative weight without the historical overreach.
  • Universal impossibility claims overstate the argument. "No quantitative evaluation method can foresee exaptations" and "Benchmark culture applied to pre-paradigmatic problems does not accelerate progress toward the goal" are categorical claims stronger than the paper's own reasoning (which supports strong tendency claims about incompleteness and lock-in, not absolute impossibility).
  • "Validated by decades of use" overstates the epistemic status of the Q1 correlation. Use is not validation, and the paper itself elsewhere cites construct-validity critiques that undercut the very correlation being asserted here.
  • Invisibility Objection rebuttal leans on survivorship examples. Backpropagation, attention, and RLHF are things that did survive despite the selection regime; they are logically the wrong direction of evidence for a tax on lost work. The conservation-biology analogy is doing more work than acknowledged; naming concrete plausible losses (or foregrounding the Klinger stagnation finding as the strongest empirical anchor) would strengthen the section.
  • Proposals do not confront the funding-structure diagnosis they inherit. Section 3.1 argues that AI funding has shifted to VC and large tech, and that benchmark culture aligns with those funders' commercial horizons. Section 5 then proposes plural venues, protected communities, and long-horizon funding without naming who would design or fund them under the shifted structure. Even a short paragraph identifying the relevant actors (TMLR-style venues, national funders, AI Safety Institutes, journal boards, ICML review-quality initiatives) and their leverage would move the proposals from aspiration to institutional design.
  • Section 5 does not engage with existing pluralism. The venue landscape is already partly plural (TMLR, position papers, workshops, ACL, journals, safety institutes). A sharper diagnosis might be "prestige and career signal concentrate on benchmark-driven venues" rather than "the community lacks plural regimes." Engaging this reframing would sharpen the proposal.
  • Normative-only scoping brackets counterexamples. Parts of interpretability (mechanistic circuits), safety (formal verification, red-teaming coverage), and alignment (reward-model calibration) do admit relatively world-side operationalization and have been benchmark-driven. The paper waves at "the normative subset" but does not engage with the fact that the world-side/human-side boundary runs through several such-that clauses rather than around them.
  • Loose analogies in §5.1. Lottery schemes (stochastic) and the Rooney Rule (deterministic constraint) are grouped under "similar logic" despite operating on different mechanisms. The randomization-against-correlated-bias argument mischaracterizes what randomization corrects — randomization primarily addresses variance, not systematic tilt.
  • Historical/attribution slips. Non-Euclidean geometry is conventionally attributed to Bolyai/Lobachevsky/Gauss; Riemann introduced Riemannian differential geometry. The GAN "reviewer quote" is presented as "precisely documented" but is grammatically spliced and uncited. The SVM-era framing is placed immediately after "two AI winters," inviting readers to associate the SVM period (an active well-funded paradigm competition) with a funding winter.
  • Bibliographic inconsistencies. The Dartmouth proposal appears as two separate entries (1955, 1956) with inconsistent author-initial formatting; the same Raji et al. paper appears as 2021a and 2021b with different first-author initials. Q2's enumeration varies between three and four clauses across the abstract, Introduction, and §4.2 restatements.
  • COI boilerplate contradicts the disclosures. The paragraph discloses Microsoft, Lila Sciences, and Google employment and then declares that the research was conducted "in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest." The two sentences cannot both be literally true.
  • Discussion is thin. Section 7 is a two-sentence coda relative to the argumentative load of §3–§4; a paragraph on tractability, falsification conditions, or how the argument revises under specific empirical outcomes would substantially strengthen the ending.

Recommended Changes

Essential

  • Stabilize the Q1/Q2 characterization. Rewrite the Introduction and §4.2 to use a single, consistent formulation ("objectives specifiable in advance vs. requiring ongoing deliberation" is the strongest candidate), and acknowledge that Q1 was itself bypassed rather than resolved by benchmarks. Fix the objective/subjective slide by either replacing or explicitly reconciling the Fisher et al. citation with §4.2's disclaimer.
  • Repair the Chouldechova inference. Either replace it with a source about deliberation-constructed goals, or add a sentence explaining why formal-tradeoff impossibility supports the pre-paradigmatic claim despite presuming crisp specification.
  • Downgrade "produced by it" and re-cite Hao et al. Change the §3.1 topic sentence to "amplified by" (matching the paragraph's own subsequent phrasing) and change "confirms" to "is consistent with" for the Hao et al. citation, or add a source that isolates benchmark culture from co-varying industrial factors.
  • Soften universal claims. Replace "No quantitative evaluation method can foresee exaptations" and "does not accelerate progress toward the goal" with tendency/conditional phrasings that match the strength of the argument that supports them. Similarly downgrade "validated by decades of use."
  • Weaken the AlphaFold2 counterfactual. Change "required having had linguists in AI" and "could not have been" to "benefited from" / "drew on." The exaptation point survives; the counterfactual overreach does not.
  • Address who implements the proposals. Add a paragraph in §5 (or a subsection) naming the actors — TMLR-style venues, national funders, AI Safety Institutes, journal boards, ICML/NeurIPS review-quality initiatives — and their leverage relative to the funding-shift diagnosis in §3.1. Engage the counterproposal that pluralism already exists at the venue level and that the real intervention is redistributing prestige across it.
  • Confront benchmark-tractable such-that work. Add engagement with mechanistic interpretability, formal-verification-of-safety, and reward-model calibration as cases where the world-side/human-side boundary runs through an individual such-that clause. Explain why the exaptation-tax argument still applies to these subareas, or scope the argument to exclude them.

Suggested

  • Strengthen the Invisibility rebuttal. Foreground the Klinger-style stagnation-in-thematic-diversity finding as the strongest empirical anchor for the tax, and explicitly acknowledge that the backprop/attention/RLHF list is a survivorship set.
  • Tighten §5.1 mechanisms. Distinguish the Rooney Rule (deterministic constraint) from lottery schemes (stochastic), and correct the "randomization corrects correlated bias" argument — randomization primarily addresses variance among near-equal candidates, not systematic tilt.
  • Fix historical claims. Change "Non-Euclidean geometry, developed by Riemann" to "Riemannian geometry, developed by Riemann." Either cite the GAN reviewer quote to a specific source or drop the "precisely documented" framing. Separate the "two AI winters" claim from the SVM-era paradigm competition.
  • Reconcile bibliographic duplicates. Consolidate the two Dartmouth proposal entries and the two Raji et al. entries into single canonical entries, and fix the author-initial inconsistencies.
  • Stabilize Q2's enumeration. Pick a canonical list ("aligned, interpretable, safe, and governable") and use it in every restatement.
  • Rewrite the COI boilerplate. Either drop the second sentence or rephrase it so it does not directly contradict the employment disclosures.
  • Narrow the mid-2026 saturation claim. Either restrict it to LLM benchmarks (matching the Akhtar et al. scope) or add supporting sources for the broader claim.
  • Expand the Discussion. Add a paragraph on which proposal is most tractable in the next 12–24 months, what would count as evidence for or against the exaptation-tax mechanism, and how the argument revises if Q2 problems turn out to admit unexpected benchmark tractability.
  • Adjust the Zollman and Werbos characterizations. "A well-known modeling result illustrates" rather than "formal models confirm"; "developed and later applied ... to neural network training" for Werbos to preserve accuracy without weakening the exaptation point.