SAI
← All ICML 2026 orals

Position: Irresponsible AI: big tech’s influence on AI research and associated impacts

Alex Hernández-García, Alexandra Volokhova, Ezekiel Williams, Dounia Shaaban Kabakibo, Mélisande Teng

OralPosition TrackReplication score 79%Paper PDFCode repoOpenReview

Position: Irresponsible AI: big tech’s influence on AI research and associated impacts

SAI replication review · Referee report

Summary

This position paper argues that big tech's outsized involvement in AI research is a structural driver of what the authors label "irresponsible AI" (iAI) — environmental harm, labour degradation, mis/disinformation, weaponisation, and inequality. The conceptual move is to reframe familiar critiques of AI harms as consequences not of the technology in the abstract but of a specific actor set whose incentives (scale thinking, general-purpose systems, cloud/data-centre capex, growth imperatives) systematically favour irresponsible development. That framing is joined to a fresh empirical exhibit — a hand-built authorship analysis of NeurIPS, ICML, and ICLR papers over 2013–2025 documenting how industry-affiliated the ML literature has become — and to a concrete call to action aimed at researchers themselves, spanning individual reflection, collective organising, cooperative alternatives, and grass-roots and union-based avenues. The Alternative Views section is genuinely engaged rather than perfunctory, and the paper connects critical STS and political-economy literatures to a technical audience that rarely reads either.

The main conceptual limitations are (i) that "big tech" is defined narrowly in the prose but operationalised as a much broader industrial keyword list in the empirical analysis, quietly widening the paper's causal claims; (ii) that the authorship exhibit measures co-authorship presence yet is used to argue about agenda-setting influence, and its fraction actually declines after 2020, sitting awkwardly with the "growing influence" thesis; and (iii) that several inferential leaps — from cloud-market concentration to environmental costs being "constitutive" of the AI agenda, from growth-oriented press releases to actual under-regulation, and from industry co-citation to strategic agenda-shaping — are stated more strongly than the cited evidence supports. Re-executing the released code reproduces the headline authorship averages within about one percentage point and confirms the reported peak years, materially raising confidence in the empirical exhibit; the shortfalls are confined to ICLR's recent years, which are gated behind OpenReview access, and to undisclosed aggregation and matching choices rather than to disagreements in the computed numbers. Overall, the paper's core position is well-motivated and its argumentative architecture is unusually deliberate; strengthening it is mainly a matter of calibrating claims to evidence and making the empirical methodology fully transparent.

Strengths

  • Timely, well-motivated position. The paper connects a rich critical literature — political economy, STS, AI ethics, environmental critique — to a technical audience that rarely encounters it, and does so in service of a clearly stated normative thesis rather than as a survey.
  • Original empirical exhibit that reproduces. The NeurIPS/ICML/ICLR authorship analysis (Figure 2, Appendix A) is a nontrivial data-engineering effort, the code is released, and re-execution recovers the headline averages within ~1 percentage point (NeurIPS 20.7% vs 21.1%, ICML 23.6% vs 23.5%) and the reported peak years exactly, so the paper rests on its own reproducible base rather than only on prior citations.
  • Structured argument architecture. The four-move structure (influence → impact → economic drivers → call to action) is uncommon in position pieces and lets the reader see how the pieces are meant to fit together, including the reflective "implicated subjects" framing that positions researchers inside the argument.
  • Concrete call to action. Section 5 spends real effort on tractable, differentiated recommendations — reading groups, sponsorship policies, cooperatives, unions, whistle-blowing, community-centred research — rather than gesturing at "systemic change" without a lever.
  • Genuine engagement with alternative views. Section 6 takes on the strongest counter-arguments (progress-enabling, aggregate-energy modesty, top-down reform primacy, capitalism-as-least-bad) with actual argument, and concedes where warranted.

Weaknesses

  • Conceptual/operational gap in "big tech". The narrative definition centres on Google, Meta, Microsoft, Amazon and OpenAI, but the Appendix A keyword list folds in Intel, Nvidia, Samsung, IBM, Apple, Tesla, Uber, Baidu, Alibaba and others. Because the headline percentages are used to argue about "big tech's" agenda-setting influence, the reader is asked to treat two quite different actor sets as one.
  • Construct gap between co-authorship and influence. The exhibit measures the share of papers with an industry-affiliated co-author, but the text slides from this to claims that big tech "steers", "moulds", and "drives" the research agenda. Co-authorship presence is not a measure of influence, still less of irresponsibility, and the leap ("increasingly strong impact on the AI research community") is not argued.
  • Declining post-2020 fraction versus the "growing influence" framing. The paper's own data show the fraction peaking around 2020 and declining to ~14–17% by 2025, yet the framing is of growing, pervasive influence. The reconciliation — a shift to non-peer-reviewed venues — is asserted on a single citation and is load-bearing.
  • Big-tech-specific thesis versus general-capitalism attribution. Section 4 attributes iAI to "fundamental features of capitalism" affecting all corporations, which undercuts the distinctive causal role assigned to big tech; the paper concedes the counterfactual ("alternative actors could have led to similar outcomes") is out of scope, weakening the strong causal framing used elsewhere.
  • Overreach in the "self-citation" framing. The claim that industry-funded work shows "higher self-citation tendency" and thereby moulds the agenda mislabels an in-group/homophily citation pattern as self-citation and treats a citation-pattern observation as an established agenda-shaping mechanism.
  • "Constitutive" leap from correlational cloud-market data. Cloud-market concentration and a 35% year-on-year growth "tracking the generative AI boom" support concentration and co-movement, not the far stronger claim that environmental costs are "constitutive" of the AI agenda.
  • Military dual-use argument at odds with its own examples. The claim that general-purpose systems exacerbate military harm because they are "inherently dual-use" is asserted against examples (Project Maven, Lavender-style targeting, Palantir/Anduril, defence cloud) that are purpose-built military tools, not repurposed general-purpose systems.
  • Regulation critique underpowered by the sections it cites. "Regulation has done little to rein in iAI" is anchored to sections that document harms, not regulatory effectiveness, and pro-growth press releases are treated as evidence of under-regulation, collapsing two distinct steps.
  • Sourcing quirks in the impact discussion. The Microsoft "water stress" and Google "high water scarcity" figures are juxtaposed despite non-comparable bands and denominators; the "half of all internet articles in English" figure rests on a single non-peer-reviewed source; a direct quotation about distance and unseen harm carries no adjacent citation.
  • Under-argued cooperative viability claim. The assertion that agricultural and electrification cooperatives "demonstrate long-term economic viability" applicable to AI is uncited and does not engage the capital-intensive compute profile the paper itself stresses would make an AI cooperative structurally different.
  • Unreconciled aggregate-energy tension. The Section 6 counter-figure (data-centre share 1%→1.5%, 2005–2024) is answered with measurement-uncertainty and localised-effects points that do not squarely reconcile it with Section 3's "soaring", high-growth framing.
  • Local wording and framing issues. "Liberal thinkers" as the sole framing for regulation-as-remedy; an ambiguous antecedent in the U.S. data-centre share sentence; the imprecise "Indian revolution" analogy; the internally implausible Stargate $500 billion / $500 million pairing; and a military-section sentence that cues a stronger causal link between cloud services and civilian deaths than the paper defends.

Reproducibility & code

Re-executing the released repository (github.com/AlexandraVolokhova/irresponsible-ai) reproduces the empirical exhibit well, with the shortfalls confined to external data access and to undisclosed methodological choices rather than to disagreements in the computed numbers (reproduction score about 0.79; both headline claims match, three supporting claims partial).

  • Headline aggregates reproduce, and are unweighted means. The reported 21.1 / 23.5 / 31.4% are recovered within ~1 percentage point (NeurIPS 20.7%, ICML 23.6%, ICLR 30.6% over available years). Execution confirms these are computed as an unweighted mean of per-year fractions, not a pooled ratio — a choice that disproportionately weights the small, high-fraction early years (especially ICLR) and is not disclosed in the manuscript.
  • Peak years reproduce; ICLR decline only partially. NeurIPS and ICML peak in 2020 and decline monotonically through 2025, and ICLR peaks in 2016 (45.8%), all matching the text. The ICLR "decreases thereafter" portion is verifiable only through 2017 because later years are unavailable.
  • ICLR 2018–2025 not reproducible; counts not shipped. No cached CSVs/PDFs are released, and OpenReview now gates anonymous access to recent ICLR venues (challenge verification for 2024/2025, rate-limiting for 2019–2023). NeurIPS and ICML reproduced end-to-end; ICLR reproduced only for 2013–2017. Neither the paper nor the README documents that the credentialed ICLR download path no longer works for recent years.
  • Coverage statistic reproduces; cause attribution is narrow. The unprocessable-paper fraction recomputes to 0.97–3.22% per year, closely overlapping the stated 1.3–3.2% band. The failure counter, however, also fires on PDF parse errors, silently caught index errors, and heuristic filters, so the manuscript's "unconventional formatting or lack of affiliations" gloss is narrower than what the counter aggregates.
  • Affiliation matching ran unchanged but is unaudited. The keyword matcher executed as shipped, but common-word/place-name keywords ("Apple", "Amazon", "Intel", "Uber", "Meta") carry a false-positive risk, and no token boundaries, disambiguation, or validation sample are reported.
  • NeurIPS track scope undocumented. The pipeline captures but does not filter the track field, so Datasets & Benchmarks papers enter numerator and denominator from 2021; this can shift both the average and the peak year. (The reproduction did confirm the pipeline's NeurIPS 2019 count of 1428, matching the paper's in-text cross-check.)
  • Runnable but fragile. Reproduction required three patches, none touching analysis logic (a missing-file guard in the PDF parser, NaN-tolerant plotting, and a metadata-guarded ICLR labelling loop). The pipeline assumes a complete local corpus and is brittle to any download gap.

Recommended Changes

Essential

  • Align the conceptual and operational definitions of "big tech". Acknowledge in Section 2 or Appendix A that the keyword list captures a broader industrial set than the narrative definition, and report a narrow-big-tech aggregate (Google, Meta, Microsoft, Amazon, OpenAI) alongside the broad-industry aggregate for the headline percentages and peak-year claims (addresses the conceptual/operational-gap weakness).
  • Distinguish co-authorship presence from influence. Qualify the influence claims so that the authorship share is presented as a proxy with stated limits, and argue the additional channels (sponsorship, board seats, talent flow) that would bridge presence to agenda-setting (addresses the construct-gap weakness).
  • Reconcile the declining post-2020 fraction with the "growing influence" thesis. Either provide direct evidence for the venue-shift explanation or acknowledge that the conference-authorship trend has reversed and rest the influence argument on other channels (addresses the declining-fraction weakness).
  • Disclose the aggregation convention and add a robustness check. State that the reported averages are unweighted means of per-year fractions, report the publication-weighted (pooled) version, and state whether the NeurIPS Datasets & Benchmarks track is included, showing robustness to excluding it (addresses the headline-aggregates and NeurIPS-track reproducibility points).
  • Ship the resolved counts CSVs and document the ICLR access gate. Publish the per-year numerator/denominator files so Figure 2 can be audited without a full re-scrape, and add a README note that recent ICLR years now require interactive OpenReview verification (addresses the ICLR-not-reproducible and counts-not-shipped points).
  • Document and audit the affiliation-matching procedure. Specify token boundaries and case handling, add any disambiguation for common-word keywords, and report a small hand-audited validation sample with a false-positive rate (addresses the affiliation-matching point).
  • Soften "constitutive" and add the missing step. Weaken to "tightly coupled to", or add a paragraph arguing why environmental costs are internal to the scaling paradigm via the physical link between training-compute demand and data-centre capex (addresses the "constitutive" weakness).
  • Reground or reframe the general-purpose military argument. Add a general-purpose-system-repurposed-to-military example, or reframe the argument as general-purpose systems lowering the marginal cost of entering military applications rather than being intrinsically more harmful once deployed (addresses the military dual-use weakness).
  • Fix the regulation-critique chain. Rewrite the "regulation has done little" and "press releases provide evidence" sentences so the inferential steps are visible, ideally citing regulatory-effectiveness studies or building from the EU Directive example already in Section 6 (addresses the regulation-underpowered weakness).

Suggested

  • Broaden the drop-rate wording. Either partition the 1.3–3.2% by cause or note that it aggregates parse errors and heuristic filters beyond "unconventional formatting or lack of affiliations" (addresses the coverage-statistic reproducibility point).
  • Harden the released pipeline. Add file-existence guards, NaN-tolerant plotting, and a documented list of required inputs so the release survives download gaps (addresses the runnable-but-fragile point).
  • Flag the water-stress vs. water-scarcity mismatch. Add a one-sentence caveat that the Microsoft 42% and Google 15% figures use different WRI Aqueduct bands and denominators and are not directly comparable (addresses the sourcing-quirks weakness).
  • Qualify the "half of all internet articles" claim. Add a note on measurement methodology or cite a peer-reviewed source with explicit sampling and classification (addresses the sourcing-quirks weakness).
  • Support the cooperative-viability claim and engage the compute disanalogy. Add citations for the "abundant precedent" of scale-operating co-ops and briefly address why AI's compute-capital profile does not defeat the analogy (addresses the cooperative-viability weakness).
  • Reconcile the aggregate-energy rebuttal quantitatively. Point to the most recent year-on-year data-centre demand trajectory, or concede that aggregate share understates near-term dynamics while foregrounding local externalities (addresses the aggregate-energy weakness).
  • Loosen the causal cue in the Israel/Gaza sentence. Split "provide cloud services … which has killed tens of thousands" into two sentences so the empirical attributions remain without overreaching the paper's hedged claim about AI's role (addresses the local-wording weakness).
  • Small fixes. Add a citation to the "makes it possible to not see harm" quotation; clarify the Stargate $500 billion / $500 million pairing; disentangle the Birhane/Jurowetzki attribution of the 167/1428 NeurIPS 2019 figure; note that the 13%→47% highly-cited figures are a different population from Figure 2; replace "liberal thinkers" with a broader framing; disambiguate the U.S. energy-production/consumption sentence; replace "the Indian revolution" with "the Indian independence movement"; and add a second citation or description for the Sriraman et al. (2017) cooperative-data claim (address the local-wording and sourcing weaknesses).