SAI
← All ICML 2026 orals

Position: Don't Just "Fix it in Post'': A Science of AI Must Study Learning Dynamics

Stella Biderman, Mohammad Aflah Khan, Niloofar Mireshghallah, Catherine Arnett, Fazl Barez, Naomi Saphra

OralPosition TrackReplication not startedPaper PDFOpenReview

Position: Don't Just "Fix it in Post'': A Science of AI Must Study Learning Dynamics

SAI paper + code review · Referee report

Summary

This position paper argues that AI research overemphasizes post-hoc analysis and post-training interventions — the "fix it in post" mentality — at the expense of studying the training dynamics that produce model behavior. The authors reframe models as snapshots of time-evolving processes and argue that a mature science of AI should be judged by progressively more demanding capabilities: predicting training outcomes, intervening on live trajectories, and designing procedures that reliably produce desired properties. They ground this vision in Kuhn's six characteristics of good scientific practice, and survey four case studies — mechanistic interpretability, fairness, memorization, and simplicity bias — identifying open problems in each.

The conceptual move is timely and intuitive: scaling laws have already shown that some training outcomes are predictable, so extending that paradigm to capabilities, biases, and safety-relevant properties is a plausible research program. The paper is at its best when it points to concrete instances where a training-dynamics lens already outperforms static analysis — induction heads, memorization forecasting from intermediate checkpoints, distributional simplicity bias — and when it presses on how seed- and data-sensitivity should discipline what counts as a stable "mechanism" at all.

Three limitations weaken the argument. First, the scaffolding is thin: the six Kuhn desiderata are named but only Empirical Accuracy and Scope really do work; the predict \to intervene \to design hierarchy is asserted as strictly nested when design-without-intervention seems plausible; and the necessity argument for those three capabilities equivocates between "some prediction" and these specific operational abilities. Second, several claims are stronger than the cited evidence supports — an absolute negative that no training method exists for controlling image-generator gender rates, cross-scale-and-architecture invariance of an n-gram phase order, and a conclusion-level "guarantee" that contradicts the paper's own hedged definitions elsewhere. Third, the case-study coverage is uneven: reinforcement-learning training dynamics, feature-learning theory and the neural-tangent-kernel literature, and grokking / generalization phase transitions are largely absent, even though they represent some of the most theoretically developed instances of the very program the paper advocates.

Strengths

  • Timely reframing. Casting models as snapshots of time-evolving processes rather than fixed artifacts is a productive lens, and the paper does a good job connecting it to concrete lines of work (induction heads, memorization dynamics, distributional simplicity bias) so it does not remain purely programmatic.
  • Operational milestones. The "predict / intervene / design" progression, whatever its logical seams, gives readers a way to talk about what a training-dynamics science would deliver at each stage — a useful vocabulary for what has otherwise been vague talk about "understanding" AI.
  • Seed and object-of-study discussion. Section 2.3 makes an unusually forceful case that seed-, data-, and language-level variation are not implementation details but choices that shape what counts as the phenomenon. The examples (SAE seed instability, seed-varying attention head layers, near-emergence loss basins) are well chosen.
  • Memorization case study. The memorization section is the strongest of the four case studies: it walks through hypothesis-driven work, an initial null result, methodological refinement, cross-model replication, and remaining open problems, which is exactly the kind of narrative the paper argues the field should aspire to.
  • Alternative-views section exists at all. Many position papers skip serious engagement with dissenting views; this one at least enumerates four (engineering sufficiency, anti-theory skepticism, automation-as-science, safety pragmatism), even if the treatment is short.

Weaknesses

  • Kuhn's desiderata are named and abandoned. Section 2 introduces six criteria (Empirical Accuracy, Internal Consistency, External Consistency, Scope, Simplicity, Fruitfulness) as the anchor for what a "science of AI" would look like, but only Empirical Accuracy and Scope get any real use in the sections that follow. Simplicity, Internal/External Consistency, and Fruitfulness never resurface — including in the case-study "progress report" at the end of §2.4 — which makes the framework feel decorative rather than load-bearing.
  • Predict \to intervene \to design is not obviously monotone. The paper asserts that "each is harder than the last because each requires everything the one before it did, and then more," but designing training procedures that reliably yield a property does not obviously require the ability to intervene mid-run: architecture search, data-mixture design, and objective design all target "design" without an in-training intervention step. The nesting claim needs an argument, not just an assertion.
  • Necessity argument for the three capabilities equivocates. The justification given — a theory predicting nothing and guiding no intervention is "hardly a theory" — supports the mild claim that some predictive content is necessary. It does not license the stronger conclusion that these three specific operational capabilities are all necessary signs of scientific understanding.
  • "Fix it in post" framing risks strawmanning. The dichotomy between post-hoc patching and dynamics-aware science is rhetorically powerful, but many post-training methods (e.g., preference optimization, activation steering, unlearning) are themselves theory-motivated, and the paper does not distinguish "post-hoc without theory" from "post-hoc informed by theory." The result is a false dichotomy that makes the target easier to hit.
  • Missing bodies of relevant theory. For a paper about the science of training dynamics, the near-total absence of feature-learning theory, neural-tangent-kernel and mean-field analyses, statistical-learning-theory perspectives on generalization, RL training dynamics, and grokking / phase-transition work is notable. These are exactly the settings where formal "training dynamics" theories already exist; excluding them makes the surveyed progress look thinner than it is and skews the sample toward the authors' own research neighborhood.
  • Analogy oversells physics. The tennis-ball opening implies that having "position, velocity, and acceleration" gives predictions "until the end of the universe." Physics achieves that for a small class of near-integrable systems; for turbulent, chaotic, or many-body systems it does not — which is arguably the more honest analogy for neural-network training. The chosen analogy sets an expectation that the paper's own material on seed variability and loss-basin diversity later undermines.
  • Data-attribution passage: recency claim mis-supports the distinction. The sentence "Data encountered later in training has a larger influence on model behavior, so the answer depends on what is held fixed" uses recency as the reason two attribution framings diverge, but the actual divergence is between marginalizing over data orderings and conditioning on a specific realized run. Recency is downstream of the choice, not the cause of it.
  • SAE-as-proxy inference overshoots. The claim that SAE feature instability across initializations "reveal[s] limitations for using them as a proxy for a full model" mixes two claims: SAEs are non-unique decompositions (supported), and SAEs misrepresent the underlying model (a stronger claim not established by seed-sensitivity alone).
  • Apparent contradiction in the Duan et al. citation. The memorization section cites Duan et al. (2024) for "most 'absorbent' around 10–20% into training" and then, a few lines later, for "later checkpoints memorize more readily even for out-of-distribution sequences." The two can be reconciled (absorbency window vs. propensity to memorize OOD content) but as written the source appears to be cited for opposite claims.
  • Membership-inference reading overstates. Duan et al.'s low-MIA-success result is presented as "most sequences are not memorized in ways that leave detectable traces, consistent with findings that substantial repetition is required." MIA failure is at best weak evidence about memorization, especially given the authors' own attribution of the result partly to "fuzzy member/non-member boundaries" — an attack-side or definitional issue rather than a memorization-side one.
  • Absolute negative on bias control. The claim in the "Designing Social Bias" box that "there is currently no training method that results in the desired behavior" asserts a universal negative that is stronger than the survey-style presentation warrants; "no established/reliable method" would be more defensible.
  • N-gram phase-order invariance is asserted, not cited. "These patterns hold across model scales and architectures" follows two specific citations that do not obviously establish cross-scale-and-architecture invariance; the claim needs either a supporting reference (Michaelov et al., 2025 arguably is the right one — then say so and specify which claim it establishes) or hedging.
  • Section reference for cross-model generalization is off. The line "These interchange interventions transfer across models, exemplifying the generalization recommendations of Section 2.3" points to §2.3 (Properly Identifying Objects of Study), but cross-model transfer of mechanisms is treated most explicitly in §3.1 and, methodologically, in §2.2. A careful reader tracking cross-references will bounce.
  • "Guarantee" contradicts the paper's own hedges. The conclusion promises "design training procedures that guarantee desired properties," while the abstract says "more reliably produce" and §2.4 says "reliably produces." §2.4 itself cautions that a working recipe might have "no account of why it works." "Guarantee" is inconsistent with those hedges.
  • Scaling-law framing conflates inputs. "Given model size and a compute budget, final loss can be predicted" is slightly loose: Kaplan et al. and Chinchilla predict loss from (size, tokens) or equivalently compute; because C6NDC \approx 6ND, size and compute are not two independent knobs but coupled ones. Minor, but avoidable.
  • No falsifiable measure of progress. The paper's own criterion — that we track predict/intervene/design capabilities — is never operationalized. There is no proposed metric, benchmark, or heuristic for judging whether the field has moved along it. This is ironic given the paper's central argument about the value of predictions and measurable success criteria.
  • Case-study selection is not justified. The four surveyed areas are useful but the reader is not told why these four, or why (for example) safety training or optimization theory are excluded. A brief note on selection criteria would help calibrate the "measured this way, progress is real but uneven" claim in §2.4.

Recommended Changes

Essential

  • Argue the predict \to intervene \to design nesting, don't just assert it. In §2.4, either give an argument that design requires the ability to intervene (and prediction), or drop the strict-nesting claim and describe the three as related-but-distinct capabilities. This is the scaffolding the whole paper hangs on.
  • Weaken the necessity argument for the three capabilities. In the paragraph containing "Abilities like these are necessary signs of understanding, since a theory that predicts nothing and guides no intervention is hardly a theory," acknowledge that this justifies necessity of some predictive content, not necessity of the three specific operational milestones.
  • Fix the "guarantee" in the conclusion. Replace "design training procedures that guarantee desired properties" with the paper's own consistent formulation ("reliably produce") to avoid contradicting the abstract and §2.4.
  • Reconcile or split the Duan et al. citation in the memorization section. Clarify that "most 'absorbent' around 10–20% into training" and "later checkpoints memorize more readily even for out-of-distribution sequences" concern distinct measurements, and either add a one-clause distinction or split the citation.
  • Soften the MIA reading. Change "This suggests that, at scale, most sequences are not memorized in ways that leave detectable traces, consistent with findings that substantial repetition is required" to make clear that low MIA success is one-way evidence: it rules in the "high repetition needed" story but does not, on its own, rule out extensive memorization the attack cannot detect.
  • Soften the bias absolute negative. In the "Designing Social Bias" box, change "there is currently no training method that results in the desired behavior" to "no established / reliably validated method," matching the survey framing.
  • Support or hedge the n-gram invariance claim. Attach a specific citation to "These patterns hold across model scales and architectures" (Michaelov et al., 2025 seems the intended anchor — specify which patterns and which architectures/scales), or rephrase as "have been observed to hold …".

Suggested

  • Use more of the Kuhn framework, or drop the rest. Either reintroduce Internal/External Consistency, Simplicity, and Fruitfulness when reviewing case-study progress in §2.4 and §3, or reduce the framework in §2 to the two desiderata (Empirical Accuracy, Scope) that actually do work.
  • Add missing case-study areas or acknowledge their exclusion. Even a brief paragraph placing neural-tangent-kernel / feature-learning theory, RL training dynamics, and grokking / generalization phase transitions on the "progress report" grid would make the coverage argument more defensible; alternatively, name the exclusions and explain them.
  • Rework the tennis-ball analogy or complicate it. Add a sentence acknowledging that physics does not deliver end-of-universe predictions for chaotic / many-body / turbulent systems, and that neural-network training may sit closer to that regime than to a projectile — this is more honest and does not weaken the paper's motivation.
  • Distinguish theory-informed from theory-free post-hoc work. In §1 and §4, separate "post-hoc without theory" from "post-hoc informed by theory" so that the paper's target is the former, not all post-training methods.
  • Fix the data-attribution "so." Rewrite the sentence "Data encountered later in training has a larger influence… so the answer depends on what is held fixed" to make clear that the divergence between the two attribution framings comes from what is held fixed, not from recency effects (which are relevant only under the fixed-run framing).
  • Qualify the SAE-as-proxy inference. Change "revealing limitations for using them as a proxy for a full model" to something like "revealing that SAE decompositions are underdetermined and thus that any single SAE gives at best a partial view of model computation."
  • Fix the §2.3 cross-reference. Point "exemplifying the generalization recommendations of Section 2.3" to §3.1 (mechanistic interpretability) or, if §2.3 is really the intended anchor, be explicit that "generalization" here means "invariance across the seed / model / data dimensions of variation §2.3 identifies."
  • Tighten the scaling-law framing. In §2.1, rephrase "given model size and a compute budget" to make explicit that size and compute (or size and training tokens) are the two coupled knobs, not independent ones.
  • Propose an operationalization of progress. Even a paragraph sketching what a "measurement of progress along predict/intervene/design" might look like (e.g., predicting a property PP on a held-out training configuration within a stated tolerance) would give the paper's central proposal more traction.
  • Justify the case-study selection. Add one or two sentences at the top of §3 explaining why mechanistic interpretability, fairness, memorization, and simplicity bias were chosen as the survey slice.