SAI
← All ICML 2026 orals

Unsupervised Partner Design Enables Robust Ad-hoc Teamwork

Constantin Ruhdorfer, Matteo Bortoletto, Victor Oei, Anna Penzkofer, Andreas Bulling

OralReplication score 0%Paper PDFCode repoOpenReview

Unsupervised Partner Design Enables Robust Ad-hoc Teamwork

SAI replication review · Referee report

Summary

The paper introduces Unsupervised Partner Design (UPD), a population-free ad-hoc teamwork (AHT) method that repurposes the machinery of unsupervised environment design — cheap online candidate generation plus learnability-based selection — to build a curriculum over partner policies instead of environment parameters. The conceptual move is neat: rather than pretraining a diverse population (as in FCP or MEP) or hand-tuning a single mixing coefficient (as in E3T), UPD samples ϵU(0,1)\epsilon\sim U(0,1) mixtures of the ego and a Dirichlet-biased random policy, scores a large candidate batch by return variance, and trains on the top-B|B| partners in a single self-play-style stage. The formulation extends cleanly to a joint partner-level curriculum (JUPD) when a procedural generator is available. Empirically, UPD is competitive with or ahead of FCP, MEP, and E3T on LBF and Overcooked-AI, matches or beats ROTATE when its own curriculum hyperparameters are tuned per layout, and is preferred by 12 human collaborators over SP/MEP/E3T on three Overcooked layouts. The analysis of why it works — buffer-averaged ϵ\epsilon varies by layout, and average bias-mask directions reverse mid-training — is one of the most interesting parts. The main conceptual limitation is that the learnability score is justified by appeal to a Foster et al. bound whose object is advantage-signal variance, not the raw trajectory-return variance UPD estimates; the paper does not distinguish them. The main empirical limitation is that the aggregate Overcooked wins are driven almost entirely by Asymmetric Advantages: on most layouts UPD is within one standard deviation of its own ablations, and against ROTATE UPD loses on three of five layouts. An independent replication of the shipped code was attempted but stalled after arming a background batch that produced no output, so the headline values remain unverified in this run — no algorithmic discrepancy was uncovered, but the numbers behind Tables 1–3 and OGC could not be regenerated. Overall the work is a genuine, well-executed conceptual step for AHT.

Strengths

  • Conceptual contribution. Extending UED-style learnability curricula from environments to partners is a clean, well-motivated idea; framing Gπp,θG_{\pi_p,\theta} as an induced single-agent game is the right abstraction and makes the extension to a joint partner-level curriculum (JUPD) very natural.
  • Simple, self-contained method. Sampling ϵU(0,1)\epsilon\sim U(0,1) and Dirichlet-biased random policies is much lighter than the population-training pipelines used by FCP/MEP/ROTATE, and works with a single set of hyperparameters across LBF and all five Overcooked layouts.
  • Ablations are informative. The UPD w/o bias and UPD w/o ℓ rows in Table 1 do real work — they show that most of the aggregate gain comes from randomised-ϵ\epsilon + bias masking, with learnability adding a smaller (but real) increment, especially on AA.
  • Analysis is candid. Figures 4 and 6 (per-layout ϵ\epsilon selection, emergent bias-mask reversal) go beyond the standard AHT-results-table pattern and give a genuine mechanism story; the Appendix J.1 comparison across four learnability functions demonstrates robustness rather than cherry-picking var\ell_\mathrm{var}.
  • Honest empirical framing. The paper notes that UPD does not dominate every layout (CRoom, for instance), reports a transparency note about an earlier corrected version, and disclaims the human study given n=12n=12.
  • Human-in-the-loop evaluation. A double-blind, IRB-approved user study with pre-registered survey items, Holm–Bonferroni-corrected non-parametric tests, and Cronbach's α=0.916\alpha=0.916 across the seven items lends credibility to the composite score.

Weaknesses

  • Load-bearing theoretical claim is loose. The Foster et al. result concerns the variance of the scalar advantage-forming learning signal, not the raw trajectory-return variance Varτ[R(τ)]\mathrm{Var}_\tau[R(\tau)] that UPD scores. In the induced game Gπp,θG_{\pi_p,\theta} these differ because the observed return mixes ego, partner, and environment stochasticity, and the partner generator here is heavily stochastic. The paper's 'Applying the result of Foster et al. … implies that … increases with Var' overstates the connection; var\ell_\mathrm{var} is a plausible heuristic, but the derivation should acknowledge the gap. Appendix J.1's finding that mean\ell_\mathrm{mean}, gauss\ell_\mathrm{gauss}, and var\ell_\mathrm{var} all work similarly reinforces this reading.
  • Aggregate Overcooked wins are AA-driven. In Table 1, UPD's +18% aggregate margin over E3T is dominated by a very large gap on AA (181.4 vs 127.9); on CRoom UPD is behind, and on CR/CC/FC the gap to its own ablations is inside one standard deviation. Table 3 makes this sharper: ROTATE wins on CRoom, CR, and FC, and UPD's average lead again comes from AA. The paper would be strengthened by discussing why AA is favourable to UPD (asymmetric layouts + convention breaking?) rather than leading with the single aggregate number.
  • FC comparison against E3T reduces to self-play. E3T on FC uses ϵ=0.0\epsilon=0.0, i.e. pure self-play. The E3T-vs-UPD improvement on FC therefore does not really test 'fixed-ϵ\epsilon mixing'; it tests self-play. Either sweep ϵ\epsilon on FC (as was done for LBF) or explicitly flag this reduction in the Table 1 caption.
  • OGC evaluation drops AA and CC. Table 2 evaluates JUPD on only three layouts (CRoom, CR, FC). AA (where UPD's fixed-environment advantage is largest) and CC (the paper's coordination-difficulty exemplar) are absent, which weakens the joint-curriculum headline. And as rendered, Table 2 is unparseable — the JUPD row that is the subject of the section is missing.
  • Human study claims outrun what n=12n=12 can support. 'Significantly more adaptive, more human-like, a better collaborator, and less frustrating' is a blanket claim; Figure 7's variable asterisks show that not every UPD-vs-baseline / item pair reaches the same significance level after Holm–Bonferroni over three comparisons. The abstract's 'more X than all evaluated baseline methods' should be aligned with the §5.3 disclaimer, and effect sizes should be reported alongside p-values.
  • Attribution among UPD's components is not cleanly separated. Table 1's ablation rows (E3T, UPD w/o bias, UPD w/o ℓ, UPD) bundle (i) randomised ϵ\epsilon and (ii) bias masking together. A fourth ablation with randomised ϵ\epsilon but no bias mask and no \ell-selection would let readers isolate the three axes and would make the §7 suggestion ('extending E3T with randomised mixture coefficients') concrete.
  • Convention-breaking narrative is suggestive but not causal. The bias-mask reversal in Figure 6 is striking, but 'consistent with learnability favouring convention-violating partners' is only one plausible reading — the ego's own preferred direction may itself be shifting. Showing that the ego's action distribution is stable while the selected bias-mask reverses, or ablating a bias-mask freeze and reporting a test-time regression, would move this from narrative to mechanism.
  • CV² correction is introduced without ablation. CV² normalisation is presented as necessary for JUPD (different levels induce different reward ranges) but no JUPD-with-vs-without-CV² comparison is shown on OGC. Given the robustness across learnability functions demonstrated for the fixed-environment setting, an OGC-side ablation would clarify whether CV² is essential or stylistic.
  • Cost model states precise crossover numbers under strong assumptions. The 'n3n\geq 3', 'CPPO/Cenv>600C_\mathrm{PPO}/C_\mathrm{env}>600', and 'n=2.5n^\star=2.5' thresholds depend on fully-parallel candidate scoring across KK simulators and on wall-clock rollout cost scaling with horizon HH rather than EHEH — both hold only in the unsaturated-hardware regime. The '<10% additional runtime' claim in the main text is stated without a wall-clock measurement. Please separate the analytical model from the empirical overhead.
  • Overstated inferences from single-example arguments. The footnote claim that the optimal ϵ\epsilon depends on the task '(as in this example)' rests on an example with only one task; the pedagogical example in Appendix C mixes explicit definitions with an implicit πpmix\pi_p^\mathrm{mix} mixture; and the universal quantifier 'across all possible testing games' inside the same appendix is not justified by the toy matrix game shown.
  • Notational and table issues. Table 5 introduces a hyperparameter ('# sampled partners') that is not in Table 4, making the fine-tuned configuration hard to map back to defaults. Table 6 reuses ϵ\epsilon for the PPO clip range, colliding with the paper's own mixing coefficient. Table 2 as rendered is unparseable. 'Converges stably' as a shared claim across Figures 10–14 rests on cross-method training-return comparisons the paper elsewhere disclaims. CEC in §6 is framed as a partner-generation baseline but is defined as self-play. The transparency note about an earlier version does not specify what changed.

Reproducibility & code

We ran the shipped repository end-to-end. No algorithmic divergence from the paper was uncovered — the codebase does ship scripts for every experiment the paper describes (SP, FCP, MEP, E3T, UPD and its ablations, ROTATE evaluation, JUPD/OGC, LBF, and evaluation harnesses). The replication attempt itself failed to complete, so none of the numeric headline values could be independently verified in this run.

  • Replication attempt stalled. The pipeline halted after arming a background batch that never produced output; 0 of 8 planned evaluation steps executed. As a consequence the +18% Overcooked-AI aggregate over E3T, the ROTATE comparison, the OGC JUPD result, the LBF-ϵ\epsilon sweep, and the '282 trained policies' bookkeeping are all unverified here. This is a repro-pipeline outcome, not a paper defect, but it means the numeric claims stand only on the paper.
  • Evaluation partners are not shipped. The Overcooked and LBF evaluation scripts hard-code eval_populations/FF_BRDiv/{layout_name}/params_*.pt paths; no BRDiv checkpoints or deterministic-generation seeds are included. Tables 1–3, and the human-study ego checkpoints they feed, therefore require regenerating BRDiv from scratch to match the paper's numbers.
  • LBF training defaults diverge from the paper. The LBF DPD training script's TrainConfig defaults (5×1075\times 10^7 timesteps, sfl_buffer_size=128, num_steps_per_env=128) do not match the paper's stated 10710^7 / B=64|B|=64 / 400-steps LBF configuration. A reproducer running the shipped script silently trains a different curriculum from the one that produced Figure 3.
  • Ablation configurations are unlabelled. UPD w/o bias corresponds to a separate script (dpd_ippo_overcooked_rnn.py vs _w_bias_rnn.py), and UPD w/o ℓ is not exposed as a config flag in the fixed-environment DPD script — only the OGC-side JUPD script has a learnability_function="none" branch. Mapping Table 1's ablation rows to script + flag requires code archaeology.
  • No orchestration / aggregation harness. The 6-seed × 5-layout × 7-method matrix behind Table 1, the OGC matrix behind Table 2, and the plots behind Figures 4/5/6/19/20 all have to be assembled by hand. Training curves live in wandb, which defaults to disabled.
  • Human-study artifacts are absent. No NiceWebRL configuration / web interface, no anonymised per-participant Likert responses, no per-game return log, and no Wilcoxon / Cronbach's-α\alpha analysis script. The significance markers in Figure 7 and the composite α=0.916\alpha=0.916 cannot be checked without re-running the 12-person study.
  • ROTATE comparison depends on external, non-shipped training with hard-coded paths. configs/heldout_ego.yaml points to /scratch locations that will not exist for external users; ROTATE training itself lives in the upstream Wang et al. repo.
  • SP-IPPO transparency note lacks the evaluation population. The 32.4 vs 41.4 improvement is reported without stating the partner set, the number of seeds, or which layouts it averages, all of which matter for taking the note at face value.

Recommended Changes

Essential

  • Tighten the theoretical justification for var\ell_\mathrm{var}. Rewrite the passage around 'Applying the result of Foster et al. …' to distinguish advantage-signal variance from raw trajectory-return variance in the induced game, and present var\ell_\mathrm{var} as a heuristic proxy motivated by (rather than derived from) the Foster et al. bound.
  • Decompose the AA-driven aggregate wins. Add a per-layout discussion showing that the +18% Overcooked-AI aggregate over E3T and the average lead over ROTATE are AA-dominated, and hypothesise why. This addresses the Aggregate Overcooked wins are AA-driven weakness.
  • Fix and complete the OGC table. Provide a corrected Table 2 with mean ±\pm std per method ×\times layout — including the JUPD row that is currently missing — and either add AA and CC to the OGC evaluation or explicitly justify why they are excluded.
  • Align the human-study language with n=12n=12. Rewrite the abstract sentence 'rated as more adaptive, more human-like, and less frustrating than all evaluated baseline methods' to match the §5.3 disclaimer, and in §5.3 spell out exactly which UPD-vs-baseline / item pairs reach p<0.05p<0.05 after Holm–Bonferroni, with effect sizes.
  • Ship the evaluation partners and a reproduction harness. Release the exact BRDiv checkpoints (or deterministic seeds + configs) used for Tables 1–3, remove the hard-coded absolute paths in heldout_ego.yaml and the AHT evaluation scripts, and add scripts/reproduce_tableN.sh + per-figure plotting notebooks. Align the LBF training-script defaults with the paper (or ship a configs/lbf_paper.yaml).
  • Release the human-study analysis pipeline. Publish anonymised per-participant Likert responses, per-game returns, and the Wilcoxon + Cronbach's-α\alpha analysis script so readers can verify the significance markers and the composite score without re-running the study.
  • Document the ablation-to-script mapping. Add a README section mapping each Table 1 row (UPD w/o bias, UPD w/o ℓ, full UPD) to the exact script + config it corresponds to, since the current split across scripts (with learnability_function="none" only exposed on the OGC side) makes the ablations hard to reproduce.
  • Document the ROTATE dependency. Either bundle the exact ROTATE checkpoints used for Table 3 or specify the upstream Wang et al. commit + hyperparameters + seeds, and remove the /scratch hard-codes from configs/heldout_ego.yaml.

Suggested

  • Add a fixed-ϵ\epsilon Overcooked sweep. The claim 'no single ϵ\epsilon would be optimal' currently rests on buffer-averaged ϵ\epsilon trajectories and the LBF grid; a fixed-ϵ\epsilon Overcooked-AI sweep would directly support it.
  • Sweep ϵ\epsilon on FC (or flag the reduction). Either sweep ϵ\epsilon for E3T on FC, or explicitly note in the Table 1/3 captions that E3T-FC is pure self-play because ϵ=0.0\epsilon=0.0.
  • Add a fourth ablation isolating randomised ϵ\epsilon. An E3T + U(0,1) ε only row (no bias, no \ell-selection) would let readers isolate the three UPD components and would make the §7 recipe suggestion concrete.
  • Test the convention-breaking hypothesis causally. Report whether the ego's own action distribution is stable while the selected bias-mask reverses in Figure 6, or ablate a bias-mask freeze and show a test-time regression.
  • Ablate the CV² correction in JUPD. A JUPD run with raw var\ell_\mathrm{var} on OGC would clarify whether CV² is essential or stylistic.
  • Report a wall-clock measurement. Back the '<10% additional runtime' claim in §5.1 with a concrete UPD-vs-E3T-vs-SP table on identical hardware, and surface the two throughput assumptions (KK-parallel scoring, EHHEH \to H wall-clock scaling) that underpin the Appendix F crossover thresholds.
  • Add a per-layout MEP configuration table. Convert the prose about MEP tuning ('10 PPO epochs, 64 minibatches, α=0.001\alpha=0.001 on FC, …') into a small table, and clearly distinguish PPO entropy from population-entropy α\alpha.
  • Fix minor notation and specification issues. Rename the PPO clip in Table 6 (e.g., ϵclip\epsilon_\mathrm{clip}) to avoid colliding with the mixing coefficient; align Table 5's '# sampled partners' hyperparameter with Table 4's $\rho$ / $|B|$ / '# generated agents' vocabulary; state πpmix\pi_p^\mathrm{mix} explicitly in Appendix C and scope 'across all possible testing games' to the toy game; state which of {4,096,16,384}\{4{,}096, 16{,}384\} is SFLE3T's vs JUPD's OGC buffer; and add a manifest for the '282 trained policies' bookkeeping claim.
  • Expand the transparency note. State which methods/layouts changed in the earlier-version correction described in Appendix B, and by roughly how much, so readers can audit the scope of the earlier issues.
  • Report per-partner rollout counts. In the Figure 5 and Appendix I learnability-vs-return analyses, state the number of rollouts per partner used to estimate var\ell_\mathrm{var}, and confirm that 'most partners have 0\ell\approx 0' is not a small-sample estimator artifact.
  • Sharpen the SP-IPPO transparency note. State the evaluation population, number of seeds, and layout(s) used for the 32.4 vs 41.4 comparison, and ship both configurations so the note can be checked against the released code.
  • Reframe CEC in §6 as a level-only baseline rather than a 'partner + level generation' one, since it plays self-play.
  • Weaken the 'converges stably' framing in §5.2.2. The training returns in Figures 10–14 are not cross-method comparable; a per-method convergence criterion would be more meaningful than the shared framing.