SAI
← All ICML 2026 orals

Incentivizing Truthfulness and Collaborative Fairness in Bayesian Learning

Rachael Hwee Ling Sim, Jue Fan, Xiao Tian, Xinyi Xu, Patrick Jaillet, Bryan Kian Hsiang Low

OralReplication not startedPaper PDFOpenReview

Incentivizing Truthfulness and Collaborative Fairness in Bayesian Learning

SAI paper + code review · Referee report

Summary

The paper studies the joint problem of eliciting truthful data from participating sources in collaborative machine learning while distributing rewards that satisfy collaborative fairness (F0–F3, Sec. 2). Its conceptual move is to insist that the data valuation function (DVF) depend on information OO that remains hidden from the sources — instantiated as the mediator's held-out validation set — and to identify the log-likelihood logp(T=TDC=DC)\log p(T=T^\ast \mid D_C=D_C) as a DVF that (i) is strictly proper as a scoring rule and (ii) plays nicely with semivalues. Proposition 3.2 formalizes the truthfulness incentive at the level of individual and coalition values; Proposition 4.1 lifts this to semivalues ϕi[νDv]\phi_i[\nu_{\mathbf D}^v] using an additional assumption (main-text A3 / appendix A.4) that each source believes other sources are truthful. Sec. 5 discusses two practically important relaxations — limited budgets and no held-out validation set — the latter via a per-source split trick and a modified symmetry condition. Empirically the paper covers four Bayesian models (GP, Bayesian logistic regression, BNN, multinomial logistic on BloodMNIST), six untruthful strategies for a target source, and Beta-Shapley/individual-value ablations.

The contribution is a clean, welcome unification of two prior threads (Zheng et al.'s pointwise-mutual-information truthfulness and the Shapley-value fairness line) whose combined coverage was missing. The technical novelty sits mostly in (a) identifying A3/A.4 as the additional coupling assumption required to move from single-source truthfulness to coalition-value truthfulness, and (b) mapping out where strict truthfulness necessarily breaks (limited budget, no validation set). The main conceptual limitation is that the theory is tightly bound to (i) Bayesian models, (ii) a shared, correctly specified prior/likelihood, and (iii) sources believing everyone else is truthful — a stack of premises that individually are standard but jointly narrow the applicability envelope substantially. The empirical envelope confirms this: truthfulness holds only under mild misspecification, and LO-BL already exhibits duplication beating truthful submission for a target source. The story is honest about these limits but sometimes advertises the guarantee more broadly than the supporting text sustains.

Strengths

  • Clean conceptual synthesis. The paper is upfront about the two missing pieces (fairness under peer-prediction, truthfulness under Shapley) and shows exactly where each side fails when the other is imposed. Identifying that the coalition-value truthfulness step requires A3/A.4 is a genuine technical contribution, not just a bookkeeping observation.
  • Fair-and-truthful DVF. Casting the DVF as the log posterior predictive density of a hidden validation set gives a natural, strictly-proper-scoring-rule justification, closed-form expressions in the exponential family, and a clean interpretation of E[v(DC)]\mathbb{E}[v(D_C)] as mutual information I(DC,T)I(D_C, T), which in turn delivers monotonicity in coalition size 'for free'.
  • Careful treatment of relaxations. Sec. 5 handles the two most practically salient constraints (limited budget, no validation set) with correctly weakened guarantees (ii-truthful rather than ii-s·truthful under clipping; modified F1 under split-based valuation) and explains why the natural attempts fail.
  • Broad empirical envelope. Four Bayesian model families, six strategies, size/noise/input-shift ablations, Beta-Shapley variants at three temperatures, and a 20-source approximate-Shapley table together cover a lot of ground for a truthfulness paper, and the paper is explicit about known failure modes (LO-BL duplication, high-noise misspecification).
  • Honest limitations section. App. G collects the load-bearing assumptions (validation set, Bayesian scope, trusted mediator, semivalue scalability, minor-misspecification empirical regime) and does not try to argue them away — a useful feature relative to related work.

Weaknesses

  • Circular justification of A3/A.4. In Sec. 4 the paper argues that because the truthful profile maximizes total expected reward and is the natural Schelling point, "all sources will prefer the easier truthful Nash equilibrium, justifying assumption A3." But A3 is the premise under which the truthful profile is derived to be optimal; using the derived conclusion to justify the premise is rhetorically circular, and the "no better equilibrium exists" clause conflates aggregate welfare with individual optimality (an individual source can prefer a smaller share of a smaller pie). This is the single load-bearing assumption behind the main truthfulness-plus-fairness result and it deserves an independent behavioral or mechanism-design justification.
  • Assumption numbering shifts between text and appendix. Prop. 3.2/4.1 cite A1–A3, but the appendix uses A.1 (No cloning) plus A.2/A.3/A.4 — so main-text "A3" (belief that others are truthful) corresponds to appendix "A.4", and appendix "A.3" is instead the validation-set assumption. Combined with the appendix's use of p(Oθ)p(O|\theta) where the main text uses p(Tθ)p(T|\theta), this makes it real work for the reader to check which assumption powers which lemma.
  • Empirical truthfulness is only claimed under minor misspecification. The limitations section itself says truthfulness holds "only … under minor model misspecification", and the LO-BL experiment exhibits duplication beating truthful submission even at low noise (attributed to prior over-uncertainty). Since the promised deployments (hospitals, banks) are precisely where misspecification is likely to be substantial and priors are hard to calibrate, the practical envelope of the guarantee is narrower than the abstract suggests.
  • Only Zheng et al. (2024) is empirically compared. The other reasonable baselines — Chen et al. (2020) peer-prediction, validation-accuracy Shapley (Ghorbani & Zou, 2019), volume-based DVFs — are excluded from Sec. 6 on the ground that they lack matching guarantees, but that is exactly why an empirical head-to-head is informative: a reader cannot judge how much practical accuracy/robustness is being paid for the extra theoretical property.
  • Restriction to Bayesian models with agreed-upon specification. The stack of assumptions — Bayesian model class (A2), agreed priors and likelihoods (A2/A.2), belief in others' truthfulness (A3/A.4) — excludes almost all deployed classifiers and quietly assumes an adversary-free specification stage. Strategic misreporting of the shared prior or per-source likelihood is a plausible attack channel that the current threat model ignores.
  • Sec. 5.2 relaxations are stronger than advertised. The modified F1 in the validation-set-free case only guarantees equal rewards for two sources with identical datasets and identical split randomness, so generically-symmetric sources with different split seeds receive different rewards. The abstract still claims the mechanism "provably ensures (F) collaborative fairness"; the qualifier should be visible earlier.
  • Volume duplication factor is understated (App. C.1). The claim that duplicating each point kk times increases the volume DVF by k\sqrt{k} holds only for f=1f=1; the general scaling is kf/2k^{f/2} where ff is the feature count. Since App. C.1 is the paper's argument that alternative DVFs are provably untruthful, this understates rather than exaggerates the vulnerability — worth fixing for correctness of the motivation.
  • Information-gain-vs-duplication argument does not cover duplication. The 'information never hurts' derivation shows adding distinct data cannot lower information gain, but duplicating an already-observed point yields H(θDi,dup)=H(θDi)H(\theta|D_i',\mathrm{dup})=H(\theta|D_i'), so the gain is unchanged, not increased. This weakens App. C.1's proof-of-untruthfulness for the information-gain DVF specifically for the duplication strategy.
  • Feasible-reward guarantee mixes pointwise and expected regimes. The Sec. 5.1 statement that min(ϕi/a,B)\min(\phi_i/a,B) is "ii-s·truthful only when ϕi/aB\phi_i/a\leq B, and ii-truthful otherwise" reads as a deterministic partition, but Def. 3.1 is over the expectation and clipping generically happens on a subset of realizations. The correct statement is subtler and should be given in expectation.
  • Recommendation 'use small validation sets to disincentivize duplication' is in tension with Fig. 6. Fig. 6 and F.2 argue larger validation sets reduce estimator variance; the App. F.5 recommendation to use small ones has no reconciled trade-off. Readers picking one direction will pay for it in the other.
  • Small-nn empirical scope. Every main-text experiment is n=3n=3; the 9- and 20-source runs are appendix-only, and the 20-source result uses an approximate Shapley (an unbiased estimator, not the exact quantity). Moving at least one larger-nn experiment into the main text would strengthen the empirical claim.
  • Minor exposition problems. Notation clash between input dimension ww and weight vector w\mathbf{w}; unparseable sentence in the "reward value of other sources" paragraph in App. F.5; Fig. 8 caption still says "increasing sizes" while the body varies input distribution; Strategy I labels ("mode, class 1" vs "class 0") disagree with the LO-BL setup text stating source 0 has more of class 2. Individually minor but they slow verification.

Reproducibility & code

Veritas inspected the released artifacts without executing them; the observations below are strictly about what the paper's methodology + released material would let a competent reader reproduce, not about any measured outcomes.

  • No public code / seeds released. The paper says only "please refer to the environment.yml file attached" and reports a 1-hour NUTS runtime for NN-CY, but there is no repository URL, no runnable script (train.py, compute_shapley.py, beta_shapley.py), and no random seeds. Every step of the pipeline — coalition-value computation, NUTS configuration per model, validation-subset sampling, split randomization — must be reconstructed from prose. Given how central the numeric orderings are to the paper's argument, this is the biggest reproducibility gap.
  • GP hyperparameters and prior underspecified. "Squared exponential kernel and different lengthscales for each feature" is the entire GP specification; the prior on lengthscales, the noise-variance prior, and whether hyperparameters are learned by marginal-likelihood maximization or fully Bayesian sampling are not stated. The 2- and 3-source numeric Shapley values in App. B.2 (0.098 vs 0.066; 0.042 vs 0.049) depend directly on these choices.
  • LO-BL preprocessing pipeline underspecified. The paper reports Table 4's exact Shapley to five decimal places but does not identify the pretrained ResNet-18 checkpoint (ImageNet-1k? MedMNIST-pretrained?), the PCA fit set (union of source data? mediator only?), or whether PCA is re-fit per coalition.
  • Ambiguous Shapley approximation. Table 3 uses "the approximation proposed by Kolpaczki et al. (2024)" with 3000 samples over 20 runs, but Kolpaczki et al. present multiple estimators (SVARM and variants) and the paper does not identify which. Since the reported std devs are of the same order as inter-strategy gaps, the estimator choice matters for whether the T-vs-others ordering reproduces.
  • Validation-subset and split procedures unspecified. "20 different subsets of the validation set" (Figs. 1–2) and "10 different training and validation set splits" (Fig. 3, App. F.5) do not say whether these are bootstrap resamples, disjoint folds, or random contiguous fractions, so CI widths are not reconstructible from the description alone.
  • Strategy specifications are dataset-dependent but not tabulated. Strategy N's "incorrect class label" for 8-class LO-BL is under-specified (uniform among other classes? adversarial?); Strategy I's synthetic-label target class ("mode, class 1" for LO-HE; "class 0" for LO-BL) contradicts LO-BL's stated class distribution; Strategy P's noise SD (0.05/0.1/0.2) is stated in absolute terms without confirming feature standardization.
  • Missing display equation for the validation-set-free characteristic function. Sec. 5.2 promises "The characteristic function νDTj\nu_{\mathbf D}^{T_j} is defined as" but the actual expression is not rendered; readers must infer it from surrounding prose, which is fragile for the load-bearing object of the section.
  • LO-BL under two configurations. The main-text LO-BL uses 3 sources with ratio [0.4, 0.3, 0.3]; Table 4 (App. F.4) uses 6 sources with ratio [0.2, 0.2, 0.1, 0.1, 0.2, 0.2]. Same acronym, different experiment — please rename or clearly distinguish.

Recommended Changes

Essential

  • Repair the A3/A.4 justification. In Sec. 4, replace "justifying assumption A3" with informal-motivation language, and either (a) delete "no better equilibrium exists" or (b) supply a genuine equilibrium-selection argument. Cross-references to the appendix's A.4 (not A.3) should also be corrected — see the numbering mismatch below.
  • Unify assumption numbering. Make main-text A1/A2/A3 either identical to appendix A.2/A.3/A.4 or add a translation table, so a reader tracing Prop. 4.1 lands on the right appendix paragraph. In the same pass, standardize the validation-set likelihood notation (p(Oθ)p(O|\theta) vs p(Tθ)p(T|\theta)) in Assumption A.2.
  • Fix the volume duplication scaling factor. Change k\sqrt{k} to kf/2k^{f/2} in App. C.1 (or state the f=1f=1 restriction), and separately correct the information-gain claim so duplication is not conflated with "adding data".
  • Restate the feasible-reward guarantee in expectation. The Sec. 5.1 partition (ii-s·truthful when ϕi/aB\phi_i/a\leq B, ii-truthful otherwise) should be phrased as a statement over E[min(ϕi/a,B)]\mathbb{E}[\min(\phi_i/a,B)] with the mixed-regime case handled explicitly.
  • Release runnable code and seeds. Add a repository with environment.yml, seeds, and scripts that regenerate Figs. 1–3, Figs. 10–19 and Tables 3–4. This is the single biggest reproducibility win.
  • Reconcile the "use small validation sets" recommendation with Fig. 6. Either quantify the trade-off (small set discourages duplication but raises variance / worsens the LO-BL prior-uncertainty exception) or drop the blanket recommendation.
  • Advertise Sec. 5.2's F1 relaxation up front. Since the modified symmetry needs identical datasets and identical split randomness, the abstract and intro should not claim unqualified (F) for the validation-set-free case.

Suggested

  • Add ablations under substantive misspecification. Beyond the current 'minor noise' regime, run at least one experiment with a wrong link function, wrong noise family, or wrong kernel class so readers can see where the DVF actually breaks and calibrate deployment risk (relates to the empirical-scope weakness).
  • Add an empirical baseline against Chen et al. (2020) and validation-accuracy Shapley. Even if they lack matching guarantees, the practical accuracy/robustness gap is what a practitioner needs.
  • Quantify the LO-BL exception. Report how frequently duplication beats truthful submission across seeds, prior settings, and class distributions, and state deployment guidance for well-specified vs. under-informed priors.
  • Discuss adversarial specification. In App. A / Sec. 3, note that A1/A.2 (agreed prior, likelihood, model) rules out strategic misreporting at the specification stage and briefly discuss which guarantees survive if that channel is opened.
  • Add a display equation for νDTj\nu_{\mathbf D}^{T_j}. Sec. 5.2 currently has the "is defined as" leader without the definition; please render it.
  • Tighten strategy specifications. Give a single per-dataset table pinning Strategy N's label-corruption distribution, Strategy I's target class (and confirm "mode" is literal), Strategy P's noise SD relative to standardized/raw features, and confirm the LO-BL feature-extraction pipeline (ResNet-18 checkpoint, PCA fit set).
  • Move at least one larger-nn experiment into the main text. With every main-text figure at n=3n=3, the fairness/truthfulness claim looks under-tested at scale; Table 3 (20 sources) or Fig. 10e (9 sources) would help.
  • Fix presentation issues. Repair the Fig. 8 caption ("increasing sizes" \to input distribution), the scrambled sentence in App. F.5's "reward value of other sources" paragraph, the notation clash between input dimension ww and weight vector w\mathbf w, and the Table 1 row/column alignment; and clearly distinguish 3-source LO-BL from 6-source LO-BL (e.g., LO-BL-6).
  • Soften the "community views the validation-set assumption as necessary" claim. 'Common and reasonable' is what the citations support; 'necessary' is contradicted by the active line of work removing the assumption (and by the paper's own Sec. 5.2).