🤖 AI Summary
This paper addresses incentive imbalance arising from asynchronous participant enrollment in multi-party collaborative data sharing. We propose the first time-aware fair incentive mechanism, departing from conventional synchronous-assumption frameworks by incorporating temporal dynamics into data value assessment. Grounded in game-theoretic modeling, the mechanism quantifies the higher risk borne by early contributors and designs a time-sensitive reward allocation principle to ensure both temporal fairness and individual rationality under dynamic enrollment. Our method jointly leverages model-output contribution scores and enrollment timing to generate implementable, differentiated rewards. Extensive experiments on synthetic and real-world datasets demonstrate its effectiveness: it significantly enhances latecomers’ willingness to contribute while satisfying key properties—including fairness, individual rationality, and computational feasibility.
📝 Abstract
In collaborative data sharing and machine learning, multiple parties aggregate their data resources to train a machine learning model with better model performance. However, as the parties incur data collection costs, they are only willing to do so when guaranteed incentives, such as fairness and individual rationality. Existing frameworks assume that all parties join the collaboration simultaneously, which does not hold in many real-world scenarios. Due to the long processing time for data cleaning, difficulty in overcoming legal barriers, or unawareness, the parties may join the collaboration at different times. In this work, we propose the following perspective: As a party who joins earlier incurs higher risk and encourages the contribution from other wait-and-see parties, that party should receive a reward of higher value for sharing data earlier. To this end, we propose a fair and time-aware data sharing framework, including novel time-aware incentives. We develop new methods for deciding reward values to satisfy these incentives. We further illustrate how to generate model rewards that realize the reward values and empirically demonstrate the properties of our methods on synthetic and real-world datasets.