A Principled Approach to Data Valuation for Federated Learning

Wang, Tianhao; Rausch, Johannes; Zhang, Ce; Jia, Ruoxi; Song, Dawn

doi:10.1007/978-3-030-63076-8_11

Tianhao Wang¹¹,
Johannes Rausch¹²,
Ce Zhang¹²,
Ruoxi Jia¹³ &
…
Dawn Song¹⁴

Part of the book series: Lecture Notes in Computer Science ((LNAI,volume 12500))

8047 Accesses
38 Citations

Abstract

Federated learning (FL) is a popular technique to train machine learning (ML) models on decentralized data sources. In order to sustain long-term participation of data owners, it is important to fairly appraise each data source and compensate data owners for their contribution to the training process. The Shapley value (SV) defines a unique payoff scheme that satisfies many desiderata for a data value notion. It has been increasingly used for valuing training data in centralized learning. However, computing the SV requires exhaustively evaluating the model performance on every subset of data sources, which incurs prohibitive communication cost in the federated setting. Besides, the canonical SV ignores the order of data sources during training, which conflicts with the sequential nature of FL. This chapter proposes a variant of the SV amenable to FL, which we call the federated Shapley value. The federated SV preserves the desirable properties of the canonical SV while it can be calculated without incurring extra communication cost and is also able to capture the effect of participation order on data value. We conduct a thorough empirical study of the federated SV on a range of tasks, including noisy label detection, adversarial participant detection, and data summarization on different benchmark datasets, and demonstrate that it can reflect the real utility of data sources for FL and has the potential to enhance system robustness, security, and efficiency. We also report and analyze “failure cases” and hope to stimulate future research.

This is a preview of subscription content, log in via an institution to check access.

Access this chapter

Log in via an institution

eBook: USD 16.99; Price excludes VAT (USA)

Softcover Book: USD 16.99; Price excludes VAT (USA)

Tax calculation will be finalised at checkout

Purchases are for personal use only

Institutional subscriptions

Notes

1.
We use preprocessing and the pretrained model as provided by PyTorch Hub.

References

WeBank and Swiss Re signed cooperation MoU (2019). https://markets.businessinsider.com/news/stocks/webank-and-swiss-re-signed-cooperation-mou-1028228738#
Bagdasaryan, E., Veit, A., Hua, Y., Estrin, D., Shmatikov, V.: How to backdoor federated learning. In: International Conference on Artificial Intelligence and Statistics, pp. 2938–2948 (2020)
Google Scholar
Brouwer, W.D.: The federated future is ready for shipping (2019). https://medium.com/@_doc_ai/the-federated-future-is-ready-for-shipping-d17ff40f43e3
Chessa, M., Loiseau, P.: A cooperative game-theoretic approach to quantify the value of personal data in networks. In: Proceedings of the 12th workshop on the Economics of Networks, Systems and Computation, p. 9. ACM (2017)
Google Scholar
Deng, X., Papadimitriou, C.H.: On the complexity of cooperative solution concepts. Math. Oper. Res. 19(2), 257–266 (1994)
Article MathSciNet Google Scholar
Ghorbani, A., Zou, J.: Data Shapley: equitable valuation of data for machine learning. arXiv preprint arXiv:1904.02868 (2019)
Gu, T., Liu, K., Dolan-Gavitt, B., Garg, S.: BadNets: evaluating backdooring attacks on deep neural networks. IEEE Access 7, 47230–47244 (2019)
Article Google Scholar
Hard, A., et al.: Federated learning for mobile keyboard prediction. arXiv preprint arXiv:1811.03604 (2018)
Heckman, J.R., Boehmer, E.L., Peters, E.H., Davaloo, M., Kurup, N.G.: A pricing model for data markets. In: iConference 2015 Proceedings (2015)
Google Scholar
Jia, Y., et al.: Efficient task-specific data valuation for nearest neighbor algorithms. Proc. VLDB Endow. 12(11), 1610–1623 (2019)
Article Google Scholar
Jia, R., et al..: Towards efficient data valuation based on the Shapley value. arXiv preprint arXiv:1902.10275 (2019)
Jia, R., et al..: Towards efficient data valuation based on the Shapley value. In: AISTATS (2019)
Google Scholar
Jia, R., Sun, X., Xu, J., Zhang, C., Li, B., Song, D.: An empirical and comparative analysis of data valuation with scalable algorithms. arXiv preprint arXiv:1911.07128 (2019)
Kang, J., Xiong, Z., Niyato, D., Xie, S., Zhang, J.: Incentive mechanism for reliable federated learning: a joint optimization approach to combining reputation and contract theory. IEEE Internet Things J. 6(6), 10700–10714 (2019)
Article Google Scholar
Kang, J., Xiong, Z., Niyato, D., Yu, H., Liang, Y.C., Kim, D.I.: Incentive design for efficient federated learning in mobile networks: a contract theory approach. In: 2019 IEEE VTS Asia Pacific Wireless Communications Symposium (APWCS), pp. 1–5. IEEE (2019)
Google Scholar
Kleinberg, J., Papadimitriou, C.H., Raghavan, P.: On the value of private information. In: Proceedings of the 8th Conference on Theoretical Aspects of Rationality and Knowledge. pp. 249–257. Morgan Kaufmann Publishers Inc. (2001)
Google Scholar
Kolesnikov, A., et al.: Big Transfer (BiT): general visual representation learning. arXiv preprint arXiv:1912.11370 (2019)
Koutris, P., Upadhyaya, P., Balazinska, M., Howe, B., Suciu, D.: Query-based data pricing. J. ACM (JACM) 62(5), 43 (2015)
Article MathSciNet Google Scholar
Krizhevsky, A., Hinton, G., et al.: Learning multiple layers of features from tiny images (2009)
Google Scholar
LeCun, Y., Cortes, C.: MNIST handwritten digit database (2010). http://yann.lecun.com/exdb/mnist/
Lee, J.S., Hoh, B.: Sell your experiences: a market mechanism based incentive for participatory sensing. In: 2010 IEEE International Conference on Pervasive Computing and Communications (PerCom), pp. 60–68. IEEE (2010)
Google Scholar
Leroy, D., Coucke, A., Lavril, T., Gisselbrecht, T., Dureau, J.: Federated learning for keyword spotting. In: ICASSP 2019–2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 6341–6345. IEEE (2019)
Google Scholar
Maleki, S.: Addressing the computational issues of the Shapley value with applications in the smart grid. Ph.D. thesis, University of Southampton (2015)
Google Scholar
McMahan, B., Moore, E., Ramage, D., Hampson, S., Arcas, B.A.: Communication-efficient learning of deep networks from decentralized data. In: Artificial Intelligence and Statistics, pp. 1273–1282 (2017)
Google Scholar
Mihailescu, M., Teo, Y.M.: Dynamic resource pricing on federated clouds. In: 2010 10th IEEE/ACM International Conference on Cluster, Cloud and Grid Computing (CCGrid), pp. 513–517. IEEE (2010)
Google Scholar
Shapley, L.S.: A value for n-person games. In: Contributions to the Theory of Games, vol. 2, no. 28, pp. 307–317 (1953)
Google Scholar
Song, T., Tong, Y., Wei, S.: Profit allocation for federated learning. In: 2019 IEEE International Conference on Big Data (Big Data), pp. 2577–2586. IEEE (2019)
Google Scholar
Upadhyaya, P., Balazinska, M., Suciu, D.: Price-optimal querying with data APIs. PVLDB 9(14), 1695–1706 (2016)
Google Scholar
Wang, G., Dang, C.X., Zhou, Z.: Measure contribution of participants in federated learning. In: 2019 IEEE International Conference on Big Data (Big Data), pp. 2597–2604. IEEE (2019)
Google Scholar
Yu, H., et al.: A fairness-aware incentive scheme for federated learning. In: Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, pp. 393–399 (2020)
Google Scholar

Download references

Author information

Authors and Affiliations

Harvard University, Cambridge, USA
Tianhao Wang
ETH Zurich, Zürich, Switzerland
Johannes Rausch & Ce Zhang
Virginia Tech, Blacksburg, USA
Ruoxi Jia
UC Berkeley, Berkeley, USA
Dawn Song

Authors

Tianhao Wang
View author publications
You can also search for this author in PubMed Google Scholar
Johannes Rausch
View author publications
You can also search for this author in PubMed Google Scholar
Ce Zhang
View author publications
You can also search for this author in PubMed Google Scholar
Ruoxi Jia
View author publications
You can also search for this author in PubMed Google Scholar
Dawn Song
View author publications
You can also search for this author in PubMed Google Scholar

Corresponding author

Correspondence to Ruoxi Jia .

Editor information

Editors and Affiliations

Hong Kong University of Science and Technology, Hong Kong, Hong Kong
Qiang Yang
WeBank, Shenzhen, China
Lixin Fan
Nanyang Technological University, Singapore, Singapore
Han Yu

Rights and permissions

Reprints and permissions

Copyright information

About this chapter

Cite this chapter

Wang, T., Rausch, J., Zhang, C., Jia, R., Song, D. (2020). A Principled Approach to Data Valuation for Federated Learning. In: Yang, Q., Fan, L., Yu, H. (eds) Federated Learning. Lecture Notes in Computer Science(), vol 12500. Springer, Cham. https://doi.org/10.1007/978-3-030-63076-8_11

Download citation

DOI: https://doi.org/10.1007/978-3-030-63076-8_11
Published: 26 November 2020
Publisher Name: Springer, Cham
Print ISBN: 978-3-030-63075-1
Online ISBN: 978-3-030-63076-8
eBook Packages: Computer ScienceComputer Science (R0)

Publish with us

Policies and ethics