Finding the Achilles Heel of the Web of Data: Using Network Analysis for Link-Recommendation
The Web of Data is increasingly becoming an important infrastructure for such diverse sectors as entertainment, government, e-commerce and science. As a result, the robustness of this Web of Data is now crucial. Prior studies show that the Web of Data is strongly dependent on a small number of central hubs, making it highly vulnerable to single points of failure. In this paper, we present concepts and algorithms to analyse and repair the brittleness of the Web of Data. We apply these on a substantial subset of it, the 2010 Billion Triple Challenge dataset. We first distinguish the physical structure of the Web of Data from its semantic structure. For both of these structures, we then calculate their robustness, taking betweenness centrality as a robustness-measure. To the best of our knowledge, this is the first time that such robustness-indicators have been calculated for the Web of Data. Finally, we determine which links should be added to the Web of Data in order to improve its robustness most effectively. We are able to determine such links by interpreting the question as a very large optimisation problem and deploying an evolutionary algorithm to solve this problem. We believe that with this work, we offer an effective method to analyse and improve the most important structure that the Semantic Web community has constructed to date.
KeywordsShort Path Betweenness Centrality Semantic Network Selective Strategy High Betweenness Centrality
- 4.Bader, D., Madduri, K.: SNAP, Small-world Network Analysis and Partitioning: an open-source parallel graph framework for the exploration of large-scale networks. In: IEEE International Symposium on Parallel and, pp. 1–12. IEEE, Los Alamitos (April 2008)Google Scholar
- 9.Gil, R., Garcia, R.: Measuring the semantic web. In: Advances in Metadata Research, Proceedings of MTSR 2005. Rinton Press (2006) ISBN 1-58949-053-3Google Scholar
- 10.Guéret, C., Wang, S., Schlobach, S.: The web of data is a complex system - first insight into its multi-scale network properties. In: Proceedings of the European Conference on Complex Systems, ECCS (2010) (to appear)Google Scholar
- 11.Jaffri, A., Glaser, H., Millard, I.: Uri identity management for semantic web data integration and linkage. In: 3rd International Workshop On Scalable Semantic Web Knowledge Base Systems. Springer, Heidelberg (2007)Google Scholar
- 13.Zhang, X., Cheng, G., Qu, Y.: Ontology summarization based on rdf sentence graph. In: Proceedings of the 16th International Conference on World Wide Web, WWW 2007, pp. 707–716. ACM, New York (2007)Google Scholar