Bias-Free Hypothesis Evaluation in Multirelational Domains
In propositional domains using a separate test set via random sampling or cross validation is generally considered to be an unbiased estimator of true error. In multirelational domains previous work has already noted that linkage of objects may cause these procedures to be biased and has proposed corrected sampling procedures. However, as we show in this paper, the existing procedures only address one particular case of bias introduced by linkage. In this paper we therefore introduce generalized subgraph sampling, a sampling procedure based on bin packing, which ensures that test sets are properly chosen to match the probability of reencountering previously seen objects and which includes previous approaches as a special case. Experiments with data from the Internet Movie Database illustrate the performance of our algorithm.
KeywordsTransductive Learning Neighbor Probability Inductive Logic Program Probabilistic Relational Model Internet Movie Database
Unable to display preview. Download preview PDF.
- 2.Körner, C., Wrobel, S.: Bias-free hypothesis evaluation in multirelational domains. Technical report, Fraunhofer Institut Autonome Intelligente Systeme (2005), http://www.ais.fraunhofer.de/~ckoerner
- 3.Getoor, L., Friedman, N., Koller, D., Pfeffer, A.: Relational data mining. In: Dzeroski, S., Lavrac, N. (eds.) Learning Probabilistic Relational Models, pp. 307–335. Springer, Berlin (2001)Google Scholar
- 4.Taskar, B., Abbeel, P., Koller, D.: Discriminative probabilistic models for relational data. In: Proc. of the 18th Conference on Uncertainty in Artificial Intelligence (2002)Google Scholar
- 5.Neville, J., Jensen, D.: Collective classification with relational dependency networks. In: Proc. of the 2nd Multi-Relational Data Mining Workshop, 9th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (2003)Google Scholar
- 6.Fürnkranz, J.: personal communicationGoogle Scholar