Knowledge Discovery from Semantically Heterogeneous Aggregate Databases Using Model-Based Clustering

  • Shuai Zhang
  • Sally McClean
  • Bryan Scotney
Conference paper

DOI: 10.1007/978-3-540-73390-4_22

Part of the Lecture Notes in Computer Science book series (LNCS, volume 4587)
Cite this paper as:
Zhang S., McClean S., Scotney B. (2007) Knowledge Discovery from Semantically Heterogeneous Aggregate Databases Using Model-Based Clustering. In: Cooper R., Kennedy J. (eds) Data Management. Data, Data Everywhere. BNCOD 2007. Lecture Notes in Computer Science, vol 4587. Springer, Berlin, Heidelberg

Abstract

When distributed databases are developed independently, they may be semantically heterogeneous with respect to data granularity, scheme information and the embedded semantics. However, most traditional distributed knowledge discovery (DKD) methods assume that the distributed databases derive from a single virtual global table, where they share the same semantics and data structures. This data heterogeneity and the underlying semantics bring a considerable challenge for DKD. In this paper, we propose a model-based clustering method for aggregate databases, where the heterogeneous schema structure is due to the heterogeneous classification schema. The underlying semantics can be captured by different clusters. The clustering is carried out via a mixture model, where each component of the mixture corresponds to a different virtual global table. An advantage of our approach is that the algorithm resolves the heterogeneity as part of the clustering process without previously having to homogenise the heterogeneous local schema to a shared schema. Evaluation of the algorithm is carried out using both real and synthetic data. Scalability of the algorithm is tested against the number of databases to be clustered; the number of clusters; and the size of the databases. The relationship between performance and complexity is also evaluated. Our experiments show that this approach has good potential for scalable integration of semantically heterogeneous databases.

Keywords

Model-based clustering Semantically heterogeneous databases EM algorithm 

Preview

Unable to display preview. Download preview PDF.

Unable to display preview. Download preview PDF.

Copyright information

© Springer-Verlag Berlin Heidelberg 2007

Authors and Affiliations

  • Shuai Zhang
    • 1
  • Sally McClean
    • 1
  • Bryan Scotney
    • 1
  1. 1.School of Computing and Information Engineering, University of Ulster, Coleraine, Northern IrelandUK

Personalised recommendations