International Semantic Web Conference

The Semantic Web - ISWC 2015 pp 270-278

Automatic Curation of Clinical Trials Data in LinkedCT

Conference paper

DOI: 10.1007/978-3-319-25010-6_16

Part of the Lecture Notes in Computer Science book series (LNCS, volume 9367)
Cite this paper as:
Hassanzadeh O., Miller R.J. (2015) Automatic Curation of Clinical Trials Data in LinkedCT. In: Arenas M. et al. (eds) The Semantic Web - ISWC 2015. Lecture Notes in Computer Science, vol 9367. Springer, Cham


The Linked Clinical Trials (LinkedCT) project started back in 2008 with the goal of providing a Linked Data source of clinical trials. The source of the data is from the XML data published on, which is an international registry of clinical studies. Since the initial release, the LinkedCT project has gone through some major changes to both improve the quality of the data and its freshness. The result is a high-quality Linked Data source of clinical studies that is updated daily, currently containing over 195,000 trials, 4.6 million entities, and 42 million triples. In this paper, we present a detailed description of the system along with a brief outline of technical challenges involved in curating the raw XML data into high-quality Linked Data. We also present usage statistics and a number of interesting use cases developed by external parties. We share the lessons learned in the design and implementation of the current system, along with an outline of our future plans for the project which include making the system open-source and making the data free for commercial use.


Clinical trials Linked data Data curation 


Unable to display preview. Download preview PDF.

Unable to display preview. Download preview PDF.

Copyright information

© Springer International Publishing Switzerland 2015

Authors and Affiliations

  1. 1.Department of Computer ScienceUniversity of TorontoTorontoCanada
  2. 2.IBM T.J. Watson Research CenterNew YorkUSA

Personalised recommendations