Skip to main content

Mining document, concept, and term associations for effective biomedical retrieval: introducing MeSH-enhanced retrieval models

Abstract

Manually assigned subject terms, such as Medical Subject Headings (MeSH) in the health domain, describe the concepts or topics of a document. Existing information retrieval models do not take full advantage of such information. In this paper, we propose two MeSH-enhanced (ME) retrieval models that integrate the concept layer (i.e. MeSH) into the language modeling framework to improve retrieval performance. The new models quantify associations between documents and their assigned concepts to construct conceptual representations for the documents, and mine associations between concepts and terms to construct generative concept models. The two ME models reconstruct two essential estimation processes of the relevance model (Lavrenko and Croft 2001) by incorporating the document-concept and the concept-term associations. More specifically, in Model 1, language models of the pseudo-feedback documents are enriched by their assigned concepts. In Model 2, concepts that are related to users’ queries are first identified, and then used to reweight the pseudo-feedback documents according to the document-concept associations. Experiments carried out on two standard test collections show that the ME models outperformed the query likelihood model, the relevance model (RM3), and an earlier ME model. A detailed case analysis provides insight into how and why the new models improve/worsen retrieval performance. Implications and limitations of the study are discussed. This study provides new ways to formally incorporate semantic annotations, such as subject terms, into retrieval models. The findings of this study suggest that integrating the concept layer into retrieval models can further improve the performance over the current state-of-the-art models.

This is a preview of subscription content, access via your institution.

Fig. 1
Fig. 2
Fig. 3
Fig. 4
Fig. 5
Fig. 6
Fig. 7
Fig. 8
Fig. 9
Fig. 10

Notes

  1. http://eutils.ncbi.nlm.nih.gov/entrez/eutils/.

  2. http://www.lemurproject.org/.

References

  • Abdou, S., Ruck, P., & Savoy, J. (2005). Evaluation of stemming, query expansion and manual indexing approaches for the genomic task. In Proceedings of TREC 2005.

  • Bacchin, M., & Melucci, M. (2005). Symbol-based query expansion experiments at TREC 2005 Genomics track. In Proceedings of TREC 2005.

  • Bai, J., Song, D., Bruza, P., Nie, J. Y., & Cao, G. (2005). Query expansion using term relationships in language models for information retrieval. In Proceedings of CIKM 2005 (pp. 688–695). Bremen: ACM.

  • Blei, D. M., Ng, A. Y., & Jordan, M. I. (2003). Latent dirichlet allocation. The Journal of Machine Learning Research, 3, 993–1022.

    MATH  Google Scholar 

  • Díaz-Galiano, M. C., García-Cumbreras, M. A., Martín-Valdivia, M. T., Montejo-Ráez, A., & Urena-López, L. A. (2008). Integrating mesh ontology to improve medical information retrieval. In Advances in multilingual and multimodal information retrieval (pp. 601–606). Berlin: Springer.

  • Fang, H., & Zhai, C. (2005). An exploration of axiomatic approaches to information retrieval. In Proceedings of SIGIR 2005 (pp. 480–487). Salvador: ACM.

  • Finkelstein, L. (2002). Placing search in context: The concept revisited. ACM Transactions on Information Systems, 20, 116–131.

    Article  Google Scholar 

  • Gabrilovich, E., & Markovitch, S. (2007). Computing semantic relatedness using Wikipedia-based explicit semantic analysis. In Proceedings of the 20th international joint conference on artifical intelligence (pp. 1606–1611). Morgan Kaufmann Publishers Inc.

  • Gauch, S., & Smith, J. B. (1991). Search improvement via automatic query reformulation. ACM Transactions on Information Systems (TOIS), 9(3), 249–280.

    Article  Google Scholar 

  • Gault, L. V., Shultz, M., & Davies, K. J. (2002). Variations in Medical Subject Headings (MeSH) mapping: From the natural language of patron terms to the controlled vocabulary of mapped lists. Journal of the Medical Library Association, 90(2), 173.

    Google Scholar 

  • Gonzalo, J., Verdejo, F., Chugur, I., & Cigarran, J. (1998). Indexing with WordNet synsets can improve text retrieval. arXiv preprint cmp-lg/9808002.

  • Griffon, N., Chebil, W., Rollin, L., Kerdelhue, G., Thirion, B., Gehanno, J. F., & Darmoni, S. J. (2012). Performance evaluation of unified medical language system®’s synonyms expansion to query PubMed. BMC Medical Informatics and Decision Making, 12(1), 12.

    Article  Google Scholar 

  • Guisado-Gámez, J., Dominguez-Sal, D., & Larriba-Pey, J. L. (2013). Massive query expansion by exploiting graph knowledge bases. arXiv preprint arXiv:1310.5698.

  • Guo, Y., Harkema, H., & Gaizauskas, R. (2004). Sheffield university and the TREC 2004 Genomics track: Query expansion using synonymous terms. In Proceedings of the thirteenth Text REtrieval conference. Gaithersburg, MD: Department of Commerce, National Institute of Standards and Technology.

  • Harman, D., & Buckley, C. (2009). Overview of the reliable information access workshop. Information Retrieval, 12(6), 615–641.

    Article  Google Scholar 

  • He, B., & Ounis, I. (2009). Finding good feedback documents. In Proceedings of the 18th ACM conference on Information and knowledge management (pp. 2011–2014). Hong Kong: ACM.

  • Hersh, W. (2008). Information retrieval: A health and biomedical perspective. Berlin: Springer.

    Google Scholar 

  • Hersh, W., & Bhupatiraju, R. T. (2003). TREC Genomics track overview. In Proceedings of the twelfth text retrieval conference, TREC 2003 (pp. 14–23). Gaithersburg, MD: Department of Commerce, National Institute of Standards and Technology.

  • Hersh, W., Bhupatiraju, R. T., & Price, S. (2003). Phrases, boosting, and query expansion using external knowledge resources for genomic information retrieval. In Proceedings of the twelfth text retrieval conference. Gaithersburg, MD: Department of Commerce, National Institute of Standards and Technology.

  • Hersh, W., Buckley, C., Leone, T. J., & Hickam, D. (1994). OHSUMED: An interactive retrieval evaluation and new large test collection for research. In Proceedings of SIGIR 1994 (pp. 192–201). London: Springer.

  • Hersh, W. R., Cohen, A. M., Roberts, P. M., & Rekapalli, H. K. (2006). TREC 2006 Genomics track overview. In TREC 2006.

  • Jelinek, F., & Mercer, R. L. (1980). Interpolated estimation of Markov source parameters from sparse data. In Proceedings of the workshop on pattern recognition in practice. Amsterdam: North-Holland.

  • Kamps, J. (2004). Improving retrieval effectiveness by reranking documents based on controlled vocabulary. In Advances in information retrieval (pp. 283–295). Berlin: Springer.

  • Korfhage, R. R. (1984). Query enhancement by user profiles. In Proceedings of SIGIR 1984 (pp. 111–121). Cambridge: British Computer Society.

  • Kurland, O. (2008). The opposite of smoothing: A language model approach to ranking query-specific document clusters. In Proceedings of SIGIR 2008 (pp. 171–178). Singapore: ACM.

  • Kurland, O. (2009). Re-ranking search results using language models of query-specific clusters. Information Retrieval, 12(4), 437–460.

    Article  Google Scholar 

  • Kurland, O., & Lee, L. (2004). Corpus structure, language models, and ad hoc information retrieval. In Proceedings of SIGIR 2004 (pp. 194–201). Sheffield: ACM.

  • Lafferty, J., & Zhai, C. (2001). Document language models, query models, and risk minimization for information retrieval. In Proceedings of SIGIR 2001 (pp. 111–119). New Orleans: ACM.

  • Lafferty, J., & Zhai, C. (2003). Probabilistic relevance models based on document and query generation. In Language modeling for information retrieval (pp. 1–10). Netherlands: Springer.

  • Lavrenko, V., & Croft, W. B. (2001). Relevance based language models. In Proceedings of SIGIR 2001 (pp. 120–127). New Orleans: ACM.

  • Lee, K. S., Croft, W. B., & Allan, J. (2008). A cluster-based resampling method for pseudo-relevance feedback. In Proceedings of SIGIR 2008 (pp. 235–242). Singapore: ACM.

  • Liu, X., & Croft, W. B. (2004). Cluster-based retrieval using language models. In Proceedings of SIGIR 2004 (pp. 186–193). Sheffield: ACM.

  • Lu, Z., Kim, W., & Wilbur, W. J. (2009). Evaluation of query expansion using MeSH in PubMed. Information Retrieval, 12(1), 69–80.

    Article  Google Scholar 

  • Lu, K., & Mao, J. (2013). Automatically infer subject terms and documents associations through text mining. In Proceedings of the 76th annual conference of association for information science and technology (ASIST’2013), Montreal, Canada.

  • Lv, Y., & Zhai, C. (2009). A comparative study of methods for estimating query language models with pseudo feedback. In Proceedings of CIKM 2009 (pp. 1895–1898). Hong Kong: ACM.

  • Lv, Y., Zhai, C., & Chen, W. (2011). A boosting approach to improving pseudo-relevance feedback. In Proceedings of SIGIR 2011 (pp. 165–174). Beijing: ACM.

  • Manning, C. D., Raghavan, P., & Schütze, H. (2008). Introduction to information retrieval. Cambridge: Cambridge University Press.

    MATH  Book  Google Scholar 

  • Mata, J., Crespo, M., & Maña, M. J. (2012). Using MeSH to expand queries in medical image retrieval. In Medical content-based retrieval for clinical decision support (pp. 36–46). Berlin: Springer.

  • Meij, E., & De Rijke, M. (2007). Integrating conceptual knowledge into relevance models: A model and estimation method. In International conference on the theory of information retrieval (ICTIR 2007). Budapest: Alma Mater Series.

  • Meij, E., Trieschnigg, D., De Rijke, M., & Kraaij, W. (2010). Conceptual language models for domain-specific retrieval. Information Processing and Management, 46(4), 448–469.

    Article  Google Scholar 

  • Metzler, D., & Croft, W. B. (2005). A Markov random field model for term dependencies. In Proceedings of SIGIR 2005 (pp. 472–479). Salvador: ACM.

  • Metzler, D., Dumais, S., & Meek, C. (2007). Similarity measures for short segments of text. In Advances in information retrieval (pp. 16–27). Berlin: Springer.

  • Montgomery, J., Si, L., Callan, J., & Evans, D. (2004). Effect of varying number of documents in blind feedback: Analysis of the 2003 NRRC RIA workshop “bf_numdocs” experiment suite. In Proceedings of SIGIR 2004 (pp. 476–477). Sheffield: ACM.

  • Plaunt, C., & Norgard, B. A. (1998). An association-based method for automatic indexing with a controlled vocabulary. Journal of the American Society for Information Science, 49(10), 888–902.

    Google Scholar 

  • Poikonen, T., & Vakkari, P. (2009). Lay persons’ and professionals’ nutrition-related vocabularies and their matching to a general and a specific thesaurus. Journal of Information Science, 35(2), 232–243.

    Article  Google Scholar 

  • Ponte, J. M., & Croft, W. B. (1998). A language modeling approach to information retrieval. In Proceedings of SIGIR 1998 (pp. 275–281). Melbourne: ACM.

  • Shin, K., & Han, S. Y. (2004). Improving information retrieval in MEDLINE by modulating MeSH term weights. In Natural language processing and information systems (pp. 388–394). Berlin: Springer.

  • Shiri, A. (2012). Powering search: The role of Thesauri in new information environments. Medford, NJ: Information Today Inc.

    Google Scholar 

  • Smucker, M. D., Allan, J., & Carterette, B. (2007). A comparison of statistical significance tests for information retrieval evaluation. In Proceedings of the sixteenth ACM conference on information and knowledge management (pp. 623–632). New York: ACM.

  • Srinivasan, P. (1996). Query expansion and MEDLINE. Information Processing and Management, 32(4), 431–443.

    Article  Google Scholar 

  • Stokes, N., Li, Y., Cavedon, L., & Zobel, J. (2009). Exploring criteria for successful query expansion in the genomic domain. Information Retrieval, 12(1), 17–50.

    Article  Google Scholar 

  • Trieschnigg, D. (2010). Proof of concept: Concept-based biomedical information retrieval. Doctoral dissertation, University of Twente.

  • Trieschnigg, D., Pezik, P., Lee, V., de Jong, F., Kraaij, W., & Rebholz-Schuhmann, D. (2009). MeSH up: Effective MeSH text classification for improved document retrieval. Bioinformatics, 25, 1412–1418.

    Article  Google Scholar 

  • van Rijsbergen, (1979). Information retrieval (2nd ed.). London: Butterworths.

    Google Scholar 

  • Vechtomova, O., Robertson, S., & Jones, S. (2003). Query expansion with long-span collocates. Information Retrieval, 6(2), 251–273.

    Article  Google Scholar 

  • Voorhees, E. M. (1994). Query expansion using lexical–semantic relations. In Proceedings of SIGIR 1994 (pp. 61–69). London: Springer.

  • Wang, L., Bennett, P. N., & Collins-Thompson, K. (2012). Robust ranking models via risk-sensitive optimization. In Proceedings of SIGIR 2012 (pp. 761–770). Portland: ACM.

  • Wei, X., & Croft, W. B. (2006). LDA-based document models for ad-hoc retrieval. In Proceedings of SIGIR 2006 (pp. 178–185). Seattle: ACM.

  • Xu, J., & Croft, W. B. (1996). Query expansion using local and global document analysis. In Proceedings of SIGIR 1996 (pp. 4–11). Zurich: ACM.

  • Zeng, Q. T., Crowell, J., Plovnick, R. M., Kim, E., Ngo, L., & Dibble, E. (2006). Assisting consumer health information retrieval with query recommendations. Journal of the American Medical Informatics Association, 13(1), 80–90.

    Article  Google Scholar 

  • Zeng, Q., Kogan, S., Ash, N., Greenes, R. A., & Boxwala, A. A. (2002). Characteristics of consumer terminology for health information retrieval. Methods of Information in Medicine, 41(4), 289–298.

    Google Scholar 

  • Zeng, Q. T., Kogan, S., Plovnick, R. M., Crowell, J., Lacroix, E. M., & Greenes, R. A. (2004). Positive attitudes and failed queries: an exploration of the conundrums of consumer health information retrieval. International Journal of Medical Informatics, 73(1), 45–55.

    Article  Google Scholar 

  • Zhai, C. (2002). Risk minimization and language modeling in text retrieval. Doctoral dissertation, University of Massachusetts, Amherst.

  • Zhai, C., & Lafferty, J. (2001a). A study of smoothing methods for language models applied to ad hoc information retrieval. In Proceedings of the SIGIR 2001 (pp. 334–342). New Orleans: ACM.

  • Zhai, C., & Lafferty, J. (2001b). Model-based feedback in the language modeling approach to information retrieval. In Proceedings of the CIKM 2001 (pp. 403–410). Atlanta: ACM.

  • Zhai, C., & Lafferty, J. (2002). Two-stage language models for information retrieval. In Proceedings of the SIGIR 2002 (pp. 49–56). Tampere: ACM.

  • Zhai, C., & Lafferty, J. (2004). A study of smoothing methods for language models applied to information retrieval. ACM Transactions on Information Systems, 22(2), 179–214.

    Article  Google Scholar 

  • Zhang, J., Wolfram, D., Wang, P., Hong, Y., & Gillis, R. (2008). Visualization of health-subject analysis based on query term co-occurrences. Journal of the American Society for Information Science and Technology, 59, 1933–1947.

    Article  Google Scholar 

  • Zielstorff, R. D. (2003). Controlled vocabularies for consumer health. Journal of Biomedical Informatics, 36, 326–333.

    Article  Google Scholar 

Download references

Acknowledgments

We thank the anonymous reviewers for their constructive comments on an earlier version of the manuscript. Their comments definitely help to improve the paper. We are also grateful for helpful comments from Professor Wei Lu and Associate Professor Lu An. This work was sponsored by the National Natural Science Foundation of China funded projects under grant No. 71273196, No. 71420107026, and No. 71373286.

Author information

Authors and Affiliations

Authors

Corresponding author

Correspondence to Kun Lu.

Rights and permissions

Reprints and Permissions

About this article

Verify currency and authenticity via CrossMark

Cite this article

Mao, J., Lu, K., Mu, X. et al. Mining document, concept, and term associations for effective biomedical retrieval: introducing MeSH-enhanced retrieval models. Inf Retrieval J 18, 413–444 (2015). https://doi.org/10.1007/s10791-015-9264-0

Download citation

  • Received:

  • Accepted:

  • Published:

  • Issue Date:

  • DOI: https://doi.org/10.1007/s10791-015-9264-0

Keywords

  • Relevance model
  • Concept
  • MeSH-enhanced retrieval models
  • Health information retrieval