A systematic method to create search strategies for emerging technologies based on the Web of Science: illustrated for ‘Big Data’
- 1.2k Downloads
Bibliometric and “tech mining” studies depend on a crucial foundation—the search strategy used to retrieve relevant research publication records. Database searches for emerging technologies can be problematic in many respects, for example the rapid evolution of terminology, the use of common phraseology, or the extent of “legacy technology” terminology. Searching on such legacy terms may or may not pick up R&D pertaining to the emerging technology of interest. A challenge is to assess the relevance of legacy terminology in building an effective search model. Common-usage phraseology additionally confounds certain domains in which broader managerial, public interest, or other considerations are prominent. In contrast, searching for highly technical topics is relatively straightforward. In setting forth to analyze “Big Data,” we confront all three challenges—emerging terminology, common usage phrasing, and intersecting legacy technologies. In response, we have devised a systematic methodology to help identify research relating to Big Data. This methodology uses complementary search approaches, starting with a Boolean search model and subsequently employs contingency term sets to further refine the selection. The four search approaches considered are: (1) core lexical query, (2) expanded lexical query, (3) specialized journal search, and (4) cited reference analysis. Of special note here is the use of a “Hit-Ratio” that helps distinguish Big Data elements from less relevant legacy technology terms. We believe that such a systematic search development positions us to do meaningful analyses of Big Data research patterns, connections, and trajectories. Moreover, we suggest that such a systematic search approach can help formulate more replicable searches with high recall and satisfactory precision for other emerging technology studies.
KeywordsSearch strategy Lexical query Citation analysis Big Data
We acknowledge support from the US National Science Foundation (Award #1527370—“Forecasting Innovation Pathways of Big Data & Analytics”). Besides, we are grateful for the scholarship provided by the China Scholarship Council (CSC Student ID 201406030005). The findings and observations contained in this paper are those of the authors and do not necessarily reflect the views of the National Science Foundation and China Scholarship Council.
- Cooper, H., Hedges, L. V., & Valentine, J. C. (Eds.). (2009). The handbook of research synthesis and meta-analysis. New York: Russell Sage Foundation.Google Scholar
- Garfield, E., Paris, S., & Stock, W. G. (2006). HistCiteTM: A software tool for informetric analysis of citation linkage. Information Wissenschaft und Praxis, 57(8), 391–400.Google Scholar
- Halevi, G., & Moed, H. (2012). The evolution of big data as a research and scientific topic: Overview of the literature. Research Trends, 30(1), 3–6.Google Scholar
- Manyika, J., Chiu, M., Brown, B., Bughin, J., Dobbs, R., Roxburgh, C., et al. (2011). Big data: The next frontier for innovation, competition, and productivity. McKinsey Global Institute.Google Scholar
- McAfee, A., & Brynjolfsson, E. (2012). Big data: The management revolution. Harvard Business Review, 90, 60–67.Google Scholar
- Miller, H. E. (2013). Big-data in cloud computing: A taxonomy of risks. Information Research, 18(1). http://InformationR.net/ir/18-1/paper571.html
- Porter, A. L., & Cunningham, S. W. (2005). Tech mining: Exploiting new technologies for competitive advantage. New York: Wiley. [Chinese edition, Tsinghua University Press, 2012].Google Scholar
- Porter, A. L., Huang, Y., Schuehle, J., & Youtie, J. (2015). MetaData: BigData research evolving across disciplines, players, and topics. New York (July): IEEE BigData Congress.Google Scholar
- Rousseau, R. (2012). A view on big data and its relation to informetrics. Chinese Journal of Library and Information Science, 5(3), 12–26.Google Scholar