Introduction

Claims about the value and benefits of diversityFootnote 1 have proliferated in recent decades, becoming a contemporary secular credo that animates the spirit of almost all elite institutions of the English-speaking world. Consider, for example, President Bill Clinton’s proclamation, in his State of the Union Address, that “we must never, ever believe that our diversity is a weakness—it is our greatest strength” (Clinton, 1997). Nowadays, such contentions about the value and benefits of diversity—what we refer to as the “Diversity Hypothesis”—are commonplace, even axiomatic.

However, in contemplating the meaning and effects of diversity, it is important to note that ‘diversity’ is not a single fixed concept but something that can be understood in two distinct ways. For instance, as argued by Mogilski et al. (2025), diversity can refer to diversity in demographic makeup and identity, with an emphasis on including underrepresented groups within positions of power, or it can refer to diversity of perspective.

In the context of workplaces and fields of inquiry, the contexts most relevant to the present review, the Diversity Hypothesis is arguably manifest most prominently in the cognitive or informational Diversity Hypothesis (Mello & Rentsch, 2015) that posits that teams whose members vary in terms of their perspectives, thinking styles, knowledge, skills, values, and beliefs (i.e., diversity of perspective or viewpoint) are more likely to generate unique ideas and find more creative solutions to complex problems than homogenous teams. Consistent with this, popular science writers such as Scott E. Page (2007, 2019) contend that groups of people with diverse perspectives, heuristics and interpretations systematically outperform groups composed of the “best” individual performers, with additional claims that this viewpoint diversity ultimately stems from identity diversity (e.g., demographic differences, such as gender and race).

Consistent with these formulations of the Diversity Hypothesis, in an article for The New York Times, Nicholas Kristoff asserted that “scholarly research suggests that the best problem-solving doesn’t come from a group of the best individual problem-solvers, but from a diverse team whose members complement each other” (Kristoff, 2013). Similarly, Phillips and colleagues argued in a highly cited article in Scientific American that decades of research “shows that socially diverse groups…are more innovative than homogeneous groups” (Phillips et al., 2014). More recently, in the specific context of science, Smith, in an op-ed for Canada’s University Affairs, asserted “the evidence is unambiguous: diversity fosters more original, widely recognized and impactful science” (Smith, 2025).

However, although it is often claimed that diversity offers genuine potential benefits for innovation and creativity in the context of work, these benefits are neither automatic nor guaranteed. Indeed, as posited by the social psychologist Alice Eagly, there is a chasm between research findings about the actual value and benefits of diversity and advocates’ claims about diversity (Eagly 2016b). Specifically, Eagly noted that advocates of diversity often invoke findings that support their objectives while ignoring findings that are unsupportive, thus highlighting results that are politically congenial but unrepresentative of the available scientific knowledge (Eagly 2016a).

In the context of the Diversity Hpothesis in science, although the scientific literature comprises numerous studies that support the claim that diversity results in better and more impactful science, it is also replete with studies that not only do not support the claim, but contradict it. Given the gap between actual value and benefits of diversity and advocates’ claims about diversity, as revealed in mainstream news, the popular press, and scholarly journals, what evidence supports the claim that diversity fosters more original, widely recognized, and impactful science? The purpose of this rapid scoping review is to address this question by focusing on two specific tasks; namely, the collation of the available evidence for the Diversity Hypothesis in science and the assessment of whether or not the findings of these studies are consistent or inconsistent with the Diversity Hypothesis. These two basic tasks are necessarily antecedent to future, methodologically rigorous studies of the types of diversity that foster science, as well as the processes and conditions through which the benefits of diversity do and do not obtain. Before we turn to the methods and results of this rapid scoping review, we set the scene by summarizing the main theoretical arguments that have been advanced for and against the Diversity Hypothesis in the context of science.

Arguments for the Diversity Hypothesis

Diversity enhances innovation in science

According to Hofstra and colleagues, underrepresented groups in science produce higher rates of novelty and innovation (Hofstra et al., 2020). In their analysis of ~ 1.2 million US doctoral recipients from 1977 to 2015, Hofstra et al. concluded that women and nonwhite scholars (i.e., demographically diverse individuals with minoritized identities) introduce more novel conceptual linkages in their dissertations than scholars categorized as members majority groups, suggesting that demographic diversity fosters scientific innovation.

Diversity enhances creativity, decision-making and problem-solving

According to Yang and colleagues (2022), gender-diverse teams produce more novel and higher-impact scientific ideas because diverse teams bring varied cognitive frameworks, analytical approaches, and problem-solving strategies that ultimately generate more comprehensive solutions. Similarly, in Stahl and colleagues’ meta-analysis of 10,632 multicultural work groups, culturally diverse teams experience more task conflict, which can improve vetting and decision-making (Stahl et al., 2010).

Diversity enhances long-term impact and cross-disciplinary influence

In their analysis of 23 million scientific publications and 4 million patents, Li and Zheng (2025) found that although teams with more diverse expertise have no significant impact advantage in the short- (2 years) or mid-term (5 years) over teams with less diverse expertise, they have substantially higher long-term impact (10 years), increasingly attracting larger cross-disciplinary influence. This pattern of findings suggests that benefits of diversity compound over time as novel approaches gain recognition.

Arguments against the Diversity Hypothesis

Diversity increases interpersonal conflict and reduces cohesion

According to Jehn and colleagues (Jehn et al., 1999), diversity increases multiple forms of conflict (relationship conflict, process conflict, task conflict) and reduces social integration and team cohesion, (cf. Putnam, 2007). Although social category or identity diversity (e.g., demographic differences) modestly positively influences group member morale and although informational diversity (e.g., differences in knowledge, expertise, background) positively influences group performance, value diversity (i.e., differences in what group members believe the group should be doing and why) negatively influences group members’ satisfaction, commitment to the group and intention to remain in the group.

Diversity effects are contingent and context-dependent

According to van Knippenberg, the effect of diversity on performance range from strongly negative to strongly positive depending on context, suggesting no generalizable “diversity benefit” exists (van Knippenberg, 2024). In many real-world scientific contexts lacking ideal supporting conditions, diversity may impair rather than enhance performance.

Diversity produces outcomes that are devalued and discounted

According to Hofstra and colleagues, although individuals belonging to underrepresented groups (i.e., identity or demographically diverse people) produce higher rates of scientific novelty, these novel contributions tend to be devalued and discounted compared to majority groups, undermining the hypothesis that diversity improves measured scientific outcomes and impact (Hofstra et al., 2020).

The main objective of this rapid scoping review is to collate and examine the evidence for the claim that diversity of scientists, as ‘diverse’ individuals or members of diverse scientific teams, enhances scientific output or impact. “Diversity” can be understood in two distinct ways: demographic (or identity) diversity, where an individual is considered "diverse" if they belong to a minority or minoritized group, and viewpoint diversity within a group. Because the concept is not fixed or unambiguous, this investigation is not a straightforward review. As such, we necessarily screen off a number of important issues. First, we restrict our analysis to the categorization of findings with respect to their consistency or inconsistency with the Diversity Hypothesis; namely, the hypothesis that diversity enhances scientific output or impact. Second, because of this focus on this categorization of results, we do not consider potential processes that mediate the effect of diversity of scientists and scientific teams on scientific output or impact or the contextual factors (e.g., supportive organizational climates) that moderate these relationships, noting that there is an established literature focusing on how diversity influences group decision making and problem solving in general. These are important questions but are beyond the scope of this study. Finally, given our focus on the categorization of findings with respect to their consistency with the Diversity Hypothesis, we do not examine the possible reasons why the Diversity Hypothesis was or was not supported in the studies included in our review. The absence of an examination of mediators, moderators, and reasons for observed outcomes does not imply uncritical endorsement of claims, but a deliberate neutrality.

Methods

Design

This rapid scoping review was guided by the Joanna Briggs Institute (JBI) approach (Peters et al., 2020) made up of nine steps: (1) defining and aligning the objectives and research questions; (2) developing and aligning the inclusion criteria with the objectives and questions; (3) describing the planned approach to evidence searching, selection, data extraction and presentation of the evidence; (4) searching the evidence; (5) selecting the evidence; (6) extracting the evidence; (7) analysis of the evidence; (8) presentation of the results; and (9) summarizing the evidence, making conclusions and noting implications. This review was reported according to the PRISMA extension for Scoping Review guidelines (PRIMA-ScR) (Tricco et al., 2018) and was pre-registered with the Open Science Framwork. The rest of the paper follows the overall structure of the JBI approach.

Research question

The broad research question that motivated this review is What evidence supports the claim that diversity of scientists or in scientific teams enhances scientific output or impact? The narrower and operationalized hypothesis to be tested in the literature review is “diversity results in higher scientific productivity and impact.” We refer to this as the Diversity Hypothesis.

Identifying relevant studies

Eligibility criteria

The focal population were ‘scientists and scientific teams/labs,’ the concept was ‘diversity of scientists or scientific teams/labs,’ and the context was ‘scientific productivity.’ Studies were included if they were peer-reviewed, empirical studies that satisfied the population, concept, context as well as other inclusion criteria as presented in Table 1.

Table 1 PCC and inclusion and exclusion criteria

Search

A systematic search of three electronic databases (Web of Science, Scopus, and Ovid PsycInfo) for articles that met our inclusion criteria was conducted. As indicated in Table 1, the search was restricted to articles published between 1973 and the time of the search, March 2025. The year 1973 was selected as the starting point of the search as this was the year that The Strength of Weak Ties by Mark Granovetter (1973), a seminal paper often quoted in the diversity literature and in the field of social networks, was published. Subject Headings (MeSH) and Boolean operators, including AND, OR, and NOT were applied to the search. Stemming using an asterisk ‘*’ and wildcards ‘?’ were used to widen the search and NEAR operators were used to refine the search. Our search strategy was developed in collaboration with a professional research assistant with expertise in conducting systematic reviews. A full accounting of the number of articles identified in the search of the various databases is given as supporting information.

Study selection

The search results were imported into Covidence®—a web platform specializing in functionality relevant to undertaking systematic reviews. Duplicate articles were identified and removed. JB and SW independently screened titles and abstracts, and ZP resolved conflicts. JB and ZP independently screened the full text of articles, and SW resolved conflicts. The PRISMA flowchart is presented in Fig. 1.

Fig. 1
Fig. 1
Full size image

PRISMA flow diagram of study selection process

Data coding

Data extraction and coding was performed by one investigator (JB) and checked for accuracy by a second investigator (ZP). All papers included in the study were added to a spreadsheet. Each of the papers was then categorized according to year of publication, indicator of productivity or impact, country of subject of analysis, unit of analysis, type of diversity considered in the paper, and whether results were “population-adjusted,” which refers to whether results were presented relative to a reference population and/or whether they were normalized. Papers were then coded according to the diversity effect examined in the paper. Finally, each paper was evaluated for the results it reported relating diversity to scientific productivity or impact. Each of these elements, apart from year of publication, are described in further detail below.

Productivity and impact indicators

Each paper was examined for a context indicator; that is, a measure of academic productivity and/or impact. The primary of these were number of publications for productivity and paper citations and h-index for impact. Many articles also concentrated on author position (i.e., whether authors were listed first, last or somewhere in between) as an indicator of output as well as prestige. Sometimes authorship was analyzed in terms of seniority (e.g. senior author) as well. These were grouped together as First, Last or Senior authorship. Three papers (e.g., Yang et al., 2022) included an output/impact measure termed “novelty”, defined as an unusual combination of concepts or disciplines. One paper used patents as an indicator of productivity (Puia & Ofori-Dankwa, 2013) and another paper used an “Altmetric” as an indicator of impact (Wang et al., 2022).

Country of subject of analysis

Many papers in the study did not have a particular geographic focus. This was the case for articles related to diversity and scientific output or impact in general (AlShebli et al., 2018) or within a given discipline or journal (Böhme et al., 2022). Some articles examined a given discipline but chose only a few flagship or representative journals (Brown et al., 2020). Because article authors come from across the world, no specific geographic focus could be attributed to them and they were coded as “multiple countries.” Another group of articles looked at particular sets of countries (e.g., Martín-Alcázar et al. 2020, 2022). These were also coded as “multiple countries.” In other cases, articles looked at particular disciplines within a given country (Deora et al., 2024) or at research more broadly within a given country (e.g., Pilkina & Lovakov, 2022). Articles of this type were attributed to the country that was the focus of the paper and further aggregated according to region.

Unit of analysis

Papers were also categorized in terms of their unit of analysis. Although, logically, the term “diversity” implies a grouping such as a team or the co-authors of a paper, the term is not always, or even most often, used in that sense. In particular, a common approach is to categorize individuals as “diverse” owing to their demographic or identity attributes (e.g., gender/sex, ethnicity, etc.). (What constitutes diversity itself is described in the following subsection.) As a result, the papers included in this analysis have also been categorized according to their unit of analysis. The most granular unit of analysis is the individual. Papers in which the individual is the unit of analysis typically involve comparing publication or citation rates of women as compared to men (e.g., Wu, 2023). At the first level of aggregation are articles that look at research teams of teams of authors of papers (e.g., Yang et al., 2022). Above the level of teams are institutions or research centers (e.g., Hackett et al., 2021). An additional level of aggregation involves what we term super-organizational units, which includes analyses at a regional or national level (e.g., Andersson & Le, 2023; Puia & Ofori-Dankwa, 2013). Finally, there are many papers that consider more than one unit of analysis, most commonly individuals and teams (e.g., AlShebli et al., 2018).

Type of diversity

Papers were also coded according to the type of diversity that they examined. A general view of diversity that summarizes the types of diversity to which authors adhere is found in a guidance document produced by the US Military Leadership Diversity Commission (Commission, 2010). (This organization no longer appears to exist, though the guidance document can still be found online.) In particular, the document describes the definitions of diversity adopted in a number of US military and civilian organizations (e.g., US Air Force, Walt Disney Company, etc.). The one that most closely matches that used in the literature to define diversity originates from the US Air Force (the original reference to this document no longer appears in the Air Force Military News, Roll Call) and includes: demographic diversity, behavioral/cognitive diversity, structural diversity, and global diversity. As noted in the introduction, the topic of diversity is complicated by the fact that ‘diversity’ is not a single fixed concept but rather a fuzzy one that can be understood in distinct ways (Mogilski et al., 2025). First, as seen through the lens of critical theory, intersectionality and identity politics, ‘diversity’ can refer to demographic diversity, with an emphasis on where underrepresented and nominally oppressed groups are held to exist within positions of power (Mogilski et al., 2025). According to this view, an individual can be categorized as ‘diverse’ if they are assigned to, or self-identify as a member of, a minority or ‘minoritized’ social category or identity group (e.g., female, gender diverse or trans-gender, queer, nonwhite, non-Western) as opposed to a majority, "privileged" social category (e.g., male, cis-gender, heterosexual, white, Western). Second, diversity can refer to such things as diversity, within a social unit such as a team or organization, of information, knowledge or expertise (e.g., informational diversity), perspective, viewpoint or worldview (e.g., viewpoint diversity) or differences in what group members believe the group should be doing and why (e.g., value diversity). It is worth noting that “diversity” is typically not explicitly defined in the articles included in this study but only referred to obliquely and used implicitly. Almost universally, the term ‘diversity’ is used to refer to individuals who are not white, are not male, and/or who originated from a non-Western country. In some cases it is also implied that diversity means non-cis-gendered, and not sexually heteronormative (e.g., Nelson et al., 2022). We coded papers roughly consistently with the Air Force guidance. With respect to demographic diversity, papers concerned with women vs. men were coded as Sex diversity. Papers looking at race or ethnicity were coded as Race/Ethnicity. Papers considering age and sometimes career stage were coded Career Stage/Age. If the diversity in question was the disciplinary background of individuals or members, they were coded as Discipline/Field of Study diversity. There were papers that considered the impact of the representation of different universities (e.g., Saá-Pérez et al., 2015), and this was coded as Institutional diversity. Geographic and national diversity was investigated in some papers (e.g., Naik et al., 2023), and was coded as such. One paper included an analysis of lesbian, gay, bisexual, transgender, asexual, or otherwise queer individuals (LGBTQA) and this was coded as LGBTQ (Nelson et al., 2022). Many papers included an analysis of more than one type of diversity (e.g., Mishra et al., 2025). Because there were so many of these papers and because we wanted to be able to break out results by type of diversity, we did not include a multiple diversity category. As such, when evaluating diversity impacts, some papers produced multiple results (for the definition and characterization of a “result,” see Section Definition of “Results”, below). For example, a paper could present a result for productivity as a function of Sex as well productivity as a function of Race.

Categorization of results according to population-adjustment

In the literature covered in this study, papers reported results in a number of different ways, but a particularly important dimension to characterize how they were reported relates to whether they were presented relative to a reference population and/or whether they were normalized. How results were presented along these dimensions matters because if they are not reported correctly, results across diversity types cannot be conclusively compared.

Results presented without comparison to a reference population make it impossible to compare between subsets of the population. A typical example of this is the presentation of the proportion of publications for women and for men without reference to the proportion of women and men in the population from which the sample is drawn. This is observed for example in the paper by Böhme et al. (2022). In particular, they write “women are under-represented because they make up 46.6% of authorships (publications).” This claim relies on the unstated assumption that the proportion of female and male researchers (in pediatrics) is 50% for each. Given that these authors do not report on the overall population of pediatric researchers, providing an affordance for the calculation of representative, population-adjusted results, the claim of under-representation cannot be made since the observed difference could be caused by differences in publication proportions for men and women, differences in proportions of male and female researchers, or a combination of both. Other papers (e.g., Xu et al., 2020) do compare the proportion of female authors when they make their comparison of the proportion of male to female publications.

The normalization of results simply involves adjusting overall results to average results. Some papers will present overall results in levels without controlling for the overall number of researchers being considered. To illustrate, and similar to the example offered in the previous paragraph, Pratama et al. (2024) report on the total number of citations for men and women in their sample. While they write “Clear gender disparities suggest that male rectors are more productive or have access to more resources for research and publication” (Pratama et al., 2024), there is no attempt to provide average citations (or citations per researcher). Thus, although Pratama et al. report “clear gender disparities,” the veracity of this claim is unclear because it is not apparent whether the differences in numbers of citations is because men publish more than women or whether there are simply more men or a combination of both. Such claims are clearer, however, when results are presented as averages (e.g. citations per author) as is done, for example, by Lerman et al. (2022). In a related vein, some papers provide the analysis of statistical models and, in particular, model coefficients linking diversity type to scientific output and impact. We refer to providing citations per author or the reporting of diversity and outcome or impact model results as normalization.

We refer to these two types (comparisons to reference population and normalization) of reporting collectively as Population-Adjustment. The results from all the papers were coded as to whether these types of Population-Adjustments were made.

Definition of “Results”

For each of the papers included in the study, an analysis was conducted to determine the relationship reported between type of diversity (the Concept in Table 1) and output and/or impact indicators (the Context in Table 1). We defined a reported finding relating diversity to productivity or impact as “a result.” Results were evaluated at the intersection of one diversity type and one productivity or impact indicator. For example, a paper might report on citation rates of female authors relative to male authors, and this would be considered one “result.” Similarly, a paper might report on citation rates of black authors relative to white authors. This would also have been considered one “result.” As such, a single paper could present multiple results. This could happen if different effects were found for the same diversity type-productivity/impact indicator (e.g. the effect for sex diversity might be positive in one model, but negative in another model in the same paper). It could also happen if the same paper examined multiple types of diversity (e.g. women and race) or multiple productivity/impact indicators (e.g. publication rates and citation rates). This explains why, when results are presented below, there are more “results” than papers.

Results themselves were categorized based on how they relate to the hypothesis introduced above, namely “diversity results in higher scientific productivity and impact.” This starting hypothesis was based on many statements to that effect in popular, scientific, and business publications, as discussed in the introduction section. If diversity was associated with higher productivity or impact, the result was labeled “hypothesis-consistent” (Hyp.-Cons. in Tables 3, 4 and 5). If diversity was associated with lower productivity or impact, the result was labeled “hypothesis-inconsistent” (Hyp.-Inc. in Tables 3, 4 and 5). If there was no difference in productivity or impact by diversity category, the result was labeled “neutral.” It was not uncommon for papers to report that, at first, an increase in diversity (e.g. if there were one female author as opposed to none) would increase productivity or impact, but that the effect would turn negative as the amount of diversity increased (e.g. only female authors). This is typically termed an “inverted U” effect in the literature and that was the label accorded to such results in this analysis. Finally, papers could report mixed results. This could arise, for example, if different effects were found in different models of a paper.

MMAT analysis

An assessment of the quality of the research methodology for each of the papers included in this study was undertaken using the revised Mixed Methods Assessment Tool (MMAT) (Hong et al., 2018). Consistent with recent research supporting the use of AI to streamline and enhance the efficiency of data extraction in systematic literature reviews (Jardim et al., 2022; Ofori-Boateng et al., 2024), AI (Anthropic’s Claude) was used to extract information pertaining to MMAT criteria. The accuracy of the data extracted was subsequently reviewed and verified by the research assistant and SW. Criteria were assessed against each study design type, and the research assistant and SW identified whether relevant quality criteria were met or not using ‘Yes,’ ‘No’ and ‘Can’t tell’ response options. Based on this, the research team made an overall quality assessment of ‘low,’ ‘medium,’ or ‘high’ for each article.

Table of included papers

The 104 papers included in this analysis are summarized in Table 2.

Table 2 Summary of the 104 papers included in this analysis

Results

Results are presented according to the same topics and order of the topics as described in the Data Coding section above. We begin with the distribution of papers across time related to diversity and scientific output and impact.

The distribution of publication years for the 104 studies included in this analysis are presented in Fig. 2. The oldest paper identified was published in 1982. From that time through to 2008, the number of papers published that met our inclusion criteria was at most one per year with a low frequency. There was then a small surge in papers from 2009 to 2013 and then a large surge starting after 2015, with the largest number of included papers (17) coming in 2024. In the five-year period of 2020–2024, there were an average of 13 publications per year included in this analysis.

Fig. 2
Fig. 2
Full size image

Studies per year (the literature search was performed mid-March, 2025)

With respect to papers according to type of diversity reported, output/impact measure, unit of analysis, and discipline, results are presented in Fig. 3. Some papers had more than one type of diversity and/or more than one output/impact measure, so the totals in those two categories sum to more than 104. As an example, Borsuk et al. (2009) published a paper titled The Influence of Author Gender, National Language, and Number of Authors on Citation Rate in Ecology. They reported on how citation rates and first authorship are related to sex, i.e. same type of diversity (sex) and two different output/imput indicators (citation rates and first authorship). Furthermore, this same paper reported on the influence of nationality on citation rates, representing a third result from one paper that has the same output/impact measure (citations) of another result, but as related to a different type of diversity (nationality).

Fig. 3
Fig. 3
Full size image

Number of papers based on output/impact, type of diversity, unit of analysis, and discipline

The most common output/impact measure was number of publications and the most common diversity measure, by a large margin, was sex, a type of demographic diversity. Most papers included in this study had individuals as the unit of analysis. A typical and recent example is from a radiology journal (Stirrat et al., 2025) where a bibliometric analysis of authorship (the ouput/impact measure) in interventional radiology journals with respect to sex (the type of diversity) was reported with the unit of analysis as the individual.

Population-adjusted results

A total of 200 “results” (as described in the subsection “Definition and Determination of Hypothesis and “Results”) were gleaned from the104 papers included in this scoping review, which represents an average of nearly two “results” per paper (e.g. the effect of sex on number of publications and number of citations). Of those results, 142 were population-adjusted (see Section “Categorization of Results According to Population-Adjustment,” above), meaning that they were somehow normalized (i.e. number of publications per individual) or compared to the population of interest (i.e., percent of female authors compared to percent of female researchers in the field). Since meaningful quantitative results can only be gleaned from population-adjusted data, the analysis in this section will only rely on the 142 results that were normalized or compared to the population of interest. The overall summary of these 142 population-adjusted results is presented in Table 3 and Fig. 4.

Table 3 Population-adjusted results
Fig. 4
Fig. 4
Full size image

Graphical representation of the results reported in Table 3 for the population-adjusted results

There are, of course, different ways to analyze these results. Those committed to the hypothesis a priori would likely group together the hypothesis-consistent (Hyp.-Cons. in Tables 3, 4 and 5), mixed, and inverted U results into a single category of “possibly supporting hypothesis,” whereas neutral and hypothesis-inconsistent (Hyp.-Inc. in Tables 3, 4 and 5) results would be placed into a “not supporting” category. Those more skeptical of the hypothesis, would likely only consider the “hypothesis-consistent” (Hyp.-Cons. in Tables 3, 4 and 5) results to be positive evidence. These two approaches of categorizing the data are presented in Fig. 5.

Fig. 5
Fig. 5
Full size image

Representation of population adjusted results based on two approaches of aggregating the data as discussed in the text

Table 4 Results from papers where the MMAT analysis indicated high quality
Table 5 Overall summary of the results from the 104 papers included in this analysis

MMAT Analysis Results

The MMAT rubric is a systematic approach at assessing the overall quality of qualitative, quantitative, or mixed methods papers (Hong et al., 2018). The 104 papers in our analysis were categorized into three quality categories (low, medium, and high) based on their MMAT assessment and scores. The number of papers and results representing each category is presented in Fig. 6.

Fig. 6
Fig. 6
Full size image

Left: number of papers per MMAT quality. Right: number of results per MMAT quality

In a similar approach to the population-adjusted data, the data for the high quality papers were analyzed with the results presented in Table 4; Fig. 7. In this case, there were 104 total results from 54 papers.

Fig. 7
Fig. 7
Full size image

Graphical representation of the results reported in Table 4 for the high quality papers

The results presented in Table 4; Fig. 7 were further categorized in a similar way as was done for the population-adjusted results above. That analysis is presented in Fig. 8.

Fig. 8
Fig. 8
Full size image

Representation of results from the high quality papers based on two approaches of binning the data as discussed in the text

Overall summary

Table 5 provides the overall summary of the results found from the 104 papers included in this analysis as they relate to the hypothesis that diversity improves scientific output and/or impact.

Discussion

In this rapid scoping review, we sought to collate and evaluate the scientific evidence for the Diversity Hypothesis; namely, that ‘diversity’ leads to more original and impactful science. Our review revealed three key findings: (1) overwhelmingly, the focal type of diversity in our set of studies pertain to demographic or identity diversity (i.e., sex, race) of individual scientists and teams with markedly less attention given to viewpoint diversity (e.g., scholarly/scientific discipline) within teams, which precludes our ability to properly assess the strongest version of the Diversity Hypothesis in science; (2) only between 15% and 28% of results reported in these articles are unambiguously consistent with the Diversity Hypothesis, with the balance not being so—a pattern that was observed in both the sample as a whole and within the smaller sample of high quality papers; and (3) overall, the results indicate that there is little empirical evidence that diversity, as defined narrowly by demographic or identity diversity, improves scientific output and impact.

With the headline results having been reported, we can now go into a more formal discussion of the results and what can be concluded from the literature on the topic. We will discuss the results in the same order as they were presented in the results section, concentrating on the consistency or inconsistency of the results with respect to the central hypothesis. Rather than just relying on the numbers and percentages presented above, it will be instructive to discuss some representative examples of the types of results that fall into the various categories.

The most prevalent dimension of diversity was sex—a type of demographic diversity—and the most prevalent productivity indicator was number of publications. There were five papers in our study that were rated as high quality from the MMAT analysis that also had a population-adjusted result for number of publications based on sex. All five of those results were coded as being hypothesis-inconsistent (Hyp.-Inc. in Tables 3, 4 and 5). Aksnes et al. (2011), in a study of Norwegian scientists, report that while women account for 35% of Norwegian scientists, they account for only 25% of articles published. In an effort to understand network effects, Li et al. (2022) report on the publications of men and women in the STEM fields of biology, chemistry, computer science, mathematics, medicine, and physics from 1989 to 2017. They find “gendered inequalities in observed measures of both career-wise productivity and prominence among mid-career STEM researchers, in which men both publish more papers and receive more citations than women”(Li et al., 2022). J. Scott Long had one of the earlier papers in this field in 1992 when he published about measures of sex differences in scientific productivity (Long, 1992). His population was a sample of biochemists who received their Ph.D. in the 1950s and 1960s. Long found that males, on average, published more than females and that this disparity grew over the career stages of the researchers (i.e., the gap between the numbers of articles published per year between males and females grew as a function of career year). A study from a decade earlier by Ray Over (1982) discovered a similar result for a population of psychologists, finding that men averaged 15 publications 12 years after dissertation and women averaged 5. Finally, Stack and Lester (2024) developed a multiple regression model to predict productivity based on a sample of prolific suicidologists. Their results show that the effect of gender is positive and significant, indicating that “adjusting for the other predictors, males had more publications than females” (Stack & Lester, 2024).

There were only two papers in our study that found population-adjusted, hypothesis-consistent results for gender. Unlike the preceding examples where the unit of analysis was the individual, in these two papers, team was the unit of analysis. Saá-Pérez et al. (2015) reported on the effect of diversity on the performance of a Spanish university research teams. Their dependent variable was scientific performance as measured by the number of published articles per team. One of their independent variables was gender diversity, as measured by the entropy index of sex on a research team. In all of their models, the coefficient between gender diversity and performance was positive, but the coefficient was not always statistically significant. This paper had an overall quality rating of “high” based on the MMAT analysis. In the field of earth science publications, Lerback et al. (2020) reported that for authorship teams of two to four, acceptance rates to American Geophysical Union journals were 4.5% higher for teams comprising members of more than one gender. This paper had an overall quality rating of “medium” based on the MMAT analysis.

The preceding examples are designed only to provide a sense of the types of evidence encountered in the literature and how that evidence was coded in this study. In the case of publications that focus on gender diversity, there is strong evidence that, on average, men are more productive than women. There seems to be some weaker evidence that teams may benefit from gender diversity. Please note that we make no assertions about why men publish more than women publish. We are only observing that the current evidence that diversity improves scientific output is very weak when gender is the focal type of diversity and the individual is the unit of analysis.

Similar results are found for diversity of nationality, race, age, LGBT, institution, and geography, which are all examples of demographic or identity-related diversity. In all of these cases, there were far more results that were inconsistent with the Diversity Hypothesis than were consistent with it. For example, there were no studies in our data set that showed that racial diversity led to increased number of publications or publication rates. To the contrary, several papers report that researchers who were categorized as racial minorities or racially diverse teams publish less (Chander et al., 2023; Duggan et al., 2024; Hauc et al., 2024; Hesli & Lee, 2011; Joarder et al., 2024; Kang-Auger et al., 2024; Lerback et al., 2020; Mehta et al., 2022). Lerback et al. (2020), for example, report that papers from racially homogeneous teams are accepted at a 5.5% higher rate than papers from racially diverse teams in the field of Earth and space science.

The one diversity type where the results were not heavily skewed against the hypothesis was disciplinary diversity, where the results were more mixed. We identified 18 papers with results on disciplinary diversity. From these papers, there were several results indicating higher publication rates (Adachi et al., 2025; Andersson & Le, 2023; Specht & Crowston, 2022) and citation rates (AlShebli et al., 2018; Chen et al., 2021; Larivière et al., 2015; Steele & Stier, 2000) for teams with disciplinary diversity. There was only one paper showing a negative correlation between disciplinary diversity and publications (Porac et al., 2004). Several papers showed an inverted U relationship between disciplinary diversity and publications (Belkhouja et al., 2021; Caner et al., 2024; Saá-Pérez et al., 2015) or citations (Krammer & Dahlin, 2024; Specht & Crowston, 2022), which suggests that either too little or too much disciplinary diversity impairs outputs and impact. In the interesting case of the article from Specht and Crowston, a positive correlation was found between disciplinary diversity and publications and an inverted U relationship was found between disciplinary diversity and number of citations (Specht & Crowston, 2022).

Overall, the results of the analysis section indicate that there is little empirical evidence that diversity improves scientific output and/or impact. In fact, with the possible exception of disciplinary diversity—a type of informational or viewpoint diversity operationalized at the team level—the majority of the evidence seems to point in the other direction. These findings are robust regardless of how the data is parsed and analyzed. The same conclusions can be drawn looking at the full data set, only the population-adjusted results, or only the results of high quality based on the MMAT analysis.

Conclusion

This rapid scoping review sought to understand the current state of the scientific literature with respect to the Diversity Hypothesis; that is, the proposition that diversity improves scientific output and impact. Based on over 100 scientific articles, we find that only between 15% and 28% of results reported in the literature are consistent with the hypothesis, with the balance of the results not being consistent with it. Based on this, the scientific evidence implies that the Diversity Hypothesis should be rejected. To solidify these conclusions, future work should investigate whether the conclusions can be quantified through the exploration of the possibility of undertaking one or several meta-analyses on the hypothesis. That said, while we did not do an in-depth and systematic analysis of whether results were suitable for a meta-analysis, our sense is that results are not reported in a sufficiently systematic or comparable manner to enable meta-analyses. As a result of that, we think that there is room for future work that would undertake research that would test the Diversity Hypothesis and do so in a manner that would enable a meta-analytic approach in the future. Given the currents and controversies in this field, future research may benefit from adversarial collaboration, which is a method of dispute resolution where researchers with opposing views or competing theories work together to jointly design and conduct experimental research. In conclusion, despite claims that evidence for the Diversity Hypothesis is unambiguous, the results of this rapid scoping review suggest that, presently, this broad assertion is unsubstantiated and that more nuanced propositions pertaining to the types and effects of diversity be advanced and enacted in the context of scientific claims, funding, and policy-making.