Introduction

Breast cancer screening reduces mortality by enabling early detection and timely treatment [1]. In several countries, population-based screening programs use mammography as the primary screening modality.

Breast density—defined as the proportion of fibroglandular tissue relative to fatty tissue on a mammogram—is a key factor affecting both mammography performance and breast cancer risk. High breast density is characterized by a greater amount of fibroglandular tissue, which appears radiographically dense. Higher breast density is linked to an increased breast cancer risk [2, 3] and lowers mammographic sensitivity by masking tumors [4, 5].

This dual impact has spurred interest in incorporating breast density into personalized screening strategies [6]. Women with extremely dense breasts may benefit from shorter screening intervals or supplemental imaging like MRI to improve detection [7, 8]. Tailoring screening based on individual density profiles could enhance early detection and further reduce breast cancer mortality [9].

Breast density has traditionally been assessed visually by radiologists using BI-RADS classification. The 4th BI-RADS edition relied on percentage estimates of density, while the 5th edition uses a semi-qualitative assessment emphasizing the potential masking effect of fibroglandular tissue [10]. In both editions, however, visual evaluation is inherently subjective and prone to substantial inter-reader variability [11].

To improve consistency, automated algorithms for breast density measurement have been developed, with recent updates aligning better with BI-RADS 5th edition criteria [12, 13]. These updated tools have yet to be compared directly, and it remains unclear which provides the most accurate assessment or how their results relate to breast cancer risk. Considering masking effects, understanding differences in screen-detected and interval cancer rates across automated density categories is also important.

Using data from over 61,000 women in a Dutch prospective screening cohort, this study aims to: (1) assess the agreement of density measurements from three commercial automated algorithms (Volpara, Quantra, and iCAD), and (2) evaluate associations between these measurements and breast cancer risk, overall and separately for screen-detected and interval cancers.

Materials and methods

Ethics statement

The personalized risk-based mammography screening (PRISMA) study was approved by the local medical ethics committee (CMO Arnhem-Nijmegen; reference no. 2014/177). All participants provided informed consent for use of their data and mammograms; those without consent were excluded.

Design and population

The PRISMA study is an observational, population-based, prospective cohort study, nested within the Dutch breast cancer screening program [14]. Women aged 50–75 undergo biennial mammography; those at strongly increased risk (genetic/familial history or prior thoracic irradiation) are monitored in hospital.

In the screening program, at each screening examination, a digital mammogram—including mediolateral oblique (MLO) and craniocaudal (CC) views—is acquired for each breast and read independently by two screening radiologists. Breast density is not routinely assessed.

PRISMA participants were enrolled between 2014 and 2019 in 20 screening units covering four screening regions. The distribution across regions was 55.9% North, 19.6% South West, 14.0% East, and 10.8% South. To prevent repeat participation in PRISMA, inclusion time per unit was limited to two years.

Figure 1 shows the inclusion flowchart. Of 78,463 participants, we excluded those with missing unprocessed mammograms (n = 12,260), breast implants (n = 860), missing density output (n = 3416), and screen-detected cancer at entry (n = 413). Screen-detected cancers were defined as those diagnosed within 12 months of recall, and interval cancers within 30 months after a negative screen; others were non-screen-detected.

Fig. 1
Fig. 1
Full size image

Flowchart showing the 61,514 PRISMA study participants who were included to compare the agreement of automated density measurements and their associations with breast cancer risk

Outcomes

The primary outcome of this study is the occurrence of breast cancer, including both in situ and invasive breast cancer. As a secondary outcome, we studied the occurrence of invasive breast cancer only. Cancer cases were identified through linkage with the Netherlands Cancer Registry, with follow-up through December 2022. The completeness of pathology-confirmed cancer diagnoses in this registry is estimated to be between 95% and 98.7% [15].

Breast density algorithms

Breast density was assessed using the screening mammograms, both the CC and MLO views from both breasts, obtained at study entry. All mammograms were unprocessed (‘for processing’) full-field digital mammograms acquired on Hologic systems (Selenia or Selenia Dimensions). Breast density was measured using three automated algorithms: Volpara, Quantra, and iCAD.

Volpara

Volpara version 1.5.5.4 (Volpara Health) was used. It is a physics-based algorithm that derives tissue volumes from pixel intensities, imaging parameters, breast compression data, and expected attenuation coefficients [16]. For each mammographic view and breast, the algorithm estimates the fibroglandular volume (cm³) and total breast volume (cm³), from which volumetric breast density (VBD) % is calculated. Based on the average VBD% of both breasts, Volpara Density Grades 4th edition (VDG4) are calculated as A (< 4.5%), B (4.5%–7.5%), C (7.5%–15.5%), or D (> 15.5%). The Volpara Density Grade 5th edition (VDG5) uses the maximum VBD % of either breast and a lower threshold for category A (< 3.5% instead of < 4.5%). VDG4 and VDG5 correspond to BI-RADS 4th and 5th edition density categories, respectively.

Quantra

Quantra version 2.2 (Hologic Inc.) applies a machine learning–based method that leverages pattern and texture features from unprocessed digital mammograms. The model was trained on a dataset annotated by expert radiologists according to BI-RADS 5th edition criteria [12]. Quantra provides a BI-RADS–like breast density category for each breast and for the overall examination, corresponding to the BI-RADS 5th edition classification.

iCAD

The iCAD density assessment tool (version 2.1.0.0.64, M-Vu CAD; iCAD) estimates breast density from unprocessed mammograms by computing the total breast area (cm²) and fibroglandular tissue area (cm²). The ratio of these areas yields the breast density percentage for each breast. The software also assigns BI-RADS–like density categories (A–D) consistent with the BI-RADS 5th edition.

Statistical methods

Agreement between methods

Because not all algorithms yield continuous breast density measures, agreement was primarily assessed using categorical outputs. For each pair of the three algorithms, agreement was quantified by absolute agreement (i.e., the proportion of mammograms categorized identically) and weighted kappa statistics with 95% confidence intervals (CIs), using quadratic weighting. Kappa values were interpreted using Altman’s modification of the Landis and Koch scale: poor (≤ 0.20), fair (0.21–0.40), moderate (0.41–0.60), good (0.61–0.80), and very good (> 0.80) [17, 18]. The weighted kappa accounts for the magnitude of disagreement when categories assigned by two algorithms differ by more than one level. Bangdiwala’s agreement charts were used for visual assessment of concordance between methods [19]. Volpara and iCAD also produce continuous density measures (VBD%, maximum VBD%, and percent dense tissue), but these are defined differently and not directly comparable. Therefore, instead of agreement statistics, we evaluated correlations using Spearman’s coefficient.

Associations with breast cancer risk and screen-detected and interval cancers

Associations between breast density and five-year breast cancer risk were evaluated using both categorical and continuous density measures. Kaplan–Meier curves were constructed to visualize cumulative breast cancer incidence across density categories for each algorithm. To understand how mammographic sensitivity would differ across density categories defined by the different algorithms, crude incidence rates of screen-detected and interval cancers were calculated separately across density categories.

Univariable Cox proportional hazards models were used to estimate hazard ratios (HRs) with 95% CIs. The proportional hazards assumption was verified via visual inspection of scaled Schoenfeld residuals and found not to be violated. Continuous density measures (VBD%, Maximum VBD% and percent dense tissue) were log-transformed and standardized before inclusion in the model. For categorical density measurements, category B served as the reference.

Density categories A and D are of particular interest for personalized screening, as they may trigger adapted screening recommendations. To assess whether disagreement between algorithms influences risk estimates, HRs for categories A and D (vs B) for each algorithm were re-estimated, restricting category A or D to women whose density was classified differently by at least one other algorithm (that is, higher than A or lower than D, respectively, regardless of the number of categories that they were discordant. All B-category women were retained as the reference group in both analyses.

To compare the algorithms with each other, the discriminatory performance of each density metric, both the categorical and continuous ones, was also summarized by calculating the 5-year area under the receiver operating characteristic curve (AUC) for breast cancer. Finally, to understand the ability of the different density metrics to discriminate between cancers that are screen-detected and those that are not, e.g., due to masking, we calculated a 5-year AUC for screen-detected vs interval cancers.

All analyses in this study are complete case analyses and performed in R version 4.4.3.

Results

Baseline characteristics and density distribution

A total of 61,514 mammographic exams from an equal number of PRISMA participants were included in the analysis (Fig. 1). The median age at the time of screening was 60.1 years (interquartile range (IQR): 54.5–66.6).

Figure 2 presents the distribution of breast density categories according to the four classification methods: VDG4, VDG5, Quantra, and iCAD. For the highest density category (category D), the number and percentage of exams classified as such were 5177 (8.4%) with VDG5, 4418 (7.2%) with VDG4, 4024 (6.5%) with Quantra, and 1916 (3.1%) with iCAD. The number and percentage of exams classified in the lowest density category (category A) were 20,400 (33.2%) with VDG4, 14,451 (23.5%) with iCAD, 8834 (14.4%) with Quantra, and 6594 (10.7%) with VDG5. The median continuous % breast density measures were 5.6% (IQR: 4.0%–8.8%) for VBD% (Volpara), 6.0% (IQR: 4.2%–9.4%) for Maximum VBD% (Volpara), and 9.5% (IQR: 3.0%–26.0%) for percent dense tissue (iCAD).

Fig. 2
Fig. 2
Full size image

Distributions of automatically measured breast density categories using three algorithms in the PRISMA cohort (N = 61,514)

Agreement between algorithms

Tables 1 and 2 quantify the agreement between the categorical breast density measures, with visual representation provided in Fig. 3. Although disagreements between methods were observed in all directions, disagreements spanning more than one category were infrequent. Weighted kappa values ranged from 0.725 to 0.826, indicating good agreement across all comparisons.

Fig. 3
Fig. 3
Full size image

Agreement charts to compare breast density categorization on mammograms for three different algorithms in the PRISMA cohort. Black areas indicate agreement, gray areas and their orientation indicate disagreement and its direction. Weighted Fleiss kappa represents overall agreement

Table 1 Cross-tabulation to compare breast density categorization on mammograms for three different algorithms in the PRISMA cohort
Table 2 Absolute agreement and weighted kappa statistics for agreement between automated breast density categories measured with three different algorithms in the PRISMA cohort

Comparing measures from different vendors, the strongest agreement was observed between VDG5 and Quantra, both calibrated to BI-RADS 5th edition. The absolute agreement between these algorithms was 75.0%, with a weighted kappa of 0.795 (95% CI: 0.791–0.797). In cases classified as category D by VDG5, a substantial proportion (n = 1941) were downgraded to category C by Quantra, while fewer (n = 779) were upgraded from C to D. At the lower end of the spectrum, more exams were classified as category B by VDG5 and as category A by Quantra (n = 3893) than the reverse (n = 1632).

iCAD showed the strongest agreement with VDG4, with an absolute agreement of 66.3% and a weighted kappa of 0.781 (95% CI: 0.778–0.784). Among the 4705 exams classified as category D by either algorithm, most discrepancies involved downgrading by iCAD to category C (n = 2801), with only 331 exams upgraded from C to D by iCAD. Similarly, more exams were classified as category B by iCAD while VDG4 classified them as category A (n = 7881), compared to the reverse situation (n = 1888).

Spearman correlations between continuous breast density measures were strong across vendors: 0.903 [95% CI: 0.901-0.904] between VBD % and percent dense tissue, and 0.898 [95% CI: 0.896-0.900] between Maximum VBD % and percent dense tissue.

Associations with breast cancer risk, screen-detected, and interval cancers

During a median follow-up of 4.2 years (IQR: 3.9–4.3), a total of 777 breast cancer cases were identified, of which 762 occurred within the first five years. Of those, 419 were screen-detected cancers, 317 were interval cancer and 26 were non-screen-detected. Figure 4 displays cumulative breast cancer incidence curves stratified by density category for each of the categorical breast density algorithms. The cumulative absolute number of breast cancer cases per density category for each algorithm is displayed in the risk tables in Fig. 4. It is important to note that the number of women in each density category varies depending on the classification method (Fig. 2). Consequently, the absolute number of breast cancer cases in the highest density category (category D) differed across algorithms, ranging from 37 cases among 1916 participants classified as D by iCAD, to 89 cases among 5177 participants classified as D by VDG5. Similarly, in the lowest density category A, the absolute number of cases ranged from 55 among 6594 participants classified as A by VDG5 to 196 among 20,793 participants classified as A by VDG4.

Fig. 4
Fig. 4
Full size image

Stratified cumulative breast cancer incidence (proportion) curve by breast density category for three different automated breast density algorithms. Risk tables display the cumulative number of events across breast density levels

Despite these differences in category size, the cumulative incidence curves revealed similar event rates across all methods (Fig. 4). Table 3 shows that breast cancer risk consistently increased with higher breast density across all categorical measures. The three continuous measures also demonstrated similar associations with breast cancer risk: the HR per one-standard-deviation increase in log-transformed density % was 1.22 (95% CI: 1.14–1.30) for VBD %, 1.22 (95% CI: 1.14–1.31) for maximum VBD %, and 1.29 (95% CI: 1.19–1.38) for percent dense tissue.

Table 3 Associations of breast density measures with breast cancer risk and five-year area under the receiver operating curves to discriminate between participants with and without breast cancer, and to discriminate between screen-detected and interval cancers in the PRISMA cohort

The HRs presented in Table 4 show that, for participants with breast density category A or D according to a given algorithm, associations with breast cancer risk were not different in case of disagreement with one or more of the other algorithms (i.e., classified as discordant).

Table 4 Associations of categorical breast density measures A and D with breast cancer, for all participants categorized as A or D by an algorithm, or considering only discordant measurements

Figure 5 shows crude incidence rates per 1000 women-years for screen-detected and interval cancers separately. For all algorithms, the interval cancer rates increase relative to the screen-detected cancer rates as density increases. These increases are of similar magnitude across the three algorithms.

Fig. 5
Fig. 5
Full size image

Crude screen-detected and interval cancer rates per 1000 women-years by density categories for three different breast density algorithms in the PRISMA cohort

Regarding discriminatory performance, the five-year area under the curve (AUC) estimates to distinguish between women with and without breast cancer were similar among categorical measures: 0.53 (95% CI: 0.52–0.55) for VDG4, 0.54 (95% CI: 0.52–0.55) for VDG5, 0.55 (95% CI: 0.53–0.57) for Quantra, and 0.54 (95% CI: 0.52–0.56) for iCAD. The continuous density measures showed slightly higher discrimination: 0.57 (95% CI: 0.55–0.59) for VBD %, 0.56 (95% CI: 0.54–0.58) for Maximum VBD %, and 0.56 (95% CI: 0.54–0.58) for percent dense tissue. Table 3 also shows AUC estimates to distinguish between women with screen-detected and interval cancer. For categorical metrics, the AUCs range from 0.57 (VDG4) to 0.60 (Quantra). Again, slightly higher AUCs are observed (up to 0.62) for the continuous density measures.

Of the 762 breast cancer cases in the first 5 years of follow-up, 657 were invasive. Table S1 shows that associations with invasive breast cancer risk for all density measures are similar to the associations with overall breast cancer risk (Table 3).

Discussion

Summary of findings

In this large, population-based cohort of over 61,000 women participating in the Dutch breast cancer screening program, we evaluated breast density using three automated algorithms. We compared their agreement, and association with breast cancer risk, as well as screen-detected and interval cancer rates. The algorithms demonstrated good agreement, with weighted kappa values ranging from 0.725 to 0.826 between categorical measures. Across vendors, the strongest agreement was observed between VDG5 and Quantra.

Breast cancer incidence rose consistently with increasing density across all algorithms, and all three continuous measures were significantly associated with five-year risk. Discrimination was modest for every measure, with only a small AUC improvement when using continuous instead of categorical density (from ~0.53–0.55 to ~0.56–0.57). Screen-detected vs interval cancer patterns were likewise similar across density algorithms.

Although overall agreement between algorithms was strong, they assigned different proportions of women to categories A and D. In personalized screening, this would influence both the size of risk strata and the number of cancers within them. Because cumulative risks were comparable, variation in absolute case counts mainly reflects differences in category distribution rather than underlying risk.

Literature comparison

To the best of our knowledge, this is the largest study to date comparing agreement between multiple automated breast density algorithms on screening mammograms. While several studies have evaluated breast density algorithms, direct comparison with our findings is limited, primarily because only the version of Volpara used in this study has been evaluated in previous research. Comparative studies including iCAD’s density measurement are currently lacking [20].

Prior research has largely focused on comparisons between Volpara and earlier versions of Quantra, which provided continuous volumetric density outputs [13, 21,22,23,24,25]. A recent meta-analysis reported strong agreement between the volumetric breast density estimates of these two algorithms, with a pooled correlation coefficient of 0.839 [20]. Other automated or semi-automated density measurement tools, such as Cumulus and LIBRA, have also shown strong agreement with each other and similar associations with breast cancer risk in a previous study [26].

Although Quantra version 2.2 adopts a different approach—using a multi-class classification model instead of volumetric estimation—our results show similarly strong agreement between its categorical output and that of Volpara and thus align with previous findings that used other versions.

Regarding breast cancer risk, the increasing HRs across breast density categories observed here align with earlier findings, reinforcing breast density as an important risk factor. Its impact on mammographic sensitivity, likely via masking, is confirmed by our results, which agree with studies showing increasing interval cancer rates as density increases [4, 5, 27, 28].

Clinical implications

The strong agreement between the three automated breast density algorithms, along with similar associations with breast cancer risk and comparable screen-detected and interval cancer rates, suggests no single algorithm clearly outperforms the others in this setting. This is reassuring, implying different tools may be used interchangeably to support breast density assessment, for example, in risk prediction for tailored screening. Moreover, our findings suggest that treating breast density as a continuous rather than a categorical measure may provide incremental value. The slightly higher discriminatory performance of continuous measures, reflected in AUCs, supports their potential utility in risk models.

Despite strong agreement and similar risk associations, notable differences in the proportion of women classified into the lowest or highest density categories (“A” and “D”) must be carefully considered, especially for personalized screening. Since category D may require the decision to intensify screening (e.g., supplemental imaging, shorter screening interval), variability in classification could lead to substantial differences in the number of women selected for additional interventions, depending on the algorithm, which influences the required capacity. Similarly, if de-intensified screening (e.g., extended intervals) is considered for category A, the choice of algorithm affects the number of women eligible. This also influences the absolute number of cases in these groups, affecting the potential benefit of different screening strategies.

Limitations

Several limitations should be acknowledged. First, although various commercial breast density algorithms exist, only three were included. Still, these represent distinct technical approaches covering most methodologies used in other tools [20]. One approach not represented is deep learning–based density classification. For example, Eun Lee et al compared an AI-based and physics-based model to radiologists’ assessments in 488 Asian women, finding similar agreement of both models and radiologists [29].

Second, all mammograms were acquired using a single vendor (Hologic), limiting generalizability across mammography systems. We could not assess if agreement varies by system. However, prior work found system choice did not significantly impact density estimates [30].

Another important limitation, not unique here, is the absence of a gold standard for breast density measurement, preventing definitive conclusions on any algorithm’s validity. Some studies compare volumetric mammographic density to MRI, especially for volumetric algorithms, while others focus on predicting BI-RADS categories [31]. The assessment of an algorithm’s validity may depend on the chosen target metric. Agreement between automated algorithms and radiologists’ BI-RADS is generally moderate, likely reflecting inter-reader variability. Interestingly, agreement between automated algorithms tends to be higher due to their inherent consistency [20].

In the absence of a gold standard, we assessed associations between each density metric and breast cancer incidence, and compared screen-detected and interval cancer rates. Although longer follow-up would strengthen conclusions, these outcomes are clinically meaningful, especially considering the use of breast density to guide personalized screening strategies.

Finally, the PRISMA cohort may not be fully representative of the entire Dutch screening population. However, it covers multiple regions with good geographic representation, and the age distribution closely reflects the target screening population [32].

Conclusion

In this large population-based screening cohort, three automated breast density algorithms showed good agreement and comparable associations with breast cancer risk, as well as with screen-detected and interval cancer rates, suggesting they may be used interchangeably to support density assessment. Differences in the proportion of women classified as entirely fatty or extremely dense, however, should be considered when implementing personalized screening as they directly affect the number of women eligible for de-intensified or supplemental screening and the capacity required to deliver it.