Artificial intelligence in BreastScreen Norway: a retrospective analysis of a cancer-enriched sample including 1254 breast cancer cases

Koch, Henrik Wethe; Larsen, Marthe; Bartsch, Hauke; Kurz, Kathinka D.; Hofvind, Solveig

doi:10.1007/s00330-023-09461-y

Cited by 14 publications

(4 citation statements)

References 35 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…Several prior studies explored AI performance in relation to breast density, noting a relative decline of standalone AI performance as breast density increases [ 40 – 42 ]. However, another study reported consistent sensitivity for an AI system with increased breast density, while radiologists’ sensitivity decreased [ 43 ]. In our study, performance metrics of the standalone AI were superior in women with non-dense breasts compared to dense breasts.…”

Section: Discussionmentioning

confidence: 99%

Screening mammography performance according to breast density: a comparison between radiologists versus standalone intelligence detection

Kwon,

Chang,

Ham

et al. 2024

Breast Cancer Res

View full text Add to dashboard Cite

Background Artificial intelligence (AI) algorithms for the independent assessment of screening mammograms have not been well established in a large screening cohort of Asian women. We compared the performance of screening digital mammography considering breast density, between radiologists and AI standalone detection among Korean women. Methods We retrospectively included 89,855 Korean women who underwent their initial screening digital mammography from 2009 to 2020. Breast cancer within 12 months of the screening mammography was the reference standard, according to the National Cancer Registry. Lunit software was used to determine the probability of malignancy scores, with a cutoff of 10% for breast cancer detection. The AI’s performance was compared with that of the final Breast Imaging Reporting and Data System category, as recorded by breast radiologists. Breast density was classified into four categories (A–D) based on the radiologist and AI-based assessments. The performance metrics (cancer detection rate [CDR], sensitivity, specificity, positive predictive value [PPV], recall rate, and area under the receiver operating characteristic curve [AUC]) were compared across breast density categories. Results Mean participant age was 43.5 ± 8.7 years; 143 breast cancer cases were identified within 12 months. The CDRs (1.1/1000 examination) and sensitivity values showed no significant differences between radiologist and AI-based results (69.9% [95% confidence interval [CI], 61.7–77.3] vs. 67.1% [95% CI, 58.8–74.8]). However, the AI algorithm showed better specificity (93.0% [95% CI, 92.9–93.2] vs. 77.6% [95% CI, 61.7–77.9]), PPV (1.5% [95% CI, 1.2–1.9] vs. 0.5% [95% CI, 0.4–0.6]), recall rate (7.1% [95% CI, 6.9–7.2] vs. 22.5% [95% CI, 22.2–22.7]), and AUC values (0.8 [95% CI, 0.76–0.84] vs. 0.74 [95% CI, 0.7–0.78]) (all P < 0.05). Radiologist and AI-based results showed the best performance in the non-dense category; the CDR and sensitivity were higher for radiologists in the heterogeneously dense category (P = 0.059). However, the specificity, PPV, and recall rate consistently favored AI-based results across all categories, including the extremely dense category. Conclusions AI-based software showed slightly lower sensitivity, although the difference was not statistically significant. However, it outperformed radiologists in recall rate, specificity, PPV, and AUC, with disparities most prominent in extremely dense breast tissue.

show abstract

Section: Discussionmentioning

confidence: 99%

Screening mammography performance according to breast density: a comparison between radiologists versus standalone intelligence detection

Kwon,

Chang,

Ham

et al. 2024

Breast Cancer Res

View full text Add to dashboard Cite

show abstract

“…However, the classification system has faced challenges due to the significant interobserver variability among radiologists, leading to inconsistencies and uncertainties in assessments [5][6][7]. Recent advancements in artificial intelligence (AI) and deep learning (DL) have demonstrated the potential to improve diagnostic accuracy in medical imaging [8][9][10]. This study investigates the efficacy of a deep learning-enhanced computer-aided diagnosis (CAD) system in evaluating breast tissue density according to the BI-RADS density classification.…”

Section: Introductionmentioning

confidence: 99%

“…Recent advancements in artificial intelligence (AI) and deep learning (DL) have demonstrated the potential to improve diagnostic accuracy in medical imaging [ 8 , 9 , 10 ]. This study investigates the efficacy of a deep learning-enhanced computer-aided diagnosis (CAD) system in evaluating breast tissue density according to the BI-RADS density classification.…”

Section: Introductionmentioning

confidence: 99%

Enhancing Accuracy in Breast Density Assessment Using Deep Learning: A Multicentric, Multi-Reader Study

Biroš,

Kvak,

Dandár

et al. 2024

Diagnostics

View full text Add to dashboard Cite

The evaluation of mammographic breast density, a critical indicator of breast cancer risk, is traditionally performed by radiologists via visual inspection of mammography images, utilizing the Breast Imaging-Reporting and Data System (BI-RADS) breast density categories. However, this method is subject to substantial interobserver variability, leading to inconsistencies and potential inaccuracies in density assessment and subsequent risk estimations. To address this, we present a deep learning-based automatic detection algorithm (DLAD) designed for the automated evaluation of breast density. Our multicentric, multi-reader study leverages a diverse dataset of 122 full-field digital mammography studies (488 images in CC and MLO projections) sourced from three institutions. We invited two experienced radiologists to conduct a retrospective analysis, establishing a ground truth for 72 mammography studies (BI-RADS class A: 18, BI-RADS class B: 43, BI-RADS class C: 7, BI-RADS class D: 4). The efficacy of the DLAD was then compared to the performance of five independent radiologists with varying levels of experience. The DLAD showed robust performance, achieving an accuracy of 0.819 (95% CI: 0.736–0.903), along with an F1 score of 0.798 (0.594–0.905), precision of 0.806 (0.596–0.896), recall of 0.830 (0.650–0.946), and a Cohen’s Kappa (κ) of 0.708 (0.562–0.841). The algorithm achieved robust performance that matches and in four cases exceeds that of individual radiologists. The statistical analysis did not reveal a significant difference in accuracy between DLAD and the radiologists, underscoring the model’s competitive diagnostic alignment with professional radiologist assessments. These results demonstrate that the deep learning-based automatic detection algorithm can enhance the accuracy and consistency of breast density assessments, offering a reliable tool for improving breast cancer screening outcomes.

show abstract

“…Cancers 2023, 15, 3069 2 of 12 Studies have reported that AI performance is comparable to or might even outperform humans in the interpretation of screening studies [4][5][6][7][8]. In these studies, performance is often assessed in terms of general outcome metrics such as cancer detection rate and recall rate [9], but information on the prognostic features of screen-detected breast cancers is frequently not provided.…”

Section: Introductionmentioning

confidence: 99%

Comparing Prognostic Factors of Cancers Identified by Artificial Intelligence (AI) and Human Readers in Breast Cancer Screening

Oberije

Sharma

James

et al. 2023

Cancers

View full text Add to dashboard Cite

Invasiveness status, histological grade, lymph node stage, and tumour size are important prognostic factors for breast cancer survival. This evaluation aims to compare these features for cancers detected by AI and human readers using digital mammography. Women diagnosed with breast cancer between 2009 and 2019 from three UK double-reading sites were included in this retrospective cohort evaluation. Differences in prognostic features of cancers detected by AI and the first human reader (R1) were assessed using chi-square tests, with significance at p < 0.05. From 1718 screen-detected cancers (SDCs) and 293 interval cancers (ICs), AI flagged 85.9% and 31.7%, respectively. R1 detected 90.8% of SDCs and 7.2% of ICs. Of the screen-detected cancers detected by the AI, 82.5% had an invasive component, compared to 81.1% for R1 (p-0.374). For the ICs, this was 91.5% and 93.8% for AI and R1, respectively (p = 0.829). For the invasive tumours, no differences were found for histological grade, tumour size, or lymph node stage. The AI detected more ICs. In summary, no differences in prognostic factors were found comparing SDC and ICs identified by AI or human readers. These findings support a potential role for AI in the double-reading workflow.

show abstract

Artificial intelligence in BreastScreen Norway: a retrospective analysis of a cancer-enriched sample including 1254 breast cancer cases

Cited by 14 publications

References 35 publications

Screening mammography performance according to breast density: a comparison between radiologists versus standalone intelligence detection

Screening mammography performance according to breast density: a comparison between radiologists versus standalone intelligence detection

Enhancing Accuracy in Breast Density Assessment Using Deep Learning: A Multicentric, Multi-Reader Study

Comparing Prognostic Factors of Cancers Identified by Artificial Intelligence (AI) and Human Readers in Breast Cancer Screening

Contact Info

Product

Resources

About