Background Emerging evidence suggests that artificial intelligence (AI) can increase cancer detection in mammography screening while reducing screen-reading workload, but further understanding of the clinical impact is needed. Methods In this randomised, controlled, parallel-group, non-inferiority, single-blinded, screening-accuracy within the Swedish national screening programme, women recruited at four screening sites in southwest (Malm & ouml;, Lund, Landskrona, and Trelleborg) who were eligible for mammography screening were randomly (1:1) to AI-supported screening or standard double reading. The AI system (Transpara version 1.7.0 ScreenPoint Nijmegen, Netherlands) was used to triage screening examinations to single or double reading and as detection highlighting suspicious findings. This is a protocol-defined analysis of the secondary outcome measures cancer detection, false-positive rates, positive predictive value of recall, type and stage of cancer detected, reading workload. This trial is registered at ClinicalTrials.gov, NCT04838756 and is closed to accrual. Findings Between April 12, 2021, and Dec 7, 2022, 105 934 women were randomly assigned to the intervention control group. 19 women were excluded from the analysis. The median age was 537 years (IQR 465-632). supported screening among 53 043 participants resulted in 338 detected cancers and 1110 recalls. Standard among 52 872 participants resulted in 262 detected cancers and 1027 recalls. Cancer-detection rates were 64 (95% CI 57-71) screened participants in the intervention group and 50 per 1000 (44-56) in the control ratio of 129 (95% CI 109-151; p=00021). AI-supported screening resulted in an increased detection cancers (270 vs 217, a proportion ratio of 124 [95% CI 104-148]), wich were mainly small lymph-node cancers (58 more T1, 46 more lymph-node negative, and 21 more non-luminal A). AI-supported screening resulted in an increased detection of in situ cancers (68 vs 45, a proportion ratio of 151 [103-219]), with of the increased detection being high-grade in situ cancer (12 more nuclear grade III, and no increase grade I). The recall and false-positive rate were not significantly higher in the intervention group (a ratio [95% CI 099-117; p=0084] and 101 [091-111; p=092], respectively). The positive predictive value of significantly higher in the intervention group compared with the control group, with a ratio of 119 (95% CI p=0012). There were 61 248 screen readings in the intervention group and 109 692 in the control group, a 442% reduction in the screen-reading workload. Interpretation The findings suggest that AI contributes to the early detection of clinically relevant breast reduces screen-reading workload without increasing false positives.
BACKGROUND:In the Malmö Breast Tomosynthesis Screening Trial (MBTST), digital breast tomosynthesis (DBT) increased the detection rate of breast cancer compared to currently used digital mammography (DM). This study evaluated the cost effectiveness of DBT versus DM in Sweden using data from MBTST. MATERIALS AND METHOD:A health economic model was used to compare a breast cancer screening programme using DBT to breast cancer screening using DM in women aged 40-74. Model input values on demographics, breast cancer incidence and mortality were derived from regional registries. The effect of DBT on detection rates and interval breast cancer was based on MBTST. Key information is not yet available and scenario analyses were used to explore the reduction in relative breast cancer mortality required to meet willingness-to-pay thresholds between €10,000 and €50,000/quality-adjusted life year (QALY), assuming three levels of DBT examination cost. RESULTS:Compared with DM, DBT-screening was estimated to detect 56 more invasive cancers and 18 fewer interval cancers per 100,000 invited women. The relative reduction in breast cancer mortality for DBT to be cost effective varied between 2.3 and 14.2% for the lowest versus highest evaluated cost and willingness-to-pay scenarios. The required reduction at a DBT cost and willingness-to-pay threshold of €61 and €50,000/QALY respectively was 3.6%. CONCLUSION:This model analysis shows that the cost effectiveness of DBT-screening is dependent on the expected relative reduction in breast cancer mortality and the willingness-to-pay for screening. Furthermore, the results are driven by the future costs for DBT-screening examinations.
Limited understanding exists regarding non-detected cancers in digital breast tomosynthesis (DBT) screening. This study aims to classify non-detected cancers into true or false negatives, compare them with true positives, and analyze reasons for non-detection. Conducted between 2010 and 2015, the prospective single-center Malmö Breast Tomosynthesis Screening Trial (MBTST) compared one-view DBT and two-view digital mammography (DM). Cancers not detected by DBT, i.e., interval cancers, those detected in the next screening round, and those only identified by DM, underwent a retrospective informed review by in total four breast radiologists. Reviewers classified cancers into true negative, false negative, or non-visible based on both DBT and DM findings and assessed radiographic appearances at screening and diagnosis, breast density, and reasons for non-detection. Statistics included the Pearson X2 test. In total, 89 cancers were not detected with DBT in the MBTST; eight cancers were solely in the DM reading mode, 59 during subsequent DM screening rounds, and 22 interval cancers. The proportion of cancers classified as false negative was 25
Background: Analysis of how digital breast tomosynthesis (DBT) screening affects consecutive screening performance is important to estimate its future value in screening. Purpose: To evaluate whether DBT contributes to early detection of breast cancer by assessing cancer detection rates (CDRs), including the fraction of invasive cancers and cancer subtypes in consecutive routine digital mammography (DM). Materials and Methods: The paired prospective Malm & ouml; Breast Tomosynthesis Screening Trial (MBTST) was performed between January 2010 and February 2015. Participating women underwent one-view DBT and two-view DM at one screening occasion. In this secondary analysis, women were followed up through their first (DM1) and second (DM2) consecutive two-view DM screening rounds after MBTST participation. Cancer diagnoses were identified by referencing records. A logistic regression model, adjusted for age, was used to calculate the odds of luminal A-like cancers with use of the MBTST as reference. Results: There were 14 848 final participants in the MBTST (median age, 57 years [IQR, 49-65 years]). Of those, 12 876 women were screened in DM1 (median age, 58 years [IQR, 50-66 years]) and 10 883 were screened in DM2 (median age, 59 years [IQR, 51-67 years]). Compared with CDRs in the trial of 6.5 of 1000 women (95% CI: 5.2, 7.9) for DM and 8.7 of 1000 women (95% CI: 7.3, 10.3) for DBT, the CDR was lower in DM1 (4.6 of 1000 women [95% CI: 3.6, 5.9]) and DM2 (5.3 of 1000 women [95% CI: 4.1, 6.9]). The proportion of invasive cancers was 84.9% (118 of 139 cancers) in the MBTST; the corresponding numbers were 66% (39 of 59 cancers) for DM1 and 83% (50 of 60 cancers) for DM2. The odds of luminal A-like cancers were lower in DM1 at 0.28 (95% CI: 0.12, 0.66 [P P = .004]) but not in DM2 at 0.80 (95% CI: 0.40, 1.58 [P P = .52]) versus screening in the MBTST. Conclusion: CDR and the fraction of invasive cancers were lower in DM1 and then increased in DM2 following the MBTST, indicating earlier cancer detection mainly due to increased detection of luminal A-like cancers in DBT screening.
Background Retrospective studies have shown promising results using artificial intelligence (AI) to improve mammography screening accuracy and reduce screen-reading workload; however, to our knowledge, a randomised trial has not yet been conducted. We aimed to assess the clinical safety of an AI-supported screen-reading protocol compared with standard screen reading by radiologists following mammography. Methods In this randomised, controlled, population-based trial, women aged 40-80 years eligible for mammography screening (including general screening with 1 & BULL;5-2-year intervals and annual screening for those with moderate hereditary risk of breast cancer or a history of breast cancer) at four screening sites in Sweden were informed about the study as part of the screening invitation. Those who did not opt out were randomly allocated (1:1) to AI-supported screening (intervention group) or standard double reading without AI (control group). Screening examinations were automatically randomised by the Picture Archive and Communications System with a pseudo-random number generator after image acquisition. The participants and the radiographers acquiring the screening examinations, but not the radiologists reading the screening examinations, were masked to study group allocation. The AI system (Transpara version 1.7.0) provided an examination-based malignancy risk score on a 10-level scale that was used to triage screening examinations to single reading (score 1-9) or double reading (score 10), with AI risk scores (for all examinations) and computer-aided detection marks (for examinations with risk score 8-10) available to the radiologists doing the screen reading. Here we report the prespecified clinical safety analysis, to be done after 80 000 women were enrolled, to assess the secondary outcome measures of early screening performance (cancer detection rate, recall rate, false positive rate, positive predictive value [PPV] of recall, and type of cancer detected [invasive or in situ]) and screen-reading workload. Analyses were done in the modified intention-to-treat population (ie, all women randomly assigned to a group with one complete screening examination, excluding women recalled due to enlarged lymph nodes diagnosed with lymphoma). The lowest acceptable limit for safety in the intervention group was a cancer detection rate of more than 3 per 1000 participants screened. The trial is registered with ClinicalTrials.gov, NCT04838756, and is closed to accrual; follow-up is ongoing to assess the primary endpoint of the trial, interval cancer rate. Findings Between April 12, 2021, and July 28, 2022, 80 033 women were randomly assigned to AI-supported screening (n=40 003) or double reading without AI (n=40 030). 13 women were excluded from the analysis. The median age was 54 & BULL;0 years (IQR 46 & BULL;7-63 & BULL;9). Race and ethnicity data were not collected. AI-supported screening among 39 996 participants resulted in 244 screen-detected cancers, 861 recalls, and a total of 46 345 screen readings. Standard screening among 40 024 participants resulted in 203 screen-detected cancers, 817 recalls, and a total of 83 231 screen readings. Cancer detection rates were 6 & BULL;1 (95% CI 5 & BULL;4-6 & BULL;9) per 1000 screened participants in the intervention group, above the lowest acceptable limit for safety, and 5 & BULL;1 (4 & BULL;4-5 & BULL;8) per 1000 in the control group-a ratio of 1 & BULL;2 (95% CI 1 & BULL;0-1 & BULL;5; p=0 & BULL;052). Recall rates were 2 & BULL;2% (95% CI 2 & BULL;0-2 & BULL;3) in the intervention group and 2 & BULL;0% (1 & BULL;9-2 & BULL;2) in the control group. The false positive rate was 1 & BULL;5% (95% CI 1 & BULL;4-1 & BULL;7) in both groups. The PPV of recall was 28 & BULL;3% (95% CI 25 & BULL;3-31 & BULL;5) in the intervention group and 24 & BULL;8% (21 & BULL;9-28 & BULL;0) in the control group. In the intervention group, 184 (75%) of 244 cancers detected were invasive and 60 (25%) were in situ; in the control group, 165 (81%) of 203 cancers were invasive and 38 (19%) were in situ. The screen-reading workload was reduced by 44 & BULL;3% using AI. Interpretation AI-supported mammography screening resulted in a similar cancer detection rate compared with standard double reading, with a substantially lower screen-reading workload, indicating that the use of AI in mammography screening is safe. The trial was thus not halted and the primary endpoint of interval cancer rate will be assessed in 100 000 enrolled participants after 2-years of follow up.
To evaluate the total number of false-positive recalls, including radiographic appearances and false-positive biopsies, in the Malmö Breast Tomosynthesis Screening Trial (MBTST). The prospective, population-based MBTST, with 14,848 participating women, was designed to compare one-view digital breast tomosynthesis (DBT) to two-view digital mammography (DM) in breast cancer screening. False-positive recall rates, radiographic appearances, and biopsy rates were analyzed. Comparisons were made between DBT, DM, and DBT + DM, both in total and in trial year 1 compared to trial years 2 to 5, with numbers, percentages, and 95
The encouraging results of modern breast cancer care builds on tremendous improvements in diagnostics and therapy during the 20th century. Scandinavian countries have made important footprints in the development of breast diagnostics regarding technical development of imaging, cell and tissue sampling methods and, not least, population screening with mammography. The multimodality approach in combination with multidisciplinary clinical work in breast cancer serve as a role model for the management of many cancer types worldwide. The development of breast radiology is well represented in the research published in this journal and this historical review will describe the most important steps.
Background: Interval cancer rates can be used to evaluate whether screening with digital breast tomosynthesis (DBT) contributes to a screening benefit. Purpose: To compare interval cancer rates and tumor characteristics in DBT screening to those in a contemporary population screened with digital mammography (DM). Materials and Methods: The prospective population-based Malmo Breast Tomosynthesis Screening Trial (MBTST) was designed to compare one-view DBT to two-view DM in breast cancer detection. The interval cancer rates and cancer characteristics in the MBTST were compared with an age-matched contemporary control group, screened with two-view DM at the same center. Conditional logistic regression was used for data analysis. Results: There were 14 848 women who were screened with DBT and DM in the MBTST between January 2010 and February 2015. The trial women were matched with two women of the same age and screening occasion at DM screening during the same period. Matches for 13 369 trial women (mean age, 56 years. 10 [standard deviation]) were found with 26 738 women in the control group (mean age, 56 years. 10). The interval cancer rate in the MBTST was 1.6 per 1000 screened women (21 of 13 369; 95% CI: 1.0, 2.4) compared with 2.8 per 1000 screened women in the control group (76 of 26 738 [95% CI: 2.2, 3.6]; conditional odds ratio, 0.6 [95% CI: 0.3, 0.9]; P =.02). The invasive interval cancers in the MBTST and in the control group showed in general high Ki-67 (63% [12 of 19] and 75% [54 of 72]), and low proportions of luminal A-like subtype (26% [five of 19] and 17% [12 of 72]), respectively. Conclusion: The reduced interval cancer rate after screening with digital breast tomosynthesis compared with a contemporary age-matched control group screened with digital mammography might translate into screening benefits. Interval cancers in the trial generally had nonfavorable characteristics. (C) RSNA, 2021
Purpose To investigate how an artificial intelligence (AI) system performs at digital mammography (DM) from a screening population with ground truth defined by digital breast tomosynthesis (DBT), and whether AI could detect breast cancers at DM that had originally only been detected at DBT. Materials and Methods In this secondary analysis of data from a prospective study, DM examinations from 14 768 women (mean age, 57 years), examined with both DM and DBT with independent double reading in the Malmӧ Breast Tomosynthesis Screening Trial (MBTST) (ClinicalTrials.gov: NCT01091545; data collection, 2010-2015), were analyzed with an AI system. Of 136 screening-detected cancers, 95 cancers were detected at DM and 41 cancers were detected only at DBT. The system identifies suspicious areas in the image, scored 1-100, and provides a risk score of 1 to 10 for the whole examination. A cancer was defined as AI detected if the cancer lesion was correctly localized and scored at least 62 (threshold determined by the AI system developers), therefore resulting in the highest examination risk score of 10. Data were analyzed with descriptive statistics, and detection performance was analyzed with receiver operating characteristics. Results The highest examination risk score was assigned to 10% (1493 of 14 786) of the examinations. With 90.8% specificity, the AI system detected 75% (71 of 95) of the DM-detected cancers and 44% (18 of 41) of cancers at DM that had originally been detected only at DBT. The majority were invasive cancers (17 of 18). Conclusion Almost half of the additional DBT-only screening-detected cancers in the MBTST were detected at DM with AI. AI did not reach double reading performance; however, if combined with double reading, AI has the potential to achieve a substantial portion of the benefit of DBT screening.Keywords: Computer-aided Diagnosis, Mammography, Breast, Diagnosis, Classification, Application DomainClinical trial registration no. NCT01091545© RSNA, 2021.
In this study we used a large previously built database of 2,892 mammograms and 31,650 single mammogram radiologists' assessments to simulate the impact of replacing one radiologist by an AI system in a double reading setting. The double human reading scenario and the double hybrid reading scenario (second reader replaced by an AI system) were simulated via bootstrapping using different combinations of mammograms and radiologists from the database. The main outcomes of each scenario were sensitivity, specificity and workload (number of necessary readings). The results showed that when using AI as a second reader, workload can be reduced by 44%, sensitivity remains similar (difference - 0.1%; 95% CI = -4.1%, 3.9%), and specificity increases by 5.3% (P<0.001). Our results suggest that using AI as a second reader in a double reading setting as in screening programs could be a strategy to reduce workload and false positive recalls without affecting sensitivity.
BACKGROUND:Artificial intelligence (AI) systems performing at radiologist-like levels in the evaluation of digital mammography (DM) would improve breast cancer screening accuracy and efficiency. We aimed to compare the stand-alone performance of an AI system to that of radiologists in detecting breast cancer in DM.METHODS:Nine multi-reader, multi-case study datasets previously used for different research purposes in seven countries were collected. Each dataset consisted of DM exams acquired with systems from four different vendors, multiple radiologists' assessments per exam, and ground truth verified by histopathological analysis or follow-up, yielding a total of 2652 exams (653 malignant) and interpretations by 101 radiologists (28 296 independent interpretations). An AI system analyzed these exams yielding a level of suspicion of cancer present between 1 and 10. The detection performance between the radiologists and the AI system was compared using a noninferiority null hypothesis at a margin of 0.05.RESULTS:The performance of the AI system was statistically noninferior to that of the average of the 101 radiologists. The AI system had a 0.840 (95% confidence interval [CI] = 0.820 to 0.860) area under the ROC curve and the average of the radiologists was 0.814 (95% CI = 0.787 to 0.841) (difference 95% CI = -0.003 to 0.055). The AI system had an AUC higher than 61.4% of the radiologists.CONCLUSIONS:The evaluated AI system achieved a cancer detection accuracy comparable to an average breast radiologist in this retrospective setting. Although promising, the performance and impact of such a system in a screening setting needs further investigation.
To study the feasibility of automatically identifying normal digital mammography (DM) exams with artificial intelligence (AI) to reduce the breast cancer screening reading workload. A total of 2652 DM exams (653 cancer) and interpretations by 101 radiologists were gathered from nine previously performed multi-reader multi-case receiver operating characteristic (MRMC ROC) studies. An AI system was used to obtain a score between 1 and 10 for each exam, representing the likelihood of cancer present. Using all AI scores between 1 and 9 as possible thresholds, the exams were divided into groups of low- and high likelihood of cancer present. It was assumed that, under the pre-selection scenario, only the high-likelihood group would be read by radiologists, while all low-likelihood exams would be reported as normal. The area under the reader-averaged ROC curve (AUC) was calculated for the original evaluations and for the pre-selection scenarios and compared using a non-inferiority hypothesis. Setting the low/high-likelihood threshold at an AI score of 5 (high likelihood > 5) results in a trade-off of approximately halving (− 47%) the workload to be read by radiologists while excluding 7% of true-positive exams. Using an AI score of 2 as threshold yields a workload reduction of 17% while only excluding 1% of true-positive exams. Pre-selection did not change the average AUC of radiologists (inferior 95% CI > − 0.05) for any threshold except at the extreme AI score of 9. It is possible to automatically pre-select exams using AI to significantly reduce the breast cancer screening reading workload. • There is potential to use artificial intelligence to automatically reduce the breast cancer screening reading workload by excluding exams with a low likelihood of cancer. • The exclusion of exams with the lowest likelihood of cancer in screening might not change radiologists’ breast cancer detection performance. • When excluding exams with the lowest likelihood of cancer, the decrease in true-positive recalls would be balanced by a simultaneous reduction in false-positive recalls.
Background: Screening accuracy can be improved with digital breast tomosynthesis (DBT). To further evaluate DBT in screening, it is important to assess the molecular subtypes of the detected cancers. Purpose: To describe tumor characteristics, including molecular subtypes, of cancers detected at DBT compared with those detected at digital mammography (DM) in breast cancer screening. Materials and Methods: The Malmo Breast Tomosynthesis Screening Trial is a prospective, population-based screening trial comparing one-view DBT with two-view DM. Tumor characteristics were obtained, and invasive cancers were classified according to St Gallen as follows: luminal A-like, luminal B-like human epidermal growth factor receptor (HER)2-negative/HER2-positive, HER2-positive, and triple-negative cancers. Tumor characteristics were compared by mode of detection: DBT alone or DM (ie, DBT and DM or DM alone). chi(2) test was used for data analysis. Results: Between January 2010 and February 2015, 14 848 women were enrolled (mean age, 57 years +/- 10; age range, 40-76 years). In total, 139 cancers were detected; 118 cancers were invasive and 21 were ductal carcinomas in situ. Thirty-seven additional invasive cancers (36 cancers with complete subtypes and stage) were detected at DBT alone, and 81 cancers (80 cancers with complete stage) were detected at DM. No differences were seen between DBT and DM in the distribution of tumor size 20 mm or smaller (86% [31 of 36] vs 85% [68 of 80], respectively; P = .88), node-negative status (75% [27 of 36] vs 74% [59 of 80], respectively; P = .89), or luminal A-like subtype (53% [19 of 36] vs 46% [37 of 81], respectively; P = .48). Conclusion: The biologic profile of the additional cancers detected at digital breast tomosynthesis in a large prospective population-based screening trial was similar to those detected at digital mammography, and the majority were early-stage luminal A-like cancers. This indicates that digital breast tomosynthesis screening does not alter the predictive and prognostic profile of screening-detected cancers. (C) RSNA, 2019
Background Digital breast tomosynthesis is an advancement of the mammographic technique, with the potential to increase detection of lesions during breast cancer screening. The main aim of the Malmo Breast Tomosynthesis Screening Trial (MBTST) was to investigate the accuracy of one-view digital breast tomosynthesis in population screening compared with standard two-view digital mammography. Methods In this prospective, population-based screening study, of women aged 40-74 years invited to attend national breast cancer screening at Skane University Hospital, Malmo, Sweden, a random sample was asked to participate in the trial (every third woman who was invited to attend regular screening was invited to participate). Participants had to be able to speak English or Swedish and were excluded from the study if they were pregnant. Participants underwent screening with two-view digital mammography (ie, craniocaudal and mediolateral oblique views) followed by one-view digital breast tomosynthesis with reduced compression in the mediolateral oblique view (with a wide tomosynthesis angle of 50 degrees) at one screening visit. Images were read with masked double reading and scoring by two separate reading groups, one for each method, made up of seven radiologists. Any cancer detected with a malignancy probability score of three or higher by any reader in either group was discussed in a consensus meeting of at least two readers, from which the decision of whether or not to recall the woman for further investigation was made. The primary outcome measures were sensitivity and specificity of breast cancer detection. Secondary outcome measures were screening performance measures of cancer detection, recall, and interval cancers (cancers clinically detected between screenings), and positive predictive value for screen recalls and negative predictive value of each method. Outcomes were analysed in the per-protocol population. Follow-up of the participants for at least 2 years allowed for identification of interval cancers. This trial is registered with ClinicalTrials.gov, number NCT01091545. Findings Between Jan 27, 2010, and Feb 13, 2015, of 21 691 women invited, 14 851 (68%) agreed to participate. Three women withdrew consent during follow-up and were excluded from the analyses. 139 breast cancers were detected in 137 (<1%) of 14 848 women. Sensitivity was higher for digital breast tomosynthesis than for digital mammography (81.1%, 95% CI 74.2-86.9, vs 60.4%, 52.3-68.0) and specificity was slightly lower for digital breast tomosynthesis than was for digital mammography (97.2%, 95% CI 97.0-97.5, vs 98.1%, 97.9-98.3). The proportion of cancers detected was significantly higher with digital breast tomosynthesis than with digital mammography (8.7 cancers per 1000 women screened, 95% CI 7.3-10.3 vs 6.5 cancers per 1000 screened, 5.2-7.9; p<0.0001). The proportion of women recalled after discussion was higher among cancers detected by digital breast tomosynthesis than for those detected by digital mammography after consensus (3.6%, 95% CI 3.3-3.9 vs 2.5%, 2.2-2.8; p<0.0001). The positive predictive value for screen recalls was 24.1% (95% CI 20.5-28.0) for digital breast tomosynthesis and 25.9% (21.6-30.7) for digital mammography, and the negative predictive value was 99.8% (99.7-99.9) and 99.6% (99.4-99.7), respectively. The proportion of women who developed interval cancers after trial screening was 1.48 cancers per 1000 women screened (95% CI 0.93-2.24). Interpretation Breast cancer screening by use of one-view digital breast tomosynthesis with a reduced compression force has higher sensitivity at a slightly lower specificity for breast cancer detection compared with two-view digital mammography and has the potential to reduce the radiation dose and screen-reading burden required by two-view digital breast tomosynthesis with two-view digital mammography. Copyright (C) 2018 Elsevier Ltd. All rights reserved.
To compare the performance of one-view digital breast tomosynthesis (1v-DBT) to that of three other protocols combining DBT and mammography (DM) for breast cancer detection.
This study aimed to investigate the effects of adding adjunct mechanical imaging to mammography breast screening. We hypothesized that mechanical imaging could detect increased local pressure caused by both malignant and benign breast lesions and that a pressure threshold for malignancy could be established. The impact of this on breast screening was investigated with regard to reductions in recall and biopsy rates.