Purpose:We aim to investigate whether a breast cancer risk model can be trained with transfer learning from a breast cancer detection model. Approach:An existing open-source breast cancer detection model was used to extract the latent space representation of a local dataset composed of images from 17,878 breast cancer screening participants. An autoencoder network was used for dimensionality reduction. A final model training step mapped the reduced representation to 4-year breast cancer risk. Two experimental studies were performed: first, a local dataset was used to investigate the impact of transfer learning to local data. Second, the training was executed for breast cancer risk by excluding cases with obvious signs of cancer in the model training process, to avoid the influence of cancer signs. In addition, the Mirai model was evaluated. Results:The open-source model achieved a baseline performance of 0.72 with 95%CI [0.66, 0.78] for 4-year risk prediction when applied to local test cases of breast cancer risk data. This performance was improved after training on local breast cancer detection data to 0.75 with 95%CI [0.69, 0.81], although not statistically significant ( p = 0.09 ). Further training with local breast cancer risk data achieved an area under the curve (AUC) of 0.77 with 95%CI [0.71, 0.82], which is higher than the original model ( p < 0.01 ). The Mirai model demonstrated a lower performance after fine-tuning ( p < 0.01 ). Conclusions:Transfer learning enabled the creation of a breast cancer risk prediction model from a previously developed breast cancer detection model, demonstrating that a publicly available model can be adapted to local clinical needs.
Neuroendocrine neoplasms (NENs) are a heterogeneous group of neoplasms in which an adequate expression of a somatostatin receptor (SSTR) serves as a target for imaging and therapy, typically using SSTR agonists. In the past decade, research in the field of NENs has focused on SSTR antagonists because of their higher and more sustained binding to a target, which may expand current diagnostic and therapeutic options, especially in patients with lower SSTR expression. The aim of this study was to perform a first-in-human injection of a novel 99mTc-labeled SSTR type 2 antagonist to establish safety, pharmacokinetics, and first assessment of SSTR status in patients with NENs through a qualitative and quantitative comparison with 68Ga-SSTR PET as a reference standard. Methods: Patients with metastatic grade 1 or 2 NENs and proven SSTR expression in primary and metastatic lesions on SSTR PET were imaged after a first-in-human injection of [99mTc]Tc-TECANT1. Assessment of safety, pharmacokinetics, and dosimetry was performed as a primary endpoint, with comparison of diagnostic performance and quantification of uptake to SSTR PET as a secondary endpoint. Results: Ten patients were enrolled. [99mTc]Tc-TECANT1 was found to be safe, with favorable pharmacokinetics and dosimetry (average total body effective dose of 3.3 ± 0.7 mSv). The diagnostic performance in comparison to that of SSTR PET was found to be superior, with a comparable or a higher number of detected lesions and superior target-to-background contrast (a factor of 2 and above for lesions with the highest uptake). Conclusion: [99mTc]Tc-TECANT1 appears to be a safe and widely utilizable radiopharmaceutical for the assessment of SSTR expression status in NENs. Validation of results in a larger cohort is warranted.
Objective. The delineation of the clinical target volume (CTV) in radiotherapy is fundamentally uncertain due to the invisibility of microscopic disease on medical images. The ICRU 83 report acknowledges this by proposing a probabilistic interpretation of the CTV, but it does not define how to compute the probability of microscopic tumor presence (MTP) in tissue. This work addresses this gap by introducing a novel stochastic model that estimates the probability of MTP at the voxel level based on local spatial correlations in the voxels’ neighborhood. Approach. We developed two first-principles stochastic models to simulate MTP under different assumptions, incorporating spatial correlation between neighboring voxels. The constant marginal probability (CMP) model assumes spatially uniform MTP and is suited for tumors without radial dependence on the distance from the gross tumor volume (GTV). The variable marginal probability (VMP) model introduces radial dependence, modeling decreasing MTP with distance from the GTV. The CMP model was evaluated on prostate cancer data, while the VMP model was assessed using breast and lung cancer data. Results. Both models accurately reproduced the fraction of times that MTP is present. In the prostate case, the CMP model estimated a marginal probability of MTP of 0.03, consistent with a literature report that indicates an average total microscopic tumor volume of approximately 583 mm 3 across patients. The VMP model successfully replicated the radial distribution of tumor islets, achieving mean absolute errors of 0.01 mm and 0.011 mm for breast and lung cancer distance distributions, respectively. However, not all MTP characteristics could be fully captured by the models, and in some cases discrepancies with population based tumor characteristics remain. Significance. This work introduces a statistically consistent framework that enables a probabilistic definition of the CTV. The proposed models provide a new way to capture key aspects of microscopic disease spread by introducing local voxel correlations.
Objective.Deep-learning-based models have achieved state-of-the-art breast cancer risk (BCR) prediction performance. However, these models are highly complex, and the underlying mechanisms of BCR prediction are not fully understood. Key questions include whether these models can detect breast morphologic changes that lead to cancer. These findings would boost confidence in utilizing BCR models in practice and provide clinicians with new perspectives. In this work, we aimed to determine when oncogenic processes in the breast provide sufficient signal for the models to detect these changes.Approach.In total, 1210 screening mammograms were collected for patients screened at different times before the cancer was screen-detected and 2400 mammograms for patients with at least ten years of follow-up. MIRAI, a BCR risk prediction model, was used to estimate the BCR. Attribution heterogeneity was defined as the relative difference between the attributions obtained from the right and left breasts using one of the eight interpretability techniques. Model reliance on the side of the breast with cancer was quantified with AUC. The Mann-Whitney U test was used to check for significant differences in median absolute Attribution Heterogeneity between cancer patients and healthy individuals.Results.All tested attribution methods showed a similar longitudinal trend, where the model reliance on the side of the breast with cancer was the highest for the 0-1 years-to-cancer interval (AUC = 0.85-0.95), dropped for the 1-3 years-to-cancer interval (AUC = 0.64-0.71), and remained above the threshold for random performance for the 3-5 years-to-cancer interval (AUC = 0.51-0.58). For all eight attribution methods, the median values of absolute attribution heterogeneity were significantly larger for patients diagnosed with cancer at one point (p< 0.01).Significance.Interpretability of BCR prediction has revealed that long-term predictions (beyond three years) are most likely based on typical breast characteristics, such as breast density; for mid-term predictions (one to three years), the model appears to detect early signs of tumor development, while for short-term predictions (up to a year), the BCR model essentially functions as a breast cancer detection model.
Objective.When it comes to the implementation of deep-learning based breast cancer risk (BCR) prediction models in clinical settings, it is important to be aware that these models could be sensitive to various factors, especially those arising from the acquisition process. In this work, we investigated how sensitive the state-of-the-art BCR prediction model is to realistic image alterations that can occur as a result of different positioning during the acquisition process.Approach.5076 mammograms (1269 exams, 650 participants) from the Slovenian and Belgium (University Hospital Leuven) Breast Cancer Screening Programs were collected. The Original MIRAI model was used for 1-5 year BCR estimation. First, BCR was predicted for the original mammograms, which were not changed. Then, a series of different image alteration techniques was performed, such as swapping left and right breasts, removing tissue below the inframammary fold, translations, cropping, rotations, registration and pectoral muscle removal. In addition, a subset of 81 exams, where at least one of the mammograms had to be retaken due to inadequate image quality, served as an approximation of a test-retest experiment. Bland-Altman plots were used to determine prediction bias and 95% limits of agreement (LOA). Additionally, the mean absolute difference in BCR (Mean AD) was calculated. The impact on the overall discrimination performance was evaluated with the AUC.Results.Swapping left and right breasts had no impact on the predicted BCR. The removal of skin tissue below the inframammary fold had minimal impact on the predicted BCR (1-5 year LOA: [-0.02, 0.01]). The model was sensitive to translation, rotation, registration, and cropping, where LOAs of up to ±0.1 were observed. Partial pectoral muscle removal did not have a major impact on predicted BCR, while complete removal of pectoral muscle introduced substantial prediction bias and LOAs (1 year LOA: [-0.07, 0.04], 5 year LOA: [-0.06, 0.03]). The approximation of a real test-retest experiment resulted in LOAs similar to those of simulated image alterations. None of the alterations impacted the overall BCR discrimination performance; the initial 1 year AUC (0.90 [0.88, 0.92]) and 5 year AUC (0.77 [0.75, 0.80]) remained unchanged.Significance.While tested image alterations do not impact overall BCR discrimination performance, substantial changes in predicted 1-5 year BCR can occur on an individual basis.
Quantitative 18F-FDG PET/CT-derived metabolic metrics are strongly associated with patient outcomes in diffuse large B-cell lymphoma (DLBCL), but the lack of consensus on optimal segmentation thresholds limits standardization. This study evaluated the prognostic value of various metabolic tumor volume (MTV) segmentation approaches in 140 stage II-IV DLBCL patients treated with standard immunochemotherapy. MTV was derived using fixed SUV (≥2.5, ≥4.0), relative (>41% SUVmax), and adaptive (liver-to-background) thresholds. Baseline MTV metrics significantly correlated with 3-year overall survival (OS3) in univariate analysis in overall cohort, with MTV41 showing the strongest association (HR: 1.27; p = 0.003). MTV25 and MTV41 remained significant in the stage 4 patient subgroup. However, in multivariate analysis, no MTV metric independently predicted OS3 when adjusted for the International Prognostic Index (IPI), which remained the dominant predictor (HR: 1.95; p < 0.0001). ROC analysis confirmed superior AUC for IPI (0.76) over PET-based metrics (0.64-0.69). Predictive models integrating IPI with PET metrics were robust but failed to improve prognostic accuracy beyond IPI alone. Although PET-derived MTV metrics provide prognostic value in univariate analysis, threshold selection has minimal impact, and their added value is limited when combined with IPI, reinforcing its role as the most reliable survival predictor in DLBCL.
BACKGROUND:A considerable proportion of metastatic melanoma (mM) patients do not respond to immune checkpoint inhibitors (ICIs). There is a great need to develop noninvasive biomarkers to detect patients, who do not respond to ICIs early during the course of treatment. The aim of this study was to evaluate the role of early [18F]2fluoro-2-deoxy-D-glucose PET/CT (18F-FDG PET/CT) at week four (W4) and other possible prognostic biomarkers of survival in mM patients receiving ICIs. PATIENTS AND METHODS:. In this prospective noninterventional clinical study, mM patients receiving ICIs regularly underwent 18F-FDG PET/CT: at baseline, at W4 after ICI initiation, at week sixteen and every 16 weeks thereafter. The tumor response to ICIs at W4 was assessed via modified European Organisation for Research and Treatment of Cancer (EORTC) criteria. Patients with progressive metabolic disease (PMD) were classified into the no clinical benefit group (no-CB), and those with other response types were classified into the clinical benefit group (CB). The primary end point was survival analysis on the basis of the W4 18F-FDG PET/CT response. The secondary endpoints were survival analysis on the basis of LDH, the number of metastatic localizations, and immune-related adverse events (irAEs). Kaplan-Meier analysis and univariate Cox regression analysis were used to assess the impact on survival. RESULTS:Overall, 71 patients were included. The median follow-up was 37.1 months (952% CI = 30.1-38.0). Three (4%) patients had only baseline scans due to rapid disease progression and death prior to W4 18F-FDG-PET/CT. Fifty-one (72%) patients were classified into the CB group, and 17 (24%) were classified into the no-CB group. There was a statistically significant difference in median overall survival (OS) between the CB group (median OS not reached [NR]; 95% CI = 17.8 months - NR) and the no-CB group (median OS 6.2 months; 95% CI = 4.6 months - NR; p = 0.003). Univariate Cox analysis showed HR of 0.4 (95% CI = 0.18 - 0.72; p = 0.004). median OS was also significantly longer in the group with normal serum LDH levels and the group with irAEs and cutaneous irAEs. CONCLUSIONS:Evaluation of mM patients with early 18F-FDG-PET/CT at W4, who were treated with ICIs, could serve as prognostic imaging biomarkers. Other recognized prognostic biomarkers were the serum LDH level and occurrence of cutaneous irAEs.
In breast cancer screening, evaluating risk prediction independently from cancer detection is challenging due to factors such as early cancer signs and missed cancers in negative mammograms. Recent developments in risk prediction include Mirai, a deep learning-based mammographic breast cancer risk prediction model. As a means of establishing which image features contribute to the Mirai risk estimate, we developed CalcMirai, a Mirai variant that is limited to calcification features. The aim of this study was to use CalcMirai to test the casual contribution of calcification features in Mirai to the resultant risk prediction. Screening mammograms from the EMory BrEast imaging Dataset (EMBED) were used in selective mirroring experiments, where mammograms from one breast were mirrored to replace the contralateral breast. Results showed that both Mirai and CalcMirai had better performance in the positive mirroring setting i.e. considering only the future cancerous side, compared to using both left and right breast views in the original models (p-values < 0.01). While negative mirroring i.e. considering the future healthy breast side of the cancerous patient, resulted in poorer performance (p-values < 0.01), both models remained discriminative (AUCs 0.61-0.62). There was no significant difference between the negative mirroring performance of Mirai and CalcMirai. Additionally, visual assessments of receptive fields confirmed that the calcification features identified in CalcMirai accurately captured the positions of calcifications in the contralateral healthy breasts of future cancer patients. This underlines the role of calcification features as strong risk factors. Our findings suggest that the predictive power of Mirai derives mainly from its ability to detect early micro-calcifications and/or identify high-risk calcifications. These results provide new insight into mammographic risk factors, implying that calcification may be underestimated in predicting breast cancer risk compared to the well-known risk factor of parenchymal patterns.
Purpose To evaluate whether features extracted by Mirai can be aligned with mammographic observations and contribute meaningfully to the prediction of breast cancer risk. Materials and Methods This retrospective study examined the correlation of 512 Mirai features with mammographic observations in terms of receptive field and anatomic location. A total of 29 374 screening examinations with mammograms (10 415 female patients; mean age at examination, 60 years ± 11 [SD]) from the EMory BrEast imaging Dataset (EMBED) (2013-2020) were used to evaluate feature importance using a feature-centric explainable artificial intelligence pipeline. Risk prediction was evaluated using only calcification features (CalcMirai) or mass features (MassMirai) against Mirai. Performance was assessed in screening and screen-negative (time to cancer, >6 months) populations using the area under the receiver operating characteristic curve (AUC). Results Eighteen calcification features and 18 mass features were selected for CalcMirai and MassMirai, respectively. Both CalcMirai and MassMirai had lower performance than Mirai in lesion detection (screening population: Mirai 1-year AUC, 0.81 [95% CI: 0.78, 0.84]; CalcMirai 1-year AUC, 0.76 [95% CI: 0.73, 0.80]; MassMirai 1-year AUC, 0.74 [95% CI: 0.71, 0.78] [P < .001]). In risk prediction, there was no evidence of a difference in performance between CalcMirai and Mirai (screen-negative population: Mirai 5-year AUC, 0.66 [95% CI: 0.63, 0.69]; CalcMirai 5-year AUC, 0.66 [95% CI: 0.64, 0.69] [P = .71]). However, MassMirai achieved lower performance than Mirai (5-year AUC, 0.57 [95% CI: 0.54, 0.60]; P < .001). Radiologist review of calcification features confirmed Mirai's use of benign calcification in risk prediction. Conclusion The explainable AI pipeline demonstrated that Mirai implicitly learned to identify mammographic lesion features, particularly calcifications, for lesion detection and risk prediction. Keywords: Breast, Mammography, Screening Supplemental material is available for this article. © The Author(s) 2025. Published by the Radiological Society of North America under a CC BY 4.0 license. See also commentary by Gichoya and Trivedi in this issue.
Objective. To determine whether pre-treatment brain metabolic network patterns measured with 18F-FDG PET are associated with treatment response and survival in cancer patients. Approach. Exploratory retrospective study of two independent cohorts: stage III breast cancer patients treated with neoadjuvant chemotherapy and stage IV melanoma patients treated with anti-PD-1 immunotherapy. Metabolic brain network scores were derived from pre-treatment 18F-FDG PET scans and evaluated for their ability to stratify good versus poor responders using ROC analysis (AUC). Longitudinal changes in network scores were assessed across follow-up, and progression-free survival (PFS) and overall survival (OS) analyses were performed in the melanoma cohort. Main results. Specific brain networks were associated with treatment outcome; the cognition/language network was the strongest predictor (AUC > 0.84 for distinguishing good vs. poor responders in both cohorts). Good responders showed lower cognition/language scores than poor responders and healthy controls. Longitudinally, cognition/language scores remained stable in good responders, while poor responders exhibited a gradual convergence toward the scores observed in good responders. In the melanoma cohort, lower cognition/language scores were significantly associated with longer PFS and OS. Significance. These findings indicate that metabolic brain network patterns, particularly the cognition/language network, may serve as noninvasive biomarkers linked to treatment efficacy and survival in oncology. The results support a possible complex interaction between brain metabolism, immune response, and clinical outcomes. Key limitations include the retrospective design and lack of direct immune-function and psychometric measures; prospective, multimodal studies are needed to validate these observations and elucidate underlying mechanisms.
Objective. State-of-the-art breast cancer risk (BCR) prediction models have been originally trained on mammograms with pectoral muscle (PM) included. This study investigated whether excluding PM during training/fine-tuning improves the model's BCR discrimination performance, calibration, and robustness. Approach. First, the Original deep learning model (MIRAI), trained on the US (Massachusetts General Hospital) data, was validated, and the relative contribution of PM to BCR predictions was evaluated using saliency maps. Additionally, 23 792 mammograms from the Slovenian screening program were collected and two datasets were created, with and without screening positive exams. The original MIRAI was then fine-tuned on the training/fine-tuning set of Slovenian mammograms with and without PM, creating Fine-tuned MIRAI models. In total, four models (Original MIRAI with PM, Original MIRAI without PM, Fine-tuned MIRAI with PM, Fine-tuned MIRAI without PM) were compared on a test set in terms of discrimination performance for 1-5 Year BCR (evaluating area under the curve), calibration performance (measured with expected calibration error-ECE) and robustness to incremental PM removals/additions, and to incremental breast tissue removals. Results. The relative contribution of PM to the BCR prediction on the Original MIRAI model was low (similar to 5%); however, there were significant outliers where the relative contribution was more than 50%. The removal of PM did not impact the 1-5 Year BCR discrimination performance of the Original MIRAI (with screening positive exams: 0.77-0.91, without screening positive exams: 0.64-0.67). Fine-tuned MIRAI on mammograms with PM removed achieved significantly higher 1-5 Year BCR discrimination performance (with screening positive exams: 0.82-0.93, without screening positive exams: 0.71-0.79). After recalibration, all models had similar ECE (with screening positive exams: 0.04-0.05, without screening positive exams: 0.02-0.03). Significance. Improved BCR discrimination performance can be achieved when the model is trained/fine-tuned on mammograms with PM removed.
Aim: This study proposes a method to use longitudinal breast cancer screening data to develop a 1- to 4-year breast cancer risk prediction model. It uses transfer learning from an open-source breast cancer detection model, an Autoencoder to perform dimensionality reduction as well as an LSTM network to incorporate the sequential data. Methods: The study utilizes a labelled dataset of 846 patients with up to five different mammography screening exams. The exams were taken on three systems from the vendor Siemens and the images are of the "FOR PRESENTATION" type. In this dataset there are 423 low risk cases and 423 high risk cases. A breast cancer detection model was used to obtain a latent representation of features extracted from the screening images. Dimensionality reduction was performed on the latent space using an Autoencoder architecture. The reduced latent space was then mapped to 1- to 4-year breast cancer risk with an LSTM model. Results: The model achieved an AUC of 0.74 for differentiating high and low risk cases, outperforming the Tyrer-Cuzick model. At the reference specificity operating point of 85.4% from the Tyrer-Cuzick model, the longitudinal model achieves a sensitivity of 60%, outperforming a similar model trained by only seeing a single exam of a given patient. Conclusions: The incorporation of longitudinal data into breast cancer risk assessment models can increase the sensitivity to underlying patterns that are correlated to breast cancer and therefore improve breast cancer screening strategies.
Background Detection of bone marrow involvement (BMI) in diffuse large B-cell lymphoma (DLBCL) typically relies on invasive bone marrow biopsy (BMB) that faces procedure limitations, while F-18-FDG PET/CT imaging offers a noninvasive alternative. The present study assesses the performance of F-18-FDG PET/CT in DLBCL BMI detection, its agreement with BMB, and the impact of BMI on survival outcomes. Patients and methods This retrospective study analyzes baseline F-18-FDG PET/CT and BMB findings in145 stage II-IV DLBCL patients, evaluating both performance of the two diagnostic procedures and the impact of BMI on survival. Results DLBCL BMI was detected in 38 patients (26.2%) using PET/CT and in 18 patients (12.4%) using BMB. Concordant results were seen in 79.3% of patients, with 20.7% showing discordant results. Combining PET/CT and BMB data, we identified 29.7% of patients with BMI. The sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), and accuracy of PET/CT for detecting DLBCL BMI were 88.4%, 100%, 100%, 95.3%, and 96.5%, respectively, while BMB showed lower sensitivity (41.9%) and NPV (46.8%). The median overall survival (OS) was not reached in any gender subgroup, with 5-year OS rates of 82% (total), 84% (female), and 80% (male) (p = 0.461), while different International Prognostic Index (IPI) groups exhibited varied 5-year OS rates: 94% for low risk (LR), 91% for low-intermediate risk (LIR), 84% for high-intermediate risk (HIR), and 65% for high risk (HR) (p = 0.0027). Bone marrow involvement did not impact OS significantly (p = 0.979). Conclusions 18F-FDG PET/CT demonstrated superior diagnostic accuracy compared to BMB. While other studies reported poorer overall and BMI 5-year OS in DLBCL, our findings demonstrated favourable survival data.
Background: This study assessed the prognostic value of tumor burden in bone marrow (BM) and total disease (TD), as depicted on 18F-FDG PET/CT in 140 DLBCL patients, for complete remission after first-line systemic treatment (iCR) and 3- and 5-year overall survival (OS3 and OS5). Methods: Baseline 18F-FDG PET/CT scans of 140 DLBCL patients were segmented to quantify metabolic tumor volume (MTV), total lesion glycolysis (TLG), and SUVmax in BMI, findings elsewhere (XL), and TD. Results: Bone marrow involvement (BMI) presented in 35 (25%) patients. Median follow-up time was 47 months; 79 patients (56%) achieved iCR. iCR was significantly associated with TD MTV, XL MTV, BM PET positivity, and International Prognostic Index (IPI). OS3 was significantly worse with TD MTV, XL MTV, IPI, and age. OS5 was significantly associated with IPI, but not with MTVs and TLGs. Univariate factors predicting OS3 were XL MTV (hazard ratio [HR] = 1.29), BMI SUVmax (HR = 0.56), and IPI (HR = 1.92). By multivariate analysis, higher IPI (HR = 2.26) and BMI SUVmax (HR = 0.91) were significant independent predictors for OS3. BMI SUVmax resulted in a negative coefficient and hence indicated a protective effect. Conclusions: Baseline 18F-FDG PET/CT MTV is significantly associated with survival. BMI identified on 18F-FDG PET/CT allows appropriate treatment that may improve survival.
Aim: This study proposes a method to bypass the requirement of large amounts of original training data to develop a 1- to 4-year breast cancer risk prediction model using transfer learning from a breast cancer detection model with digital mammography images as input. Methods: The study utilizes a labelled dataset of 423 low risk cases and 423 high risk cases, which is considered a small amount of data in terms of AI development, but from the viewpoint of a regional screening organization this represents a large number of high risk cases, given the rarity of such cases compared to the large number of low risk cases available. A breast cancer detection model was used to obtain a latent representation of features extracted from 'FOR PRESENTATION' screening mammography images from three systems from a single vendor (Siemens). Dimensionality reduction was performed on the latent space using an Autoencoder architecture. The reduced latent space was then mapped to 1- to 4-year breast cancer risk with a fully-connected model. Results: The resulting model achieved an AUC of 0.77 for differentiating high and low risk cases, outperforming the Tyrer-Cuzick model and achieving state-of-the-art performance. Conclusions: The use of transfer learning from breast cancer detection models can produce image-based breast cancer risk prediction models that are comparable to the state-of-the-art, while requiring only moderate amounts of data.