Multiple sclerosis is a chronic autoimmune disease that affects the central nervous system. Understanding multiple sclerosis progression and identifying the implicated brain structures is crucial for personalized treatment decisions. Deformation-based morphometry utilizes anatomical magnetic resonance imaging to quantitatively assess volumetric brain changes at the voxel level, providing insight into how each brain region contributes to clinical progression with regards to neurodegeneration. Utilizing such voxel-level data from a relapsing multiple sclerosis clinical trial, we extend a model-agnostic feature importance metric to identify a robust and predictive feature set that corresponds to clinical progression. These features correspond to brain regions that are clinically meaningful in MS disease research, demonstrating their scientific relevance. When used to predict progression using classical survival models and 3D convolutional neural networks, the identified regions led to the best-performing models, demonstrating their prognostic strength. We also find that these features generalize well to other definitions of clinical progression and can compensate for the omission of highly prognostic clinical features, underscoring the predictive power and clinical relevance of deformation-based morphometry as a regional identification tool.
Abstract Background: Diffuse large B-cell lymphoma (DLBCL) is a heterogeneous disease and accurate outcome prediction is essential for optimizing treatment. Baseline total metabolic tumor volume (TMTV; Vercellino et al. Blood 2020), and interim and end-of-treatment (EOT) FDG-PET responses are established prognostic indicators in frontline DLBCL treatment (Kostakoglu et al. Blood Adv 2021; Jemaa et al. J Clin Oncol 2024). However, their relevance in regimens other than rituximab (R) plus cyclophosphamide, doxorubicin, vincristine, and prednisolone (CHOP) remains unclear (Kostakoglu et al. Haematologica 2022). Polatuzumab vedotin (Pola) with R-CHP is a new standard-of-care therapy based on the POLARIX study (NCT03274492; Phase III Pola-R-CHP vs R-CHOP). This analysis evaluates the prognostic value of quantitative PET parameters, including baseline TMTV, total lesion glycolysis (TLG), and interim (iPET) and EOT PET responses in patients with previously untreated DLBCL in the POLARIX study. Methods: Available baseline FDG-PET/CT scans were centrally reviewed for TMTV by an Independent Review Committee, using a standardized uptake value (SUV) threshold of 1.5 times the liver reference plus 2 standard deviations, with a minimum lesion volume of 1mL. After 4 cycles, iPET (iPET4) positivity was defined as a reduction in maximum SUV (SUVmax) of <70%. At EOT, metabolic response was assessed according to the 2014 Lugano criteria. Prognosis of FDG-PET-derived metrics, and treatment effects stratified by FDG-PET-derived metrics and other predictive covariates, were analyzed using the Kaplan–Meier method and Cox proportional hazards model in the POLARIX population with 5-years’ follow-up. Results: In the POLARIX global population,835 patients had evaluable baseline FDG-PET/CT scans, 661 had evaluable iPET4, and 760 had evaluable EOT FDG-PET/CT scans. Median TMTV was 411mL (range: 0–10,881). In the R-CHOP arm, high (>411mL) baseline TMTV (HR: 2.05 [95% CI: 1.49–2.83]) and high TLG (HR: 1.55 [95% CI: 1.13–2.12]) were strong predictors of progression-free survival (PFS). Notably, high TMTV and TLG were similarly predictive in the Pola-R-CHP arm, HR: 1.85 (95% CI: 1.30–2.63) and HR: 1.60 (95% CI: 1.13–2.27), respectively. There was no significant evidence of association between Pola-R-CHP benefit and TMTV or TLG (treatment interaction: TMTV p=0.67; TLG p=0.92), or association between baseline TMTV and cell-of-origin (COO [by Nanostring]; activated B-cell [ABC] interaction: TMTV p=0.38; TLG p=0.99). In the ABC COO population with high TMTV, 5-year PFS and OS rates were significantly higher (PFS HR: 0.27 [95% CI: 0.14–0.50]; OS HR: 0.28 [95% CI: 0.13–0.63]) with Pola-R-CHP (PFS: 74% [95% CI: 63–87]; OS: 84% [95% CI: 75–95]) compared with R-CHOP (PFS: 28% [95% CI: 17–45]; OS: 54% [95% CI: 42–69]). In 661 evaluable patients from POLARIX, 583 (88%) had negative iPET4 (Pola-R-CHP: 312/340 [92%]; R-CHOP: 271/321 [84%]); of 760 patients with evaluable EOT scans, 627 (83%) achieved complete metabolic response (CMR) at EOT (Pola-R-CHP: 327/387 [84%]; R-CHOP: 300/373 [80%]). In the Pola-R-CHP arm, negative iPET4 and CMR at EOT differentiated patients at high risk of relapse (5-year PFS rates from iPET4 positive vs negative: 69% [95% CI: 63–75] vs 46% [95% CI: 29–71], respectively; HR:0.38 [95% CI: 0.22–0.68]) and high risk of death (5-year OS rates from EOT CMR vs non-CMR: 91% [95% CI: 88–95] vs 58% [95% CI: 47–73], respectively; HR:0.17 [95% CI: 0.10–0.29]). Similar results were also observed in the R-CHOP arm (iPET4 PFS HR: 0.42 [95% CI: 0.27–0.65]; CMR at EOT OS HR: 0.41 [95% CI: 0.25–0.67]). No significant evidence of association between EOT or iPET4 responses and COO was found (ABC interaction: EOT response p=0.76; iPET4 p=0.75). Among patients with CMR at EOT, significant OS and PFS benefits were observed with Pola-R-CHP relative to R-CHOP (OS HR: 0.55 [95% CI: 0.35–0.88]; PFS HR:0.66 [95% CI: 0.48–0.91]).Conclusions: In the POLARIX study, baseline TMTV was a strong prognostic factor of patient outcomes and consistent across treatment arms, suggesting its relevance regardless of DLBCL biological subtype or therapy received. FDG-PET responses at both interim and EOT assessments were also prognostic. These findings confirm that well-established PET prognostic factors remain relevant in patients treated with Pola-R-CHP. In addition, Pola-R-CHP led to more sustained responses and longer survival among patients with EOT CMR compared with R-CHOP.
18-Fluoro-deoxyglucose positron emission tomography/computed tomography (FDG-PET/CT) is a valuable imaging tool widely used in the management of cancer patients. Deep learning models excel at segmenting highly metabolic tumors but face challenges in regions with complex anatomy and normal cell uptake, such as the gastro-intestinal tract. Despite these challenges, it remains important to achieve accurate segmentation of gastro-intestinal tumors. Here, we present an international multicenter comparative study between a novel organ-focused approach and a whole-body training method to evaluate the effectiveness of training data homogeneity in accurately identifying gastro-intestinal tumors. In the organ-focused method, the training data is limited to cases with intestinal tumors which makes the network trained with more homogeneous data and with stronger presence of intestinal tumor signals. The whole body approach extracts the intestinal tumors from the results of a model trained on the whole-body scans. Both approaches were trained using diffuse large B cell (DLBCL) patients from a large multi-center clinical trial (NCT01287741). We report an improved mean(±std) Dice score of 0.78(±0.21) for the organ-based approach on the hold-out set, compared to 0.63(±0.30) for the whole-body approach, with the p-value of less than 0.0001. At the lesion level, the proposed organ-based approach also shows increased precision, recall, and F1-score. An independent trial was used to evaluate the generalizability of the proposed method to non-Hodgkin’s lymphoma (NHL) patients with follicular lymphoma (FL). Given the variability in structure and metabolism across tissues in the body, our quantitative findings suggest organ-focused training enhances intestinal tumor segmentation by leveraging tissue homogeneity in the training data, contrasting with the whole-body training approach, which, by its very nature, is a more heterogeneous data set.
Introduction Artificial intelligence-based tools are increasingly used to assist radiologists with assessing a patient's disease status and evaluating treatment effectiveness. However, before these tools can be applied in clinical trials or routine practice, robust measures must be established to evaluate their performance. Total metabolic tumor volume (TMTV) is a volumetric tumor burden assessment derived from positron emission tomography/computed tomography (PET/CT) scans that has demonstrated prognostic value in patients with B-cell non-Hodgkin lymphoma. In this study, we investigate the variability among human readers in quantifying TMTV from PET/CT scans of patients with lymphoma, then compare their performance with a fully automated deep-learning-based algorithm (Jemaa et al. J Digit Imaging 2020). Methods Diagnostic quality PET/CT scans at screening or at first on-treatment tumor response assessment (FOT) visits were obtained from 125 patients enrolled in an ongoing phase 1/2 multicenter, open-label, dose escalation and expansion study in patients with relapsed or refractory B-cell non-Hodgkin lymphoma (ClinicalTrials.gov identifier, NCT02500407). One scan per patient was selected at random from either time point. Manual TMTV (mTMTV) was assessed individually by 3 radiologists using a semiautomated reference region-based threshold method (1.5 × mean liver standardized uptake value + 2 standard deviations [SDs]). A consolidated ground truth read was created for each scan from the 3 mTMTV reads using a majority-vote approach. Automated TMTV (aTMTV) was applied to the same images, and aTMTV and mTMTV values were compared. All TMTV values were cubic root-transformed for statistical analysis. Bland-Altman analysis and Dice similarity coefficients (DSCs) were used to assess concordance among mTMTV reads and between aTMTV and mTMTV reads, including ground truth. Correlation between aTMTV and the ground truth mTMTV was assessed using Pearson's correlation coefficient (r). Subgroup analyses compared aTMTV reads with ground truth mTMTV reads in patients of different age, sex, cancer subtype and prior therapy lines at baseline. Results In total, 110 images comprising 78 screening and 32 FOT scans from patients with diverse lymphoma types met image quality requirements and were included (aggressive lymphomas, n = 67; indolent lymphomas, n = 43). Bland-Altman analyses revealed good agreement between individual mTMTV reads of TMTV: the mean differences between mTMTV readers 1 versus 2, 2 versus 3, and 3 versus 1 were −0.02 (limit of agreement [LoA]: −1.89, 1.84), −0.03 (LoA: −1.65, 1.59) and 0.05 (LoA: −2.14, 2.24), respectively. A high degree of overlap was also observed between lesion segmentations, with DSCs of 0.89 (SD: 0.17) for reader 1 versus 2, 0.90 (SD: 0.16) for reader 2 versus 3, and 0.89 (SD: 0.19) for reader 3 versus 1. Slightly greater measurement variability was observed when individual mTMTV reads were compared with aTMTV. The mean differences between methods were 0.10 (LoA: −1.53, 1.73) for reader 1 versus aTMTV, 0.12 (LoA: −1.75, 1.99) for reader 2 versus aTMTV, and 0.15 (LoA: −1.97, 2.28) for reader 3 versus aTMTV. DSCs for readers 1, 2 and 3 versus aTMTV were 0.70 (SD: 0.21), 0.70 (SD: 0.22) and 0.71 (SD: 0.22), respectively. Good agreement and a strong correlation were observed between the consolidated ground truth mTMTV read and aTMTV (mean difference between methods: 0.05 [95% confidence interval (CI): −0.12, 0.22; LoA: −1.71, 1.80]; r = 0.97 [95% CI: 0.97, 0.98]). However, slight differences were observed in the spatial overlap of lesion segmentations between aTMTV and ground truth mTMTV (DSC: 0.71; SD: 0.22). Consistent results were observed across all patient subgroups. Conclusion The findings of this study provide valuable insights into the comparative performance of manual readers and aTMTV when assessing TMTV from PET/CT scans. aTMTV demonstrated similar performance to expert radiologists using semiautomated software for TMTV estimation, although slightly different regions were identified as lesions by the algorithm. Further research is warranted to explore the clinical implications of these findings. Understanding the degree of measurement variability among manual readers and how algorithm-based approaches differ may ultimately contribute to improved accuracy of automated TMTV assessments and support clinical adoption.
Accurate quantification of multiple sclerosis (MS) lesions using multi-contrast magnetic resonance imaging (MRI) plays a crucial role in disease assessment. While many methods for automatic MS lesion segmentation in MRI are available, these methods typically require a fixed set of MRI modalities as inputs. Such full multi-contrast inputs are not always acquired, limiting their utility in practice. To address this issue, a training strategy known as modality dropout (MD) has been widely adopted in the literature. However, models trained via MD still underperform compared to dedicated models trained for particular modality configurations. In this work, we hypothesize that the poor performance of MD is the result of an overly constrained multi-task optimization problem. To reduce harmful task interference, we propose to incorporate task-conditional mixture-of-expert layers into our segmentation model, allowing different tasks to leverage different parameters subsets. Second, we propose a novel online self-distillation loss to help regularize the model and to explicitly promote model invariance to input modality configuration. Compared to standard MD training, our method demonstrates improved results on a large proprietary clinical trial dataset as well as on a small publicly available dataset of T2 lesions.
PURPOSE:Artificial intelligence can reduce the time used by physicians on radiological assessments. For 18F-fluorodeoxyglucose-avid lymphomas, obtaining complete metabolic response (CMR) by end of treatment is prognostic. METHODS:Here, we present a deep learning-based algorithm for fully automated treatment response assessments according to the Lugano 2014 classification. The proposed four-stage method, trained on a multicountry clinical trial (ClinicalTrials.gov identifier: NCT01287741) and tested in three independent multicenter and multicountry test sets on different non-Hodgkin lymphoma subtypes and different lines of treatment (ClinicalTrials.gov identifiers NCT02257567, NCT02500407; 20% holdout in ClinicalTrials.gov identifier NCT01287741), outputs the detected lesions at baseline and follow-up to enable focused radiologist review. RESULTS:The method's response assessment achieved high agreement with the adjudicated radiologic responses (eg, agreement for overall response assessment of 93%, 87%, and 85% in ClinicalTrials.gov identifiers NCT01287741, NCT02500407, and NCT02257567, respectively) similar to inter-radiologist agreement and was strongly prognostic of outcomes with a trend toward higher accuracy for death risk than adjudicated radiologic responses (hazard ratio for end of treatment by-model CMR of 0.123, 0.054, and 0.205 in ClinicalTrials.gov identifiers NCT01287741, NCT02500407, and NCT02257567, compared with, respectively, 0.226, 0.292, and 0.272 for CMR by the adjudicated responses). Furthermore, a radiologist review of the algorithm's assessments was conducted. The radiologist median review time was 1.38 minutes/assessment, and no statistically significant differences were observed in the level of agreement of the radiologist with the model's response compared with the level of agreement of the radiologist with the adjudicated responses. CONCLUSION:These results suggest that the proposed method can be incorporated into radiologic response assessment workflows in cancer imaging for significant time savings and with performance similar to trained medical experts.
F-Fluorodeoxyglucose-positron emission tomography (FDG-PET) imaging is a valuable diagnostic tool in oncology with a wide range of clinical applications for cancer diagnosis, staging, and monitoring treatment response. Accurate tumor segmentation from these images is vital for understanding the biochemical and physiological alterations within the tumors. End-to-end deep learning approaches enable rapid and reproducible tumor identification and extraction, surpassing manual and semi-automatic methods. Compared to other organs, intestinal tumor segmentation poses a significant challenge due to its complex anatomical shape and acute non-malignant findings. This study aims to investigate the impact of training data homogeneity on the segmentation results of intestinal tumors using Convolutional Neural Networks (CNNs). To achieve this, we propose an organ-based approach where the training data is limited to the small intestine region. We will compare the results obtained by the organ-based approach with those from a model trained on the whole-body PET/CT data. In the whole-body approach, tumor segmentation predictions for the intestine are extracted from the results obtained by training on the whole-body data. Quantitative results show that the organ-based approach outperform the whole-body method in segmentation of intestinal tumors. Whole-body and organ-based approaches generated a dice score (mean±std) of 0.63±0.30 and 0.78±0.21 for the whole-body and organ-based approaches respectively with p-value less than 0.0001. The lesion level analysis yielded F1 scores of 0.79 for the whole-body approach and 0.86 for the organ-based approach.
Deep neural networks (DNNs) have recently showed remarkable performance in various computer vision tasks, including classification and segmentation of medical images. Deep ensembles (an aggregated prediction of multiple DNNs) were shown to improve a DNN's performance in various classification tasks. Here we explore how deep ensembles perform in the image segmentation task, in particular, organ segmentations in CT (Computed Tomography) images. Ensembles of V-Nets were trained to segment multiple organs using several in-house and publicly available clinical studies. The ensembles segmentations were tested on images from a different set of studies, and the effects of ensemble size as well as other ensemble parameters were explored for various organs. Compared to single models, Deep Ensembles significantly improved the average segmentation accuracy, especially for those organs where the accuracy was lower. More importantly, Deep Ensembles strongly reduced occasional "catastrophic" segmentation failures characteristic of single models and variability of the segmentation accuracy from image to image. To quantify this we defined the "high risk images": images for which at least one model produced an outlier metric (performed in the lower 5% percentile). These images comprised about 12% of the test images across all organs. Ensembles performed without outliers for 68%-100% of the "high risk images" depending on the performance metric used.
PDF file, 869K, Proliferation and apoptosis analysis by IF staining in EL4 tumor sections.
PDF file - 47K, Anti-proliferative EC50's for GNE-317 against a Panel of Nine Glioma Cell lines
Supplementary Figure 1 from MetMAb, the One-Armed 5D5 Anti-c-Met Antibody, Inhibits Orthotopic Pancreatic Tumor Growth and Improves Survival
Supplementary Figure S2. Number of histogram analyses of tumor imaging data between January 1990 and October 2014.
PDF file - 58K, Bioluminescence Analysis of U87 Intracranial Tumors in Mice at the Start of Treatment (Day 7 post-surgery) and at the End of Study (Day 28)
Introduction: Total metabolic tumor volume (TMTV) holds promise as a method for quantifying tumor burden in patients with F-18 fluorodeoxyglucose (FDG)-avid lymphomas and, with further validation, has potential as a prognostic biomarker. We have developed a deep learning model for the automatic detection of lesions and quantification of TMTV from FDG-positron emission tomography/computed tomography (FDG-PET/CT) scans (Jemaa et al., 2020, 2022). We aimed to further evaluate the model and identify factors that may influence the performance of lesion detection and TMTV quantification in patients with diffuse large B-cell lymphoma (DLBCL) and follicular lymphoma (FL). Methods: The test data set was compiled using baseline and post-treatment FDG-PET/CT scans from the phase 3 GOYA (NCT01287741) and GALLIUM (NCT01332968) clinical trials, which included 166 patients with DLBCL and 201 patients with FL, respectively. The model was trained using an independent data set from GOYA (n = 836). FDG-PET/CT images were assessed manually (mTMTV) and by the algorithm (aTMTV). Pearson’s correlation coefficient (r) was used to determine the overall performance of aTMTV versus mTMTV. Bias was assessed using the slope and intercept from a weighted Deming regression. Lesion detection performance was evaluated based on sensitivity (the proportion of mTMTV-detected lesions identified by aTMTV) and positive predictive value (PPV; the proportion of aTMTV-detected lesions identified by mTMTV). Performance was compared among patient populations with different demographics and clinical characteristics, and across images from different PET/CT scanner manufacturers. Results: aTMTV quantification highly correlated with mTMTV in the test data set (n = 367; Figure A). No systematic bias was observed between aTMTV and mTMTV (slope, 1.06 [95% CI: 1.02, 1.09]; intercept, –0.27 [95% CI: –0.52, –0.03]; mean difference between methods, 0.10 [standard deviation (SD): 1.15]). Agreement between aTMTV and mTMTV was consistent among patients with different baseline demographics and clinical characteristics, and across scans from different PET/CT scanner manufacturers. Overall mean sensitivity and PPV for lesion detection were both >0.8 (Figure B). Performance was lower for lesions ≤10 mL (mean sensitivity, 0.67; mean PPV, 0.72) than for lesions >10 mL (mean sensitivity and PPV >0.95). Conclusions: The aTMTV algorithm demonstrated good performance for the measurement of TMTV in patients with non-Hodgkin lymphoma and good generalizability across patient subpopulations and PET/CT scanner manufacturers. Reduced algorithm performance for small lesions (≤10 mL) may be the result of higher variability among readers in the determination of small lesions; future work aims to refine and optimize performance of the algorithm for use as a prognostic tool in clinical practice. The research was funded by: This study was funded by F. Hoffmann-La Roche AG, Basel, Switzerland. Medical writing support was provided by Liv Hayward PhD of PharmaGenesis Cardiff, Cardiff, UK and funded by F. Hoffmann-La Roche AG. Keywords: Diagnostic and Prognostic Biomarkers, Indolent non-Hodgkin lymphoma, PET-CT Conflicts of interests pertinent to the abstract. T. Xu Employment or leadership position: F. Hoffmann-La Roche AG Stock ownership: F. Hoffmann-La Roche AG S. Jemaa Employment or leadership position: Genentech, Inc Stock ownership: F. Hoffmann-La Roche AG Other remuneration: Genentech, Inc M. Kumar Employment or leadership position: Genentech, Inc S. Balasubramanian Employment or leadership position: Genentech, Inc. Stock ownership: F. Hoffman-La Roche AG and Genentech, Inc. A. Shamas-Din Employment or leadership position: F. Hoffmann-La Roche AG S. Lyalina Employment or leadership position: F. Hoffmann-La Roche AG and Genentech, Inc. Stock ownership: F. Hoffmann-La Roche AG S. Ounadjela Employment or leadership position: Genentech, Inc. B. Malik Employment or leadership position: Genentech, Inc. Stock ownership: F. Hoffmann-La Roche AG J. Lee Employment or leadership position: Genentech, Inc. Stock ownership: F. Hoffmann-La Roche AG, YungShin Global Holding and Yung Zip Chemical S. Figueroa Morales Employment or leadership position: F. Hoffmann-La Roche AG Stock ownership: F. Hoffmann-La Roche AG T. Nielsen Employment or leadership position: F. Hoffmann-La Roche AG Stock ownership: F. Hoffmann-La Roche AG R. A. D. Carano Employment or leadership position: Genentech, Inc. Stock ownership: F. Hoffmann-La Roche AG Other remuneration: F. Hoffmann-La Roche AG L. Kostakoglu Consultant or advisory role: F. Hoffmann-La Roche AG and Genentech, Inc. W. Capra Employment or leadership position: Genentech, Inc. Stock ownership: F. Hoffmann-La Roche.
PDF file - 60K, Plasma Concentrations, Brain Concentrations and Brain-to-Plasma Ratio Measured at the End of Study (42 days) 2 hours Following the Last Daily PO Administration of GNE-317 (40 mg/kg), GDC-0941 (250 mg/kg) or GDC-0980 (10 mg/kg) to GS2 Tumor-Bearing Mice
Supplementary Figure 3 from MetMAb, the One-Armed 5D5 Anti-c-Met Antibody, Inhibits Orthotopic Pancreatic Tumor Growth and Improves Survival
PDF file, 28K, Effects of 30D8 on CXCL12 and VEGF serum levels in HM7 tumor bearing mice.