PURPOSE:Neurofibromatosis type 2 (NF2) is a tumor predisposition syndrome characterized by bilateral vestibular schwannomas (VSs) resulting in deafness and brainstem compression. This study evaluated efficacy and biomarkers of bevacizumab activity for NF2-associated progressive and symptomatic VSs.PATIENTS AND METHODS:Bevacizumab 7.5 mg/kg was administered every 3 weeks for 46 weeks, followed by 24 weeks of surveillance after treatment with the drug. The primary end point was hearing response defined by word recognition score (WRS). Secondary end points included toxicity, tolerability, imaging response using volumetric magnetic resonance imaging analysis, durability of response, and imaging and blood biomarkers.RESULTS:Fourteen patients (estimated to yield > 90% power to detect an alternative response rate of 50% at alpha level of 0.05) with NF2, with a median age of 30 years (range, 14 to 79 years) and progressive hearing loss in the target ear (median baseline WRS, 60%; range 13% to 82%), were enrolled. The primary end point, confirmed hearing response (improvement maintained ≥ 3 months), occurred in five (36%) of 14 patients (95% CI, 13% to 65%; P < .001). Eight (57%) of 14 patients had transient hearing improvement above the 95% CI for WRS. No patients experienced hearing decline. Radiographic response was seen in six (43%) of 14 target VSs. Three grade 3 adverse events, hypertension (n = 2) and immune-mediated thrombocytopenic purpura (n = 1), were possibly related to bevacizumab. Bevacizumab treatment was associated with decreased free vascular endothelial growth factor (not bound to bevacizumab) and increased placental growth factor in plasma. Hearing responses were inversely associated with baseline plasma hepatocyte growth factor (P = .019). Imaging responses were associated with high baseline tumor vessel permeability and elevated blood levels of vascular endothelial growth factor D and stromal cell-derived factor 1α (P = .037 and .025, respectively).CONCLUSION:Bevacizumab treatment resulted in durable hearing response in 36% of patients with NF2 and confirmed progressive VS-associated hearing loss. Imaging and plasma biomarkers showed promising associations with response that should be validated in larger studies.
Deep reinforcement learning(DRL) is increasingly being explored in medical imaging. However, the environments for medical imaging tasks are constantly evolving in terms of imaging orientations, imaging sequences, and pathologies. To that end, we developed a Lifelong DRL framework, SERIL to continually learn new tasks in changing imaging environments without catastrophic forgetting. SERIL was developed using selective experience replay based lifelong learning technique for the localization of five anatomical landmarks in brain MRI on a sequence of twenty-four different imaging environments. The performance of SERIL, when compared to two baseline setups: MERT(multi-environment-best-case) and SERT(single-environment-worst-case) demonstrated excellent performance with an average distance of $9.90\pm7.35$ pixels from the desired landmark across all 120 tasks, compared to $10.29\pm9.07$ for MERT and $36.37\pm22.41$ for SERT($p<0.05$), demonstrating the excellent potential for continuously learning multiple tasks across dynamically changing imaging environments.
INTRODUCTION:Although magnetic resonance imaging (MRI) has been used to selectively characterize patterns of muscle involvement in myotonic dystrophy type 2 (DM2), the relationship between quantitative MRI-derived muscle composition and functional performance remains incompletely defined. This study aimed to characterize the imaging phenotype in DM2 using whole-body MRI scans and to evaluate the association between MRI-based muscle measurements and clinical endpoints. METHODS:Clinical endpoints collected included strength, physical function, and patient-reported outcomes. AI-based algorithms were utilized to segment 38 muscles bilaterally from Dixon MRI scans collected in a prospective, cross-sectional study. The segmentations were used to calculate muscle fat fraction (MFF) and contractile volume (CV) for each muscle. Associations between MRI measurements and clinical outcomes were examined using Spearman correlations. RESULTS:Adult patients (6 females, 6 males) with DM2 and no contraindications to MRI were enrolled at a single center. The mean age was 52.5 ± 15.4 years and mean disease duration was 12.5 ± 7.8 years. The gluteus maximus and erector spinae muscles had the highest average MFF. Mean MFF in fully captured muscles ranged from 7.4% to 36.5%. Higher MFFs were moderately associated with shorter 6-min walk distances (ρ = -0.62, p = 0.03) and slower 10-m walk/run times (ρ = 0.61, p = 0.04). Higher CVs were moderately associated with longer 6-min walk distance (ρ = 0.67, p = 0.02) and faster 10-m walk/run (ρ = -0.67, p = 0.02). DISCUSSION:These results identify a pattern of proximal muscle involvement and moderate associations between MFFs and CVs with physical function. Larger longitudinal studies are needed to validate the utility of imaging biomarkers in DM2.
Self-supervised learning has revolutionized medical imaging by enabling efficient and generalizable feature extraction from large-scale unlabeled datasets. Recently, self-supervised foundation models have been extended to three-dimensional (3D) computed tomography (CT) data, generating compact, information-rich embeddings with 1408 features that achieve state-of-the-art performance on downstream tasks such as intracranial hemorrhage detection and lung cancer risk forecasting. However, these embeddings have been shown to encode demographic information, such as age, sex, and race, which poses a significant risk to the fairness of clinical applications. In this work, we propose a Variation Autoencoder (VAE) based adversarial debiasing framework to transform these embeddings into a new latent space where demographic information is no longer encoded, while maintaining the performance of critical downstream tasks. We validated our approach on the NLST lung cancer screening dataset, demonstrating that the debiased embeddings effectively eliminate multiple encoded demographic information and improve fairness without compromising predictive accuracy for lung cancer risk at 1-year and 2-year intervals. Additionally, our approach ensures the embeddings are robust against adversarial bias attacks. These results highlight the potential of adversarial debiasing techniques to ensure fairness and equity in clinical applications of self-supervised 3D CT embeddings, paving the way for their broader adoption in unbiased medical decision-making.
In addition to focal lesions, diffusely abnormal white matter (DAWM) is seen on brain MRI of multiple sclerosis (MS) patients and may represent early or distinct disease processes. The role of MRI-observed DAWM is understudied due to a lack of automated assessment methods. Supervised deep learning (DL) methods are highly capable in this domain, but require large sets of labeled data. To overcome this challenge, a DL-based network (DAWM-Net) was trained using semi-supervised learning on a limited set of labeled data for segmentation of DAWM, focal lesions, and normal-appearing brain tissues on multiparametric MRI. DAWM-Net segmentation performance was compared to a previous intensity thresholding-based method on an independent test set from expert consensus (N = 25). Segmentation overlap by Dice Similarity Coefficient (DSC) and Spearman correlation of DAWM volumes were assessed. DAWM-Net showed DSC > 0.93 for normal-appearing brain tissues and DSC > 0.81 for focal lesions. For DAWM-Net, the DAWM DSC was 0.49 +/- 0.12 with a moderate volume correlation (rho = 0.52, p < 0.01). The previous method showed lower DAWM DSC of 0.26 +/- 0.08 and lacked a significant volume correlation (rho = 0.23, p = 0.27). These results demonstrate the feasibility of DL-based DAWM auto-segmentation with semi-supervised learning. This tool may facilitate future investigation of the role of DAWM in MS.
Purpose To describe the design, conduct, and results of the Breast Multiparametric MRI for prediction of neoadjuvant chemotherapy Response (BMMR2) challenge. Materials and Methods The BMMR2 computational challenge opened on May 28, 2021, and closed on December 21, 2021. The goal of the challenge was to identify image-based markers derived from multiparametric breast MRI, including diffusion-weighted imaging (DWI) and dynamic contrast-enhanced (DCE) MRI, along with clinical data for predicting pathologic complete response (pCR) following neoadjuvant treatment. Data included 573 breast MRI studies from 191 women (mean age [±SD], 48.9 years ± 10.56) in the I-SPY 2/American College of Radiology Imaging Network (ACRIN) 6698 trial (ClinicalTrials.gov: NCT01042379). The challenge cohort was split into training (60%) and test (40%) sets, with teams blinded to test set pCR outcomes. Prediction performance was evaluated by area under the receiver operating characteristic curve (AUC) and compared with the benchmark established from the ACRIN 6698 primary analysis. Results Eight teams submitted final predictions. Entries from three teams had point estimators of AUC that were higher than the benchmark performance (AUC, 0.782 [95% CI: 0.670, 0.893], with AUCs of 0.803 [95% CI: 0.702, 0.904], 0.838 [95% CI: 0.748, 0.928], and 0.840 [95% CI: 0.748, 0.932]). A variety of approaches were used, ranging from extraction of individual features to deep learning and artificial intelligence methods, incorporating DCE and DWI alone or in combination. Conclusion The BMMR2 challenge identified several models with high predictive performance, which may further expand the value of multiparametric breast MRI as an early marker of treatment response. Clinical trial registration no. NCT01042379 Keywords: MRI, Breast, Tumor Response Supplemental material is available for this article. © RSNA, 2024.
One vision of a future artificial intelligence (AI) is where many separate units can learn independently over a lifetime and share their knowledge with each other. The synergy between lifelong learning and sharing has the potential to create a society of AI systems, as each individual unit can contribute to and benefit from the collective knowledge. Essential to this vision are the abilities to learn multiple skills incrementally during a lifetime, to exchange knowledge among units via a common language, to use both local data and communication to learn, and to rely on edge devices to host the necessary decentralized computation and data. The result is a network of agents that can quickly respond to and learn new tasks, that collectively hold more knowledge than a single agent and that can extend current knowledge in more diverse ways than a single agent. Open research questions include when and what knowledge should be shared to maximize both the rate of learning and the long-term learning performance. Here we review recent machine learning advances converging towards creating a collective machine-learned intelligence. We propose that the convergence of such scientific and technological advances will lead to the emergence of new types of scalable, resilient and sustainable AI systems. An emerging research area in AI is developing multi-agent capabilities with collections of interacting AI systems. Andrea Soltoggio and colleagues develop a vision for combining such approaches with current edge computing technology and lifelong learning advances. The envisioned network of AI agents could quickly learn new tasks in open-ended applications, with individual AI agents independently learning and contributing to and benefiting from collective knowledge.
Facioscapulohumeral muscular dystrophy (FSHD) affects roughly 1 in 7500 individuals. While at the population level there is a general pattern of affected muscles, there is substantial heterogeneity in muscle expression across- and within-patients. There can also be substantial variation in the pattern of fat and water signal intensity within a single muscle. While quantifying individual muscles across their full length using magnetic resonance imaging (MRI) represents the optimal approach to follow disease progression and evaluate therapeutic response, the ability to automate this process has been limited. The goal of this work was to develop and optimize an artificial intelligence-based image segmentation approach to comprehensively measure muscle volume, fat fraction, fat fraction distribution, and elevated short-tau inversion recovery signal in the musculature of patients with FSHD. Intra-rater, inter-rater, and scan-rescan analyses demonstrated that the developed methods are robust and precise. Representative cases and derived metrics of volume, cross-sectional area, and 3D pixel-maps demonstrate unique intramuscular patterns of disease. Future work focuses on leveraging these AI methods to include upper body output and aggregating individual muscle data across studies to determine best-fit models for characterizing progression and monitoring therapeutic modulation of MRI biomarkers.
Self-supervised foundation models have recently been successfully extended to encode three-dimensional (3D) computed tomography (CT) images, with excellent performance across several downstream tasks, such as intracranial hemorrhage detection and lung cancer risk forecasting. However, as self-supervised models learn from complex data distributions, questions arise concerning whether these embeddings capture demographic information, such as age, sex, or race. Using the National Lung Screening Trial (NLST) dataset, which contains 3D CT images and demographic data, we evaluated a range of classifiers: softmax regression, linear regression, linear support vector machine, random forest, and decision tree, to predict sex, race, and age of the patients in the images. Our results indicate that the embeddings effectively encoded age and sex information, with a linear regression model achieving a root mean square error (RMSE) of 3.8 years for age prediction and a softmax regression model attaining an AUC of 0.998 for sex classification. Race prediction was less effective, with an AUC of 0.878. These findings suggest a detailed exploration into the information encoded in self-supervised learning frameworks is needed to help ensure fair, responsible, and patient privacy-protected healthcare AI.
Objectives: Glioblastomas (GBM) are the most common primary invasive neoplasms of the brain. Distinguishing between lesion recurrence and different types of treatment related changes in patients with GBM remains challenging using conventional MRI imaging techniques. Therefore, accurate and precise differentiation between true progression or pseudoresponse is crucial in deciding on the appropriate course of treatment. This retrospective study investigated the potential of apparent diffusion coefficient (ADC) map values derived from diffusion-weighted imaging (DWI) as a noninvasive method to increase diagnostic accuracy in treatment response. Methods: A cohort of 21 glioblastoma patients (mean age: 59.2 ± 11.8, 12 Male, 9 Female) that underwent treatment with bevacizumab were selected. The ADC values were calculated from the DWI images obtained from a standardized brain protocol across 1.5-T and 3-T MRI scanners. Ratios were calculated for rADC values. Lesions were classified as bevacizumab-induced cytotoxicity based on characteristic imaging features (well-defined regions of restricted diffusion with persistent diffusion restriction over the course of weeks without tissue volume loss and absence of contrast enhancement). The rADC value was compared to these values in radiation necrosis and recurrent lesions, which were concluded in our prior study. The nonparametric Wilcoxon signed rank test with p < 0.05 was used for significance. Results: The mean ± SD age of the selected patients was 59.2 ± 11.8. ADC values and corresponding mean rADC values for bevacizumab-induced cytotoxicity were 248.1 ± 67.2 and 0.39 ± 0.10, respectively. These results were compared to the ADC values and corresponding mean rADC values of tumor progression and radiation necrosis. Significant differences between rADC values were observed in all three groups (p < 0.001). Bevacizumab-induced cytotoxicity had statistically significant lower ADC values compared to both tumor recurrence and radiation necrosis. Conclusion: The study demonstrates the potential of ADC values as noninvasive imaging biomarkers for differentiating recurrent glioblastoma from radiation necrosis and bevacizumab-induced cytotoxicity.
Background: Identifying active lesions in magnetic resonance imaging (MRI) is crucial for the diagnosis and treatment planning of multiple sclerosis (MS). Active lesions on MRI are identified following the administration of Gadolinium-based contrast agents (GBCAs). However, recent studies have reported that repeated administration of GBCA results in the accumulation of Gd in tissues. In addition, GBCA administration increases health care costs. Thus, reducing or eliminating GBCA administration for active lesion detection is important for improved patient safety and reduced healthcare costs. Current state-of-the-art methods for identifying active lesions in brain MRI without GBCA administration utilize data-intensive deep learning methods. Objective: To implement nonlinear dimensionality reduction (NLDR) methods, locally linear embedding (LLE) and isometric feature mapping (Isomap), which are less data-intensive, for automatically identifying active lesions on brain MRI in MS patients, without the administration of contrast agents. Materials and Methods: Fluid-attenuated inversion recovery (FLAIR), T2-weighted, proton density-weighted, and pre- and post-contrast T1-weighted images were included in the multiparametric MRI dataset used in this study. Subtracted pre- and post-contrast T1-weighted images were labeled by experts as active lesions (ground truth). Unsupervised methods, LLE and Isomap, were used to reconstruct multiparametric brain MR images into a single embedded image. Active lesions were identified on the embedded images and compared with ground truth lesions. The performance of NLDR methods was evaluated by calculating the Dice similarity (DS) index between the observed and identified active lesions in embedded images. Results: LLE and Isomap, were applied to 40 MS patients, achieving median DS scores of 0.74 ± 0.1 and 0.78 ± 0.09, respectively, outperforming current state-of-the-art methods. Conclusions: NLDR methods, Isomap and LLE, are viable options for the identification of active MS lesions on non-contrast images, and potentially could be used as a clinical decision tool.
Federated learning is an exciting area within machine learning that allows cross-silo training of large-scale machine learning models on disparate or similar tasks in a privacy-preserving manner. However, conventional federated learning frameworks require a synchronous training schedule and are incapable of lifelong learning. To that end, we propose an asynchronous decentralized federated lifelong learning (ADFLL) method that allows agents in the system to asynchronously and continually learn from their own previous experiences and others', thus overcoming the potential drawbacks of conventional federated learning. We evaluate the ADFLL framework in two experimental setups for deep reinforcement learning (DRL) based landmark localization across different imaging modalities, orientations, and sequences. The ADFLL was compared to central aggregation and conventional lifelong learning for upper-bound comparison and with a conventional DRL model for lower-bound comparison. Across all the experiments, ADFLL demonstrated excellent capability to collaboratively learn all tasks across all the agents compared to the baseline models in in-distribution and out-of-distribution test sets. In conclusion, we provide a flexible, efficient, and robust federated lifelong learning framework that can be deployed in real-world applications.
Federated learning is a recent development in the machine learning area that allows a system of devices to train on one or more tasks without sharing their data to a single location or device. However, this framework still requires a centralized global model to consolidate individual models into one, and the devices train synchronously, which both can be potential bottlenecks for using federated learning. In this paper, we propose a novel method of asynchronous decentralized federated lifelong learning (ADFLL) method that inherits the merits of federated learning and can train on multiple tasks simultaneously without the need for a central node or synchronous training. Thus, overcoming the potential drawbacks of conventional federated learning. We demonstrate excellent performance on the brain tumor segmentation (BRATS) dataset for localizing the left ventricle on multiple image sequences and image orientation. Our framework allows agents to achieve the best performance with a mean distance error of 7.81, better than the conventional all-knowing agent's mean distance error of 11.78, and significantly (p=0.01) better than a conventional lifelong learning agent with a distance error of 15.17 after eight rounds of training. In addition, all ADFLL agents have comparable or better performance than a conventional LL agent. In conclusion, we developed an ADFLL framework with excellent performance and speed-up compared to conventional RL agents.
We introduce tumor connectomics, a novel MRI-based complex graph theory framework that describes the intricate network of relationships within the tumor and surrounding tissue, and combine this with multiparametric radiomics (mpRad) in a machine-learning approach to distinguish radiation necrosis (RN) from true progression (TP). Pathologically confirmed cases of RN vs. TP in brain metastases treated with SRS were included from a single institution. The region of interest was manually segmented as the single largest diameter of the T1 post-contrast (T1C) lesion plus the corresponding area of T2 FLAIR hyperintensity. There were 40 mpRad features and 6 connectomics features extracted, as well as 5 clinical and treatment factors. We developed an Integrated Radiomics Informatics System (IRIS) based on an Isomap support vector machine (IsoSVM) model to distinguish TP from RN using leave-one-out cross-validation. Class imbalance was resolved with differential misclassification weighting during model training using the IRIS. In total, 135 lesions in 110 patients were analyzed, including 43 cases (31.9%) of pathologically proven RN and 92 cases (68.1%) of TP. The top-performing connectomics features were three centrality measures of degree, betweenness, and eigenvector centralities. Combining these with the 10 top-performing mpRad features, an optimized IsoSVM model was able to produce a sensitivity of 0.87, specificity of 0.84, AUC-ROC of 0.89 (95% CI: 0.82-0.94), and AUC-PR of 0.94 (95% CI: 0.87-0.97).
While Deep Reinforcement Learning has been widely researched in medical imaging, the training and deployment of these models usually require powerful GPUs. Since imaging environments evolve rapidly and can be generated by edge devices, the algorithm is required to continually learn and adapt to changing environments, and adjust to low-compute devices. To this end, we developed three image coreset algorithms to compress and denoise medical images for selective experience replayed-based lifelong reinforcement learning. We implemented neighborhood averaging coreset, neighborhood sensitivity-based sampling coreset, and maximum entropy coreset on full-body DIXON water and DIXON fat MRI images. All three coresets produced 27x compression with excellent performance in localizing five anatomical landmarks: left knee, right trochanter, left kidney, spleen, and lung across both imaging environments. Maximum entropy coreset obtained the best performance of $11.97\pm 12.02$ average distance error, compared to the conventional lifelong learning framework's $19.24\pm 50.77$.
COVID-19 is an ongoing global health pandemic. Although COVID-19 can be diagnosed with various tests such as PCR, these tests do not establish pulmonary disease burden. Whereas point-of-care lung ultrasound (POCUS) can directly assess the severity of characteristic pulmonary findings of COVID-19, the advantage of using US is that it is inexpensive, portable, and widely available for use in many clinical settings. For automated assessment of pulmonary findings, we have developed an unsupervised learning technique termed the calculated lung ultrasound (CLU) index. The CLU can quantify various types of lung findings, such as A or B lines, consolidations, and pleural effusions, and it uses these findings to calculate a CLU index score, which is a quantitative measure of pulmonary disease burden. This is accomplished using an unsupervised, patient-specific approach that does not require training on a large dataset. The CLU was tested on 52 lung ultrasound examinations from several institutions. CLU demonstrated excellent concordance with radiologist findings in different pulmonary disease states. Given the global nature of COVID-19, the CLU would be useful for sonographers and physicians in resource-strapped areas with limited ultrasound training and diagnostic capacities for more accurate assessment of pulmonary status.