Accurate segmentation of pediatric brain tumors in MRI is essential for diagnosis, treatment planning, and response assessment. In this study, we investigate uncertainty-aware segmentation of multi-subregion pediatric gliomas using the BraTS-PEDs 2025 dataset. Leveraging the nnUNet v2 framework, we establish a strong baseline and conduct a series of ablation experiments to assess the impact of technical modifications in the design of the pipelines. Key contributions include the use of skull stripping (SynthStrip), atlas-based brain subregion masking (SynthSeg), and an ensemble-based cropping approach guided by whole tumor segmentation. We also evaluate synthetic channel augmentation and multi-task learning with auxiliary skull stripping, though these did not yield performance gains. A voxel-wise ensemble framework is used to identify spatial uncertainty. Results show that whole tumor segmentation and region-specific cropping significantly improve subregion Dice scores, particularly for enhancing and non-enhancing tumor regions. All code, exploratory data analysis outputs, and experiment results are made publicly available to support reproducibility and further research.
INTRODUCTION:Nondiagnostic results after thyroid fine-needle aspiration (FNA) are common and may delay diagnosis. PURPOSE:To identify demographic and sonographic predictors of nondiagnostic cytology at repeat thyroid FNA and to report malignancy rates in this group. MATERIALS AND METHODS:Single-center retrospective cohort including consecutive adult patients who underwent repeat ultrasound-guided FNA of thyroid nodules with prior nondiagnostic cytology (Bethesda I) between 2015 and 2023. Nodule-level ultrasound features were extracted from structured reports. The primary outcome was nondiagnostic cytology at repeat FNA. Multivariable logistic regression with generalized estimating equations was used to account for clustering of nodules. RESULTS:A total of 208 patients with 242 thyroid nodules were included. On repeat FNA, 77 of 242 nodules (31.8%) remained nondiagnostic. In multivariable analysis, older age (odds ratio [OR] per year, 1.03) and nodule composition were independently associated with a nondiagnostic result at repeat FNA. Solid nodules had lower odds of a nondiagnostic result than cystic or mixed cystic-solid nodules (OR, 0.30). Patient sex, maximum diameter, echogenicity, echogenic foci, calcifications, shape, and margins showed no significant association. Cytology suspicious for or diagnostic of malignancy (Bethesda V-VI) was found in 4 nodules (1.6%); however, this estimate may be affected by verification bias. CONCLUSION:Approximately one-third of nodules remained nondiagnostic at repeat FNA. Older age and cystic or mixed nodule composition were independently associated with a higher risk of repeat nondiagnostic cytology. Alternative diagnostic strategies (eg, rapid on-site evaluation, core-needle biopsy, or surveillance) can be considered early in the diagnostic workup of these nodules.
Cerebrovascular disease (CVD) is a leading cause of mortality with a strong link to cognitive impairment and dementia. White matter lesions (WML) are prevalent in CVD and are early markers of vascular compromise, particularly in relation to intraplaque hemorrhage (IPH), an indicator of carotid artery plaque instability. As vascular disease represents a possible treatment window for dementia subjects, this study explores the relationship between hemispheric WML asymmetry and IPH utilizing a large multicenter cohort to find novel biomarkers of disease. FLAIR MRI scans of 264 subjects from the Canadian Atherosclerosis Imaging Network were categorized as IPH positive (IPH+) or IPH negative (IPH-) and WML biomarkers were automatically computed (Figure 1). Biomarkers related to WML prevalence (volume) and WML ischemia and progression (intensity) were extracted: ICV-normalized WML volume (WML-ICV), WML mean intensity (WML-Intensity), and WML intensity ratio (WML-IR). WML asymmetry was assessed via an asymmetry index measure (AIM). Linear mixed models and regression analyses were conducted, with adjustments for age, sex, scanner manufacturer, and stenosis, to evaluate associations between WML biomarkers and IPH status. IPH+ patients exhibited significant rightward asymmetry in WML-ICV (0.0032 ± 0.002, p < 0.05), WML-Intensity (7.26 ± 5.41, p < 0.05), and WML-IR (0.0271 ± 0.0204, p < 0.05); Table 1. IPH+ subjects (left, right or bilateral) had more lesions that were brighter in the right hemisphere. This trend was most pronounced in younger male patients (<65 years), suggesting a high-risk demographic. Regression analysis revealed IPH as a significant predictor of WML asymmetry, with stronger effects observed in subjects with IPH in the right carotid artery. Previous studies suggest more injury in the right hemisphere for subjects with small vessel disease, and this work supports this finding. With rightward WML asymmetry being strongly associated with IPH, this could be reflecting a surrogate marker for overall vascular disease and its contribution to brain health and dementia. Automated WML biomarkers can be used to identify these high-risk patients and guide early interventions for subjects with vascular disease and dementia. Future work should validate these findings in larger, longitudinal datasets to enhance clinical applications.
Pathology provides the definitive diagnosis, and Artificial Intelligence (AI) tools are poised to improve accuracy, inter-rater agreement, and turn-around time (TAT) of pathologists, leading to improved quality of care. A high value clinical application is the grading of Lymph Node Metastasis (LNM) which is used for breast cancer staging and guides treatment decisions. A challenge of implementing AI tools widely for LNM classification is domain shift, where Out-of-Distribution (OOD) data has a different distribution than the In-Distribution (ID) data used to train the model, resulting in a drop in performance in OOD data. This work proposes a novel clustering and sampling method to automatically curate training datasets in an unsupervised manner with the aim of improving model generalization abilities. To evaluate the generalization performance of the proposed models, we applied a novel use of the Two One-sided Tests (TOST) method. This method examines whether the performance on ID and OOD data is equivalent, serving as a proxy for generalization. We provide the first evidence for computing equivalence margins that are data-dependent, which reduces subjectivity. The proposed framework shows the ensembled models constructed from models that generalized across both tumor and normal patches enhanced performance, achieving an F1 score of 0.81 for LNM classification on unseen ID and OOD samples. Interactive viewing of slide-level segmentations can be accessed on PathcoreFlow™ through https://web.pathcore.com/folder/18555?s=QTJVHJuhrfe5 . Segmentation models are available at https://github.com/IAMLAB-Ryerson/OOD-Generalization-LNM .
This study investigates the feasibility of replacing expert ratings with those from trained students in medical imaging research, specifically focusing on adult knee recess distension in ultrasound images. Expert radiologists provide high precision, but their limited availability and high cost can restrict scalability for large datasets. Ten university students with limited prior experience received structured training from a professional sonographer. The training included foundational knowledge of knee recess distension, segmentation practice using Label Studio software, and guidance on submission protocols. Students segmented ultrasound images from 25 patients diagnosed with recess distension, with each student performing three segmentations per image. An expert radiologist segmented the same images to establish a benchmark. Metrics such as Coefficient of Variation (CV), Intraclass Correlation Coefficient (ICC), and Root Mean Square Error (RMSE) were used to assess the accuracy and consistency of the students’ ratings compared to the expert. A simulation further evaluated the impact of variability and aggregation size on accuracy. Most students achieved ICC values above 0.80, indicating good to excellent agreement with the expert. Some students showed moderate agreement (ICC = 0.76) and higher RMSE values, reflecting variability in performance. The simulation revealed that aggregation reduces RMSE, though it eventually reaches a saturation point. For low-variability student ratings (CV = 0.55), accuracy comparable to expert benchmarks was achievable with smaller groups. Higher variability required larger groups, and in some cases, the lowest expert CV benchmarks (CV = 0.2) were unattainable. Trained student raters have the potential to serve as cost-effective alternatives to expert radiologists in large-scale imaging studies. By optimizing training and leveraging aggregation, student raters can achieve accuracy approaching expert levels, offering a scalable solution for resource-limited settings in medical imaging research.
Current methods for performing 3D reconstruction and novel view synthesis (NVS) in ultrasound imaging data often face severe artifacts when training NeRF-based approaches. The artifacts produced by current approaches differ from NeRF floaters in general scenes because of the unique nature of ultrasound capture. Furthermore, existing models fail to produce reasonable 3D reconstructions when ultrasound data is captured or obtained casually in uncontrolled environments, which is common in clinical settings. Consequently, existing reconstruction and NVS methods struggle to handle ultrasound motion, fail to capture intricate details, and cannot model transparent and reflective surfaces. In this work, we introduced NeRF-US, which incorporates 3D-geometry guidance for border probability and scattering density into NeRF training, while also utilizing ultrasound-specific rendering over traditional volume rendering. These 3D priors are learned through a diffusion model. Through experiments conducted on our new "Ultrasound in the Wild" dataset, we observed accurate, clinically plausible, artifact-free reconstructions.
ABSTRACT Background: Prognosticating outcomes for traumatic brain injury (TBI) patients is challenging due to the required specialized skills and variability among clinicians. Recent attempts to standardize TBI prognosis have leveraged machine learning (ML) methodologies. This study evaluates the necessity and influence of ML-assisted TBI prognostication through healthcare professionals’ perspectives via focus group discussions. Methods: Two virtual focus groups included ten key TBI care stakeholders (one neurosurgeon, two emergency clinicians, one internist, two radiologists, one registered nurse, two researchers in ML and healthcare and one patient representative). They answered six open-ended questions about their perceptions and potential ML use in TBI prognostication. Transcribed focus group discussions were thematically analyzed using qualitative data analysis software. Results: The study captured diverse perceptions and interests in TBI prognostication across clinical specialties. Notably, certain clinicians who currently do not prognosticate expressed an interest in doing so independently provided they had access to ML support. Concerns included ML’s accuracy and the need for proficient ML researchers in clinical settings. The consensus suggested using ML as a secondary consultation tool and promoting collaboration with internal or external research resources. Participants believed ML prognostication could enhance disposition planning and standardize care regardless of clinician expertise or injury severity. There was no evidence of perceived bias or interference during the discussions. Conclusion: Our findings revealed an overall positive attitude toward ML-based prognostication. Despite raising multiple concerns, the focus group discussions were particularly valuable in underscoring the potential of ML in democratizing and standardizing TBI prognosis practices.
Computed tomography (CT) is an important imaging modality for guiding prognostication in patients with traumatic brain injury (TBI). However, because of the specialized expertise necessary, timely and dependable TBI prognostication based on CT imaging remains challenging. This study aimed to enhance the efficiency and reliability of TBI prognostication by employing machine learning (ML) techniques on CT images. A retrospective analysis was conducted on the Collaborative European NeuroTrauma Effectiveness Research in TBI (CENTER-TBI) data set (n = 1016). An ML-driven binary classifier was developed to predict favorable or unfavorable outcomes at 6 months post-injury. The prognostic performance was assessed using the area under the curve (AUC) over fivefold cross-validation and compared with conventional models that depend on clinical variables and CT scoring systems. An external validation was performed using the Comparative Indian Neurotrauma Effectiveness Research in Traumatic Brain Injury (CINTER-TBI) data set (n = 348). The developed model achieved superior performance without the necessity for manual CT assessments (AUC = 0.846 [95% CI: 0.843-0.849]) compared with the model based on the clinical and laboratory variables (AUC = 0.817 [95% CI: 0.814-0.820]) and established CT scoring systems requiring manual interpretations (AUC = 0.829 [95% CI: 0.826-0.832] for Marshall and 0.838 [95% CI: 0.835-0.841] for International Mission for Prognosis and Analysis of Clinical Trials in TBI [IMPACT]). The external validation demonstrated the prognostic capacity of the developed model to be significantly better (AUC = 0.859 [95% CI: 0.857-0.862]) than the model using clinical variables (AUC = 0.809 [95% CI: 0.798-0.820]). This study established an ML-based model that provides efficient and reliable TBI prognosis based on CT scans, with potential implications for earlier intervention and improved patient outcomes.
Background: Recurrent hemarthrosis and resultant hemophilic arthropathy are significant causes of morbidity in persons with hemophilia, despite the marked evolution of hemophilia care. Prevention, timely diagnosis, and treatment of bleeding episodes are key. However, a physical examination or a patient’s assessment of musculoskeletal pain may not accurately identify a joint bleed. This difficulty is compounded as hemophilic arthropathy progresses. Objectives: Our system aims to utilize artificial intelligence and ultrasonography (US; point-of-care and handheld) to enable providers, and ultimately patients, to detect joint bleeds at the bedside and at home. We aimed to develop and assess the reliability of artificial intelligence algorithms in detecting and segmenting synovial recess distension (SRD; an indicator of disease activity) on US images of adult and pediatric knee, elbow, and ankle joints. Methods: A total of 12,145 joint exams, comprising 61,501 US images from 7 international healthcare centers, were collected. The dataset included healthy participants and adult and pediatric persons with hemophilia, with and without SRD. Images were manually labeled by 2 experts and used to train binary convolutional neural network classifiers and segmentation models. Metrics to evaluate performance included accuracy, sensitivity, specificity, and area under the curve. Results: The algorithms exhibited high performance across all joints and all cohorts. Specifically, the knee model showed an accuracy of 97%, sensitivity of 96%, specificity of 97%, and an area under the curve of 0.97 in SRD. High Dice coefficients (80%-85%) were achieved in segmentation tasks across all joints. Conclusion: This technology could assist with the early detection and management of hemarthrosis in hemophilia.
Background: Machine learning models can provide quick and reliable assessments in place of medical practitioners. With over 50 million adults in the United States suffering from osteoarthritis, there is a need for models capable of interpreting musculoskeletal ultrasound images. However, machine learning requires lots of data, which poses significant challenges in medical imaging. Therefore, we explore two strategies for enriching a musculoskeletal ultrasound dataset independent of these limitations: traditional augmentation and diffusion-based image synthesis. Methods: First, we generate augmented and synthetic images to enrich our dataset. Then, we compare the images qualitatively and quantitatively, and evaluate their effectiveness in training a deep learning model for detecting thickened synovium and knee joint recess distension. Results: Our results suggest that synthetic images exhibit some anatomical fidelity, diversity, and help a model learn representations consistent with human opinion. In contrast, augmented images may impede model generalizability. Finally, a model trained on synthetically enriched data outperforms models trained on un-enriched and augmented datasets. Conclusions: We demonstrate that diffusion-based image synthesis is preferable to traditional augmentation. Our study underscores the importance of leveraging dataset enrichment strategies to address data scarcity in medical imaging and paves the way for the development of more advanced diagnostic tools.
In recent years, Artificial Intelligence has been used to assist healthcare professionals in detecting and diagnosing neurodegenerative diseases. In this study, we propose a methodology to analyze functional Magnetic Resonance Imaging signals and perform classification between Parkinson’s disease patients and healthy participants using Machine Learning algorithms. In addition, the proposed approach provides insights into the brain regions affected by the disease. The functional Magnetic Resonance Imaging from the PPMI and 1000-FCP datasets were pre-processed to extract time series from 200 brain regions per participant, resulting in 11,600 features. Causal Forest and Wrapper Feature Subset Selection algorithms were used for dimensionality reduction, resulting in a subset of features based on their heterogeneity and association with the disease. We utilized Logistic Regression and XGBoost algorithms to perform PD detection, achieving 97.6% accuracy, 97.5% F1 score, 97.9% precision, and 97.7%recall by analyzing sets with fewer than 300 features in a population including men and women. Finally, Multiple Correspondence Analysis was employed to visualize the relationships between brain regions and each group (women with Parkinson, female controls, men with Parkinson, male controls). Associations between the Unified Parkinson’s Disease Rating Scale questionnaire results and affected brain regions in different groups were also obtained to show another use case of the methodology. This work proposes a methodology to (1) classify patients and controls with Machine Learning and Causal Forest algorithm and (2) visualize associations between brain regions and groups, providing high-accuracy classification and enhanced interpretability of the correlation between specific brain regions and the disease across different groups.
Background: Primary small vessel CNS vasculitis (sv-cPACNS) is a challenging inflammatory brain disease in children. Brain biopsy is mandatory to confirm the diagnosis. This study aims to develop and validate a histological scoring tool for diagnosing small vessel CNS vasculitis. Methods: A standardized brain biopsy scoring instrument was developed and applied to consecutive full-thickness brain biopsies of pediatric cases and controls at a single center. Stains included immunohistochemistry and Hematoxylin & Eosin. Nine North American neuropathologists, blinded to patients’ presentation, diagnosis, and therapy, scored de-identified biopsies independently. Results: A total of 31 brain biopsy specimens from children with sv-cPACNS, 11 with epilepsy, and 11 with non-vasculitic inflammatory brain disease controls were included. Angiocentric inflammation in the cortex or white matter increases the likelihood of sv-cPACNS, with odds ratios (ORs) of 3.231 (95CI: 0.914-11.420, p = 0.067) and 3.923 (95CI: 1.13-13.6, p = 0.031). Moderate to severe inflammation in these regions is associated with a higher probability of sv-cPACNS, with ORs of 5.56 (95CI: 1.02-29.47, p = 0.046) in the cortex and 6.76 (95CI: 1.26-36.11, p = 0.025) in white matter. CD3, CD4, CD8, and CD20 cells predominated the inflammatory infiltrate. Reactive endothelium was strongly associated with sv-cPACNS, with an OR of 8.93 (p = 0.001). Features reported in adult sv-PACNS, including granulomas, necrosis, or fibrin deposits, were absent in all biopsies. The presence of leptomeningeal inflammation in isolation was non-diagnostic. Conclusion: Distinct histological features were identified in sv-cPACNS biopsies, including moderate to severe angiocentric inflammatory infiltrates in the cortex or white matter, consisting of CD3, CD4, CD8, and CD20 cells, alongside reactive endothelium with specificity of 95%. In the first study of its kind proposing histological criteria for evaluating brain biopsies, we aim to precisely characterize the type and severity of the inflammatory response in patients with sv-cPACNS; this can enable consolidation of this population to assess outcomes and treatment methodologies comprehensively.
Medical image analysis has become a prominent area where machine learning has been applied. However, high quality, publicly available data is limited either due to patient privacy laws or the time and cost required for experts to annotate images. In this retrospective study, we designed and evaluated a pipeline to generate synthetic labeled polyp images for augmenting medical image segmentation models with the aim of reducing this data scarcity. In particular, we trained diffusion models on the HyperKvasir dataset, comprising 1000 images of polyps in the human GI tract from 2008 to 2016. Qualitative expert review, Fréchet Inception Distance (FID), and Multi-Scale Structural Similarity (MS-SSIM) were tested for evaluation. Additionally, various segmentation models were trained with the generated data and evaluated using Dice score and Intersection over Union. We found that our pipeline produced images more akin to real polyp images based on FID scores, and segmentation performance also showed improvements over GAN methods when trained entirely, or partially, with synthetic data, despite requiring less compute for training. Moreover, the improvement persists when tested on different datasets, showcasing the transferability of the generated images.
This study explores the use of deep generative models to create synthetic ultrasound images for the detection of hemarthrosis in hemophilia patients. Addressing the challenge of sparse datasets in rare disease diagnostics, the study aims to enhance AI model robustness and accuracy through the integration of domain knowledge into the synthetic image generation process. The study employed two ultrasound datasets: a base dataset (Db) of knee recess distension images from non-hemophiliac patients and a target dataset (Dt) of hemarthrosis images from hemophiliac patients. The synthetic generation framework included a content generator (Gc) trained on Db and a context generator (Gs) to adapt these images to match Dt’s context. This approach generated a synthetic target dataset (Ds), primed for AI training in rare disease research. The assessment of synthetic image generation involved expert evaluations, statistical analysis, and the use of domain-invariant perceptual distance and Fréchet inception distance for quality measurement. Expert evaluation revealed that images produced by our synthetic generation framework were comparable to real ones, with no significant difference in overall quality or anatomical accuracy. Additionally, the use of synthetic data in training convolutional neural networks demonstrated robustness in detecting hemarthrosis, especially with limited sample sizes. This study presents a novel approach for generating synthetic ultrasound images for rare disease detection, such as hemarthrosis in hemophiliac knees. By leveraging deep generative models and integrating domain knowledge, the proposed framework successfully addresses the limitations of sparse datasets and enhances AI model training and robustness. The synthetic images produced are of high quality and contribute significantly to AI-driven diagnostics in rare diseases, highlighting the potential of synthetic data in medical imaging.
Neuroimaging has a key role in identifying small-vessel vasculitis from common diseases it mimics, such as multiple sclerosis. Oftentimes, a multitude of these conditions present similarly, and thus diagnosis is difficult. To date, there is no standardized method to differentiate between these diseases. This review identifies and presents existing scoring tools that could serve as a starting point for integrating artificial intelligence/machine learning (AI/ML) into the clinical decision-making process for these rare diseases. A scoping literature review of EMBASE and MEDLINE included 114 articles to evaluate what criteria exist to diagnose small-vessel vasculitis and common mimics. This paper presents the existing criteria of small-vessel vasculitis conditions and mimics them to guide the future integration of AI/ML algorithms to aid in diagnosing these conditions, which present similarly and non-specifically.
Supervised machine learning classification is the most common example of artificial intelligence (AI) in industry and in academic research. These technologies predict whether a series of measurements belong to one of multiple groups of examples on which the machine was previously trained. Prior to real-world deployment, all implementations need to be carefully evaluated with hold-out validation, where the algorithm is tested on different samples than it was provided for training, in order to ensure the generalizability and reliability of AI models. However, established methods for performing hold-out validation do not assess the consistency of the mistakes that the AI model makes during hold-out validation. Here, we show that in addition to standard methods, an enhanced technique for performing hold-out validation—that also assesses the consistency of the sample-wise mistakes made by the learning algorithm—can assist in the evaluation and design of reliable and predictable AI models. The technique can be applied to the validation of any supervised learning classification application, and we demonstrate the use of the technique on a variety of example biomedical diagnostic applications, which help illustrate the importance of producing reliable AI models. The validation software created is made publicly available, assisting anyone developing AI models for any supervised classification application in the creation of more reliable and predictable technologies.
The field of healthcare has witnessed remarkable advancements in recent years, particularly in the care and management of persons with hemophilia. With the advent of artificial intelligence (AI) and the growing importance of ultrasonography in medicine, the future of hemophilia care holds great promise. These developments not only improve the delivery of healthcare but also empower patients, leading to better outcomes and enhanced quality of life. In this article, I will explore the recent developments in hemophilia care, the impact of AI on healthcare delivery, and the significance of ultrasonography in digital health for persons with hemophilia.
Abstract Background Although MRI is a radiation-free imaging modality, it has historically been limited in lung imaging due to inherent technical restrictions. The aim of this study is to explore the performance of lung MRI in detecting solid and subsolid pulmonary nodules using T1 gradient-echo (GRE) (VIBE, Volumetric interpolated breath-hold examination), ultrashort time echo (UTE) and T2 Fast Spin Echo (HASTE, Half fourier Single-shot Turbo spin-Echo). Methods Patients underwent a lung MRI in a 3 T scanner as part of a prospective research project. A baseline Chest CT was obtained as part of their standard of care. Nodules were identified and measured on the baseline CT and categorized according to their density (solid and subsolid) and size (> 4 mm/ ≤ 4 mm). Nodules seen on the baseline CT were classified as present or absent on the different MRI sequences by two thoracic radiologists independently. Interobserver agreement was determined using the simple Kappa coefficient. Paired differences were compared using nonparametric Mann-Whitney U tests. The McNemar test was used to evaluate paired differences in nodule detection between MRI sequences. Results Thirty-six patients were prospectively enrolled. One hundred forty-nine nodules (100 solid/49 subsolid) with mean size 10.8 mm (SD = 9.4) were included in the analysis. There was substantial interobserver agreement (k = 0.7, p = 0.05). Detection for all nodules, solid and subsolid nodules was respectively; UTE: 71.8%/71.0%/73.5%; VIBE: 61.6%/65%/55.1%; HASTE 72.4%/72.2%/72.7%. Detection rate was higher for nodules > 4 mm in all groups: UTE 90.2%/93.4%/85.4%, VIBE 78.4%/88.5%/63.4%, HASTE 89.4%/93.8%/83.8%. Detection of lesions ≤4 mm was low for all sequences. UTE and HASTE performed significantly better than VIBE for detection of all nodules and subsolid nodules (diff = 18.4 and 17.6%, p = < 0.01 and p = 0.03, respectively). There was no significant difference between UTE and HASTE. There were no significant differences amongst MRI sequences for solid nodules. Conclusions Lung MRI shows adequate performance for the detection of solid and subsolid pulmonary nodules larger than 4 mm and can serve as a promising radiation-free alternative to CT.