Background:Pancreatic cancer (PC) is frequently missed on unenhanced CT examinations performed for unrelated clinical indications, where the pancreas is included incidentally and clinical suspicion is low. Purpose:To develop and validate a deep learning-based tool for PC diagnosis and risk stratification on unenhanced CT. Materials and Methods:This retrospective study included 3080 unenhanced CT studies of Taiwanese patients with PC, other pancreatic diseases and normal pancreas between 2004 and 2019 from a tertiary hospital, randomly divided into training, validation, and internal test sets. Unenhanced CT studies from United States institutions were used for external testing. A hybrid convolutional neural network-transformer model was trained for PC diagnosis and risk stratification. Performance was evaluated using sensitivity, specificity, and area under the curve (AUC), with comparisons to 2 radiologists by McNemar's test and exploratory decision curve analysis. Results:The internal dataset included 713 PCs (mean age, 64.6 ± 12.0 years; 384 men), 1661 normal pancreas and 706 other pancreatic diseases. In an exploratory comparison restricted to unenhanced CT (29 PCs, 31 controls), the sensitivity of computer-aided diagnosis (CAD) tool (89.7%, 72.6-97.8) seemed comparable with that of 1 radiologist (86.2%, 68.3-96.1) and higher than another (41.4%, 23.5-61.1); but wide confidence intervals and inter-radiologist variability warrant cautious interpretation. In the internal test set (142 PCs, 474 controls), sensitivity was 90.8% (84.9-95.0) and specificity 93.0% (90.4-95.2) (AUC: 0.98), with sensitivity comparable to radiologist reports based on enhanced and unenhanced CT (95.4%, 90.2-98.3; P = .21). In the external set (42 PCs, 22 controls), sensitivity was 76.2% (60.5-87.9) and specificity 86.4% (65.1-97.1) (AUC: 0.89). The tool stratified cases into 7 risk levels with likelihood ratios ranging from <0.01 to 173.46. Exploratory decision curve analysis suggested potential net benefit across threshold probabilities. Conclusion:This tool may assist in the opportunistic detection and risk stratification of PC on unenhanced CT.
BACKGROUND AND PURPOSE:Accurate MRI-based target delineation for hypopharyngeal squamous cell carcinoma (HPSCC) is clinically important but expertise dependent. We aimed to develop a multicenter-validated tri-sequence deep-learning model, determine whether AI assistance narrows contouring expertise gap, and explore quality-aware low-overlap risk modeling to inform deployment support. MATERIALS AND METHODS:This retrospective study included 727 HPSCC patients from three institutions. A tri-sequence 3D nnU-Net trained on the development cohort (n = 530) was evaluated in the internal test cohort (n = 37), external cohort 1 (n = 109), and external cohort 2 (n = 51) using Dice similarity coefficient (DSC), surface DSC, average symmetric surface distance (ASSD), and mean surface distance (MSD). Clinical utility was assessed in a randomized double-crossover study of the 51-case external cohort 2 involving three junior and three senior radiation oncologists, comparing manual with AI-assisted contouring by DSC, contouring time, Fleiss' κ, and 5-point Likert scores. For exploratory deployment-support analysis, MRI-quality features and auto-segmentation-derived tumor volume were used to characterize domain shift and perform XGBoost-based low-overlap classification (DSC < 0.75). RESULTS:Tri-sequence mean DSC was 0.87 ± 0.11 internally and 0.85 ± 0.14 and 0.82 ± 0.16 in external cohorts. In the reader study, AI assistance increased mean DSC in juniors from 0.73 ± 0.16 to 0.86 ± 0.14 and in seniors from 0.79 ± 0.13 to 0.84 ± 0.15, reduced contouring time by 55%, and improved Fleiss' κ from 0.69 ± 0.12 to 0.86 ± 0.12 (all p < 0.01). The multivariable low-overlap risk model achieved an area under the receiver operating characteristic curve of 0.89 internally and 0.71-0.78 externally. CONCLUSION:Deep-learning-assisted tri-sequence MRI segmentation enabled robust multicenter HPSCC delineation, improved contouring efficiency and consistency, and supports quality-aware analysis in radiotherapy planning.
PURPOSE:The purpose of this study was to develop and validate a computer-aided detection (CAD) tool for the detection of pancreatic cancer (PC) on diagnostic and prediagnostic computed tomography (CT) examinations. MATERIALS AND METHODS:A CAD tool was developed using 2496 contrast-enhanced CT images (596 PCs, 1335 normal pancreas, 565 other pancreatic diseases) from a referral center (October 2004-December 2019) and underwent external validation at two independent institutions (January 2018-December 2020) in a retrospective case-control design. Prediagnostic CT images obtained one to 12 months before the clinical diagnosis of PC, representing clinically challenging or missed images, were collected (November 2004-August 2022) from three referral centers to further evaluate the performance of the CAD tool. Classification performance of the CAD tool was assessed using sensitivity, specificity, and area under the receiver operating characteristic curve (AUC), RESULTS: From internal and external datasets, the diagnostic test sets included 200 PCs and 4998 controls of 4744 patients with normal pancreas and 254 patients with other pancreatic diseases (2448 women and 2750 men; median age, 63 years; age range: 18-101). The CAD tool achievedan AUC of 0.950 (95 % confidence interval [CI]: 0.932-0.968), 90.0 % sensitivity (180 out of 200; 95 % CI: 85.0-93.8), and 87.8 % specificity (4389 out of 4998; 95 % CI: 86.9-88.7) in the diagnosis of PC. For prediagnostic test sets, which included 54 PCs and 118 controls of 89 patients with normal pancreas and 19 patients with other pancreatic diseases (63 women and 99 men; median age, 61 years; age range: 18-99), the sensitivity was 66.7 % (36 out of 54; 95 % CI: 52.5-78.9). Sensitivities for PCs ≤ 2 cm were 77.1 % (27 out of 35; 95 % CI: 59.9-89.6) and 66.7 % (14 out of 21; 95 % CI: 43.0-85.4) in diagnostic and prediagnostic test sets, respectively. CONCLUSION:This CAD tool demonstrates high diagnostic performance for the detection of PC, including for small PC or clinically unrecognized patients.
Contrast-enhanced computed tomography (CT) is recommended in clinical guidelines for managing hepatocellular carcinoma (HCC), yet complete multi-phase acquisition is often unavailable. Here, we present DuoProto, a novel dualbranch, prototype-guided framework that leverages limited multi-phase CT during training to enhance single-phase ER prediction in HCC. DuoProto aligns class-level representations across single- and multi-phase inputs via prototype learning, incorporates clinically informed ranking, and infers with single-phase only. Under conditions of class imbalance and missing phase, DuoProto consistently outperforms existing methods, achieving absolute gains of $3-9 \%$ across all metrics. The proposed framework provides a clinically aligned solution, enabling more informed decision-making. Code is available at https://github.com/idssplab/DuoProto.
INTRODUCTION:Limited linkage between epidermal growth factor receptor (EGFR) mutations and recurrence-predictive radiomic signatures restricts the application of radiomics-guided therapy for brain metastases (BMs) from non-small-cell lung cancer (NSCLC). This study aimed to establish an EGFR-associated radiomic signature (EGFR-RS), compare its consistency with that of conventional whole radiomic features-based radiomic signature (WF-RS), and evaluate its efficacy in predicting local recurrence for BMs treated with radiosurgery. METHODS:Brain magnetic resonance (MR) and computed tomography (CT) images of NSCLC patients with BMs undergoing radiosurgery between 2008 and 2020 were examined. The least absolute shrinkage and selection operator was utilized to select features and develop signatures. Discriminative abilities were assessed using the area under the curve, while univariable and multivariable competing risk regression determined predictors and established a clinical-radiomic model. RESULTS:In total, 318 patients with 759 BMs were enrolled. The EGFR-RS, incorporating 11 MR and six CT EGFR-associated prognostic radiomic features, displayed better consistency, and superior predictive performance than the WF-RS, with C-indices of 0.746 (95 %CI 0.616, 0.876) in the test cohort, compared with 0.655 (95 %CI 0.527, 0.784) for the WF-RS. Multivariable analysis indicated EGFR-RS as the sole significant predictor of local recurrence in both the discovery and test sets (P < 0.001, hazard ratio [HR] = 2.75; and P = 0.01, HR = 2.13, respectively). The clinical-radiomic model (EGFR-RS + EGFR mutation status + BM size) outperformed the clinical model in identifying high-risk lesions with local recurrence (discovery: P < 0.001; HR = 4.54; test: P = 0.002; HR = 5.1). CONCLUSION:The multimodal EGFR-RS, demonstrating better consistency than the WF-RS, effectively predicted the local recurrence of NSCLC BMs.
Computational competitions are the standard for benchmarking medical image analysis algorithms, but they typically use small curated test datasets acquired at a few centers, leaving a gap to the reality of diverse multicentric patient data. To this end, the Federated Tumor Segmentation (FeTS) Challenge represents the paradigm for real-world algorithmic performance evaluation. The FeTS challenge is a competition to benchmark (i) federated learning aggregation algorithms and (ii) state-of-the-art segmentation algorithms, across multiple international sites. Weight aggregation and client selection techniques were compared using a multicentric brain tumor dataset in realistic federated learning simulations, yielding benefits for adaptive weight aggregation, and efficiency gains through client sampling. Quantitative performance evaluation of state-of-the-art segmentation algorithms on data distributed internationally across 32 institutions yielded good generalization on average, albeit the worst-case performance revealed data-specific modes of failure. Similar multi-site setups can help validate the real-world utility of healthcare AI algorithms in the future.
BACKGROUND:Pancreatic ductal adenocarcinoma (PDAC) has the worst prognosis among major cancer types, primarily due to late diagnosis on contrast-enhanced CT. Artificial intelligence (AI) can improve diagnostic performance, but robust benchmarks and reliable comparison to radiologists' performance are scarce. We established an open-source benchmark with the aim of investigating AI systems for PDAC detection on CT and compared them to radiologists' performance, at scale. METHODS:In this international, paired, non-inferiority, confirmatory, observational study (PANORAMA), the AI system was trained and externally validated within an international benchmark, with a cohort of 2310 patients from four tertiary care centres in the Netherlands and the USA for training (n=2224) and tuning (n=86), and a sequestered cohort of 1130 patients from five tertiary care centres (the Netherlands, Sweden, and Norway) for testing. A multi-reader, multi-case observer study with 68 radiologists (40 centres, 12 countries; median 9·0 [IQR 6·0-14·5] years of experience) was conducted on a subset of 391 patients from the testing cohort. The reference standard was established with histopathology and at least 3 years of clinical follow-up. The primary endpoint was the mean area under the receiver operating characteristic curve (AUROC) of the AI system compared to that of radiologists at PDAC detection on CT. The study protocol and statistical plan were prespecified to test non-inferiority (considering a margin of 0·05), followed by superiority towards the AI system. This study is registered with Zenodo (https://doi.org/10.5281/zenodo.10599559) and is complete. FINDINGS:Of the 3440 (1511 [44%] female, 1929 [56%] male; median age 67 [IQR 58-74] years) included patients (Jan 1, 2004 to Dec 31, 2023), 1103 (32%) received a positive PDAC diagnosis. In the sequestered testing cohort of 1130 patients (406 with histologically confirmed PDAC), AI achieved an AUROC of 0·92 (95% CI 0·90-0·93). In the subset of 391 patients (144 [37%] with histologically confirmed PDAC) used for the reader study, AI achieved statistically non-inferior (p<0·0001) and superior (p=0·001) performance with an AUROC of 0·92 (95% CI 0·89-0·94), compared to the pool of 68 participating radiologists, with an AUROC of 0·88 (0·85-0·91). INTERPRETATION:AI demonstrated substantially improved PDAC detection on routine CT scans compared to radiologists on average, showing potential to detect cancer earlier and improve patient outcomes. FUNDING:European Union's Horizon 2020 research and innovation programme.
In medical imaging, developing generalized segmentation models that can handle multiple organs and lesions is crucial. However, the scarcity of fully annotated datasets and strict privacy regulations present significant barriers to data sharing. Federated Learning (FL) allows decentralized model training, but existing FL methods often struggle with partial labeling, leading to model divergence and catastrophic forgetting. We propose ConDistFL, a novel FL framework incorporating conditional distillation to address these challenges. ConDistFL enables effective learning from partially labeled datasets, significantly improving segmentation accuracy across distributed and non-uniform datasets. In addition to its superior segmentation performance, ConDistFL maintains computational and communication efficiency, ensuring its scalability for real-world applications. Furthermore, ConDistFL demonstrates remarkable generalizability, significantly outperforming existing FL methods in out-of-federation tests, even adapting to unseen contrast phases (e.g., non-contrast CT images) in our experiments. Extensive evaluations on 3D CT and 2D chest X-ray datasets show that ConDistFL is an efficient, adaptable solution for collaborative medical image segmentation in privacy-constrained settings.
Deep learning has revolutionized medical imaging, offering advanced methods for accurate diagnosis and treatment planning. The BCLC staging system is crucial for staging Hepatocellular Carcinoma (HCC), a high-mortality cancer. An automated BCLC staging system could significantly enhance diagnosis and treatment planning efficiency. However, we found that BCLC staging, which is directly related to the size and number of liver tumors, aligns well with the principles of the Multiple Instance Learning (MIL) framework. To effectively achieve this, we proposed a new preprocessing technique called Masked Cropping and Padding(MCP), which addresses the variability in liver volumes and ensures consistent input sizes. This technique preserves the structural integrity of the liver, facilitating more effective learning. Furthermore, we introduced Re ViT, a novel hybrid model that integrates the local feature extraction capabilities of Convolutional Neural Networks (CNNs) with the global context modeling of Vision Transformers (ViTs). Re ViT leverages the strengths of both architectures within the MIL framework, enabling a robust and accurate approach for BCLC staging. We will further explore the trade-off between performance and interpretability by employing TopK Pooling strategies, as our model focuses on the most informative instances within each bag.
Purpose To develop two deep learning-based systems for diagnosing and localizing pneumothorax on portable supine chest X-rays (SCXRs). Methods For this retrospective study, images meeting the following inclusion criteria were included: (1) patient age ≥ 20 years; (2) portable SCXR; (3) imaging obtained in the emergency department or intensive care unit. Included images were temporally split into training (1571 images, between January 2015 and December 2019) and testing (1071 images, between January 2020 to December 2020) datasets. All images were annotated using pixel-level labels. Object detection and image segmentation were adopted to develop separate systems. For the detection-based system, EfficientNet-B2, DneseNet-121, and Inception-v3 were the architecture for the classification model; Deformable DETR, TOOD, and VFNet were the architecture for the localization model. Both classification and localization models of the segmentation-based system shared the UNet architecture. Results In diagnosing pneumothorax, performance was excellent for both detection-based (Area under receiver operating characteristics curve [AUC]: 0.940, 95% confidence interval [CI]: 0.907–0.967) and segmentation-based (AUC: 0.979, 95% CI: 0.963–0.991) systems. For images with both predicted and ground-truth pneumothorax, lesion localization was highly accurate (detection-based Dice coefficient: 0.758, 95% CI: 0.707–0.806; segmentation-based Dice coefficient: 0.681, 95% CI: 0.642–0.721). The performance of the two deep learning-based systems declined as pneumothorax size diminished. Nonetheless, both systems were similar or better than human readers in diagnosis or localization performance across all sizes of pneumothorax. Conclusions Both deep learning-based systems excelled when tested in a temporally different dataset with differing patient or image characteristics, showing favourable potential for external generalizability.
OBJECTIVES: We aimed to develop a computer-aided detection (CAD) system to localize and detect the malposition of endotracheal tubes (ETTs) on portable supine chest radiographs (CXRs). DESIGN: This was a retrospective diagnostic study. DeepLabv3+ with ResNeSt50 backbone and DenseNet121 served as the model architecture for segmentation and classification tasks, respectively. SETTING: Multicenter study. PATIENTS: For the training dataset, images meeting the following inclusion criteria were included: 1) patient age greater than or equal to 20 years; 2) portable supine CXR; 3) examination in emergency departments or ICUs; and 4) examination between 2015 and 2019 at National Taiwan University Hospital (NTUH) (NTUH-1519 dataset: 5,767 images). The derived CAD system was tested on images from chronologically (examination during 2020 at NTUH, NTUH-20 dataset: 955 images) or geographically (examination between 2015 and 2020 at NTUH Yunlin Branch [YB], NTUH-YB data set: 656 images) different datasets. All CXRs were annotated with pixel-level labels of ETT and with image-level labels of ETT presence and malposition. INTERVENTIONS: None. MEASUREMENTS AND MAIN RESULTS: For the segmentation model, the Dice coefficients indicated that ETT would be delineated accurately (NTUH-20: 0.854; 95% CI, 0.824-0.881 and NTUH-YB: 0.839; 95% CI, 0.820-0.857). For the classification model, the presence of ETT could be accurately detected with high accuracy (area under the receiver operating characteristic curve [AUC]: NTUH-20, 1.000; 95% CI, 0.999-1.000 and NTUH-YB: 0.994; 95% CI, 0.984-1.000). Furthermore, among those images with ETT, ETT malposition could be detected with high accuracy (AUC: NTUH-20, 0.847; 95% CI, 0.671-0.980 and NTUH-YB, 0.734; 95% CI, 0.630-0.833), especially for endobronchial intubation (AUC: NTUH-20, 0.991; 95% CI, 0.969-1.000 and NTUH-YB, 0.966; 95% CI, 0.933-0.991). CONCLUSIONS: The derived CAD system could localize ETT and detect ETT malposition with excellent performance, especially for endobronchial intubation, and with favorable potential for external generalizability.
Abstract Introduction: Reliability of machine and deep learning models on medical images can be compromised by the presence of image artifacts of variability in image acquisition. This is further exacerbated when validating model performance across multiple institutions, where multiple acquisition protocols and scanner equipment may result in significant numbers of outlier scans. We developed a privacy-preserving outlier identification algorithm via federated MR imaging quality evaluation of a distributed multi-institutional cohort of prostate MRI scans. Methods: This retrospective study utilized T2-weighted prostate MRI scans from 1 public collection (used as a template) and 3 federated institutions. An open-source MR image quality evaluation tool (MRQy) was implemented via Docker within the RhinoHealth Federated Computing Platform (FCP). This enabled remote extraction of 9 specific metadata values and remote computation of 15 quality measurements from each scan, while ensuring confidentiality of patient data by avoiding direct access to patient images and information. Centralized analysis of extracted image quality attributes utilized a rule-based classifier to determine whether individual metadata or measurement values for a given MRI scan fell within or outside predefined bounds: (i) ±20% of template metadata values, (ii) 5th - 95th percentile range of template measurement values. Scans were considered outliers if more than 60% of their image quality attributes fell outside predefined bounds. Results: In total, 630 MRI scans from 3 institutions were analyzed in a federated manner. 344 MRI scans from 1 public collection formed a benchmark cohort to calibrate the rule-based classifier and determine predefined bounds for outlier detection. Centralized analysis of federated image quality attributes showed marked differences in the number of outliers per institution (3%, 19%, 79% for Institutions 1, 2, 3, respectively). Further analysis of Institution 3 determined significant intensity and resolution differences compared to other institutions, requiring extensive post-processing prior to downstream analysis. Conclusion: Federated image quality assessment can allow for distributed, privacy preserving analysis of outliers and thus curate multi-institutional cohorts of consistent quality for downstream machine learning tasks. Further analysis of the impact of quality-based curation on federated machine learning applications is currently underway. Citation Format: Mohsen Hariri, Prathyush Chirra, Malhar Patel, Tal Tiano Einat, Ittai Dayan, Alex Tonetti, Yuval Baror, Tristan Barrett, Nikita Sushentsev, Joshua D. Kaggie, Shuaiyu Yuan, Dufan Wu, Baihui Yu, Zhiliang Lyu, Cheyu Hsu, Weichung Wang, Smitha Krishnamurthi, Satish E. Viswanath. Federated image quality assessment of prostate MRI scans in a multi-institutional setting [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2024; Part 1 (Regular Abstracts); 2024 Apr 5-10; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2024;84(6_Suppl):Abstract nr 2344.
Using the digital image correlation method for on-site measurement, limited by continuous operation conditions and environment, two inevitable challenges are introduced. First, the reference images at the balanced positions needed for DIC analysis are always unknown; and environmental disturbances contaminate the continuously recording images and distort the measurement results. To overcome the two challenges, a preliminary attempt is to adopt the optical flow algorithm to redefine the reference image and reject the outlier video images to suppress the possible significant departures introduced by unknown reference images and environmental disturbances on the DIC-determined dynamic displacement and strain field. In this study, the balloon inflation experiment evaluates the improvement by adopting the optical flow method. The results show that the optical flow can provide additional information that helps redefine the reference image and reject unwanted image frames. After adopting the redefined reference image for DIC calculation, the averaged displacement and strain intuitively show that the balloon is in extension. Meanwhile, the overall frames' averaged displacement and averaged strain can conduct a 3.6% difference for displacement and 5.6% for strain in magnitude as the redefined one is used. According to the results, selecting reference images needs further study to improve the measurement accuracy, particularly for continuous displacement and strain measurement using DIC.
Prompt and correct detection of pulmonary tuberculosis (PTB) is critical in preventing its spread. We aimed to develop a deep learning-based algorithm for detecting PTB on chest X-ray (CXRs) in the emergency department. This retrospective study included 3498 CXRs acquired from the National Taiwan University Hospital (NTUH). The images were chronologically split into a training dataset, NTUH-1519 (images acquired during the years 2015 to 2019; n = 2144), and a testing dataset, NTUH-20 (images acquired during the year 2020; n = 1354). Public databases, including the NIH ChestX-ray14 dataset (model training; 112,120 images), Montgomery County (model testing; 138 images), and Shenzhen (model testing; 662 images), were also used in model development. EfficientNetV2 was the basic architecture of the algorithm. Images from ChestX-ray14 were employed for pseudo-labelling to perform semi-supervised learning. The algorithm demonstrated excellent performance in detecting PTB (area under the receiver operating characteristic curve [AUC] 0.878, 95% confidence interval [CI] 0.854-0.900) in NTUH-20. The algorithm showed significantly better performance in posterior-anterior (PA) CXR (AUC 0.940, 95% CI 0.912-0.965, p-value < 0.001) compared with anterior-posterior (AUC 0.782, 95% CI 0.644-0.897) or portable anterior-posterior (AUC 0.869, 95% CI 0.814-0.918) CXR. The algorithm accurately detected cases of bacteriologically confirmed PTB (AUC 0.854, 95% CI 0.823-0.883). Finally, the algorithm tested favourably in Montgomery County (AUC 0.838, 95% CI 0.765-0.904) and Shenzhen (AUC 0.806, 95% CI 0.771-0.839). A deep learning-based algorithm could detect PTB on CXR with excellent performance, which may help shorten the interval between detection and airborne isolation for patients with PTB.
Chiou-Shann Fuh (傅楸善)合作论文数Department of Computer Science and Information Engineering, National Taiwan University4