Background: Complex oncological procedures pose various surgical challenges including dissection in distinct tissue planes and preservation of vulnerable anatomical structures throughout different surgical phases. In rectal surgery, a violation of dissection planes increases the risk of local recurrence and autonomous nerve damage resulting in incontinence and sexual dysfunction. While deep learning-based identification of target structures has been described in basic laparoscopic procedures, feasibility of artificial intelligence-based guidance has not yet been investigated in complex abdominal surgery. Methods: A dataset of 57 robot-assisted rectal resection (RARR) videos was split into a pre-training dataset of 24 temporally non-annotated videos and a training dataset of 33 temporally annotated videos. Based on phase annotations and pixel-wise annotations of randomly selected image frames, convolutional neural networks were trained to distinguish surgical phases and phase-specifically segment anatomical structures and tissue planes. To evaluate model performance, F1 score, Intersection-over-Union (IoU), precision, recall, and specificity were determined. Results: We demonstrate that both temporal (average F1 score for surgical phase recognition: 0.78) and spatial features of complex surgeries can be identified using machine learning-based image analysis. Based on analysis of a total of 8797 images with pixel-wise target structure segmentations, mean IoUs for segmentation of anatomical target structures range from 0.09 to 0.82 and from 0.05 to 0.32 for dissection planes and dissection lines throughout different phases of RARR in our analysis. Conclusions: Image-based recognition is a promising technique for surgical guidance in complex surgical procedures. Future research should investigate clinical applicability, usability, and therapeutic impact of a respective guidance system.
Graph neural networks (GNNs) are becoming increasingly popular in the medical domain for the tasks of disease classification and outcome prediction. Since patient data is not readily available as a graph, most existing methods either manually define a patient graph, or learn a latent graph based on pairwise similarities between the patients. There are also hypergraph neural network (HGNN)-based methods that were introduced recently to exploit potential higher order associations between the patients by representing them as a hypergraph. In this work, we propose a patient hypergraph network (PHGN), which has been investigated in an inductive learning setup for binary outcome prediction in oropharyngeal cancer (OPC) patients using computed tomography (CT)-based radiomic features for the first time. Additionally, the proposed model was extended to perform time-to-event analyses, and compared with GNN and baseline linear models.
Abstract Clinically relevant postoperative pancreatic fistula (CR-POPF) can significantly affect the treatment course and outcome in pancreatic cancer patients. Preoperative prediction of CR-POPF can aid the surgical decision-making process and lead to better perioperative management of patients. In this retrospective study of 108 pancreatic head resection patients, we present risk models for the prediction of CR-POPF that use combinations of preoperative computed tomography (CT)-based radiomic features, mesh-based volumes of annotated intra- and peripancreatic structures and preoperative clinical data. The risk signatures were evaluated and analysed in detail by visualising feature expression maps and by comparing significant features to the established CR-POPF risk measures. Out of the risk models that were developed in this study, the combined radiomic and clinical signature performed best with an average area under receiver operating characteristic curve (AUC) of 0.86 and a balanced accuracy score of 0.76 on validation data. The following pre-operative features showed significant correlation with outcome in this signature ( $$p < 0.05$$ p < 0.05 ) - texture and morphology of the healthy pancreatic segment, intensity volume histogram-based feature of the pancreatic duct segment, morphology of the combined segment, and BMI. The predictions of this pre-operative signature showed strong correlation (Spearman correlation co-efficient, $$\rho = 0.7$$ ρ = 0.7 ) with the intraoperative updated alternative fistula risk score (ua-FRS), which is the clinical gold standard for intraoperative CR-POPF risk stratification. These results indicate that the proposed combined radiomic and clinical signature developed solely based on preoperatively available clinical and routine imaging data can perform on par with the current state-of-the-art intraoperative models for CR-POPF risk stratification.
Table S1 Analyzed genes of the hypothesis-driven gene set. Table S2 Programs and packages used for evaluation. Figure S1 Results of repeated 3-fold cross validation for LRC. Figure S2 Mean oob ci depending of the signature size for LRC. Figure S3 Importance score for genes associated with LRC. Table S3 Increasing robustness by including highly correlated genes. Figure S4 Patient stratification by the 7-gene signature for LRC. Figure S5 Patient stratification by the 7-gene signature and clinical parameters for LRC. Table S4 Multivariable Cox regression of OS. Figure S6 Patient stratification by the 7-gene signature and clinical parameters for OS. Table S5 Multivariable Cox regression of DM. Figure S7 Patient stratification by the 7-gene signature and clinical parameters for DM. Table S6 Means and standard deviations of the 7-gene signature gene expressions. Table S7 Comparison of gene expressions between patient groups. Table S8 Hyper-parameters for feature selection algorithms and predictive models. Table S9 Patient characteristics of all patients. Figure S8 Identification of the final gene signature.
Background: Lack of anatomy recognition represents a clinically relevant risk in abdominal surgery. Machine learning (ML) methods can help identify visible patterns and risk structures; however, their practical value remains largely unclear. Materials and methods: Based on a novel dataset of 13 195 laparoscopic images with pixel-wise segmentations of 11 anatomical structures, we developed specialized segmentation models for each structure and combined models for all anatomical structures using two state-of-the-art model architectures (DeepLabv3 and SegFormer) and compared segmentation performance of algorithms to a cohort of 28 physicians, medical students, and medical laypersons using the example of pancreas segmentation. Results: Mean Intersection-over-Union for semantic segmentation of intra-abdominal structures ranged from 0.28 to 0.83 and from 0.23 to 0.77 for the DeepLabv3-based structure-specific and combined models, and from 0.31 to 0.85 and from 0.26 to 0.67 for the SegFormer-based structure-specific and combined models, respectively. Both the structure-specific and the combined DeepLabv3-based models are capable of near-real-time operation, while the SegFormer-based models are not. All four models outperformed at least 26 out of 28 human participants in pancreas segmentation. Conclusions: These results demonstrate that ML methods have the potential to provide relevant assistance in anatomy recognition in minimally invasive surgery in near-real-time. Future research should investigate the educational value and subsequent clinical impact of the respective assistance systems.
Clinically relevant postoperative pancreatic fistula (CR-POPF) is a common severe surgical complication after pancreatic surgery. Current risk stratification systems mostly rely on intraoperatively assessed factors like manually determined gland texture or blood loss. We developed a preoperatively available image-based risk score predicting CR-POPF as a complication of pancreatic head resection. Frequency of CR-POPF and occurrence of salvage completion pancreatectomy during the hospital stay were associated with an intraoperative surgical (sFRS) and image-based preoperative CT-based (rFRS) fistula risk score, both considering pancreatic gland texture, pancreatic duct diameter and pathology, in 195 patients undergoing pancreatic head resection. Based on its association with fistula-related outcome, radiologically estimated pancreatic remnant volume was included in a preoperative (preFRS) score for POPF risk stratification. Intraoperatively assessed pancreatic duct diameter ( p < 0.001), gland texture ( p < 0.001) and high-risk pathology ( p < 0.001) as well as radiographically determined pancreatic duct diameter ( p < 0.001), gland texture ( p < 0.001), high-risk pathology ( p = 0.001), and estimated pancreatic remnant volume ( p < 0.001) correlated with the risk of CR-POPF development. PreFRS predicted the risk of CR-POPF development (AUC = 0.83) and correlated with the risk of rescue completion pancreatectomy. In summary, preFRS facilitates preoperative POPF risk stratification in patients undergoing pancreatic head resection, enabling individualized therapeutic approaches and optimized perioperative management.
Background and purpose: Radiomics analyses have been shown to predict clinical outcomes of radiotherapy based on medical imaging-derived biomarkers. However, the biological meaning attached to such image features often remains unclear, thus hindering the clinical translation of radiomics analysis. In this manuscript, we describe a preclinical radiomics trial, which attempts to establish correlations between the expression of histological tumor microenvironment (TME)- and magnetic resonance imaging (MRI)-derived image features. Materials & Methods: A total of 114 mice were transplanted with the radioresistant and radiosensitive head and neck squamous cell carcinoma cell lines SAS and UT-SCC-14, respectively. The models were irradiated with five fractions of protons or photons using different doses. Post-treatment T1-weighted MRI and histopathological evaluation of the TME was conducted to extract quantitative features pertaining to tissue hypoxia and vascularization. We performed radiomics analysis with leave-one-out cross validation to identify the features most strongly associated with the tumor's phenotype. Performance was assessed using the area under the curve (AUC(Valid)) and F1-score. Furthermore, we analyzed correlations between TME- and MRI features using the Spearman correlation coefficient rho. Results: TME and MRI-derived features showed good performance (AUC(Valid, TME) = 0.72, AUC(Valid, MRI) = 0.85, AUC(Valid, Combined) = 0.85) individual tumor phenotype prediction. We found correlation coefficients of rho = -0.46 between hypoxia-related TME features and texture-related MRI features. Tumor volume was a strong confounder for MRI feature expression. Conclusion: We demonstrated a preclinical radiomics implementation and notable correlations between MRI- and TME hypoxia-related features. Developing additional TME features may help to further unravel the underlying biology. (c) 2022 The Authors. Published by Elsevier B.V.
Zielsetzung Die Studie evaluiert den Stellenwert der präoperativen CT-Bildgebung in der Vorhersage einer postoperativen Pancreasfistel (POPF) nach Pankreaskopfresektion mit Hilfe einer Radiomics-basierten Analyse.
Background and purpose: Hypoxia Positron-Emission-Tomography (PET) as well as Computed Tomography (CT) radiomics have been shown to be prognostic for radiotherapy outcome. Here, we investigate the stratification potential of CT-radiomics in head and neck cancer (HNC) patients and test if CT-radiomics is a surrogate predictor for hypoxia as identified by PET. Materials and methods: Two independent cohorts of HNC patients were used for model development and validation, HN1 (n = 149) and HN2 (n = 47). The training set HN1 consisted of native planning CT data whereas for the validation cohort HN2 also hypoxia PET/CT data was acquired using [18F]-Fluoromisonidazole (FMISO). Machine learning algorithms including feature engineering and classifier selection were trained for two-year loco-regional control (LRC) to create optimal CT-radiomics signatures.Secondly, a pre-defined [18F]FMISO-PET tumour-to-muscle-ratio (TMRpeak ≥ 1.6) was used for LRC prediction. Comparison between risk groups identified by CT-radiomics or [18F]FMISO-PET was performed using area-under–the-curve (AUC) and Kaplan-Meier analysis including log-rank test. Results: The best performing CT-radiomics signature included two features with nearest-neighbour classification (AUC = 0.76 ± 0.09), whereas AUC was 0.59 for external validation. In contrast, [18F]FMISO TMRpeak reached an AUC of 0.66 in HN2. Kaplan-Meier analysis of the independent validation cohort HN2 did not confirm the prognostic value of CT-radiomics (p = 0.18), whereas for [18F]FMISO-PET significant differences were observed (p = 0.02). Conclusions: No direct correlation of patient stratification using [18F]FMISO-PET or CT-radiomics was found in this study. Risk groups identified by CT-radiomics or hypoxia PET showed only poor overlap. Direct assessment of tumour hypoxia using PET seems to be more powerful to stratify HNC patients.
For treatment individualisation of patients with locally advanced head and neck squamous cell carcinoma (HNSCC) treated with primary radiochemotherapy, we explored the capabilities of different deep learning approaches for predicting loco-regional tumour control (LRC) from treatment-planning computed tomography images. Based on multicentre cohorts for exploration (206 patients) and independent validation (85 patients), multiple deep learning strategies including training of 3D- and 2D-convolutional neural networks (CNN) from scratch, transfer learning and extraction of deep autoencoder features were assessed and compared to a clinical model. Analyses were based on Cox proportional hazards regression and model performances were assessed by the concordance index (C-index) and the model's ability to stratify patients based on predicted hazards of LRC. Among all models, an ensemble of 3D-CNNs achieved the best performance (C-index 0.31) with a significant association to LRC on the independent validation cohort. It performed better than the clinical model including the tumour volume (C-index 0.39). Significant differences in LRC were observed between patient groups at low or high risk of tumour recurrence as predicted by the model (p = 0.001). This 3D-CNN ensemble will be further evaluated in a currently ongoing prospective validation study once follow-up is complete.
Imaging features for radiomic analyses are commonly calculated from the entire gross tumour volume (GTVentire). However, tumours are biologically complex and the consideration of different tumour regions in radiomic models may lead to an improved outcome prediction. Therefore, we investigated the prognostic value of radiomic analyses based on different tumour sub-volumes using computed tomography imaging of patients with locally advanced head and neck squamous cell carcinoma. The GTVentire was cropped by different margins to define the rim and the corresponding core sub-volumes of the tumour. Subsequently, the best performing tumour rim sub-volume was extended into surrounding tissue with different margins. Radiomic risk models were developed and validated using a retrospective cohort consisting of 291 patients in one of the six Partner Sites of the German Cancer Consortium Radiation Oncology Group treated between 2005 and 2013. The validation concordance index (C-index) averaged over all applied learning algorithms and feature selection methods using the GTVentire achieved a moderate prognostic performance for loco-regional tumour control (C-index: 0.61 ± 0.04 (mean ± std)). The models based on the 5 mm tumour rim and on the 3 mm extended rim sub-volume showed higher median performances (C-index: 0.65 ± 0.02 and 0.64 ± 0.05, respectively), while models based on the corresponding tumour core volumes performed less (C-index: 0.59 ± 0.01). The difference in C-index between the 5 mm tumour rim and the corresponding core volume showed a statistical trend (p = 0.10). After additional prospective validation, the consideration of tumour sub-volumes may be a promising way to improve prognostic radiomic risk models.
Background Radiomic features may quantify characteristics present in medical imaging. However, the lack of standardized definitions and validated reference values have hampered clinical use. Purpose To standardize a set of 174 radiomic features. Materials and Methods Radiomic features were assessed in three phases. In phase I, 487 features were derived from the basic set of 174 features. Twenty-five research teams with unique radiomics software implementations computed feature values directly from a digital phantom, without any additional image processing. In phase II, 15 teams computed values for 1347 derived features using a CT image of a patient with lung cancer and predefined image processing configurations. In both phases, consensus among the teams on the validity of tentative reference values was measured through the frequency of the modal value and classified as follows: less than three matches, weak; three to five matches, moderate; six to nine matches, strong; 10 or more matches, very strong. In the final phase (phase III), a public data set of multimodality images (CT, fluorine 18 fluorodeoxyglucose PET, and T1-weighted MRI) from 51 patients with soft-tissue sarcoma was used to prospectively assess reproducibility of standardized features. Results Consensus on reference values was initially weak for 232 of 302 features (76.8%) at phase I and 703 of 1075 features (65.4%) at phase II. At the final iteration, weak consensus remained for only two of 487 features (0.4%) at phase I and 19 of 1347 features (1.4%) at phase II. Strong or better consensus was achieved for 463 of 487 features (95.1%) at phase I and 1220 of 1347 features (90.6%) at phase II. Overall, 169 of 174 features were standardized in the first two phases. In the final validation phase (phase III), most of the 169 standardized features could be excellently reproduced (166 with CT; 164 with PET; and 164 with MRI). Conclusion A set of 169 radiomics features was standardized, which enabled verification and calibration of different radiomics software. © RSNA, 2020 Online supplemental material is available for this article. See also the editorial by Kuhl and Truhn in this issue.
Our contribution to the BraTS 2019 challenge consisted of a deep learning based approach for segmentation of brain tumours from MR images using cross validation ensembles of 2D-UNet models. Furthermore, different approaches for the prediction of patient survival time using clinical as well as imaging features were investigated. A simple linear regression model using patient age and tumour volumes outperformed more elaborate approaches like convolutional neural networks or radiomics-based analysis with an accuracy of 0.55 on the validation cohort and 0.51 on the test cohort.
Intraoperative tracking of laparoscopic instruments is often a prerequisite for computer and robotic-assisted interventions. While numerous methods for detecting, segmenting and tracking of medical instruments based on endoscopic video images have been proposed in the literature, key limitations remain to be addressed: Firstly, robustness, that is, the reliable performance of state-of-the-art methods when run on challenging images (e.g. in the presence of blood, smoke or motion artifacts). Secondly, generalization; algorithms trained for a specific intervention in a specific hospital should generalize to other interventions or institutions. In an effort to promote solutions for these limitations, we organized the Robust Medical Instrument Segmentation (ROBUST-MIS) challenge as an international benchmarking competition with a specific focus on the robustness and generalization capabilities of algorithms. For the first time in the field of endoscopic image processing, our challenge included a task on binary segmentation and also addressed multi-instance detection and segmentation. The challenge was based on a surgical data set comprising 10,040 annotated images acquired from a total of 30 surgical procedures from three different types of surgery. The validation of the competing methods for the three tasks (binary segmentation, multi-instance detection and multi-instance segmentation) was performed in three different stages with an increasing domain gap between the training and the test data. The results confirm the initial hypothesis, namely that algorithm performance degrades with an increasing domain gap. While the average detection and segmentation quality of the best-performing algorithms is high, future research should concentrate on detection and segmentation of small, crossing, moving and transparent instrument(s) (parts).
In 2015 we began a sub-challenge at the EndoVis workshop at MICCAI in Munich using endoscope images of ex-vivo tissue with automatically generated annotations from robot forward kinematics and instrument CAD models. However, the limited background variation and simple motion rendered the dataset uninformative in learning about which techniques would be suitable for segmentation in real surgery. In 2017, at the same workshop in Quebec we introduced the robotic instrument segmentation dataset with 10 teams participating in the challenge to perform binary, articulating parts and type segmentation of da Vinci instruments. This challenge included realistic instrument motion and more complex porcine tissue as background and was widely addressed with modifications on U-Nets and other popular CNN architectures. In 2018 we added to the complexity by introducing a set of anatomical objects and medical devices to the segmented classes. To avoid over-complicating the challenge, we continued with porcine data which is dramatically simpler than human tissue due to the lack of fatty tissue occluding many organs.
Non-rigid registration is a key component in soft-tissue navigation. We focus on laparoscopic liver surgery, where we register the organ model obtained from a preoperative CT scan to the intraoperative partial organ surface, reconstructed from the laparoscopic video. This is a challenging task due to sparse and noisy intraoperative data, real-time requirements and many unknowns - such as tissue properties and boundary conditions. Furthermore, establishing correspondences between pre- and intraoperative data can be extremely difficult since the liver usually lacks distinct surface features and the used imaging modalities suffer from very different types of noise. In this work, we train a convolutional neural network to perform both the search for surface correspondences as well as the non-rigid registration in one step. The network is trained on physically accurate biomechanical simulations of randomly generated, deforming organ-like structures. This enables the network to immediately generalize to a new patient organ without the need to re-train. We add various amounts of noise to the intraoperative surfaces during training, making the network robust to noisy intraoperative data. During inference, the network outputs the displacement field which matches the preoperative volume to the partial intraoperative surface. In multiple experiments, we show that the network translates well to real data while maintaining a high inference speed. Our code is made available online.
Purpose: To develop and validate a CT-based radiomics signature for the prognosis of loco-regional tumour control (LRC) in patients with locally advanced head and neck squamous cell carcinoma (HNSCC) treated by primary radiochemotherapy (RCTx) based on retrospective data from 6 partner sites of the German Cancer Consortium - Radiation Oncology Group (DKTK-ROG). Material and methods: Pre-treatment CT images of 318 patients with locally advanced HNSCC were collected. Four-hundred forty-six features were extracted from each primary tumour volume and then filtered through stability analysis and clustering. First, a baseline signature was developed from demographic and tumour-associated clinical parameters. This signature was then supplemented by CT imaging features. A final signature was derived using repeated 3-fold cross-validation on the discovery cohort. Performance in external validation was assessed by the concordance index (C-Index). Furthermore, calibration and patient stratification in groups with low and high risk for loco-regional recurrence were analysed. Results: For the clinical baseline signature, only the primary tumour volume was selected. The final signature combined the tumour volume with two independent radiomics features. It achieved moderately good discriminatory performance (C-Index [95% confidence interval]: 0.66 [0.55–0.75]) on the validation cohort along with significant patient stratification (p = 0.005) and good calibration. Conclusion: We identified and validated a clinical-radiomics signature for LRC of locally advanced HNSCC using a multi-centric retrospective dataset. Prospective validation will be performed on the primary cohort of the HNprädBio trial of the DKTK-ROG once follow-up is completed.
Intraoperative tracking of laparoscopic instruments is often a prerequisite for computer and robotic-assisted interventions. While numerous methods for detecting, segmenting and tracking of medical instruments based on endoscopic video images have been proposed in the literature, key limitations remain to be addressed: Firstly, robustness, that is, the reliable performance of state-of-the-art methods when run on challenging images (e.g. in the presence of blood, smoke or motion artifacts). Secondly, generalization; algorithms trained for a specific intervention in a specific hospital should generalize to other interventions or institutions. In an effort to promote solutions for these limitations, we organized the Robust Medical Instrument Segmentation (ROBUST-MIS) challenge as an international benchmarking competition with a specific focus on the robustness and generalization capabilities of algorithms. For the first time in the field of endoscopic image processing, our challenge included a task on binary segmentation and also addressed multi-instance detection and segmentation. The challenge was based on a surgical data set comprising 10,040 annotated images acquired from a total of 30 surgical procedures from three different types of surgery. The validation of the competing methods for the three tasks (binary segmentation, multi-instance detection and multi-instance segmentation) was performed in three different stages with an increasing domain gap between the training and the test data. The results confirm the initial hypothesis, namely that algorithm performance degrades with an increasing domain gap. While the average detection and segmentation quality of the best-performing algorithms is high, future research should concentrate on detection and segmentation of small, crossing, moving and transparent instrument(s) (parts).