Introduction: Deep Learning has been proposed as promising tool to classify malignant nodules. Our aim was to retrospectively validate our Lung Cancer Prediction Convolutional Neural Network (LCP-CNN), which was trained on US screening data, on an independent dataset of indeterminate nodules in an European multicentre trial, to rule out benign nodules maintaining a high lung cancer sensitivity. Methods: The LCP-CNN has been trained to generate a malignancy score for each nodule using CT data from the U.S. National Lung Screening Trial (NLST), and validated on CT scans containing 2106 nodules (205 lung cancers) detected in patients from from the Early Lung Cancer Diagnosis Using Artificial Intelligence and Big Data (LUCINDA) study, recruited from three tertiary referral centers in the UK, Germany and Netherlands. We pre-defined a benign nodule rule-out test, to identify benign nodules whilst maintaining a high sensitivity, by calculating thresholds on the malignancy score that achieve at least 99 % sensitivity on the NLST data. Overall performance per validation site was evaluated using Area-Under-the-ROC-Curve analysis (AUC). Results: The overall AUC across the European centers was 94.5 % (95 %CI 92.6-96.1). With a high sensitivity of 99.0 %, malignancy could be ruled out in 22.1 % of the nodules, enabling 18.5 % of the patients to avoid followup scans. The two false-negative results both represented small typical carcinoids. Conclusion: The LCP-CNN, trained on participants with lung nodules from the US NLST dataset, showed excellent performance on identification of benign lung nodules in a multi-center external dataset, ruling out malignancy with high accuracy in about one fifth of the patients with 5-15 mm nodules.
Improving stratification of patients with indeterminate pulmonary nodules (IPNs) can lead both to earlier diagnosis of lung cancer and to reduced scanning and reduced intervention in cases of benign disease. AI-based decision support software has been shown to outperform conventional risk models at classifying IPNs as low or high risk, but its performance in addition to clinician assessment has yet to be investigated. We report the results of a Multiple-Reader Multiple-Case reader evaluation comparing reader performance for both radiologists and pulmonologists on an IPN risk stratification task with and without AI assistance from the previously-published Lung Cancer Prediction Convolutional Neural Network (LCP-CNN).
Many of the species in decline around the world are subject to different environmental stressors across their range, so replicated large-scale monitoring programmes, are necessary to disentangle the relative impacts of these threats. At the same time as funding for long-term monitoring is being cut, studies are increasingly being criticised for lacking statistical power. For those taxa or environments where a single vantage point can observe individuals or ecological processes, time-lapse cameras can provide a cost-effective way of collecting time series data replicated at large spatial scales that would otherwise be impossible. However, networks of time-lapse cameras needed to cover the range of species or processes create a problem in that the scale of data collection easily exceeds our ability to process the raw imagery manually. Citizen science and machine learning provide solutions to scaling up data extraction (such as locating all animals in an image). Crucially, citizen science, machine learning-derived classifiers, and the intersection between them, are key to understanding how to establish monitoring systems that are sensitive to – and sufficiently powerful to detect –changes in the study system. Citizen science works relatively ‘out of the box’, and we regard it as a first step for many systems until machine learning algorithms are sufficiently trained to automate the process. Using Penguin Watch (www.penguinwatch.org) data as a case study, we discuss a complete workflow from images to parameter estimation and interpretation: the use of citizen science and computer vision for image processing, and parameter estimation and individual recognition for investigating biological questions. We discuss which techniques are easily generalizable to a range of questions, and where more work is needed to supplement ‘out of the box’ tools. We conclude with a horizon scan of the advances in camera technology, such as on-board computer vision and decision making.
Background Estimation of the risk of malignancy in pulmonary nodules detected by CT is central in clinical management. The use of artificial intelligence (AI) offers an opportunity to improve risk prediction. Here we compare the performance of an AI algorithm, the lung cancer prediction convolutional neural network (LCP-CNN), with that of the Brock University model, recommended in UK guidelines. Methods A dataset of incidentally detected pulmonary nodules measuring 5–15 mm was collected retrospectively from three UK hospitals for use in a validation study. Ground truth diagnosis for each nodule was based on histology (required for any cancer), resolution, stability or (for pulmonary lymph nodes only) expert opinion. There were 1397 nodules in 1187 patients, of which 234 nodules in 229 (19.3%) patients were cancer. Model discrimination and performance statistics at predefined score thresholds were compared between the Brock model and the LCP-CNN. Results The area under the curve for LCP-CNN was 89.6% (95% CI 87.6 to 91.5), compared with 86.8% (95% CI 84.3 to 89.1) for the Brock model (p≤0.005). Using the LCP-CNN, we found that 24.5% of nodules scored below the lowest cancer nodule score, compared with 10.9% using the Brock score. Using the predefined thresholds, we found that the LCP-CNN gave one false negative (0.4% of cancers), whereas the Brock model gave six (2.5%), while specificity statistics were similar between the two models. Conclusion The LCP-CNN score has better discrimination and allows a larger proportion of benign nodules to be identified without missing cancers than the Brock model. This has the potential to substantially reduce the proportion of surveillance CT scans required and thus save significant resources.
Time-lapse cameras facilitate remote and high-resolution monitoring of wild animal and plant communities, but the image data produced require further processing to be useful. Here we publish pipelines to process raw time-lapse imagery, resulting in count data (number of penguins per image) and ‘nearest neighbour distance’ measurements. The latter provide useful summaries of colony spatial structure (which can indicate phenological stage) and can be used to detect movement – metrics which could be valuable for a number of different monitoring scenarios, including image capture during aerial surveys. We present two alternative pathways for producing counts: (1) via the Zooniverse citizen science project Penguin Watch and (2) via a computer vision algorithm ( Pengbot ), and share a comparison of citizen science-, machine learning-, and expert- derived counts. We provide example files for 14 Penguin Watch cameras, generated from 63,070 raw images annotated by 50,445 volunteers. We encourage the use of this large open-source dataset, and the associated processing methodologies, for both ecological studies and continued machine learning and computer vision development.
Rationale: The management of indeterminate pulmonary nodules (IPNs) remains challenging, resulting in invasive procedures and delays in diagnosis and treatment. Strategies to decrease the rate of unnecessary invasive procedures and optimize surveillance regimens are needed. Objectives: To develop and validate a deep learning method to improve the management of IPNs. Methods: A Lung Cancer Prediction Convolutional Neural Network model was trained using computed tomography images of IPNs from the National Lung Screening Trial, internally validated, and externally tested on cohorts from two academic institutions. Measurements and Main Results: The areas under the receiver operating characteristic curve in the external validation cohorts were 83.5% (95% confidence interval [CI], 75.4-90.7%) and 91.9% (95% CI, 88.7-94.7%), compared with 78.1% (95% CI, 68.7-86.4%) and 81.9 (95% CI, 76.1-87.1%), respectively, for a commonly used clinical risk model for incidental nodules. Using 5% and 65% malignancy thresholds defining low- and high-risk categories, the overall net reclassifications in the validation cohorts for cancers and benign nodules compared with the Mayo model were 0.34 (Vanderbilt) and 0.30 (Oxford) as a rule-in test, and 0.33 (Vanderbilt) and 0.58 (Oxford) as a rule-out test. Compared with traditional risk prediction models, the Lung Cancer Prediction Convolutional Neural Network was associated with improved accuracy in predicting the likelihood of disease at each threshold of management and in our external validation cohorts. Conclusions: This study demonstrates that this deep learning algorithm can correctly reclassify IPNs into low- or high-risk categories in more than a third of cancers and benign nodules when compared with conventional risk models, potentially reducing the number of unnecessary invasive procedures and delays in diagnosis.
The implementation of video-based non-contact technologies to monitor the vital signs of preterm infants in the hospital presents several challenges, such as the detection of the presence or the absence of a patient in the video frame, robustness to changes in lighting conditions, automated identification of suitable time periods and regions of interest from which vital signs can be estimated. We carried out a clinical study to evaluate the accuracy and the proportion of time that heart rate and respiratory rate can be estimated from preterm infants using only a video camera in a clinical environment, without interfering with regular patient care. A total of 426.6 h of video and reference vital signs were recorded for 90 sessions from 30 preterm infants in the Neonatal Intensive Care Unit (NICU) of the John Radcliffe Hospital in Oxford. Each preterm infant was recorded under regular ambient light during daytime for up to four consecutive days. We developed multi-task deep learning algorithms to automatically segment skin areas and to estimate vital signs only when the infant was present in the field of view of the video camera and no clinical interventions were undertaken. We propose signal quality assessment algorithms for both heart rate and respiratory rate to discriminate between clinically acceptable and noisy signals. The mean absolute error between the reference and camera-derived heart rates was 2.3 beats/min for over 76% of the time for which the reference and camera data were valid. The mean absolute error between the reference and camera-derived respiratory rate was 3.5 breaths/min for over 82% of the time. Accurate estimates of heart rate and respiratory rate could be derived for at least 90% of the time, if gaps of up to 30 seconds with no estimates were allowed.
Lung cancer diagnostic pathway guidelines promote the use of risk stratification models. Artificial Intelligence (AI)-based risk models have been shown to achieve better diagnostic accuracy than clinical models like Mayo Clinic (Mayo) for particular clinical populations. The aim of this study is to examine whether this could translate into faster diagnosis for high-risk cancer patients. 116 patients (116 nodules) have been collected from a retrospective consecutive cohort acquired at Vanderbilt University Hospital. Time to diagnosis (TTD) was defined as the number of days between the CT scan and diagnosis date. Mean TTD was calculated on the cohort on which TTD could be defined, and on a reduced group comprising of TTD >31 days only. Risk scores for each nodule were found using the Mayo model and an AI-based Lung Cancer Prediction model (LCP) based on CT images alone. A 65% risk of cancer was taken to be the threshold at which surgical intervention is indicated (according to ACCP guidelines). Seven patients were dropped due to negative TTD, and six for having no definitive diagnosis date. The eventual cohort contained 61 cancer patients and 42 controls. Mean TTD is 140 days (Interquartile Range – IQR 1-77 days). 25 patients have TTD=0, 60 are within 31 days (28 cancers, 32 controls) and 43 (33 cancers, 10 controls) are above 32 days. On the full cohort: Mayo risk score is ≥65% for 15 cancers and 4 controls (sensitivity, 24.6%, specificity 90.5%), with a mean cancer TTD of 75 days. The LCP scores ≥65% in 43 cancers and 10 controls (sensitivity, 70.5%, specificity 76.2%), mean cancer TTD 81 days. On the reduced group: Mayo ≥65% for 7 cancers and 2 controls (sensitivity, 21.2%, specificity 80.0%) with mean cancer TTD 150 days. The LCP scores ≥65% in 21 cancers and 4 controls (sensitivity, 63.6%, specificity 60.0%), with mean cancer TTD 156 days. The LCP risk model could potentially accelerate the diagnosis in 40% more cancer patients who were not worked up fully in the month following a scan (the jump in sensitivity going from Mayo to LCP risk ≥65% is 42.4%). For these patients, time to a cancer diagnosis and treatment could be shortened by up to 156 days compared to recommendations if applying the Mayo risk model.
Automated time-lapse cameras can facilitate reliable and consistent monitoring of wild animal populations. In this report, data from 73,802 images taken by 15 different Penguin Watch cameras are presented, capturing the dynamics of penguin (Spheniscidae; Pygoscelis spp. ) breeding colonies across the Antarctic Peninsula, South Shetland Islands and South Georgia (03/2012 to 01/2014). Citizen science provides a means by which large and otherwise intractable photographic data sets can be processed, and here we describe the methodology associated with the Zooniverse project Penguin Watch , and provide validation of the method. We present anonymised volunteer classifications for the 73,802 images, alongside the associated metadata (including date/time and temperature information). In addition to the benefits for ecological monitoring, such as easy detection of animal attendance patterns, this type of annotated time-lapse imagery can be employed as a training tool for machine learning algorithms to automate data extraction, and we encourage the use of this data set for computer vision development.
Deep neural networks (DNN) have been shown to offer a viable alternative for risk cancer prediction of indeterminate pulmonary nodules (IPNs). While the type of data used for training is known to impact performance, this issue has not been extensively studied. We present, for the first time, a study of the effect of including training data that matches the clinical pathway of the independent validation dataset, a nodule clinic of incidental findings. Two identical DNNs were trained on the task of diagnosis prediction of pulmonary nodules from CT images. The first one (DNNnlst) used purely screening data from the US National Lung Screening Trial (922 cancer and 14733 benign nodules), while the second one (DNNnlst+incidental) included data of incidentally detected nodules from European hospitals (1064 cancer and 7207 benign nodules). Both models were evaluated in an independent validation set of nodules coming from a referral center in the UK (Royal Brompton and Harefield Hospital, London) consisting of baseline scans of 406 cancer and 325 benign nodules. The models were compared in terms of AUC, as well as their ability to reclassify cancer patients with intermediate risk nodules. The Intermediate risk sub-population was defined by selecting patients with nodules in the size range of 8 to 15mm, and who were followed-up within a year with CT, referred to PET-CT, or referred to biopsy. Within this sub-population, a cancer prevalence of 30% was assumed. The operating points of the cancer prediction models were chosen by setting a cancer risk of 70%, corresponding to high-risk nodules in the guidelines of the British Thoracic Society. The DNNnlst and DNNnlst+incidental models achieved an AUC of 84.33 (95%CI: 81.49, 87.15) and 87.43 (95%CI: 84.79, 89.82) respectively on the entire validation set, showing an improvement in the discrimination capabilities (p <0.01). For reference, using the nodule’s maximum axial diameter as a predictor led to an AUC of 79.07 (95%CI: 75.73, 82.73). Additionally, considering only the intermediate risk population of the data, all of which would require workup according to guidelines, the DNNnlst+incidental model correctly classified as high risk 34.34% more cases than the DNNnlst model (sensitivity 59.39% (95%CI: 47.68, 75.17) vs. 44.21% (95%CI: 20.98, 64.42)), an improvement significant at p <0.05. Although a DNN trained only on the US lung cancer screening data could have clinical utility in an incidental setting, exposing it to further incidental data can not only increase its discriminability, as expected, but also make it a potentially more effective tool for speeding up the diagnosis of cancer patients with intermediate risk nodules and reducing unnecessary workups.
Non-contact vital sign monitoring enables the estimation of vital signs, such as heart rate, respiratory rate and oxygen saturation (SpO 2 ), by measuring subtle color changes on the skin surface using a video camera. For patients in a hospital ward, the main challenges in the development of continuous and robust non-contact monitoring techniques are the identification of time periods and the segmentation of skin regions of interest (ROIs) from which vital signs can be estimated. We propose a deep learning framework to tackle these challenges. Approach : This paper presents two convolutional neural network (CNN) models. The first network was designed for detecting the presence of a patient and segmenting the patient’s skin area. The second network combined the output from the first network with optical flow for identifying time periods of clinical intervention so that these periods can be excluded from the estimation of vital signs. Both networks were trained using video recordings from a clinical study involving 15 pre-term infants conducted in the high dependency area of the neonatal intensive care unit (NICU) of the John Radcliffe Hospital in Oxford, UK. Main results : Our proposed methods achieved an accuracy of 98.8% for patient detection, a mean intersection-over-union (IOU) score of 88.6% for skin segmentation and an accuracy of 94.5% for clinical intervention detection using two-fold cross validation. Our deep learning models produced accurate results and were robust to different skin tones, changes in light conditions, pose variations and different clinical interventions by medical staff and family visitors. Significance : Our approach allows cardio-respiratory signals to be continuously derived from the patient’s skin during which the patient is present and no clinical intervention is undertaken.
The crystallization of solidifying Al-Cu alloys over a wide range of conditions was studied in situ by synchrotron x-ray radiography, and the data were analyzed using a computer vision algorithm trained using machine learning. The effect of cooling rate and solute concentration on nucleation undercooling, crystal formation rate, and crystal growth rate was measured automatically for thousands of separate crystals, which was impossible to achieve manually. Nucleation undercooling distributions confirmed the efficiency of extrinsic grain refiners and gave support to the widely assumed free growth model of heterogeneous nucleation. We show that crystallization occurred in temporal and spatial bursts associated with a solute-suppressed nucleation zone.
Non-contact vital-sign estimation allows the monitoring of physiological parameters (such as heart rate, res-demonstrated that a convolutional neural network (CNN) can be used to detect the presence of a patient and segment the patient's skin area for vital-sign estimation, thus enabling the automatic continuous monitoring of vital signs in a hospital environment. In a study approved by the local Research Ethical Committee, we made video recordings of pre-term infants nursed in a Neonatal Intensive Care Unit (NICU) at the John Radcliffe Hospital in Oxford, UK. We extended the CNN model to detect the head, torso and diaper of the infants. We extracted multiple photoplethysmographic imaging (PPGi) signals from each body part, analysed their signal quality, and compared them with the PPGi signal derived from the entire skin area. Our results demonstrated the benefits of estimating heart rate combined from multiple regions of interest using data fusion. In the test dataset, we achieved a mean absolute error of 2.4 beats per minute for 80% (31.1 hours) from a total recording time of 38.5 hours for which both reference heart rate and video data were valid.
Artificial Intelligence (AI) based malignancy prediction of indeterminate pulmonary nodules has been previously demonstrated to perform well on screen-detected nodules imaged with low-dose, non-contrast CT. This study aimed to assess the impact of contrast media on the classification performance of such a system. A Convolutional Neural Network (CNN) was trained on the US National Lung Screening Trial (NLST), which contained only low-dose non-contrast screening images, selecting all nodules 6mm and greater in size (14761 benign nodules from 5972 patients; 932 cancer from 575 patients). A CNN classifier was trained using Deep Learning on this data to produce a malignancy score per nodule. For validation, an independent retrospective dataset of incidentally detected solid nodules was used. None of the patients had a cancer diagnosis within the past 5 years, and all had fewer than 5 nodules. The dataset contained 571 nodules from 505 patients, including 42 cancer from 39 patients. A CT was considered to be contrasted if the mode HU value within an ROI placed at the aortic arch was greater than 60HU. This resulted in two groups: the non-contrast group had 313 nodules from 276 patients (16 cancer from 14 patients); the contrast group had 258 nodules from 229 patients (26 cancer from 25 patients). The overall efficacy was assessed using Area-Under-the-ROC-Curve analysis (AUC) for each group. The AUC on the non-contrast CT group was 0.96 (95% CI 0.93 to 0.98) and 0.95 (95% CI 0.90 to 0.99) on the contrast CT group. Further analysis revealed that excluding high contrast cases, where the HU in the aortic arch was greater than 300HU (40 nodules from 38 patients; 4 cancer), resulted in an AUC of 0.97 (95% CI 0.92 to 1.00). The CNN classifier seems to be robust to the presence of contrast media with only a moderate reduction in performance. Excluding cases with high contrast restored the performance, although, with only 38 nodules excluded by this, the result may be not statistically significant. These results indicate that a CNN developed to predict pulmonary nodule malignancy, that has been trained on low-dose, non-contrast enhanced CT images, may be used with CT images with moderate levels of contrast without retraining.