Biomedical engineering has been targeted as a potential research candidate for machine learning applications, with the purpose of detecting or diagnosing pathologies. However, acquiring relevant, high-quality, and heterogeneous medical datasets is challenging due to privacy and security issues and the effort required to annotate the data. Generative models have recently gained a growing interest in the computer vision field due to their ability to increase dataset size by generating new high-quality samples from the initial set, which can be used as data augmentation of a training dataset. This study aimed to synthesize artificial lung images from corresponding positional and semantic annotations using two generative adversarial networks and databases of real computed tomography scans: the Pix2Pix approach that generates lung images from the lung segmentation maps; and the conditional generative adversarial network (cCGAN) approach that was implemented with additional semantic labels in the generation process. To evaluate the quality of the generated images, two quantitative measures were used: the domain-specific Frechet Inception Distance and Structural Similarity Index. Additionally, an expert assessment was performed to measure the capability to distinguish between real and generated images. The assessment performed shows the high quality of synthesized images, which was confirmed by the expert evaluation. This work represents an innovative application of GAN approaches for medical application taking into consideration the pathological findings in the CT images and the clinical evaluation to assess the realism of these features in the generated images.
Advancements in the development of computer-aided decision (CAD) systems for clinical routines provide unquestionable benefits in connecting human medical expertise with machine intelligence, to achieve better quality healthcare. Considering the large number of incidences and mortality numbers associated with lung cancer, there is a need for the most accurate clinical procedures; thus, the possibility of using artificial intelligence (AI) tools for decision support is becoming a closer reality. At any stage of the lung cancer clinical pathway, specific obstacles are identified and "motivate" the application of innovative AI solutions. This work provides a comprehensive review of the most recent research dedicated toward the development of CAD tools using computed tomography images for lung cancer-related tasks. We discuss the major challenges and provide critical perspectives on future directions. Although we focus on lung cancer in this review, we also provide a more clear definition of the path used to integrate AI in healthcare, emphasizing fundamental research points that are crucial for overcoming current barriers.
Lung diseases affect the lives of billions of people worldwide, and 4 million people, each year, die prematurely due to this condition. These pathologies are characterized by specific imagiological findings in CT scans. The traditional Computer-Aided Diagnosis (CAD) approaches have been showing promising results to help clinicians; however, CADs normally consider a small part of the medical image for analysis, excluding possible relevant information for clinical evaluation. Multiple Instance Learning (MIL) approach takes into consideration different small pieces that are relevant for the final classification and creates a comprehensive analysis of pathophysiological changes. This study uses MIL-based approaches to identify the presence of lung pathophysiological findings in CT scans for the characterization of lung disease development. This work was focus on the detection of the following: Fibrosis, Emphysema, Satellite Nodules in Primary Lesion Lobe, Nodules in Contralateral Lung and Ground Glass, being Fibrosis and Emphysema the ones with more outstanding results, reaching an Area Under the Curve (AUC) of 0.89 and 0.72, respectively. Additionally, the MIL-based approach was used for EGFR mutation status prediction - the most relevant oncogene on lung cancer, with an AUC of 0.69. The results showed that this comprehensive approach can be a useful tool for lung pathophysiological characterization.
The evolution of personalized medicine has changed the therapeutic strategy from classical chemotherapy and radiotherapy to a genetic modification targeted therapy, and although biopsy is the traditional method to genetically characterize lung cancer tumor, it is an invasive and painful procedure for the patient. Nodule image features extracted from computed tomography (CT) scans have been used to create machine learning models that predict gene mutation status in a noninvasive, fast, and easy-to-use manner. However, recent studies have shown that radiomic features extracted from an extended region of interest (ROI) beyond the tumor, might be more relevant to predict the mutation status in lung cancer, and consequently may be used to significantly decrease the mortality rate of patients battling this condition. In this work, we investigated the relation between image phenotypes and the mutation status of Epidermal Growth Factor Receptor (EGFR), the most frequently mutated gene in lung cancer with several approved targeted-therapies, using radiomic features extracted from the lung containing the nodule. A variety of linear, nonlinear, and ensemble predictive classification models, along with several feature selection methods, were used to classify the binary outcome of wild-type or mutant EGFR mutation status. The results show that a comprehensive approach using a ROI that included the lung with nodule can capture relevant information and successfully predict the EGFR mutation status with increased performance compared to local nodule analyses. Linear Support Vector Machine, Elastic Net, and Logistic Regression, combined with the Principal Component Analysis feature selection method implemented with 70% of variance in the feature set, were the best-performing classifiers, reaching Area Under the Curve (AUC) values ranging from 0.725 to 0.737. This approach that exploits a holistic analysis indicates that information from more extensive regions of the lung containing the nodule allows a more complete lung cancer characterization and should be considered in future radiogenomic studies.
Lung cancer is still the leading cause of cancer death in the world. For this reason, novel approaches for early and more accurate diagnosis are needed. Computer-aided decision (CAD) can be an interesting option for a noninvasive tumour characterisation based on thoracic computed tomography (CT) image analysis. Until now, radiomics have been focused on tumour features analysis, and have not considered the information on other lung structures that can have relevant features for tumour genotype classification, especially for epidermal growth factor receptor (EGFR), which is the mutation with the most successful targeted therapies. With this perspective paper, we aim to explore a comprehensive analysis of the need to combine the information from tumours with other lung structures for the next generation of CADs, which could create a high impact on targeted therapies and personalised medicine. The forthcoming artificial intelligence (AI)-based approaches for lung cancer assessment should be able to make a holistic analysis, capturing information from pathological processes involved in cancer development. The powerful and interpretable AI models allow us to identify novel biomarkers of cancer development, contributing to new insights about the pathological processes, and making a more accurate diagnosis to help in the treatment plan selection.
Statistics have demonstrated that one of the main factors responsible for the high mortality rate related to lung cancer is the late diagnosis. Precision medicine practices have shown advances in the individualized treatment according to the genetic profile of each patient, providing better control on cancer response. Medical imaging offers valuable information with an extensive perspective of the cancer, opening opportunities to explore the imaging manifestations associated with the tumor genotype in a non-invasive way. This work aims to study the relevance of physiological features captured from Computed Tomography images, using three different 2D regions of interest to assess the Epidermal growth factor receptor (EGFR) mutation status: nodule, lung containing the main nodule, and both lungs. A Convolutional Autoencoder was developed for the reconstruction of the input image. Thereafter, the encoder block was used as a feature extractor, stacking a classifier on top to assess the EGFR mutation status. Results showed that extending the analysis beyond the local nodule allowed the capture of more relevant information, suggesting the presence of useful biomarkers using the lung with nodule region of interest, which allowed to obtain the best prediction ability. This comparative study represents an innovative approach for gene mutations status assessment, contributing to the discussion on the extent of pathological phenomena associated with cancer development, and its contribution to more accurate Artificial Intelligence-based solutions, and constituting, to the best of our knowledge, the first deep learning approach that explores a comprehensive analysis for the EGFR mutation status classification.
Artificial intelligence (AI)-based solutions have revolutionized our world, using extensive datasets and computational resources to create automatic tools for complex tasks that, until now, have been performed by humans. Massive data is a fundamental aspect of the most powerful AI-based algorithms. However, for AI-based healthcare solutions, there are several socioeconomic, technical/infrastructural, and most importantly, legal restrictions, which limit the large collection and access of biomedical data, especially medical imaging. To overcome this important limitation, several alternative solutions have been suggested, including transfer learning approaches, generation of artificial data, adoption of blockchain technology, and creation of an infrastructure composed of anonymous and abstract data. However, none of these strategies is currently able to completely solve this challenge. The need to build large datasets that can be used to develop healthcare solutions deserves special attention from the scientific community, clinicians, all the healthcare players, engineers, ethicists, legislators, and society in general. This paper offers an overview of the data limitation in medical predictive models; its impact on the development of healthcare solutions; benefits and barriers of sharing data; and finally, suggests future directions to overcome data limitations in the medical field and enable AI to enhance healthcare. This perspective is dedicated to the technical requirements of the learning models, and it explains the limitation that comes from poor and small datasets in the medical domain and the technical options that try or can solve the problem related to the lack of massive healthcare data.
Poster: ECR 2018 / C-3086 / Osteomyelitis- What radiologists should know by: B. S. D. Flor de Lima, E. F. M. P. Negrao, C. Sousa, J. Rebelo, M. J. Leite, F. Duarte, J. N. C. Lobo, M. Pimenta; Porto/PT