The progress of deep learning architectures, machine learning models and pathology slide digitization is an encouraging step toward meeting the growing demand for more precise classification and prediction diagnosis for the breast tumours. The BreakHis dataset with four magnification factors (40X, 100X, 200X and 400X), as well as seven deep learning architectures used for feature extraction (DenseNet 201, Inception ResNet V2, Inception V3, ResNet 50, MobileNet V2,VGG16 and VGG19), four machine learning models for classification (MLP, SVM, DT, and KNN), and two combination rules (hard and weighted voting) were investigated in this paper to design and evaluate a new proposed approach consisting of building deep hybrid homogenous ensemble. Additionally, the best proposed models were compared to deep stacked, deep bagging, deep boosting, and deep hybrid heterogenous ensemble to choose the best strategy in building deep ensemble learning techniques. The four performance measures accuracy, precision, recall, and F1-score were used in the empirical evaluations, as well as 5-fold cross-validation, the Scott Knott statistical test, and the Borda Count voting method. The results demonstrated the new approach's potential since it outscored both singles and other deep ensemble learning strategies, achieving accuracy values of 98.3% and 97.7% for the MFs 40X, 100X and 200X, 400X, respectively. The empirical results demonstrated that the proposed ensembles are impactful for histopathological breast cancer images classification, and they provided a promising tool to assist pathologists in the diagnosis of breast cancer.
One of the most consequential public health issues in the world and a major factor in women's mortality is breast cancer. Early detection and diagnosis can significantly improve the likelihood of survival. Therefore, this study suggests a deep end-to-end heterogeneous ensemble approach by using deep convolutional neural networks models for breast histological images classification tested on the BreakHis dataset. The proposed approach showed a significant increase of performances compared to their base learners. Thus, seven deep learning architectures (VGG16, VGG19, ResNet50, Inception V3, Inception ResNet V2, Xception, and MobileNet V2) were trained using fivefold cross-validation. Thereafter, deep end-to-end heterogeneous ensembles of two up to seven models were constructed based on three selection criteria’s (by accuracy, by diversity, and by both accuracy and diversity) and combined with two voting methods: majority voting by tacking the mode of the distribution of the predicted labels, and weighted voting by taking the average of predicted probabilities. Results showed the effectiveness of deep end-to-end ensemble learning techniques for histopathological breast cancer images classification since the ensembles designed using weighted voting with the selection by accuracy strategy method exceeded the ones designed using the selection by diversity or by accuracy and diversity strategies. The accuracy values of the proposed approach have shown a significant amelioration compared to the least performing base learner used as a baseline ResNet 50 with an accuracy increased from 78.14%, 78.57%, 82.80 and 79.43% to 93.8%, 93.4%, 93.3%, and 91.8% through the BreakHis dataset's four magnification factors: 40X, 100X, 200X, and 400X respectively.
Nuclei semantic segmentation is a key component for advancing machine learning and deep learning applications in digital pathology. However, most existing segmentation models are trained and tested on high-quality data acquired with expensive equipment, such as whole slide scanners, which are not accessible to most pathologists in developing countries. These pathologists rely on low-resource data acquired with low-precision microscopes, smartphones, or digital cameras, which have different characteristics and challenges than high-resource data. Therefore, there is a gap between the state-of-the-art segmentation models and the real-world needs of low-resource settings. This work aims to bridge this gap by presenting the first fully annotated African multi-organ dataset for histopathology nuclei semantic segmentation acquired with a low-precision microscope. We also evaluate state-of-the-art segmentation models, including spectral feature extraction encoder and vision transformer-based models, and stain normalization techniques for color normalization of Hematoxylin and Eosin-stained histopathology slides. Our results provide important insights for future research on nuclei histopathology segmentation with low-resource data. Code and dataset: https://github.com/zerouaoui/AMONUSEG.
Breast cancer (BC) is considered one of the major public health issues and a leading cause of death among women in the world. Its early diagnosis using histological imaging modalities can significantly help to increase the chances of survival rate by choosing the adequate treatment. This research work is a comparative study that assesses and compares the performances of: (1) seven deep end-to-end architectures using seven deep learning models including VGG16, VGG19, ResNet 50, MobileNet V2, Inception V3, Inception ResNet V2 and DenseNet 201, (2) a deep hybrid architectures combining the deep learning model DenseNet 201 for feature extraction and MLP for classification and (3) a deep end-to-end heterogeneous ensembles of seven deep learning models used as base learners including VGG16, VGG19, ResNet50, Inception V3, Inception ResNet V2, Xception, and MobileNet V2 and combined with weighted voting over a hematoxylin and eosin (H E)-stained pathological images. The experiments were conducted using the public BreakHis dataset, 5-fold cross validation method, four metrics of performances, the Scott Knott statistical test and the Borda Count voting method. The results showed that the deep end-to-end architecture DenseNet 201 is outperforming the others architectures with an accuracy value of 94,9
Purpose Histopathology biopsy imaging is currently the gold standard for the diagnosis of breast cancer in clinical practice. Pathologists examine the images at various magnifications to identify the type of tumor because if only one magnification is taken into account, the decision may not be accurate. This study explores the performance of transfer learning and late fusion to construct multi-scale ensembles that fuse different magnification-specific deep learning models for the binary classification of breast tumor slides. Design/methodology/approach Three pretrained deep learning techniques (DenseNet 201, MobileNet v2 and Inception v3) were used to classify breast tumor images over the four magnification factors of the Breast Cancer Histopathological Image Classification dataset (40×, 100×, 200× and 400×). To fuse the predictions of the models trained on different magnification factors, different aggregators were used, including weighted voting and seven meta-classifiers trained on slide predictions using class labels and the probabilities assigned to each class. The best cluster of the outperforming models was chosen using the Scott–Knott statistical test, and the top models were ranked using the Borda count voting system. Findings This study recommends the use of transfer learning and late fusion for histopathological breast cancer image classification by constructing multi-magnification ensembles because they perform better than models trained on each magnification separately. Originality/value The best multi-scale ensembles outperformed state-of-the-art integrated models and achieved an accuracy mean value of 98.82 per cent, precision of 98.46 per cent, recall of 100 per cent and F 1-score of 99.20 per cent.
: Artificial intelligence (AI)-assisted cervical cytology is poised to enhance sensitivity whilst lessening bias, labor, and time expenses. It typically involves image processing and deep learning to automatically recognize pre-cancerous lesions on a given whole-slide image (WSI) prior to lethal invasive cancer development. Here, we introduce autoencoder (AE)-based hybrid models for cervical carcinoma prediction on the Mendeley-liquid-based cytology dataset. This is built on fourteen combinations of AE, DenseNet-201, and six state-of-the-art classifiers: adaptive boosting (AdaBoost), support vector machine (SVM), multilayer perceptron (MLP), decision tree (DT), k-nearest neighbors (k-NN), and random forest (RF). As empirical evaluations, four performance metrics, Scott-Knott (SK), and Borda count voting scheme, were performed. The AE-based hybrid models integrating AdaBoost, MLP, and RF as classifiers are among the top-ranked architectures, with respective accuracy values of 99.30, 99.20, and 98.48%. Yet, DenseNet-201 remains a solid option when adopting an end-to-end training strategy.
Computational Fluid Dynamics (CFD) simulation of multiphase industrial flows is a significant research concern for studying the performance and efficiency of chemical processes. Within the last years, CFD had growth interest from researchers with the significant increase in computational resources capacities and high-performance calculations. The trend toward focusing on machine learning (ML) techniques to solve complex industrial problems is observed. In contrast, ML has found encouraging and promising applications in that research field by offering a wealth of techniques to extract unreachable knowledge from data that can improve the chemical processes understanding and allows efficient optimization and intensification. This study aims to present the different uses of ML in the chemical processes' field, particularly the CFD modeling, and highlight the main and variate use of ML for complex geometries design and mesh optimization. We have paid particular attention to the ML trend in turbulence modeling to generate data-driven physical models from CFD simulations at different fidelity levels. Then, the extended application of ML to accelerate CFD calculations, and the development of surrogate models are provided. This work reveals that the application of ML to chemical engineering is a promising way for developing this research area.
: This paper proposes the use of transfer learning and ensemble learning for binary classification of breast cancer histological images over the four magnification factors of the BreakHis dataset: 40×, 100×, 200× and 400×. The proposed bagging ensembles are implemented using a set of hybrid architectures that combine pre-trained deep learning techniques for feature extraction with machine learning classifiers as base learners (MLP, SVM and KNN). The study evaluated and compared: (1) bagging ensembles with their base learners, (2) bagging ensembles with a different number of base learners (3, 5, 7 and 9), (3) single classifiers with the best bagging ensembles, and (4) best bagging ensembles of each feature extractor and magnification factor. The best cluster of the outperforming models was chosen using the Scott Knott (SK) statistical test, and the top models were ranked using the Borda Count voting system. The best bagging ensemble achieved a mean accuracy value of 93.98%, and was constructed using 3 base learners, 200× as a magnification factor, MLP as a classifier, and DenseNet201 as a feature extractor. The results demonstrated that bagging hybrid deep learning is an effective and a promising approach for the automatic classification of histopathological breast cancer images.
PurposeHundreds of thousands of deaths each year in the world are caused by breast cancer (BC). An early-stage diagnosis of this disease can positively reduce the morbidity and mortality rate by helping to select the most appropriate treatment options, especially by using histological BC images for the diagnosis.Design/methodology/approachThe present study proposes and evaluates a novel approach which consists of 24 deep hybrid heterogenous ensembles that combine the strength of seven deep learning techniques (DenseNet 201, Inception V3, VGG16, VGG19, Inception-ResNet-V3, MobileNet V2 and ResNet 50) for feature extraction and four well-known classifiers (multi-layer perceptron, support vector machines, K-nearest neighbors and decision tree) by means of hard and weighted voting combination methods for histological classification of BC medical image. Furthermore, the best deep hybrid heterogenous ensembles were compared to the deep stacked ensembles to determine the best strategy to design the deep ensemble methods. The empirical evaluations used four classification performance criteria (accuracy, sensitivity, precision and F1-score), fivefold cross-validation, Scott–Knott (SK) statistical test and Borda count voting method. All empirical evaluations were assessed using four performance measures, including accuracy, precision, recall and F1-score, and were over the histological BreakHis public dataset with four magnification factors (40×, 100×, 200× and 400×). SK statistical test and Borda count were also used to cluster the designed techniques and rank the techniques belonging to the best SK cluster, respectively.FindingsResults showed that the deep hybrid heterogenous ensembles outperformed both their singles and the deep stacked ensembles and reached the accuracy values of 96.3, 95.6, 96.3 and 94 per cent across the four magnification factors 40×, 100×, 200× and 400×, respectively.Originality/valueThe proposed deep hybrid heterogenous ensembles can be applied for the BC diagnosis to assist pathologists in reducing the missed diagnoses and proposing adequate treatments for the patients.
Purpose Breast cancer (BC) is the most common diagnosed cancer type and one of the top leading causes of death in women worldwide. This paper aims to investigate ensemble learning and transfer learning for binary classification of BC histological images over the four-magnification factor (MF) values of the BreakHis dataset: 40X, 100X, 200X, and 400X. Methods The proposed homogeneous ensembles are implemented using a hybrid architecture that combines: (1) three of the most recent deep learning (DL) techniques for feature extraction: DenseNet_201, MobileNet_V2, and Inception_V3, and (2) four of the most popular boosting methods for classification: AdaBoost (ADB), Gradient Boosting Machine (GBM), LightGBM (LGBM) and XGBoost (XGB) with Decision Tree (DT) as a base learner. The study evaluated and compared: (1) a set of boosting ensembles designed with the same hybrid architecture and different number of trees (50, 100, 150 and 200); (2) different boosting methods, and (3) the single DT classifier with the best boosting ensembles. The empirical evaluations used: four classification performance criteria (accuracy, recall, precision and F1-score), the fivefold cross-validation, Scott Knott statistical test to select the best cluster of the outperforming models, and Borda Count voting system to rank the best performing ones. Results The best boosting ensemble achieved an accuracy value of 92.52% and it was constructed using XGB with 200 trees and Inception_V3 as feature extractor (FE). Conclusions The results showed the potential of combining DL techniques for feature extraction and boosting ensembles to classify BC in malignant and benign tumors.
Breast cancer (BC) is the most common diagnosed cancer type and one of the top leading causes of death in women worldwide. The early diagnosis of this type of cancer is the main driver of high survival rate. This paper aims to use homogenous ensemble learning and transfer learning for binary classification of BC histological images over the four-magnification factor (MF) values of the BreakHis dataset: 40X, 100X, 200X, and 400X. The proposed ensembles are implemented using a hybrid architecture (HA) that combines: (1) three of the most recent deep learning (DL) techniques as feature extractors (FE): DenseNet_201, Inception_V3, and MobileNet_V2, and (2) the boosting method AdaBoost with Decision Tree (DT) as a base learner. The study evaluated and compared: the ensembles designed with the same HA but with different number of trees (50, 100, 150 and 200), the single DT classifiers with the best AdaBoost ensembles and the best AdaBoost ensembles of each FE over each MF. The empirical evaluations used: four classification performance criteria (accuracy, recall, precision and F1-score), 5-fold cross-validation, Scott Knott (SK) statistical test to select the best cluster of the outperforming models, and Borda Count voting system to rank the best performing ones. Results showed the potential of combining DL techniques for FE and AdaBoost boosting method to classify BC in malignant and benign tumors, furthermore the AdaBoost ensemble constructed using 200 trees, DenseNet_201 as FE and MF 200X achieved the best mean accuracy value with 90.36
Breast cancer affects thousands of people worldwide each year, artificial intelligence used for digital pathological computer-aided diagnosis for breast cancer classification is a valuable domain. The advancement of machine learning and deep learning methods, as well as pathology slide digitization for primary diagnosis, is a major development toward meeting the demand for more accurate breast tumor diagnosis, classification, and prediction. This paper proposes and evaluates a new approach consisting of deep hybrid homogenous ensemble method based on seven deep learning models for feature extraction (DenseNet 201, Inception V3, Inception ResNet V2, MobileNet V2, ResNet 50, VGG16, and VGG19), a multi-layer perceptron (MLP) for classification and two combination rules (hard and weighted voting) for histological classification using the BreakHis dataset four magnification factors: 40X, 100X, 200X and 400X. The outcomes proved the potential of the proposed new approach since it outperformed its singles achieving an accuracy value of 98,3% for the MFs 40X and 100X and 97,7% for the MFs 200X and 400X.
Breast cancer is the most common cancer in women worldwide. While the early diagnosis and treatment can significantly reduce the mortality rate, it is a challenging task for pathologists to accurately estimate the cancerous cells and tissues. Therefore, machine learning techniques are playing a significant role in assisting pathologists and improving the diagnosis results. This paper proposes a hybrid architecture that combines: three of the most recent deep learning techniques for feature extraction (DenseNet_201, Inception_V3, and MobileNet_V2) and random forest to classify breast cancer histological images over the BreakHis dataset with its four magnification factors: 40X, 100X, 200X and 400X. The study evaluated and compared: (1) the developed random forest models with their base learners, (2) the designed random forest models with the same architecture but with a different number of trees, (3) the decision tree classifiers with the best random forest models and (4) the best random forest models of each feature extractor. The empirical evaluations used: four classification performance criteria (accuracy, sensitivity, precision and F1-score), 5-fold cross-validation, Scott Knott statistical test, and Borda Count voting method. The best random forest model achieved an accuracy mean value of 85.88%, and was constructed using 9 trees, 200X as a magnification factor, and Inception_V3 as a feature extractor. The experimental results demonstrated that combining random forest with deep learning models is effective for the automatic classification of malignant and benign tumors using histopathological images of breast cancer.
Breast cancer is considered one of the major public health issues and a leading cause of death among women in the world. Its early diagnosis can significantly help to increase the chances of survival rate. Therefore, this study proposes a deep stacking ensemble technique for binary classification of breast histopathological images over the BreakHis dataset. Initially, to form the base learners of the deep stacking ensemble, we trained seven deep learning (DL) techniques based on pre-trained VGG16, VGG19, ResNet50, Inception_V3, Inception_ResNet_V2, Xception, and MobileNet with a 5-fold cross-validation method. Then, a meta-model was built, a logistic regression algorithm that learns how to best combine the predictions of the base learners. Furthermore, to evaluate and compare the performance of the proposed technique, we used: (1) four classification performance criteria (accuracy, precision, recall, and F1-score), and (2) Scott Knott (SK) statistical test to cluster and identify the outperforming models. Results showed the potential of the stacked deep learning techniques to classify breast cancer images into malignant or benign tumor. The proposed deep stacking ensemble reports an overall accuracy of 93.8%, 93.0%, 93.3%, and 91.8% over the four magnification factors (MF) values of the BreakHis dataset: 40X, 100X, 200X and 400X, respectively.
: Cervical cancer (CxCa) is heavily swerved toward low- and middle- income countries (LMICs). Without prompt actions, the burden is anticipated to worsen by 50% from 2020 to 2040 - nearly 90% of deaths to occur in sub-Saharan Africa (SSA). Yet, uterine cervix neoplasms are readily avoidable due to a protracted latent cancer period. As it stands, deep learning (DL) is a potent solution for enhancing the early detection of cervical cancer. This work assesses and compares the performance of seven end-to-end learning architectures to automatically recognize cervical lesions and carcinoma histotypes upon hematoxylin and eosin (H&E)-stained pathology images. Pre-trained VGG16, VGG19, InceptionV3, ResNet50, MobileNetV2, InceptionResNetV2, and DenseNet201 were the implemented deep convolutional neural networks (dCNNs) throughout the present empirical analysis. Experiments are conducted on two datasets: (i) Mendeley liquid-based cytology (LBC) and (ii) The Cancer Genome Atlas (TCGA) Cervical Squamous Cell Carcinoma and Endocervical Adenocarcinoma diagnostic slides. All tests were validated under a 5-fold cross-validation, with four key metrics, Scott-Knott (SK), and Borda count schemes. Both pathology data appear to promote InceptionV3 and DenseNet201. Yet, while VGG16 is a weak-performing approach for liquid-based cytology, it evinces promise in histopathology yielding 99.33% accuracy, 98.85% precision, 99.83% recall, and 99.34% F-measure.
The diagnosis of breast cancer in the early stages significantly decreases the mortality rate by allowing the choice of adequate treatment. This study developed and evaluated twenty-eight hybrid architectures combining seven recent deep learning techniques for feature extraction (DenseNet 201, Inception V3, Inception ReseNet V2, MobileNet V2, ResNet 50, VGG16, and VGG19), and four classifiers (MLP, SVM, DT, and KNN) for a binary classification of breast pathological images over the BreakHis and FNAC datasets. The designed architectures were evaluated using: (1) four classification performance criteria (accuracy, precision, recall, and F1-score), (2) Scott Knott (SK) statistical test to cluster the proposed architectures and identify the best cluster of the outperforming architectures, and (3) the Borda Count voting method to rank the best performing architectures. The results showed the potential of combining deep learning techniques for feature extraction and classical classifiers to classify breast cancer in malignant and benign tumors. The hybrid architecture using the MLP classifier and DenseNet 201 for feature extraction (MDEN) was the top performing architecture with higher accuracy values reaching 99% over the FNAC dataset, 92.61%, 92%, 93.93%, and 91.73% over the four magnification factor values of the BreakHis dataset: 40X, 100X, 200X, and 400X, respectively. The results of this study recommend the use of hybrid architectures using DenseNet 201 for the feature extraction of the breast cancer histological images because it gave the best results for both datasets BreakHis and FNAC, especially when combined with the MLP classifier.
Breast cancer (BC) is the leading cause of death among women worldwide. It affects in general women older than 40 years old. Medical images analysis is one of the most promising research areas since it provides facilities for diagnosis and decision-making of several diseases such as BC. This paper conducts a Structured Literature Review (SLR) of the use of Machine Learning (ML) and Image Processing (IP) techniques to deal with BC imaging. A set of 530 papers published between 2000 and August 2019 were selected and analyzed according to ten criteria: year and publication channel, empirical type, research type, medical task, machine learning techniques, datasets used, validation methods, performance measures and image processing techniques which include image pre-processing, segmentation, feature extraction and feature selection. Results showed that diagnosis was the most used medical task and that Deep Learning techniques (DL) were largely used to perform classification. Furthermore, we found out that classification was the most ML objective investigated followed by prediction and clustering. Most of the selected studies used Mammograms as imaging modalities rather than Ultrasound or Magnetic Resonance Imaging with the use of public or private datasets with MIAS as the most frequently investigated public dataset. As for image processing techniques, the majority of the selected studies pre-process their input images by reducing the noise and normalizing the colors, and some of them use segmentation to extract the region of interest with the thresholding method. For feature extraction, we note that researchers extracted the relevant features using classical feature extraction techniques (e.g. Texture features, Shape features, etc.) or DL techniques (e. g. VGG16, VGG19, ResNet, etc.), and finally few papers used feature selection techniques in particular the filter methods.
Diagnosis of breast cancer in the early stages allows to significantly decrease the mortality rate by allowing to choose the adequate treatment. This paper develops and evaluates twenty-eight hybrid architectures combining seven recent deep learning techniques for feature extraction (DenseNet 201, Inception V3, Inception ReseNet V2, MobileNet V2, ResNet 50, VGG16 and VGG19), and four classifiers (MLP, SVM, DT and KNN) for binary classification of breast cytological images over the FNAC dataset. To evaluate the designed architectures, we used: (1) four classification performance criterias (accuracy, precision, recall and F1-score), (1) Scott Knott (SK) statistical test to cluster the developed architectures and identify the best cluster of the outperforming architectures, and (2) Borda Count voting method to rank the best performing architectures. Results showed the potential of combining deep learning techniques for feature extraction and classical classifiers to classify breast cancer in malignant and benign tumors. The hybrid architectures using MLP classifier and DenseNet 201 for feature extraction were the top performing architectures with a higher accuracy value reaching 99% over the FNAC dataset. As results, the findings of this study recommend the use of the hybrid architectures using DenseNet 201 for the feature extraction of the breast cancer cytological images since it gave the best results for the FNAC data images, especially if combined with the MLP classifier.