Accurate classification of cancer subtypes is crucial for personalised therapies and targeted interventions. In this study, we propose BioGAT-LGG, a deep learning framework that integrates multi-omics data, including mRNA, miRNA, and DNA methylation, using a correlation-based Graph Attention Network version 2 (GATv2) for biomarker discovery and Lower-Grade Glioma (LGG) subtype classification. Unlike existing methodologies that rely on external biological priors, such as protein-protein interaction networks or reference graphs, BioGAT-LGG constructs gene-driven correlation graphs, enabling the model to learn biologically meaningful molecular interactions. To improve feature interpretability and reduce dimensionality, LASSO regression is performed during model training. The model achieved 98.03% accuracy, with precision (98.12%), recall (97.74%), and F1-score (97.87%) in a stratified 10-fold cross-validation. Extensive analysis and enrichment of known cancer-related pathways, including PI3K-Akt signalling, Small Cell Lung Cancer, and Transcriptional Misregulation in Cancer, identified the biomarkers hsa-mir-3936, MTCO1P40, and CCND2, which were subsequently validated. These results indicate that BioGAT-LGG effectively captures biologically validated mechanisms and can enable clinically significant subtype classification and biomarker-guided decision-making. This framework thus lays a scalable foundation for multi-omics integration in oncology, which can be further adopted in other tumour types.
Background: Alzheimer disease (AD) is a neurodegenerative condition that progressively develops structural changes in the brain, resulting in different stages of severity, which makes accurate multiclass classification from magnetic resonance imaging (MRI) challenging. Despite promising outcomes of deep learning models, a great number of current methods disregard disease progression, suffer from evaluation leakage, or lack interpretability. Objectives: This paper introduces DeepAttentionADNet, a lightweight hybrid CNN–Transformer framework designed for multiclass staging of Alzheimer’s disease using MRI images. Methods: The proposed model integrates convolutional feature extraction with transformer-based global context modeling. To capture the ordered nature of disease severity, a progression-aware ordinal learning objective is proposed. Moreover, consistency regularization is utilized to enhance robustness by imposing consistent prediction with spatial perturbation. A leakage-free k-fold cross-validation protocol is adopted, in which data splitting is performed prior to augmentation. Also, to promote interpretability, token-level importance maps based on transformer embeddings are utilized to visualize spatial regions that were used to make classification decisions. Results: The experimental findings on a multiclass MRI dataset of Alzheimer disease demonstrate consistent and high performance across cross-validation folds (mean F1-score (0.991 ± 0.003) and AUROC (0.9998 ± 0.0002)), without losing transparency and progress awareness. Conclusions: The proposed framework provided a robust and interpretable method of Alzheimer disease severity classification using MRI.
The integration of multi-omics data presents a major challenge in precision medicine, requiring advanced computational methods for accurate disease classification and biological interpretation. This study introduces the Multi-Omics Graph Kolmogorov-Arnold Network (MOGKAN), a deep learning model that integrates messenger RNA, micro RNA sequences, and DNA methylation data with Protein-Protein Interaction (PPI) networks for accurate and interpretable cancer classification across 31 cancer types. MOGKAN employs a hybrid approach combining differential expression with DESeq2, Linear Models for Microarray (LIMMA), and Least Absolute Shrinkage and Selection Operator (LASSO) regression to reduce multi-omics data dimensionality while preserving relevant biological features. The model architecture is based on the Kolmogorov-Arnold theorem principle, using trainable univariate functions to enhance interpretability and feature analysis. MOGKAN achieves classification accuracy of 96.28 percent and demonstrates low experimental variability with a standard deviation that is reduced by 1.58 to 7.30 percents compared to Convolutional Neural Networks (CNNs) and Graph Neural Networks (GNNs). The biomarkers identified by MOGKAN have been validated as cancer-related markers through Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analysis. The proposed model presents an ability to uncover molecular oncogenesis mechanisms by detecting phosphoinositide-binding substances and regulating sphingolipid cellular processes. By integrating multi-omics data with graph-based deep learning, our proposed approach demonstrates superior predictive performance and interpretability that has the potential to enhance the translation of complex multi-omics data into clinically actionable cancer diagnostics.
Recent advances in deep learning have expanded the potential for predictive modeling in survival analysis, particularly in high-dimensional datasets with time-varying covariates. This paper applies deep learning approaches, DeepSurv, DeepHit, and Dynamic DeepHit, to model HIV incidence (time-to-event outcome) using high-dimensional longitudinal data, incorporating time-varying cytokine profiles alongside baseline covariates. We employ the time-dependent concordance index (C-index) and Brier scores to assess the models’ predictive accuracy. We also address missing data using missForest, evaluating model performance on imputed and complete-case datasets. Different strategies for integrating cytokine profiles were explored: DeepSurv and DeepHit utilized derived variables, mean, and difference between the first and last measurements, while Dynamic DeepHit preserved the original time-varying nature of the cytokine data. Our findings demonstrate that retaining the dynamic nature of cytokine covariates, rather than relying on derived summary measures, underscores the robustness and suitability of Dynamic DeepHit as a clinical prediction model, particularly in scenarios where key variables evolve over time.
(1) Background: Circular RNAs (circRNAs) are covalently closed single-stranded molecules that play crucial roles in gene regulation, while microRNAs (miRNAs), specifically mature microRNAs, are naturally occurring small molecules of non-coding RNA with 17-25-nucleotide sizes. Understanding circRNA–miRNA interactions (CMIs) can reveal new approaches for diagnosing and treating complex human diseases. (2) Methods: In this paper, we propose a novel approach for predicting CMIs based on a graph attention network (GAT). We utilized DNABERT to extract molecular features of the circRNA and miRNA sequences and role-based graph embeddings generated by Role2Vec to extract the CMI features. The GAT’s ability to learn complex node dependencies in biological networks provided enhanced performance over the existing methods and the traditional deep neural network models. (3) Results: Our simulation studies showed that our GAT model achieved accuracies of 0.8762 and 0.8837 on the CMI-9905 and CMI-9589, respectively. These accuracies were the highest among the other existing CMI prediction methods. Our GAT method also achieved the highest performance as measured by the precision, recall, F1-score, area under the receiver operating characteristic (AUROC) curve, and area under the precision–recall curve (AUPR). (4) Conclusions: These results reflect the GAT’s ability to capture the intricate relationships between circRNAs and miRNAs, thus offering an efficient computational approach for prioritizing potential interactions for experimental validation.
This study addresses the pressing need for improved lung cancer diagnosis and treatment by leveraging computational methods and omics data analysis. Lung cancer remains a leading cause of cancer-related deaths globally, highlighting the urgency for more effective diagnostic and therapeutic approaches. Current diagnostic methods, such as imaging and biopsies, suffer from limitations in sensitivity, specificity, and accessibility, often due to factors such as poor data quality, small sample sizes, and variability in data sources. These limitations highlight the necessity for the development of advanced noninvasive techniques. Computational methods utilizing omics data have shown promise in overcoming these challenges by comprehensively understanding the molecular pathways involved in lung cancer. We propose a novel approach that utilizes RNA-Seq data and employs LASSO regression with attention mechanisms to identify lung cancer biomarkers. Our results demonstrate the effectiveness of this approach in identifying potential biomarkers for lung cancer, including well-known genes such as TP53, EGFR, KRAS, ALK, and PIK3CA, validating the model's ability to uncover key genes associated with lung cancer development and progression. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analyses revealed significant associations of the identified genes with critical biological processes and pathways, including protein synthesis, folding, cell adhesion, gene regulation, and immune responses. The PPI network analysis, constructed using the STRING database and Cytoscape application, highlighted a highly interconnected interaction landscape, with central hub genes playing pivotal roles in lung cancer progression. RPSA emerged as a crucial hub gene, consistently identified across different centrality measures. This study sheds light on the potential of computational methods and omics data analysis in improving lung cancer diagnosis and treatment, offering new insights for future research directions and personalized medicine strategies.
The integration of heterogeneous multi-omics datasets at a systems level remains a central challenge for developing analytical and computational models in precision cancer diagnostics. This paper introduces Multi-Omics Graph Kolmogorov-Arnold Network (MOGKAN), a deep learning framework that utilizes messenger-RNA, micro-RNA sequences, and DNA methylation samples together with Protein-Protein Interaction (PPI) networks for cancer classification across 31 different cancer types. The proposed approach combines differential gene expression with DESeq2, Linear Models for Microarray (LIMMA), and Least Absolute Shrinkage and Selection Operator (LASSO) regression to reduce multi-omics data dimensionality while preserving relevant biological features. The model architecture is based on the Kolmogorov-Arnold theorem principle and uses trainable univariate functions to enhance interpretability and feature analysis. MOGKAN achieves classification accuracy of 96.28% and exhibits low experimental variability in comparison to related deep learning-based models. The biomarkers identified by MOGKAN were validated as cancer-related markers through Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analysis. By integrating multi-omics data with graph-based deep learning, our proposed approach demonstrates robust predictive performance and interpretability with potential to enhance the translation of complex multi-omics data into clinically actionable cancer diagnostics.
Breast cancer remains a global health burden, with an increase in deaths related to this particular cancer. Accurately predicting and diagnosing breast cancer is important for treatment development and survival of patients. This study aimed to accurately predict breast cancer using a dataset comprising 1208 observations and 3602 genes. The study employed feature selection techniques to identify the most influential predictive genes for breast cancer using machine learning (ML) models. The study used K-nearest Neighbors (KNN), random forests (RF), and a support vector machine (SVM). Furthermore, the study employed feature- and model-based importance and explainable ML methods, including Shapley values, Partial dependency (PDPS), and Accumulated Local Effects (ALE) plots, to explain the genes’ importance ranking from the ML methods. Shapley values highlighted the significance of some of the genes in predicting cancer presence. Model-based feature ranking techniques, particularly the Leaving-One-Covariate-In (LOCI) method, identified the ten most critical genes for predicting tumor cases. The LOCI rankings from the SVM and RF methods were aligned. Additionally, visualization methods such as PDPS and ALE plots demonstrated how individual feature changes affect predictions and interactions with other genes. By combining feature selection techniques and explainable ML methods, this study has demonstrated the interpretability and reliability of machine learning models for breast cancer prediction, emphasizing the importance of incorporating explainable ML approaches for medical decision-making.
Recent studies on integrating multiple omics data highlighted the potential to advance our understanding of the cancer disease process. Computational models based on graph neural networks and attention-based architectures have demonstrated promising results for cancer classification due to their ability to model complex relationships among biological entities. However, challenges related to addressing the high dimensionality and complexity in integrating multi-omics data, as well as in constructing graph structures that effectively capture the interactions between nodes, remain active areas of research. This study evaluates graph neural network architectures for multi-omics (MO) data integration based on graph-convolutional networks (GCN), graph-attention networks (GAT), and graph-transformer networks (GTN). Differential gene expression and LASSO (Least Absolute Shrinkage and Selection Operator) regression are employed for reducing the omics data dimensionality and feature selection; hence, the developed models are referred to as LASSO-MOGCN, LASSO-MOGAT, and LASSO-MOGTN. Graph structures constructed using sample correlation matrices and protein-protein interaction networks are investigated. Experimental validation is performed with a dataset of 8,464 samples from 31 cancer types and normal tissue, comprising messenger-RNA, micro-RNA, and DNA methylation data. The results show that the models integrating multi-omics data outperformed the models trained on single omics data, where LASSO-MOGAT achieved the best overall performance, with an accuracy of 95.9%. The findings also suggest that correlation-based graph structures enhance the models’ ability to identify shared cancer-specific signatures across patients in comparison to protein-protein interaction networks-based graph structures. The code and data used in this study are available in the link (https://github.com/FadiAlharbi2024/Graph_Based_Architecture.git).
Background: Cancer survival prediction is vital in improving patients’ prospects and recommending therapies. Understanding the molecular behavior of cancer can be enhanced through the integration of multi-omics data, including mRNA, miRNA, and DNA methylation data. In light of these multi-omics data, we proposed a graph attention network (GAT) model in this study to predict the survival of non-small cell lung cancer (NSCLC). Methods: The different omics data were obtained from The Cancer Genome Atlas (TCGA) and preprocessed and combined into a single dataset using the sample ID. We used the chi-square test to select the most significant features to be used in our model. We used the synthetic minority oversampling technique (SMOTE) to balance the dataset and the concordance index (C-index) to measure the performance of our model on different combinations of omics data. Results: Our model demonstrated superior performance, with the highest value of the C-index obtained when we used both mRNA and miRNA data. This demonstrates that the multi-omics approach could be effective in predicting survival. Further pathway analysis conducted with KEGG showed that our GAT model provided high weights to the features that are associated with the viral entry pathways, such as the Epstein–Barr virus and Influenza A pathways, which are involved in lung cancer development. From our findings, it can be observed that the proposed GAT model leads to a significantly improved prediction of survival by exploiting the strengths of multiple omics datasets and the findings from the enriched pathways. Our GAT model outperforms other state-of-the-art methods that are used for NSCLC prediction. Conclusions: In this study, we developed a new model for the survival prediction of NSCLC using the GAT based on multi-omics data. Our model showed outstanding predictive values, and the KEGG analysis of the selected significant features showed that they were implicated in pivotal biological processes underlying pathways such as Influenza A and the Epstein–Barr virus infection, which are linked to lung cancer progression.
HIV remains a critical global health issue, with an estimated 39.9 million people living with the virus worldwide by the end of 2023 (according to WHO). Although the epidemic’s impact varies significantly across regions, Africa remains the most affected. In the past decade, considerable efforts have focused on developing preventive measures, such as vaccines and pre-exposure prophylaxis, to combat sexually transmitted HIV. Recently, cytokine profiles have gained attention as potential predictors of HIV incidence due to their involvement in immune regulation and inflammation, presenting new opportunities to enhance preventative strategies. However, the high-dimensional, time-varying nature of cytokine data collected in clinical research, presents challenges for traditional statistical methods like the Cox proportional hazards (PH) model to effectively analyze survival data related to HIV. Machine learning (ML) survival models offer a robust alternative, especially for addressing the limitations of the PH model’s assumptions. In this study, we applied survival support vector machine (SSVM) and random survival forest (RSF) models using changes or means in cytokine levels as predictors to assess their association with HIV incidence, evaluate variable importance, measure predictive accuracy using the concordance index (C-index) and integrated Brier score (IBS) and interpret the model’s predictions using Shapley additive explanations (SHAP) values. Our results indicated that RSFs models outperformed SSVMs models, with the difference covariate model performing better than the mean covariate model. The highest C-index for SSVM was 0.7180 under the difference covariate model, while for RSF, it reached 0.8801 under the difference covariate model using the log-rank split rule. Key cytokines identified as positive predictors of HIV incidence included TNF-A, BASIC-FGF, IL-5, MCP-3, and EOTAXIN, while 29 cytokines were negative predictors. Baseline factors such as condom use frequency, treatment status, number of partners, and sexual activity also emerged as significant predictors. This study underscored the potential of cytokine profiles for predicting HIV incidence and highlighted the advantages of RSFs models in analyzing high-dimensional, time-varying data over SSVMs. It further through ablation studies emphasized the importance of selecting key features within mean and difference based covariate models to achieve an optimal balance between model complexity and predictive accuracy.
The application of machine learning methods to analyze changes in gene expression patterns has recently emerged as a powerful approach in cancer research, enhancing our understanding of the molecular mechanisms underpinning cancer development and progression. Combining gene expression data with other types of omics data has been reported by numerous works to improve cancer classification outcomes. Despite these advances, effectively integrating high-dimensional multi-omics data and capturing the complex relationships across different biological layers remains challenging. This paper introduces LASSO-MOGAT (LASSO-Multi-Omics Gated ATtention), a novel graph-based deep learning framework that integrates messenger RNA, microRNA, and DNA methylation data to classify 31 cancer types. Utilizing differential expression analysis with LIMMA and LASSO regression for feature selection, and leveraging Graph Attention Networks (GATs) to incorporate protein-protein interaction (PPI) networks, LASSO-MOGAT effectively captures intricate relationships within multi-omics data. Experimental validation using five-fold cross-validation demonstrates the method's precision, reliability, and capacity for providing comprehensive insights into cancer molecular mechanisms. The computation of attention coefficients for the edges in the graph by the proposed graph-attention architecture based on protein-protein interactions proved beneficial for identifying synergies in multi-omics data for cancer classification.
IntroductionUnderstanding and identifying the immunological markers and clinical information linked with HIV acquisition is crucial for effectively implementing Pre-Exposure Prophylaxis (PrEP) to prevent HIV acquisition. Prior analysis on HIV incidence outcomes have predominantly employed proportional hazards (PH) models, adjusting solely for baseline covariates. Therefore, models that integrate cytokine biomarkers, particularly as time-varying covariates, are sorely needed.MethodsWe built a simple model using the Cox PH to investigate the impact of specific cytokine profiles in predicting the overall HIV incidence. Further, Kaplan-Meier curves were used to compare HIV incidence rates between the treatment and placebo groups while assessing the overall treatment effectiveness. Utilizing stepwise regression, we developed a series of Cox PH models to analyze 48 longitudinally measured cytokine profiles. We considered three kinds of effects in the cytokine profile measurements: average, difference, and time-dependent covariate. These effects were combined with baseline covariates to explore their influence on predictors of HIV incidence.ResultsComparing the predictive performance of the Cox PH models developed using the AIC metric, model 4 (Cox PH model with time-dependent cytokine) outperformed the others. The results indicated that the cytokines, interleukin (IL-2, IL-3, IL-5, IL-10, IL-16, IL-12P70, and IL-17 alpha), stem cell factor (SCF), beta nerve growth factor (B-NGF), tumor necrosis factor alpha (TNF-A), interferon (IFN) alpha-2, serum stem cell growth factor (SCG)-beta, platelet-derived growth factor (PDGF)-BB, granulocyte macrophage colony-stimulating factor (GM-CSF), tumor necrosis factor-related apoptosis-inducing ligand (TRAIL), and cutaneous T-cell-attracting chemokine (CTACK) were significantly associated with HIV incidence. Baseline predictors significantly associated with HIV incidence when considering cytokine effects included: age of oldest sex partner, age at enrollment, salary, years with a stable partner, sex partner having any other sex partner, husband's income, other income source, age at debut, years lived in Durban, and sex in the last 30 days.DiscussionOverall, the inclusion of cytokine effects enhanced the predictive performance of the models, and the PrEP group exhibited reduced HIV incidences compared to the placebo group.
Lung cancer, a life-threatening disease primarily affecting lung tissue, remains a significant contributor to mortality in both developed and developing nations. Accurate biomarker identification is imperative for effective cancer diagnosis and therapeutic strategies. This study introduces the Voting-Based Enhanced Binary Ebola Optimization Search Algorithm (VBEOSA), an innovative ensemble-based approach combining binary optimization and the Ebola optimization search algorithm. VBEOSA harnesses the collective power of the state-of-the-art classification models through soft voting. Moreover, our research applies VBEOSA to an extensive lung cancer gene expression dataset obtained from TCGA, following essential preprocessing steps including outlier detection and removal, data normalization, and filtration. VBEOSA aids in feature selection, leading to the discovery of key hub genes closely associated with lung cancer, validated through comprehensive protein-protein interaction analysis. Notably, our investigation reveals ten significant hub genes-ADRB2, ACTB, ARRB2, GNGT2, ADRB1, ACTG1, ACACA, ATP5A1, ADCY9, and ADRA1B-each demonstrating substantial involvement in the domain of lung cancer. Furthermore, our pathway analysis sheds light on the prominence of strategic pathways such as salivary secretion and the calcium signaling pathway, providing invaluable insights into the intricate molecular mechanisms underpinning lung cancer. We also utilize the weighted gene co-expression network analysis (WGCNA) method to identify gene modules exhibiting strong correlations with clinical attributes associated with lung cancer. Our findings underscore the efficacy of VBEOSA in feature selection and offer profound insights into the multifaceted molecular landscape of lung cancer. Finally, we are confident that this research has the potential to improve diagnostic capabilities and further enrich our understanding of the disease, thus setting the stage for future advancements in the clinical management of lung cancer. The VBEOSA source codes is publicly available at https://github.com/TEHNAN/VBEOSA-A-Novel-Feature-Selection-Algorithm-for-Identifying-hub-Genes-in-Lung-Cancer .
Background Intimate partner violence (IPV) remains a global public health concern for both men and women. Spatial mapping and clustering analysis can reveal subtle patterns in IPV occurrences but are yet to be explored in Rwanda, especially at a lower small-area scale. This study seeks to examine the spatial distribution, patterns, and associated factors of IPV among men and women in Rwanda.Methods This was a secondary data analysis of the 2019/2020 Rwanda Demographic and Health Survey (RDHS) individual-level data set for 1947 women aged 15-49 years and 1371 men aged 15-59 years. A spatially structured additive logistic regression model was used to assess risk factors for IPV while adjusting for spatial effects. The district-level spatial model was adjusted for fixed covariate effects and was implemented using a fully Bayesian inference within the generalized additive mixed effects framework.Results IPV prevalence amongst women was 45.9% (95% Confidence interval (CI): 43.4-48.5%) while that for men was 18.4% (95% CI: 16.2-20.9%). Using a bivariate choropleth, IPV perpetrated against women was higher in the North-Western districts of Rwanda whereas for men it was shown to be more prevalent in the Southern districts. A few districts presented high IPV for both men and women. The spatial structured additive logistic model revealed higher odds for IPV against women mainly in the North-western districts and the spatial effects were dominated by spatially structured effects contributing 64%. Higher odds of IPV were observed for men in the Southern districts of Rwanda and spatial effects were dominated by district heterogeneity accounting for 62%. There were no statistically significant district clusters for IPV in both men or women. Women with partners who consume alcohol, and with controlling partners were at significantly higher odds of IPV while those in rich households and making financial decisions together with partners were at lower odds of experiencing IPV.Conclusion Campaigns against IPV should be strengthened, especially in the North-Western and Southern parts of Rwanda. In addition, the promotion of girl-child education and empowerment of women can potentially reduce IPV against women and girls. Furthermore, couples should be trained on making financial decisions together. In conclusion, the implementation of policies and interventions that discourage alcohol consumption and control behaviour, especially among men, should be rolled out.
Breast cancer (BC) is the most incident cancer type among women. BC is also ranked as the second leading cause of death among all cancer types. Therefore, early detection and prediction of BC are significant for prognosis and in determining the suitable targeted therapy. Early detection using morphological features poses a significant challenge for physicians. It is therefore important to develop computational techniques to help determine informative genes, and hence help diagnose cancer in its early stages. Eight common hub genes were identified using three methods: the maximal clique centrality (MCC), the maximum neighborhood component (MCN), and the node degree. The hub genes obtained were CDK1, KIF11, CCNA2, TOP2A, ASPM, AURKB, CCNB2, and CENPE. Enrichment analysis revealed that the differentially expressed genes (DEGs) influenced multiple pathways. The most significant identified pathways were focal adhesion, ECM-receptor interaction, melanoma, and prostate cancer pathways. Additionally, survival analysis using Kaplan–Meier was conducted, and the results showed that the obtained eight hub genes are promising candidate genes to serve as prognostic and diagnostic biomarkers for BC. Furthermore, a correlation study between the clinicopathological factors in BC and the eight hub genes was performed. The results showed that all eight hub genes are associated with the clinicopathological variables of BC. Using an integrated analysis of RNASeq and microarray data, a protein-protein interaction (PPI) network was developed. Eight hub genes were identified in this study, and they were validated using previous studies. Additionally, Kaplan-Meier was used to verify the prognostic value of the obtained hub genes.
Cancer diagnosis and treatment depend on accurate cancer-type prediction. A prediction model can infer significant cancer features (genes). Gene expression is among the most frequently used features in cancer detection. Deep Learning (DL) architectures, which demonstrate cutting-edge performance in many disciplines, are not appropriate for the gene expression data since it contains a few samples with thousands of features. This study presents an approach that applies three feature selection techniques (Lasso, Random Forest, and Chi-Square) on gene expression data obtained from Pan-Cancer Atlas through the TCGA Firehose Data using R statistical software version 4.2.2. We calculated the feature importance of each selection method. Then we calculated the mean of the feature importance to determine the threshold for selecting the most relevant features. We constructed five models with a simple convolutional neural networks (CNNs) architecture, which are trained using the selected features and then selected the winning model. The winning model achieved a precision of 94.11%, a recall of 94.26%, an F1-score of 94.14%, and an accuracy of 96.16% on a test set.
Breast cancer is considered one of the significant health challenges and ranks among the most prevalent and dangerous cancer types affecting women globally. Early breast cancer detection and diagnosis are crucial for effective treatment and personalized therapy. Early detection and diagnosis can help patients and physicians discover new treatment options, provide a more suitable quality of life, and ensure increased survival rates. Breast cancer detection using gene expression involves many complexities, such as the issue of dimensionality and the complicatedness of the gene expression data. This paper proposes a bio-inspired CNN model for breast cancer detection using gene expression data downloaded from the cancer genome atlas (TCGA). The data contains 1208 clinical samples of 19,948 genes with 113 normal and 1095 cancerous samples. In the proposed model, Array-Array Intensity Correlation (AAIC) is used at the pre-processing stage for outlier removal, followed by a normalization process to avoid biases in the expression measures. Filtration is used for gene reduction using a threshold value of 0.25. Thereafter the pre-processed gene expression dataset was converted into images which were later converted to grayscale to meet the requirements of the model. The model also uses a hybrid model of CNN architecture with a metaheuristic algorithm, namely the Ebola Optimization Search Algorithm (EOSA), to enhance the detection of breast cancer. The traditional CNN and five hybrid algorithms were compared with the classification result of the proposed model. The competing hybrid algorithms include the Whale Optimization Algorithm (WOA-CNN), the Genetic Algorithm (GA-CNN), the Satin Bowerbird Optimization (SBO-CNN), the Life Choice-Based Optimization (LCBO-CNN), and the Multi-Verse Optimizer (MVO-CNN). The results show that the proposed model determined the classes with high-performance measurements with an accuracy of 98.3%, a precision of 99%, a recall of 99%, an f1-score of 99%, a kappa of 90.3%, a specificity of 92.8%, and a sensitivity of 98.9% for the cancerous class. The results suggest that the proposed method has the potential to be a reliable and precise approach to breast cancer detection, which is crucial for early diagnosis and personalized therapy.
Continual application of nitrogen (N), phosphorous (P) and potassium (K) fertilizer may not return a profit to farmers due to the costs of application and the loss of NPK from soil in various ways. Thus, a combination of NPK granule with a porous biochar (termed here as BNPK) appears to offer multiple benefits resulting from the excellent properties of biochar. Given the lack of information on the properties of NPK and BNPK fertilizers, it is necessary to investigate the characteristics of both to achieve a good understanding of why BNPK granule is superior to NPK granule. Therefore, this study aims to investigate the characteristics of a maize straw biochar mixed with NPK granule, before and after application to soil, and compare them to those for a commercial NPK granule. The BNPK granule, with a greater surface area and porosity, showed a higher capacity to store and donate electrons than the NPK granule. Relatively lower concentrations of Ca, P, K, Si and Mg were dissolved from the BNPK, indicating the ability of the BNPK granule to maintain these mineral elements and reduce dissolution rate. To study the nutrient storage mechanism of the BNPK granule in the soil, short- and long-term leaching experiments were conducted. During the experiments, organo-mineral clusters, comprising C, P, K, Si, Al and Fe, were formed on the surface and inside the biochar pores. However, BNPK was not effective in reducing N leaching, in the absence of plants, in a red chromosol soil.
Background Cancer remains a major public health problem, especially in Sub-Saharan Africa (SSA) where the provision of health care is poor. This scoping review mapped evidence in the literature regarding the burden of cervical, breast and prostate cancers in SSA. Methods We conducted this scoping review using the Arksey and O'Malley framework, with five steps: identifying the research question; searching for relevant studies; selecting studies; charting the data; and collating, summarizing, and reporting the data. We performed all the steps independently and resolved disagreements through discussion. We used Endnote software to manage references and the Rayyan software to screen studies. Results We found 138 studies that met our inclusion criteria from 2,751 studies identified through the electronic databases. The majority were retrospective studies of mostly registries and patient files ( n = 77, 55.8%), followed by cross-sectional studies ( n = 51, 36.9%). We included studies published from 1990 to 2021, with a sharp increase from 2010 to 2021. The quality of studies was overall satisfactory. Most studies were done in South Africa ( n = 20) and Nigeria ( n = 17). The majority were on cervical cancer ( n = 93, 67.4%), followed by breast cancer (67, 48.6%) and the least were on prostate cancer (48, 34.8%). Concerning the burden of cancer, most reported prevalence and incidence. We also found a few studies investigating mortality, disability-adjusted life years (DALYs), and years of life lost (YLL). Conclusions We found many retrospective record review cross-sectional studies, mainly in South Africa and Nigeria, reporting the prevalence and incidence of cervical, breast and prostate cancer in SSA. There were a few systematic and scoping reviews. There is a scarcity of cervical, breast and prostate cancer burden studies in several SSA countries. The findings in this study can inform policy on improving the public health systems and therefore reduce cancer incidence and mortality in SSA.