Abstract Triple-negative breast cancer (TNBC) is an aggressive subtype characterized by the absence of estrogen, progesterone, and HER2 receptors, limiting effective targeted therapies. Increasing evidence suggests that metabolic reprogramming, a hallmark of TNBC progression, is driven by underlying epigenetic mechanisms such as DNA methylation. The represented study performed an integrative analysis of transcriptomic (RNA-seq) and methylome data to uncover the metabolic–epigenetic interplay in TNBC. Differential gene expression analysis using DESeq2 revealed significant dysregulation of key metabolic genes, including upregulation of genes encoding glycolytic and serine biosynthesis enzymes and downregulation of metabolic tumor suppressors. Genome-wide methylation profiling identified extensive cytosine–phosphate–guanine (CpG) hypermethylation events associated with transcriptional repression, particularly in promoter regions. Integrative analysis pinpointed a subset of metabolism-related genes exhibiting both differential expression and methylation, such as FBP1 , RASSF1A , and PHGDH . Pathway enrichment analysis highlighted aberrations in glycolysis/gluconeogenesis, fatty acid metabolism, and one-carbon pathways (adjusted p < 0.01). Importantly, TNBC patients with hypermethylated metabolic gene signatures displayed significantly shorter overall survival (log-rank p < 0.05). These findings reveal that DNA methylation-driven metabolic dysregulation contributes to TNBC aggressiveness and may provide novel biomarkers and therapeutic targets at the metabolic–epigenetic interface.
Abstract Parkinson’s disease (PD) is a complex neurodegenerative disorder with a substantial genetic component. Over the past decade, genome-wide association studies (GWAS) have identified numerous loci associated with PD risk; however, interpretation of these findings and their broader applicability remain challenging. In this systematic review, we synthesize results from 35 GWAS published between 2015 and 2025, encompassing diverse study designs and ancestries. Recurrent risk loci, including SNCA, LRRK2, MAPT, and GBA1, were consistently replicated across multiple studies, while several ancestry-specific associations were reported, particularly in East Asian and African ancestry cohorts. Nevertheless, representation of African, South Asian, and Latino populations remains limited, constraining the global generalizability of current findings. We also discuss methodological extensions beyond single-variant GWAS, including rare variant analyses, polygenic risk scores, and machine learning–based approaches, which have been applied to complement traditional analyses but remain primarily research tools due to limited validation and interpretability. Together, this review outlines the current genetic landscape of PD and identifies key methodological and population-based gaps that must be addressed to support robust and equitable translation of GWAS discoveries.
Genomic language models (gLMs) are rapidly becoming important tools for learning biological information directly from sequence data. By adapting concepts from natural language processing, these models aim to capture contextual dependencies, regulatory grammar, evolutionary constraint, and sequence-level functional patterns that may be difficult to detect using alignment-based, motif-based, or conventional supervised methods alone. This systematic review evaluates recent model-development studies of genomic, RNA, nucleotide, codon-level, and regulatory DNA language models, with emphasis on model architecture, tokenization, training objective, biological task, benchmarking strategy, and reported limitations. A structured search of PubMed, Scopus, and Web of Science identified 469 records. After duplicate removal, screening, and full-text eligibility assessment, 58 studies met the strict inclusion criteria for primary model development or substantial model adaptation. The included studies covered diverse applications, including regulatory sequence prediction, variant-effect modeling, genome annotation, microbial and viral genome analysis, RNA splicing and regulation, codon optimization, mRNA design, and generative design of regulatory or RNA sequences. Across the included studies, stronger evidence for gLM utility was generally associated with biologically informed or task-aligned model design, including evolutionary alignments, motif-aware objectives, long-context architectures, RNA structural priors, population-aware representations, and domain-specific pretraining. However, the evidence was heterogeneous and did not support a general claim of superiority over established bioinformatics tools or specialized supervised models. In several regulatory genomics tasks, specialized supervised models, k-mer-based approaches, or conventional deep-learning baselines remained competitive or superior to pretrained language-model representations. Generative models showed growing promise for RNA, codon, viral genome, and cis-regulatory element design, although many were evaluated mainly in silico. Overall, the field is advancing quickly, but broader impact will require standardized benchmarks, clearer reporting, stronger external validation, improved interpretability, and experimental confirmation of predicted or generated biological functions.
Alzheimer’s disease (AD) is a complicated disease that attacks the brain’s neurons. Numerous efforts have been directed towards identifying interactions of single-nucleotide polymorphisms (SNPs). Nevertheless, the large volume of SNP data leads to an explosion of high-order SNP combinations, which substantially limits the effectiveness of interaction detection. Consequently, investigating SNP-SNP interactions is crucial in precision medicine (PM). This paper presents two frameworks for identifying and visualizing SNP-SNP interactions associated with AD risk. The first framework aims to integrate ensemble learning techniques and multifactor dimensionality reduction (MDR). The promising outcomes of this framework presented significant risk genes and SNP-SNP interactions with high accuracy. The achieved classification accuracy of 5-way interaction models was 0.874. The accuracy of the 2-way, 3-way, and 4-way models was 0.6648, 0.7169, and 0.7878, respectively. In the second framework, a deep neural network (DNN) is employed with SHapley Additive exPlanations (SHAP) to identify the most highly ranked SNPs that suggest significant SNP-SNP interactions that could aid in interpreting AD risk. The classification accuracy of 5-way interaction models was 0.8. The classification accuracy of the pairwise, 3-way, and 4-way models was 0.65, 0.7, and 0.77, respectively. This study identifies potential risk genes and SNP-SNP interactions associated with AD risk. This work shows that LOC105374292, NSUN7, LOC101929507, and LINC01482 are four novel genes associated with AD by both proposed frameworks.
Cancer progression triggers various molecular changes that disrupt cellular functions, impacting DNA, proteins, and other biomolecules. To gain deeper insights into the development of breast cancer, a multi-omics strategy is crucial for integrating diverse molecular data. In this study, we explore the use of variational autoencoder (VAE) and maximum mean discrepancy VAE (MMD-VAE) to integrate and analyze multi-omics data related to breast cancer. This deep learning framework supports supervised feature extraction, dimensionality reduction, and classification of breast cancer samples. Standard VAEs may fall short in capturing meaningful latent representations due to the restrictive nature of the KL divergence, which can hinder the effective integration of complex datasets. MMD-VAE replaces the KL divergence with maximum mean discrepancy, offering a more adaptable and informative latent space that improves classification and prognosis performance. We evaluated the performance of VAE and MMD-VAE in three different sample sizes and demonstrated their robustness to breast cancer classification, molecular subtype grouping, and survival analysis. In the current work, we employed a VAE to tackle the imbalance of combined di-omics and tri-omics statistics in breast cancer studies. Our evaluation included random forests, partial least squares, naive Bayes, decision trees, neural networks, and lasso regularization. Our results show that the proposed framework can effectively integrate multi-omics data and improve breast cancer analysis. Compared to VAE, MMD-VAE performs better at clustering molecular subtypes, and integrating T-SNE with MMD-VAE enhances data preservation and separation of molecular subtypes. MMD-VAE has shown superior accuracy rates to standard VAE, capturing complex associations at multi-omics levels and improving interpretability. In large-scale studies, it has good predictive power and is computationally efficient. MMD-VAE captures the interactions of multiple omics datasets, increasing accuracy and interpretability in cancer research. The random forest and lasso regularization classifiers demonstrated maximum prediction accuracy, achieving AUC values of 0.99. The study provides a novel approach to breast cancer research and has potential applications in precision medicine.
Imaging genetics is one of the important keys to precision medicine that leads to personalized treatment based on a patient's genetics, phenotype, or psychosocial characteristics. It deepens the understanding of the mechanisms through which genetic variations contribute to neurological and psychiatric disorders. This systematic review overviews the methods and applications of imaging genetics in the context of neurological diseases, mentioning its potential role in personalized medicine. Following PRISMA guidelines, this review systematically analyzes 28 studies integrating genetic and neuroimaging data to explore disease mechanisms and their implications for precision medicine. Selected research included multiple neurological disorders, including frontotemporal dementia, Alzheimer's disease, bipolar disorder, schizophrenia, Parkinson's disease, and others. Voxel-based morphometry was the most common imaging technique, while frequently examined genetic variants included APOE, C9orf72, MAPT, GRN, COMT, and BDNF. Associations between these variants and regional gray matter loss (e.g., frontal, temporal, or subcortical regions) suggest that genetic risk factors play a key role in disease pathophysiology. Integrating genetic and neuroimaging analyses enhances our understanding of disease mechanisms and supports advancements in precision medicine.
Hepatocellular carcinoma (HCC), the most prevalent form of liver cancer, remains a major global health concern due to challenges in early detection and limited treatment options. Multi-omics technologies-such as genomics, proteomics, and metabolomics-enable comprehensive insights into the disease's molecular complexity. This systematic review explores how these approaches contribute to biomarker discovery, molecular classification, and personalized treatment in HCC research. METHODS:We conducted a structured review of 32 eligible studies, categorizing their computational methodologies into five primary analytical frameworks: survival analysis, unsupervised clustering, supervised machine learning, differential expression analysis, and pathway/network analysis. Notably, unsupervised clustering and supervised machine learning approaches, such as support vector machines, random forests, and deep learning models, were frequently used for subtype classification, feature selection, and predictive modeling. RESULTS:The review identified that multi-omics approaches are widely used to discover biomarkers, classify HCC subtypes, and predict treatment responses. Common methods include clustering and machine learning. However, clinical validation remains limited, highlighting a gap in translational applicability. CONCLUSION:From a clinical perspective, multi-omics integration coupled with machine learning holds immense potential for improving early diagnosis, patient stratification, and therapeutic targeting. However, challenges related to data integration, interpretability, and cohort diversity must be addressed to realize this potential. This review underscores the transformative role of machine learning-enhanced multi-omics in reshaping liver cancer diagnosis and treatment and outlines future directions to bridge the gap between computational advances and clinical application.
Recently, antibodies have been considered essential therapeutics for combating diseases, especially viral infections. However, limited information on antibody structures has impeded their development. The interactions promoting antigen binding are determined by the structures of the six loops that comprise the complementarity-determining regions (CDRs). Artificial Intelligence (AI) technologies are overcoming many limitations by using co-evolution information from homologous proteins to predict protein function and structure and develop drug discovery. This study introduces an overview of computational methods for artificial intelligence for producing antibody structure and design and covers databases used, CDR loops, structural elements crucial to binding, and computer predictors of antibody structure and features; it highlights that Deep Learning (DL) techniques are essential for improving antibody structure prediction. Using various principles and methods, these techniques have significantly advanced the ability to predict novel antibody structures for binding.
The rapid spread of the novel coronavirus disease 2019 (COVID-19) pandemic underscores the need for early and accurate detection to enable effective treatment and containment. In this context, computational methods have shown considerable promise. This study presents an efficient end-to-end deep learning approach for detecting COVID-19, among other human coronavirus (HCoV) diseases. The proposed approach employs genomic image processing (GIP) techniques to convert complete and partial HCoV genome sequences into grayscale genomic images using the frequency chaos game representation method. These images are analyzed using a novel deep learning model, DeepCOVID-19, in conjunction with the pre-trained AlexNet model. Evaluation on a comprehensive dataset of diverse HCoV variants demonstrates that the DeepCOVID-19 model achieves an accuracy of 99.84
Hepatitis C virus (HCV) is a major global health concern, affecting millions of individuals worldwide. While existing literature predominantly focuses on disease classification using clinical data, there exists a critical research gap concerning HCV genotyping based on genomic sequences. Accurate HCV genotyping is essential for patient management and treatment decisions. While the neural models excel at capturing complex patterns, they still face challenges, such as data scarcity, that exist a lot in computational genomics. To overcome this challenges, this paper introduces an advanced deep learning approach for HCV genotyping based on the graphical representation of nucleotide sequences that outperforms classical approaches. Notably, it is effective for both partial and complete HCV genomes and addresses challenges associated with imbalanced datasets. In this work, ten HCV genotypes: 1a, 1b, 2a, 2b, 2c, 3a, 3b, 4, 5, and 6 were used in the analysis. This study utilizes Chaos Game Representation for 2D mapping of genomic sequences, employing self-supervised learning using convolutional autoencoder for deep feature extraction, resulting in an outstanding performance for HCV genotyping compared to various machine learning and deep learning models. This baseline provides a benchmark against which the performance of the proposed approach and other models can be evaluated. The experimental results showcase a remarkable classification accuracy of over 99%, outperforming traditional deep learning models. This performance demonstrates the capability of the proposed model to accurately identify HCV genotypes in both partial and complete sequences and in dealing with data scarcity for certain genotypes. The results of the proposed model are compared to NCBI genotyping tool.
Navigating a busy cityscape with a fleet of autonomous vehicles requires each to seamlessly maneuver through traffic with split-second decisions. Path planning is the backbone of such advanced machinery applications, from mobile robots to unmanned ground vehicles, where the choice of data structure plays a pivotal role in determining memory usage, planning time, and algorithm reliability. This research rigorously evaluates Grassfire, Dijkstra, A *, and RRT * algorithms based on key metrics using real GPS readings, across diverse environment representations and obstacle conditions. Our findings provide guidance for selecting the optimal algorithms and data structures tailored to specific environmental complexities. By evaluating the performance of these algorithms under various environmental conditions, the study offers insights that can help researchers and practitioners choose the most suitable algorithms and data structures for their autonomous vehicle applications. The ability to match the algorithm and data structure to the specific environmental challenges faced by autonomous vehicles is crucial for ensuring efficient and reliable path planning, which is the backbone of advanced machinery applications.
A brain computer interface (BCI) records the activities of the brain and classifies it into different classes. BCIs can be used by both severely motor disabled as well as healthy people to control devices. In this work we have concentrated on the development and application of a novel medical technology to measure the patientâs brain activity, translated it with intelligent software, and used the translated signals to drive patient-specific effectors. In this work, we deal with the EEG pattern recognition approach based on brain computer interfaces. Electroencephalographic (EEG) signals produced by the brain are used as input to our BCI system. Both offline and online BCI approaches are introduced where the offline approach was done using Dataset IA motor imagery EEG recordings and the online approach was done using our own BCI system. We have described our BCI system and its efficiency for moving the hands to right or left online. First, the measurement of the EEG and the components of a BCI system are explained. Second, the data acquisition system we developed is described in detail. Lastly, our BCI system, including all different techniques used for artifact removal, feature extraction, and classification is presented. Our results give an ideal solution for people with severe neuromuscular disorders, such as Amyotrophic Lateral Sclerosis (ALS) or spinal cord injury, people who are totally paralyzed, or âlocked-inâ, helping them to have a communication channel with others
Lung cancer has a high incidence rate and is considered highly fatal because of its low survival rate at early stages compared to other cancers. Computed tomography (CT) scans can reveal pulmonary nodules of different shapes and volumes in two dimensional (2D) slices. Three-dimensional (3D) reconstruction of pulmonary nodules can assist the radiologist in early treatment appropriate for the 3D nodule volume screened. In this research, we present a 3D reconstruction algorithm that uses 2D CT slices to reconstruct a 3D lung nodule. The equivalent diameters of small nodules ranged from 3 to 30 mm. A segmentation approach (based on bounding boxes and maximum intensity projection) was applied. Extracting the lung nodules from the 2D candidate masses was performed via a rule-based classifier. Surface rendering was used to reconstruct 3D pulmonary nodules which were visualized on the 3D Slicer software. The 3D nodule volume, as well as the accuracy rate and error of volume estimation were calculated. The proposed methodology was validated against the actual volumes of 14 3D nodules from the Lung Image Database Consortium (LIDC) database. The proposed algorithm achieved a maximum accuracy of 99.6627 % for lung nodule volume estimation. The corresponding average accuracy rate and average percentage error were 97.34 % and 2.66 %, respectively. The screening of 3D lung nodules can support surgery planning via nodule volume estimation. The average accuracy and error rates of the 3D reconstruction algorithm showed promising results in comparison with other published studies.
Alzheimer's disease (AD) is a complex disorder with strong genetic factors. The proposed framework is applied to Alzheimer's Disease Neuroimaging Initiative (ADNI) dataset. We present a novel framework integrating ensemble learning and MDR constructive induction algorithm to discover epistasis interactions associated with AD in a computationally efficient method. Discovering epistasis interactions is a big challenge and significantly impacts personalized medicine (PM). The applied ensemble learning algorithms are random forests (RF) with Gini index and permutation importance, Extreme Gradient Boosting (XGBoost), and classification and regression trees (CART). The classification accuracy of 5-way models varied between (0.8674–0.8758), whereas the accuracy of 2-way, 3-way, and 4-way models varied between (0.6515–0.6649), (0.7071–0.7170), and (0.7811–0.7878) respectively. The promising results of this proposed framework show high-ranked risk genes and up to 5-way epistasis models that contribute to the disease risk efficiently and at higher accuracy.
Coronavirus (CoV) disease 2019 (COVID-19) is a severe pandemic affecting millions worldwide. Due to its rapid evolution, researchers have been working on developing diagnostic approaches to suppress its spread. This study presents an effective automated approach based on genomic image processing (GIP) techniques to rapidly detect COVID-19, among other human CoV diseases, with high acceptable accuracy. The GIP technique was applied as follows: first, genomic graphical mapping techniques were used to convert the genome sequences into genomic grayscale images. The frequency chaos game representation (FCGR) and single gray-level representation (SGLR) techniques were used in this investigation. Then, several statistical features were obtained from the images to train and test many classifiers, including the k-nearest neighbors (KNN). This study aimed to determine the efficacy of the FCGR (with different orders) and SGLR images for accurately detecting COVID-19, using a dataset containing both partial and complete genome sequences. The results recommended the fourth-order FCGR image as a proper genomic image for extracting statistical features and achieving accurate classification. Furthermore, the results showed that KNN achieved an overall accuracy of 99.39% in detecting COVID-19, among other human CoV diseases, with 99.48% precision, 99.31% sensitivity, 99.47% specificity, 0.99 F1-score, and 0.99 Matthew's correlation coefficient.
The coronavirus disease 2019 (COVID-19) pandemic has been spreading quickly, threatening the public health system. Consequently, positive COVID-19 cases must be rapidly detected and treated. Automatic detection systems are essential for controlling the COVID-19 pandemic. Molecular techniques and medical imaging scans are among the most effective approaches for detecting COVID-19. Although these approaches are crucial for controlling the COVID-19 pandemic, they have certain limitations. This study proposes an effective hybrid approach based on genomic image processing (GIP) techniques to rapidly detect COVID-19 while avoiding the limitations of traditional detection techniques, using whole and partial genome sequences of human coronavirus (HCoV) diseases. In this work, the GIP techniques convert the genome sequences of HCoVs into genomic grayscale images using a genomic image mapping technique known as the frequency chaos game representation. Then, the pre-trained convolution neural network, AlexNet, is used to extract deep features from these images using the last convolution (conv5) and second fully-connected (fc7) layers. The most significant features were obtained by removing the redundant ones using the ReliefF and least absolute shrinkage and selection operator (LASSO) algorithms. These features are then passed to two classifiers: decision trees and k-nearest neighbors (KNN). Results showed that extracting deep features from the fc7 layer, selecting the most significant features using the LASSO algorithm, and executing the classification process using the KNN classifier is the best hybrid approach. The proposed hybrid deep learning approach detected COVID-19, among other HCoV diseases, with 99.71% accuracy, 99.78% specificity, and 99.62% sensitivity.
The Hospital Information System is a huge, integrated system that meets these needs. The primary objective of any hospital information system is to produce high-quality and timely reporting for evidence-based decisions and interventions. The electronic health record is essential in creating uniform data formats and terminology for both corporate and public sector reporting needs. In this study, we design a hospital information system that focuses on the tools, and techniques necessary to improve the gathering, archiving, retrieval, and application of information in biomedicine and the healthcare industry. With the intention of bringing about progress, the idea of a paperless hospital is presented as the embodiment of future health information systems. The outcomes of this study succeeded in proposing an integrative system for hospital information, that has the ability to enhance the capability of medical professionals to coordinate patient care.
DNA microarray data sets have been widely explored and used to analyze data without any previous biological background. However, analyzing them becomes challenging if data are missing. Thus, machine learning techniques are applied because microarray technology is promising in genomics, especially in the analysis of gene expression data. Furthermore, gene expression data can describe the transcription and translation processes of each genetic information in detail. In this study, a new system was proposed to impute more realizable values for missing data in a microarray dataset. This system was validated and evaluated on 42 samples of rectal cancer. Several evaluation tests were also conducted to confirm the effectiveness of the new system and compare it with highly known imputing algorithms. The proposed clustering column-mean quantile median technique could predict highly informative missing genes, thereby reducing the difference between the original and imputed datasets and demonstrating its efficiency.
Background Rheumatoid arthritis (RA) is an autoimmune disease in which the immune system attacks the tissues of the joints by mistake. Different factors—either genetic or environmental—affect the development of the RA disease in patients. A lot of studies aimed to examine the genetic associations with this disease in different populations. This research aspires to perform a genetic association study between six single-nucleotide polymorphisms (SNPs) and RA disease in the Egyptian population with 49 controls and 52 patients. The SNPs that are included in this study are MIR146A rs2910164 (C:G), MIR499/MIR499A rs3746444 (T:C), MTMR3 rs12537(C:T), MIR155HG rs767649 (A:T), IRAK1 rs3027898 (A:C) and PADI4 rs1748033 (C:T). Methods Real-time PCR with TaqMan allelic discrimination assay were both used to perform the genotyping. The Odds ratio models with 95% confidence interval were used to test the associations. The used models are multiplicative, recessive, dominant and co-dominant. Result The demonstrated results indicated that rs2910164 and rs12537 were associated with RA, while rs3746444 showed no association in all the tested models. The remaining SNPs were excluded as they didn't pass the Hardy–Weinberg equilibrium test. Conclusion The MIR146A and MTMR3 polymorphisms showed susceptibility to RA. Moreover, MIR499/MIR499A had no role in the disease.
Lung cancer is one of the most serious cancers in the world with the minimum survival rate after the diagnosis as it appears in Computed Tomography scans. Lung nodules may be isolated from (solitary) or attached to (juxtapleural) other structures such as blood vessels or the pleura. Diagnosis of lung nodules according to their location increases the survival rate as it achieves diagnostic and therapeutic quality assurance. In this paper, a Computer Aided Diagnosis (CADx) system is proposed to classify solitary nodules and juxtapleural nodules inside the lungs. Two main auto-diagnostic schemes of supervised learning for lung nodules classification are achieved. In the first scheme, (bounding box + Maximum intensity projection) and (Thresholding + K-means clustering) segmentation approaches are proposed then first- and second-order features are extracted. Fisher score ranking is also used in the first scheme as a feature selection method. The higher five, ten, and fifteen ranks of the feature set are selected. In the first scheme, Support Vector Machine (SVM) classifier is used. In the second scheme, the same segmentation approaches are used with Deep Convolutional neural networks (DCNN) which is a successful tool for deep learning classification. Because of the limited data sample and imbalanced data, tenfold cross-validation and random oversampling are used for the two schemes. For diagnosis of the solitary nodule, the first scheme with SVM achieved the highest accuracy and sensitivity 91.4% and 89.3%, respectively, with radial basis function and applying the (Thresholding + Kmeans clustering) segmentation approach and the higher 15 ranks of the feature set. In the second scheme, DCNN achieved the highest accuracy and sensitivity 96% and 95%, respectively, to detect the solitary nodule when applying the bounding box and maximum intensity projection segmentation approach. Receiver operating characteristic curve is used to evaluate the classifier's performance. The max. AUC = 90.3% is achieved with DCNN classifier for detecting solitary nodules. This CAD system acts as a second opinion for the radiologist to help in the early diagnosis of lung cancer. The accuracy, sensitivity, and specificity of scheme I (SVM) and scheme II (DCNN) showed promising results in comparison to other published studies.