
Genome-scale models (GEMs) are structured representations of a target organism’s metabolism based on existing genetic, biochemical, and physiological information. These models store the available knowledge of the physiology and metabolic behaviour of organisms and summarise this knowledge in a mathematical description. Flux balance analysis uses GEMs to make predictions about cellular metabolism through the solution of a constrained optimisation problem. The gene inactivity moderated by metabolism and expression (GIMME) approach further constrains FBA by means of transcriptomics data. The underlying idea is to deactivate those reactions for which transcriptomics is below a given threshold. GIMME uses a unique threshold for the entire cell. Therefore, non-essential reactions can be deactivated, even if they are required to meet the production of a certain external metabolite, because of their low associated transcript expression values. Here, we propose a new approach to enable the selection of different transcriptomics thresholds for different cell compartments or modules, such as cellular organelles and specific metabolic pathways. The approach was compared with the original GIMME in the analysis of a number of examples related to yeast batch fermentation for the production of ethanol from glucose or xylose. In some cases, the original GIMME results in biological unfeasibility, while the compartmentalised version successfully recovered flux distributions. The method is implemented in the python-based toolbox MEWpy and can be applied to other metabolic studies, opening the opportunity to obtain more refined and realistic flux distributions, which explain the connections between genotypes, environment and phenotypes.
Due to global warming, coral reefs have been directly impacted with heat stress, resulting in mass coral bleaching. Within the coral species, some are more heat resistant, which calls for an investigation towards interventions that can enhance coral resilience for other heat-susceptible species. Studying heat-resistant corals’ microbial communities can provide a potential insight to the composition of heat-susceptible corals and how their resilience is achieved. So far, techniques to efficiently classify such vast microbiome data are not sufficient. In this paper, we present an optimal machine learning based pipeline for identifying the biomarker bacterial composition of heat-tolerant coral species versus heat-susceptible ones. Through steps of feature extraction, feature selection/engineering, and machine leaning training, we apply this pipeline on publicly available 16S rRNA sequences of corals. As a result, we have identified the correlation based feature selection filter and the Random Forest classifier to be the optimal pipeline, and determined biomarkers that are indicators of thermally sensitive corals.
COVID-19 can mutate rapidly, resulting in new variants which could be more malignant. To recognize the new variant, we must identify the mutation parts by locating the nucleotide changes in the DNA sequence of COVID-19. The identification is by processing sequence alignment. In this work, we propose a method to perform multiple sequence alignment via deep reinforcement learning effectively. The proposed method integrates a progressive alignment approach by aligning each pairwise sequence center to deep Q networks. We designed the experiment by evaluating the proposed method on five COVID-19 variants: alpha, beta, delta, gamma, and omicron. The experiment results showed that the proposed method was successfully applied to align multiple COVID-19 DNA sequences by demonstrating that pairwise alignment processes can precisely locate the sequence mutation up to 90% . Moreover, we effectively identify the mutation in multiple sequence alignments fashion by discovering around 10.8% conserved region of nitrogenous bases.
Neoantigen detection is the most critical step in developing personalized vaccines in cancer immunology. However, neoantigen detection depends on the correct pMHC binding and presentation prediction. Furthermore, transformers and transfer learning have a high impact on NLP tasks. Since amino acids and proteins are like words and sentences, the pMHC binding and presentation prediction problem could be considered an NLP task. Thus, this work proposed using a BERT architecture pre-trained in 250 million proteins (ESM-1b), then we will use a BiLSTM in cascade. Our preliminary results evaluated a small BERT (TAPE) model achieving 0.80 of AUC on the netMHCpanII3.2 dataset.
Cancer immunology is a new alternative to traditional cancer treatments like radiotherapy and chemotherapy. There are some strategies, but neoantigen detection for developing cancer vaccines are methods with a high impact in recent years. However, neoantigen detection depends on the correct prediction of peptide-MHC binding. Furthermore, transformers are considered a revolution in artificial intelligence with a high impact on NLP tasks. Since amino acids and proteins could be considered like words and sentences, the peptide-MHC binding prediction problem could be seen as a NLP task. Therefore, in this work, we performed a systematic literature review of deep learning and transformer methods used in peptide-MHC binding and presentation prediction. We analyzed how ANNs, CNNs, RNNs, and Transformer are used.
The determination of protein structure has been facilitated using deep learning models, which can predict protein folding from protein sequences. In some cases, the predicted structure can be compared to the already-known distribution if there is information from classic methods such as nuclear magnetic resonance (NMR) spectroscopy, X-ray crystallography, or electron microscopy (EM). However, challenges arise when the proteins are not abundant, their structure is heterogeneous, and protein sample preparation is difficult. To determine the level of confidence that supports the prediction, different metrics are provided. These values are important in two ways: they offer information about the strength of the result and can supply an overall picture of the structure when different models are combined. This work provides an overview of the different deep-learning methods used to predict protein folding and the metrics that support their outputs. The confidence of the model is evaluated in detail using two proteins that contain four domains of unknown function.
Inferences on the evolutionary history of a gene can provide insight into whether the findings made for a given gene in a given species can be extrapolated to other species, including humans, help explain morphological evolution or give an explanation for unexpected findings regarding gene expression suppression experiments, among others. The large amount of sequence data that is already available, and that is predicted to dramatically increase in the next few years, means that life science researchers need efficient automated ways of analyzing such data. Moreover, especially when dealing with divergent sequences, inferences can be affected by the chosen alignment and tree building algorithms, and thus the same dataset should be analyzed in different ways, reinforcing the need for the availability of efficient automated ways of analyzing the sequencing data. Therefore, here, we present auto-phylo, a simple pipeline maker for phylogenetic studies, and provide two examples of its utility: one involving a small already formatted sequenced dataset (41 CDS) to determine the impact of the use of different alignment and tree building algorithms in an automated way, and another one involving the automated identification and processing of the sequences of interest, starting from 16550 bacterial CDS FASTA files downloaded from the NCBI Assembly RefSeq database, and subsequent alignment and tree building inferences.
Features selection of high-dimensional data is desirable, mainly when extensive data is used and generated more often. Currently considered research problems are related to the appropriate feature selection in a multidimensional space allowing the selection of only those relevant to the analyzed problem. The implemented and applied machine learning approach made it possible to recognize feature profiles to distinguish two classes of observations. Regarding methods based on logistic regression, 21 features were selected, and 10 features related to the examined problem were identified for neural networks. This made it possible to significantly reduce the dimensionality of the data from as many as 406 original dimensions. Moreover, the feature selection approaches allowed for consistent results; as many as eight features were common to both utilized methods. The application of the recognized profiles also made it possible to obtain very high classification quality metrics, which in the case of logistic regression both for feature selection and classification, amounted to almost 94
Schizophrenia is a complex disease with severely disabling symptoms. A consistent leading causal gene for the disease onset has not been found. There is also a lack of consensus on the disease etiology and diagnosis. Sweden poses a paradigmatic case, where relatively high misdiagnosis rates (19
As the connection between the gut and brain is further researched, more data has become available, allowing for the utilization of machine learning (ML) in such analysis. In this paper, we explore the relationship between Alzheimer’s disease (AD) and the gut microbiome and how it can be utilized for AD screening. Our main goal is to produce a reliable, noninvasive screening tool for AD. Several ML algorithms are examined separately with and without feature selection/engineering. According to the experimental results, the Naive Bayes (NB) model performs best when trained on a feature set selected by the correlation-based feature selection method, which significantly outperforms the baseline model trained on the original full feature space.
Transmembrane transport proteins (transporters) serve a crucial role for the transport of hydrophilic molecules across hydrophobic membranes in every living cell. The structures and functions of many membrane proteins are unknown due to the enormous effort required to characterize them. This article proposes TooT-BERT-T, a technique that employs the BERT representation to analyze and discriminate between transporters and non-transporters using a Logistic Regression classifier. Additionally, we evaluate frozen and fine-tuned representations from two different BERT models. Compared to state-of-the-art prediction methods, TooT-BERT-T achieves the highest accuracy of 93.89% and MCC of 0.86.
In this paper, we first present a new dataset of NDM-1 biological activities that is compiled by a cleaned version of the NMDI database. A literature review enriched the former database by 741 new compounds, comprising activities against NDM-1 classified in three classes (inactive, weakly and strongly active compounds) by specifying a unifying procedure for the labeling, which covers a range of different activity properties. Second, we restate the classification problem in the Multiple Instance Learning (MIL) setting by representing the compounds as a collection of Mol2vec vectors, each of them corresponding to a specific substructure (either atom or atom including their firsts neighbors). We observe an amelioration up to 45.7% and 38.47% in respect to balanced accuracy and F1-score, respectively, for the strongly active class in the MIL approach when compared to the classical Machine Learning paradigm. Finally, we present a classification and ranking framework based on classifiers learned by a k-fold CV procedure, which possess different hyper-parameters per fold, learnt by a Bayes optimization procedure. We observe that the top-3 and top-5 ranked accuracies of the strongly active classified compounds yield 100% for the MIL setting.
The COVID-19 pandemic remains a concrete challenge, especially in communities and rural areas where health resources are scarce. We recently developed several classifiers, useful to predict safe discharge, disease severity, and mortality risk from COVID-19, fed by routine analyses collected in the Emergency Department. In this paper, we discuss a system, made up of an app and a server, that enables doctors to use these models during the management of COVID-19 patients. The app has been developed involving the doctors since the early phases of the app design, then revised in the light of two usability cycles. We report its main features and its ease of use. So far, it has been used during the fourth wave, producing accurate results with patients that did not complete the vaccination protocol (i.e., up to the second dose).
Epilepsy is a neurological disorder (the third most common, following stroke and migraines). A key aspect of its diagnosis is the presence of seizures that occur without a known cause and the potential for new seizures to occur. Machine learning has shown potential as a cost-effective alternative for rapid diagnosis. In this study, we review the current state of machine learning in the detection and prediction of epileptic seizures. The objective of this study is to portray the existing machine learning methods for seizure prediction. Internet bibliographical searches were conducted to identify relevant literature on the topic. Through cross-referencing from key articles, additional references were obtained to provide a comprehensive overview of the techniques. As the aim of this paper aims is not a pure bibliographical review of the subject, the publications here cited have been selected among many others based on their number of citations. To implement accurate diagnostic and treatment tools, it is necessary to achieve a balance between prediction time, sensitivity, and specificity. This balance can be achieved using deep learning algorithms. The best performance and results are often achieved by combining multiple techniques and features, but this approach can also increase computational requirements.
In the last decade, miRNAs have attracted noticeable interest as potential biomarkers of neuropsychiatric conditions. However, a standard methodology for miRNA-Seq analysis does not yet exist, raising concerns about the reproducibility of the in-silico results and limiting their usefulness. This situation motivated us to design a miRNA-Seq pipeline specialized in the analysis of neuropsychiatric data, aiming to integrate the results of several bioinformatics tools in a highly reproducible workflow. In this study, we performed an initial test of the usefulness of our new pipeline, named myBrain-Seq, by reanalyzing four recent miRNA-Seq studies of neuropsychiatric conditions. We then compared the myBrain-Seq results with the original results and with an additional reanalysis done with another pipeline in order to make an estimation of the overall replicability. We found one of the three myBrain-Seq methodologies to be the one with best replicability, although the heterogeneity of the results and the absence of an experimental validation limits our conclusions. Further work is required to assess myBrain-Seq’ performance using a bigger dataset of studies with experimental validation data available.
Xylella fastidiosa is a gram-negative phytopathogenic bacterium able to infect over 500 plant species, with devastating consequences for agricultural and forest-based economies. In the last decade, genome-scale metabolic (GSM) models have become important systems biology tools for studying the metabolic behaviour of different organisms. In this work, a GSM model of X. fastidiosa subsp. pauca De Donno is presented, comprising 1164 reactions, 1379 metabolites, and 508 genes. The model was validated by comparing in silico simulations with available experimental data. The GSM model allowed identifying potential drug targets using a pipeline based on a gene essentiality analysis of the model.
The understanding of the molecular basis of cellular processes and ultimately disease, requires knowledge on protein structures, interactions, and functions. Protein–protein interaction data (PPI) is available in the publicly available main PPI databases that show little overlap due to the use of different criteria. Therefore, web platforms that aggregate the data from multiple sources, such as EvoPPI ( http://evoppi.i3s.up.pt ), where the existing databases have been updated and new ones were added, as here described, and APID ( http://cicblade.dep.usal.es:8080/APID/init.action ) are useful. Still, in both EvoPPI 1.0 and 2, here presented, we have made a special effort to make it flexible in what concerns the choice of the databases to be compared. Moreover, interacting protein pairs tend to be evolutionarily conserved, and thus the information available for one species might be used to predict the incompleteness of the network in another one, and identify putative missing interactions. This approach is now available in EvoPPI 2 for Homo sapiens and the model species Mus musculus, Caenorhabditis elegans, and Drosophila melanogaster, using either Ensembl ( https://www.ensembl.org ) or DIOPT Ortholog Finder ( https://www.flyrnai.org/cgi-bin/DRSC_orthologs.pl ) orthologies/paralogies. Moreover, since not all available PPI data is present in the main databases (e.g. PPI observed in patient tissues and mutant animal species, where PPI might be aberrant, are usually not included in the main databases, although in several studies this has been shown not to be the case), we provide the needed tools (including a Lubuntu-based virtual machine where all software is already installed and ready-to-run) to run a local EvoPPI 2 instance and create a custom database from the existing ones. This way the user can add new data for any species and from any source database, creating custom interactomes. Administrator tools are provided to help in the automatic processing and conversion of files from various sources into the custom EvoPPI database format.
Nicotinamide adenine dinucleotide (NAD) is an essential metabolite in normal cellular physiology and its deregulation may lead to several pathological conditions. NAD interacts with a vast number of proteins, acting as a coenzyme, as a substrate and regulating the interaction between proteins. The goals of this study were to characterize the proteins involved in NAD metabolism and to identify putative new NAD regulated proteins. Using an in silico approach, we first defined a NAD-binding dataset, that we characterized through pathway enrichment analysis and protein structural domains analysis. We then screened the full human proteome and further analyzed a selection of potential NAD-binding proteins. This global study of the NAD interactome resulted in the identification of new potentially NAD-binding proteins (NADPBs), including TRPC3 and a few isoforms of DGA kinases, which are involved in calcium signaling. NADBPs participate in several metabolic pathways and signaling processes in the cell, while proteins interacting with NADPBs are mostly involved in signaling pathways, including pathways related to disease, namely three major neurodegenerative diseases, Alzheimer’s, Huntington’s, and Parkinson’s.
Hairpin/cruciform structures, as well as other non-B DNA structures, are important regulators for biological processes and gene function. The formation of these structures require that the DNA sequence contains adequately spaced inverted repeats. To study the potential of DNA regions to form hairpin/cruciform structures, we developed a new procedure to analyse the variation of the concentration of occurrence of inverted repeats at different spacings along the human genome. We apply the method to the human genome and identify regions with atypical high concentration of inverted repeats when compared to a control scenario based on a Markov model of order 7. We found that the potential to form hairpin/cruciform structures is very heterogeneous across different human genome regions. Also, different regions display strikingly different patterns of enrichment of concentration depending on inverted repeats spacing.
In biological image analysis, 3D instance segmentation is a crucial step towards extracting information on objects of interest from microscopy datasets. Existing instance segmentation pipelines are frequently affected by errors such as missing boundary layer cells or poorly segmented regions. In this study, we propose several ensembles as post-processing methods for improving the quality of outputs obtained from deep learning and classical 3D segmentation pipelines. These methods take as input the results from two independent 3D segmentation pipelines and combine them using different fusion algorithms. The first algorithm uses label set intersection, the second one involves adjacency graph composition and the third one works through segmented object boundary fusion followed by 3D watershed. These 3 algorithms are tested on a dataset of 3D confocal microscopy images of floral tissues. The third fusion algorithm is found to perform best and has better global and local accuracies compared to its input segmentations. The specialty of the proposed ensemble methods is that these are model agnostic, i.e., they can be used to combine segmentation results from deep learning as well as non-deep learning or classical pipelines. These methods could be highly beneficial in correcting segmentation errors arising from missing cells in the boundary layer or under segmentation in the inner tissue layers and ultimately provide us robust segmentation results in presence of variable image qualities in biological datasets.