
This paper analyzes the application of surrogate models to improve the efficiency of Gene Regulatory Network (GRN) inference from time-series data. A Radial Basis Function (RBF) surrogate model was integrated with the Penalized mAximum LikeLihood and pArticle Swarms (PALLAS) using a Mixed Fish School Search (MFSS) algorithm to reduce the computational cost associated with evaluating the penalized log-likelihood (PLL) fitness function. Experimental results on the p53-MDM2 negative-feedback loop GRN dataset demonstrate that the surrogate-assisted approach significantly reduced fitness function calls by 50% and 89% while maintaining the quality of the PLL metric, with this showing the potential of surrogate models to accelerate GRN inference.
This study aims to analyze the genomic characteristics of a plant growth-promoting bacterium, focusing on genome assembly and annotation, as well as phylogenetic and metabolic analyses. The primary objective is to identify and validate genes related to potassium and phosphate solubilization, as well as other primary and secondary metabolic pathways. The research involves the selection of a strain, cultivation, DNA extraction, sequencing, genome assembly and annotation, followed by phylogenetic and comparative analyses. The analyzed Bacillus nitratireducensLABIM48 strain shows high genomic conservation, indicating a common origin and conserved evolution. Phylogenetic analyses confirmed that the strain belongs to the same mentioned species, with the potential to produce bioactive compounds of interest identified in silico. The ability to solubilize phosphate and potassium was validated through in vitro tests.
This paper investigates convolutional neural networks (CNN) for predicting cancer types by integrating protein-protein interaction (PPI) networks with omics data. While [Chuang et al. 2021] employed a single 3-layer CNN, we explore ten different architectures, including a custom model developed by our team (CNN2Layers), following their methodology. By evaluating the strengths and weaknesses of these models, we aim to identify the most effective CNN for accurately predicting various human cancers. Our proposed model achieved state-of-the-art performance using fewer layers. Interestingly, the simpler architectures achieved superior results, indicating their effectiveness in handling the specific characteristics of the dataset.
This study develops and evaluates pan-cancer (PC) models for cohort-specific (CS) predictions using neural networks (NNs). We adopt a dual approach, including a method inspired by few-shot learning, aiming at improving the models’ ability to distinguish between normal and tumorous tissues across diverse cohorts. The first approach trains a NN with comprehensive PC datasets containing 16 cancer types, comparing it against CS models on a target cohort, while the second analyzes whether PC models could generalize to smaller and unseen cohorts by training on 15 cohorts and evaluating on the excluded cohort. Our experiments show that PC models generally outperform CS models, even with limited sample sizes and class imbalances. Moreover, the few-shot approach successfully generalizes to other cancer types, highlighting its potential to advance personalized cancer diagnosis and treatment.
COVID-19, caused by the SARS-CoV-2 virus, has led to a global pandemic since 2020, resulting in nearly 7 million deaths. The virus’s rapid spread is due to more transmissible variants, many with spike glycoprotein mutations, which are key for cell invasion and a vaccine target. Understanding these mutations is crucial for preventing more dangerous variants. This study developed a computational method to predict the impact of mutations on the spike protein. Using data from 23,472 mutations, molecular modeling, graph-based structural signatures, and a machine-learning approach based on neural networks, the model analyzed 318 proteins, showing the methodology’s effectiveness in assessing the potential of new variants.
With the increasing volume of biological and medical data, the application of efficient data science techniques has become essential for analysis. However, healthcare data scientists often need to integrate and analyze multiple datasets simultaneously. Although these analyses share similarities, they require adjustments to various parameters, delaying development and further hindering knowledge discovery. In this paper, we propose a framework that encapsulates all stages of typical data science analyses, from data pre-processing, execution, and evaluation to the interpretation of models. In addition, the framework includes XAI analyses. In tests involving a clinical dataset, the framework achieved a reduction of 92% in lines of code.
Contacts, defined as interand intramolecular interactions predicted computationally, are typically detected using Euclidean distance and atom types. However, traditional methods can be computationally expensive and limit scalability. We introduce COCαDA (Contact Optimization by Cα Distance Analysis), a novel method that incorporates domain knowledge of amino acids to optimize distance cutoffs, simplifying implementation and enhancing efficiency. COCαDA outperforms traditional methods such as all-against-all, static cutoff (SC), and Biopython’s NeighborSearch (NS), averaging 2.5x faster than SC and 6x faster than NS. COCαDA is well-suited for exploratory and large-scale analyses and is freely available at https://github.com/LBS-UFMG/COCaDA.
Computational semantics for molecular biology was introduced to assign activities to biomolecules based on their interactions with environments and other biomolecules. We distinguish between activity and biological function towards an epistemologically neutral semantics that aligns with computational processes. Object Petri Nets (OPNs) represent complex biomolecular activities as compositions of interactions at the nucleotide level. This article introduces a way to transform the networks that shows the equivalence between OPNs and simple Place/Transition (P/T) nets while preserving their semantics. It gives an intuitive understanding of this equivalence and shows OPNs implemented on P/T PNs.
Bioinformatics requires professionals with knowledge in computing and biological sciences, but teaching it to young people remains a challenge. This article reports on a programming course focused on bioinformatics for high school students. The pilot project, launched in 2024 in Belo Horizonte, Brazil, aimed to integrate programming and Molecular Biology. Using Inquiry-Based Learning (IBL) and gamification, the course engaged students effectively. Activities were divided into quarterly stages, teaching programming through Scratch and projects involving Molecular Biology, such as DNA transcription. The initiative successfully motivated and engaged students in learning Molecular Biology. We hope that the strategies presented here can be adopted by teachers and help inspire a new generation of bioinformaticians.
Mutations in the PRKAG2 gene, which encodes the γ2 subunit of AMP-activated protein kinase (AMPK), are linked to a rare cardiomyopathy involving glycogen accumulation, left ventricular hypertrophy, and sudden death. This study investigates a novel His401Gln missense mutation in the PRKAG2 gene and its effects on AMPK γ2 subunit dynamics. Through molecular simulations and free energy analyses, we compared AMP and ATP binding affinities between the wild-type and mutant γ2 subunits. Structural modeling and simulations revealed a significant change in ATP binding at site 3 in the mutant AMPK, suggesting that the His401Gln mutation impacts protein binding behavior. This alteration may contribute to the pathological mechanisms of PRKAG2 cardiomyopathy.
We present a pipeline for exploring genomic diversity in metagenomic datasets at the species and strain levels. To achieve accurate classifications independent of taxonomy labels, we introduce the concept of Genome Reference Set (GRS), modeled using the Maximal Independent Set problem for undirected graphs. For a given user-defined target genus, we build its GRS from GenBank genomes and use it for metagenomic contig classification using BLASTn. Additional phylogenetic processing allows the identification of putative novel species. We show that our pipeline can achieve better results than general-purpose tools, and apply the pipeline to the MetaSUB dataset, identifying two putative novel strains and one putative new species of Acinetobacter.
Gene fusions are abnormal genetic events often correlated with oncogenesis. Hence, detecting them from RNA-seq data using bioinformatics methods is an important task in cancer research. Several tools have been developed for this task, but current benchmarks are inconclusive regarding their accuracy and are difficult to reproduce with new data. In this paper, we propose a computational pipeline that gathers fusion detection tools and compares them using standard classification metrics. It can also be used as an ensemble method to detect gene fusions using several tools. This pipeline was applied to simulated and real data, and supplements current benchmarks in the literature towards aiding the users in choosing the tools for their analyses.
In this work, we explore heuristics for the Adjacency Graph Packing problem, which can be applied to the Double Cut and Join (DCJ) Distance Problem. The DCJ is a rearrangement operation and the distance problem considering it is a well established method for genome comparison. Our heuristics will use the structure called adjacency graph adapted to include information about intergenic regions, multiple copies of genes in the genomes, and multiple circular or linear chromosomes. The only required property from the genomes is that it must be possible to turn one into the other with DCJ operations. We propose one greedy heuristic and one heuristic based on Genetic Algorithms. Our experimental tests in artificial genomes show that the use of heuristics is capable of finding good results that are superior to a simpler random strategy.
Various approaches utilizing Transformer architectures have achieved state-of-the-art results in Natural Language Processing (NLP). Based on this success, numerous architectures have been proposed for other types of data, such as in biology, particularly for protein sequences. Notably among these are the ESM2 architectures, pre-trained on billions of proteins, which form the basis of various state-of-the-art approaches in the field. However, the ESM2 architectures have a limitation regarding input size, restricting it to 1,022 amino acids, which necessitates the use of preprocessing techniques to handle sequences longer than this limit. In this paper, we present the long and quantized versions of the ESM2 architectures, doubling the input size limit to 2,048 amino acids.
This study aims to develop and evaluate optimized neural networks, including Multilayer Perceptrons (MLP) and Convolutional Neural Networks (CNN), by employing deep learning techniques to classify breast cancer subtypes, based on gene expression data. By implementing different neural network architectures and optimization strategies, this research seeks to determine the accuracy and efficiency of these classification methods. Data is sourced from The Cancer Genome Atlas (TCGA) repository and undergoes preprocessing, including dimensionality reduction, to prepare it for analysis. The contribution is to enhance diagnostic tools, as well as assess the predictive performance of the approaches. The comparison of networks performance presents a promising pathway to enhancing the precision of medical diagnostics and personalize treatment strategies in breast cancer.
Histone deacetylases (HDACs) are enzymes that play an essential role in regulating gene expression, with recent studies linking their inhibition to autism spectrum disorders (ASD). As a result, there is growing interest in understanding the effects of HDAC inhibition. In this paper, we used molecular docking to investigate the binding between HDACs and small ligands, focusing on two enzymes involved in embryonic development: Histone deacetylase 1 (H1) and Histone deacetylase 2 (H2). Using a graph-based structural signature algorithm, we extracted features from the resulting complexes and employed machine learning algorithms to distinguish natural ligands from decoys, achieving 72% of accuracy in the classification test.
This study aims to investigate the biological diversity of 21 genomes of Cylindrospermopsis and 6 genomes of Sphaerospermopsis belonging to the family Aphanizomenonaceae using bioinformatics methods in comparative genomics. The comparative analysis of the lineages of the groups From the Americas, Non-Americas and Sphaerospermospsis revealed conserved central genome but with different accessory genomes by geographic region. The variations observed in the organization of saxitoxin and cylindrospermopsin genes in the accessory genome of Cylindrospermopsis raciborskii from the Americas and NonAmericas, suggests in the future the formation of a new taxonomic group, due to independent evolutionary trajectories.
RNA secondary structures, determined by non-crossing base pairs, capture many of the salient features of RNA molecules, explain the free energy of structure formation very accurately, and can be computed efficiently given the sequence information only. G-quadruplexes are compact local structures that have been shown to have important biological function and can be integrated into secondary structure prediction. In recent years circular RNAs have gained considerable interest as a biological relevant subclass of RNAs. While algorithms and tools are available that extend secondary structure prediction from linear to circular RNAs, no support is provided for G-quadruplexes. In this contribution we close this gap and describe how the ViennaRNA package has been extended to include this increasingly relevant case.
Predicting the binding mode and affinity of small molecules to proteins is key to understanding their interaction. Empirical scoring functions are commonly used by docking programs, but accurately predicting them remains challenging. Docking programs can generate ligand conformations similar to crystallographic structures, yet scoring functions often struggle to identify the correct pose. This study employs Graph Attention Networks (GAT) to learn ligand-protein contact information and re-rank docking poses. Using PDBbindcore data, docking calculations with AutoDock Vina generate binding poses, evaluated by contacts and RMSD. Close contacts are mapped using BINANA, and bipartite graphs are created with atomic descriptors using RDKit.