Metal-binding sites (MBSs) are critical determinants of protein stability and biological function, yet methods for comparing their local binding environments lag behind those for whole-structure alignment. Here, we represent MBSs as atomic point clouds surrounding bound metal ligands and align them with a fine-tuned iterative closest point algorithm. Applying this framework to a redundancy-reduced collection of MBSs derived from all metalloproteins in the Protein Data Bank (PDB), we perform pairwise alignments across 23,342 sites to construct a similarity network of metal-binding environments. The resulting network topology recapitulates metal coordination chemistry and enzyme function: links are strongly enriched within metal types and across shared EC subclasses. Conserved metalloenzyme families form cohesive subnetworks; for example, the binuclear ureohydrolase domain appears as two tightly connected components that also capture atypical members such as the dinickel metformin hydrolase. We observe only a moderate global association between protein sequence and MBS geometry, yet many network links connect near-identical binding-site architectures across proteins with low sequence identity, consistent with either divergent evolution with local MBS conservation or candidate cases of molecular convergent evolution. Integrating network proximity with structural evidence of drug binding identifies drugs with enriched connectivity among their targets and predicts 528 drug-off-target combinations across 88 drugs and 151 human proteins, recovering both known off-targets (e.g., ADAM/ADAMTS for matrix metalloproteinase inhibitors) and proposing novel ones. The MBS network thus provides a scalable resource for probing metalloprotein evolution, functional convergence, and the structural basis of drug cross-reactivity.
Genetic research into atrial fibrillation (AF) and myocardial infarction (MI) has predominantly focused on comparing afflicted individuals with their healthy counterparts. However, this approach lacks granularity, thus overlooking subtleties within patient populations. In this study, we explore the distinction between AF and MI patients who experience only a single disease event and those experiencing recurrent events. Integrating hospital records, questionnaire data, clinical measurements, and genetic data from more than 500,000 HUNT and United Kingdom Biobank participants, we compare both clinical and genetic characteristics between the two groups using genome-wide association studies (GWAS) meta-analyses, phenome-wide association studies (PheWAS) analyses, and gene co-expression networks. We found that the two groups of patients differ in both clinical characteristics and genetic risks. More specifically, recurrent AF patients are significantly younger and have better baseline health, in terms of reduced cholesterol and blood pressure, than single AF patients. Also, the results of the GWAS meta-analysis indicate that recurrent AF patients seem to be at greater genetic risk for recurrent events. The PheWAS and gene co-expression network analyses highlight differences in the functions associated with the sets of single nucleotide polymorphisms (SNPs) and genes for the two groups. However, for MI patients, we found that those experiencing single events are significantly younger and have better baseline health than those with recurrent MI, yet they exhibit higher genetic risk. The GWAS meta-analysis mostly identifies genetic regions uniquely associated with single MI, and the PheWAS analysis and gene co-expression networks support the genetic differences between the single MI and recurrent MI groups. In conclusion, this work has identified novel genetic regions uniquely associated with single MI and related PheWAS analyses, as well as gene co-expression networks that support the genetic differences between the patient subgroups of single and recurrent occurrence for both MI and AF.
Adaptive Laboratory Evolution (ALE) of microorganisms can improve the efficiency of sustainable industrial processes important to the global economy. However, stochasticity and genetic background effects often lead to suboptimal outcomes during laboratory evolution. Here we report an ALE platform to circumvent these shortcomings through parallelized clonal evolution at an unprecedented scale. Using this platform, we evolved 10(4) yeast populations in parallel from many strains for eight desired wine fermentation-related traits. Expansions of both ALE replicates and lineage numbers broadened the evolutionary search spectrum leading to improved wine yeasts unencumbered by unwanted side effects. At the genomic level, evolutionary gains in metabolic characteristics often coincided with distinct chromosome amplifications and the emergence of side-effect syndromes that were characteristic of each selection niche. Several high-performing ALE strains exhibited desired wine fermentation kinetics when tested in larger liquid cultures, supporting their suitability for application. More broadly, our high-throughput ALE platform opens opportunities for rapid optimization of microbes which otherwise could take many years to accomplish.
Disease networks offer a potential road map of connections between diseases. Several studies have created disease networks where diseases are connected either based on shared genes or Single Nucleotide Polymorphism (SNP) associations. However, it is still unclear to which degree SNP-based networks map to empirical, co-observed diseases within a different, general, adult study population spanning over a long time period. We created a SNP-based phenome-wide association network (PheNet) from a large population using the UK biobank phenome-wide association studies. Importantly, the SNP-associations are unbiased towards much studied diseases, adjusted for linkage disequilibrium, case/control imbalances, as well as relatedness. We map the PheNet to significantly co-occurring diseases in the Norwegian HUNT study population, and further, identify consecutively occurring diseases with significant ordering in occurrence, independent of age and gender in the PheNet. Our analysis reveals an overlap far larger than expected by chance between the two disease networks, with diseases typically connecting within their own category. Upon examining the sequential occurrence of diseases in the HUNT dataset, we find a giant component consisting of mostly cardiovascular disorders. This allows us to identify sequentially occurring diseases that are genetically linked and co-occur frequently, while also highlighting non-sequential diseases. Furthermore, we observe that survivors of severe cardiovascular diseases subsequently often face less severe conditions, but with a reduced time until their next fatal illness. The HUNT sub-PheNet showing both genetically and co-observed diseases offers an interesting framework to study groups of diseases and examine if they, in fact, are comorbidities. We find that the HUNT sub-PheNet offers the possibility to pinpoint exactly which mutation(s) constitute shared cause of the diseases. This could be of great benefit to both researchers and clinicians studying relationships between diseases.
Flux balance analysis (FBA) remains one of the most used methods for modeling the entirety of cellular metabolism, and a range of applications and extensions based on the FBA framework have been generated. Dynamic flux balance analysis (dFBA), the expansion of FBA into the time domain, still has issues regarding accessibility limiting its widespread adoption and application, such as a lack of a consistently rigid formalism and tools that can be applied without expert knowledge. Recent work has combined dFBA with enzyme-constrained flux balance analysis (decFBA), which has been shown to greatly improve accuracy in the comparison of computational simulations and experimental data, but such approaches generally do not take into account the fact that altering the enzyme composition of a cell is not an instantaneous process. Here, we have developed a decFBA method that explicitly takes enzyme change constraints (ecc) into account, decFBAecc. The resulting software is a simple yet flexible framework for using genome-scale metabolic modeling for simulations in the time domain that has full interoperability with the COBRA Toolbox 3.0. To assess the quality of the computational predictions of decFBAecc, we conducted a diauxic growth fermentation experiment with Escherichia coli BW25113 in glucose minimal M9 medium. The comparison of experimental data with dFBA, decFBA and decFBAecc predictions demonstrates how systematic analyses within a fixed constraint-based framework can aid the study of model parameters. Finally, in explaining experimentally observed phenotypes, our computational analysis demonstrates the importance of non-linear dependence of exchange fluxes on medium metabolite concentrations and the non-instantaneous change in enzyme composition, effects of which have not previously been accounted for in constraint-based analysis.
Use of alternative non-Saccharomyces yeasts in wine and beer brewing has gained more attention the recent years. This is both due to the desire to obtain a wider variety of flavours in the product and to reduce the final alcohol content. Given the metabolic differences between the yeast species, we wanted to account for some of the differences by using in silico models. We created and studied genome-scale metabolic models of five different non-Saccharomyces species using an automated processes. These were: Metschnikowia pulcherrima, Lachancea thermotolerans, Hanseniaspora osmophila, Torulaspora delbrueckii and Kluyveromyces lactis. Using the models, we predicted that M. pulcherrima, when compared to the other species, conducts more respiration and thus produces less fermentation products, a finding which agrees with experimental data. Complex I of the electron transport chain was to be present in M. pulcherrima, but absent in the others. The predicted importance of Complex I was diminished when we incorporated constraints on the amount of enzymatic protein, as this shifts the metabolism towards fermentation. Our results suggest that Complex I in the electron transport chain is a key differentiator between Metschnikowia pulcherrima and the other yeasts considered. Yet, more annotations and experimental data have the potential to improve model quality in order to increase fidelity and confidence in these results. Further experiments should be conducted to confirm the in vivo effect of Complex I in M. pulcherrima and its respiratory metabolism.
The metabolism of all living organisms is dependent on temperature, and therefore, having a good method to predict temperature effects at a system level is of importance. A recently developed Bayesian computational framework for enzyme and temperature constrained genome-scale models (etcGEM) predicts the temperature dependence of an organism’s metabolic network from thermodynamic properties of the metabolic enzymes, markedly expanding the scope and applicability of constraint-based metabolic modelling. Here, we show that the Bayesian calculation method for inferring parameters for an etcGEM is unstable and unable to estimate the posterior distribution. The Bayesian calculation method assumes that the posterior distribution is unimodal, and thus fails due to the multimodality of the problem. To remedy this problem, we developed an evolutionary algorithm which is able to obtain a diversity of solutions in this multimodal parameter space. We quantified the phenotypic consequences on six metabolic network signature reactions of the different parameter solutions resulting from use of the evolutionary algorithm. While two of these reactions showed little phenotypic variation between the solutions, the remainder displayed huge variation in flux-carrying capacity. This result indicates that the model is under-determined given current experimental data and that more data is required to narrow down the model predictions. Finally, we made improvements to the software to reduce the running time of the parameter set evaluations by a factor of 8.5, allowing for obtaining results faster and with less computational resources.
ABSTRACT The field of metabolic modelling at the genomescale continues to grow with more models being created and curated. This comes with an increasing demand for adopting common principles regarding transparency and versioning, in addition to standardisation efforts regarding file formats, annotation and testing. Here, we present a standardised template for git-based and GitHub-hosted genome-scale metabolic models (GEMs) supporting both new models and curated ones, following FAIR principles (findability, accessibility, interoperability, and reusability), and incorporating bestpractices. standard-GEM facilitates the reuse of GEMs across web services and platforms in the metabolic modelling field and enables automatic validation of GEMs. The use of this template for new models, and its adoption for existing ones, paves the way for increasing model quality, openness, and accessibility with minimal effort. Availability standard-GEM is available from github.com/MetabolicAtlas/standard-GEM under the conditions of the CC BY 4.0 licence along with additional supporting material.
Excessive usage of antibiotics threatens the bacterial diversity in the microbiota of animals. An alternative to antibiotics that has been suggested to not disturb the microbiota is (bacterio)phage therapy. In this study, we challenged germ-free and microbially colonized yolk sac fry of Atlantic salmon with Flavobacterium columnare and observed that the mere presence of a microbiota protected the fish against lethal infection. We then investigated the effect of phage- or oxytetracycline treatment on fish survival and rearing water bacterial community characteristics using 16S rRNA gene amplicon sequencing. Phage treatment led to an increased survival of F. columnare-challenged fish and reduced the relative amounts of the pathogen in the water microbiota. In the absence of F. columnare, phage treatment did not affect the composition or the α-diversity of the rearing water microbiota. In the presence of the phage's host, phage treatment induced minor changes to the bacterial community composition, without affecting the α-diversity. Surprisingly, oxytetracycline treatment had no observable effect on the water microbiota and did not reduce the relative abundance of F. columnare in the water. In conclusion, we showed that phage treatment prevents mortality while not negatively affecting the rearing water microbiota, thus suggesting that phage treatment may be a suitable alternative to antibiotics. We also demonstrated a protective effect of the microbiota in Atlantic salmon yolk sac fry.
Genome-scale metabolic models (GEMs) are mathematical representations of metabolism that allow for in silico simulation of metabolic phenotypes and capabilities. A prerequisite for these predictions is an accurate representation of the biomolecular composition of the cell necessary for replication and growth, implemented in GEMs as the so-called biomass objective function (BOF). The BOF contains the metabolic precursors required for synthesis of the cellular macro- and micromolecular constituents (e.g. protein, RNA, DNA), and its composition is highly dependent on the particular organism, strain, and growth condition. Despite its critical role, the BOF is rarely constructed using specific measurements of the modeled organism, drawing the validity of this approach into question. Thus, there is a need to establish robust and reliable protocols for experimental condition-specific biomass determination. Here, we address this challenge by presenting a general pipeline for biomass quantification, evaluating its performance on Escherichia coli K-12 MG1655 sampled during balanced exponential growth under controlled conditions in a batch-fermentor set-up. We significantly improve both the coverage and molecular resolution compared to previously published workflows, quantifying 91.6% of the biomass. Our measurements display great correspondence with previously reported measurements, and we were also able to detect subtle characteristics specific to the particular E. coli strain. Using the modified E. coli GEM i ML1515a, we compare the feasible flux ranges of our experimentally determined BOF with the original BOF, finding that the changes in BOF coefficients considerably affect the attainable fluxes at the genome-scale.
With limited availability of vaccines, an efficient use of the limited supply of vaccines in order to achieve herd immunity will be an important tool to combat the wide-spread prevalence of COVID-19. Here, we compare a selection of strategies for vaccine distribution, including a novel targeted vaccination approach (EHR) that provides a noticeable increase in vaccine impact on disease spread compared to age-prioritized and random selection vaccination schemes. Using high-fidelity individual-based computer simulations with Oslo, Norway as an example, we find that for a community reproductive number in a setting where the base pre-vaccination reproduction number R = 2.1 without population immunity, the EHR method reaches herd immunity at 48% of the population vaccinated with 90% efficiency, whereas the common age-prioritized approach needs 89%, and a population-wide random selection approach requires 61%. We find that age-based strategies have a substantially weaker impact on epidemic spread and struggle to achieve herd immunity under the majority of conditions. Furthermore, the vaccination of minors is essential to achieving herd immunity, even for ideal vaccines providing 100% protection.
Many breast cancer patients are diagnosed with small, well-differentiated, hormone receptor-positive tumors. Risk of relapse is not easily identified in these patients, resulting in overtreatment. To identify metastasis-related gene expression patterns, we compared the transcriptomes of the non-metastatic 67NR and metastatic 66cl4 cell lines from the murine 4T1 mammary tumor model. The transcription factor nuclear factor, erythroid 2-like 2 (NRF2, encoded by NFE2L2) was constitutively activated in the metastatic cells and tumors, and correspondingly a subset of established NRF2-regulated genes was also upregulated. Depletion of NRF2 increased basal levels of reactive oxygen species (ROS) and severely reduced ability to form primary tumors and lung metastases. Consistently, a set of NRF2-controlled genes was elevated in breast cancer biopsies. Sixteen of these were combined into a gene expression signature that significantly improves the PAM50 ROR score, and is an independent, strong predictor of prognosis, even in hormone receptor-positive tumors.
Adaptive evolution under controlled laboratory conditions has been highly effective in selecting organisms with beneficial phenotypes such as stress tolerance. The evolution route is particularly attractive when the organisms are either difficult to engineer or the genetic basis of the phenotype is complex. However, many desired traits, like metabolite secretion, have been inaccessible to adaptive selection due to their trade‐off with cell growth. Here, we utilize genome‐scale metabolic models to design nutrient environments for selecting lineages with enhanced metabolite secretion. To overcome the growth‐secretion trade‐off, we identify environments wherein growth becomes correlated with a secondary trait termed tacking trait. The latter is selected to be coupled with the desired trait in the application environment where the trait manifestation is required. Thus, adaptive evolution in the model‐designed selection environment and subsequent return to the application environment is predicted to enhance the desired trait. We experimentally validate this strategy by evolving Saccharomyces cerevisiae for increased secretion of aroma compounds, and confirm the predicted flux‐rerouting using genomic, transcriptomic, and proteomic analyses. Overall, model‐designed selection environments open new opportunities for predictive evolution. EvolveX, a new algorithm enabling model‐guided design of chemical environments for targeted adaptive evolution, is applied to evolve a wine yeast strain for increased aroma secretion. EvolveX, a new algorithm enabling model‐guided design of chemical environments for targeted adaptive evolution, is applied to evolve a wine yeast strain for increased aroma secretion.
Adaptive evolution under controlled laboratory conditions has been highly effective in selecting organisms with beneficial phenotypes such as stress tolerance. The evolution route is particularly attractive when the organisms are either difficult to engineer or the genetic basis of the phenotype is complex. However, many desired traits, like metabolite secretion, have been inaccessible to adaptive selection due to their trade-off with cell growth. Here, we utilize genome-scale metabolic models to design nutrient environments for selecting lineages with enhanced metabolite secretion. To overcome the growth-secretion trade-off, we identify environments wherein growth becomes correlated with a secondary trait termed tacking trait. The latter is selected to be coupled with the desired trait in the application environment where the trait manifestation is required. Thus, adaptive evolution in the model-designed selection environment and subsequent return to the application environment is predicted to enhance the desired trait. We experimentally validate this strategy by evolving Saccharomyces cerevisiae for increased secretion of aroma compounds, and confirm the predicted flux-rerouting using genomic, transcriptomic, and proteomic analyses. Overall, model-designed selection environments open new opportunities for predictive evolution.
Abstract The metabolism of all living organisms is dependent on temperature, and therefore, having a good method to predict temperature effects at a systems level is of importance. Here, we investigate the properties and robustness of a recently developed computational framework for enzyme and temperature constrained genome-scale models (etcGEM). The approach predicts the temperature dependence of an organism's metabolic network from thermodynamic properties of the metabolic enzymes, markedly expanding the scope and applicability of constraint-based metabolic modelling. We show that the existing Bayesian algorithm for fitting an etcGEM converges under a range of different priors and random seeds, yet the solutions differ both in parameter space and on a phenotypic level. We quantified the phenotypic consequences by studying the impact of different solutions on six metabolic network signature reactions. While two of these reactions showed little phenotype variation between the solutions, the remainder displayed huge variation in flux-carrying capacity. Furthermore, we made software improvements to reduce the running time of the compute-intensive Bayesian algorithm by a factor of 8.5. In addition, we develop an evolutionary algorithm as an alternative to the Bayesian parameter fitting algorithm. Our analysis shows that it also gives decent convergence, but it yields solutions with lower variability.
Genome-scale metabolism can best be described as a highly interconnected network of biochemical reactions and metabolites. The flow of metabolites, i.e., flux, throughout these networks can be predicted and analyzed using approaches such as flux balance analysis (FBA). By knowing the network topology and employing only a few simple assumptions, FBA can efficiently predict metabolic functions at the genome scale as well as microbial phenotypes. The network topology is represented in the form of genome-scale metabolic models (GEMs), which provide a direct mapping between network structure and function via the enzyme-coding genes and corresponding metabolic capacity. Recently, the role of protein limitations in shaping metabolic phenotypes have been extensively studied following the reconstruction of enzyme-constrained GEMs. This framework has been shown to significantly improve the accuracy of predicting microbial phenotypes, and it has demonstrated that a global limitation in protein availability can prompt the ubiquitous metabolic strategy of overflow metabolism. Being one of the most abundant and differentially expressed proteome sectors, metabolic proteins constitute a major cellular demand on proteinogenic amino acids. However, little is known about the impact and sensitivity of amino acid availability with regards to genome-scale metabolism. Here, we explore these aspects by extending on the enzyme-constrained GEM framework by also accounting for the usage of amino acids in expressing the metabolic proteome. Including amino acids in an enzyme-constrained GEM of Saccharomyces cerevisiae, we demonstrate that the expanded model is capable of accurately reproducing experimental amino acid levels. We further show that the metabolic proteome exerts variable demands on amino acid supplies in a condition-dependent manner, suggesting that S. cerevisiae must have evolved to efficiently fine-tune the synthesis of amino acids for expressing its metabolic proteins in response to changes in the external environment. Finally, our results demonstrate how the metabolic network of S. cerevisiae is robust towards perturbations of individual amino acids, while simultaneously being highly sensitive when the relative amino acid availability is set to mimic a priori distributions of both yeast and non-yeast origins.
A central driver for the field of systems biology is to develop an understanding of how interactions between components affect the functioning of a system as a whole. Network analysis is an approach that is uniquely suited to uncover patterns and organizing principles in a wide variety of complex systems. In this chapter, we will give a detailed description of basic concepts for characterizing empirical networks, frequently used random network models, and how to compute properties of networks using Python packages. We will demonstrate the application of network analysis by investigating several biological networks.
Background Differential co-expression network analysis has become an important tool to gain understanding of biological phenotypes and diseases. The CSD algorithm is a method to generate differential co-expression networks by comparing gene co-expressions from two different conditions. Each of the gene pairs is assigned conserved (C), specific (S) and differentiated (D) scores based on the co-expression of the gene pair between the two conditions. The result of the procedure is a network where the nodes are genes and the links are the gene pairs with the highest C-, S-, and D-scores. However, the existing CSD-implementations suffer from poor computational performance, difficult user procedures and lack of documentation. Results We created the R-package csdR aimed at reaching good performance together with ease of use, sufficient documentation, and with the ability to play well with other tools for data analysis. csdR was benchmarked on a realistic dataset with 20, 645 genes. After verifying that the chosen number of iterations gave sufficient robustness, we tested the performance against the two existing CSD implementations. csdR was superior in performance to one of the implementations, whereas the other did not run. Our implementation can utilize multiple processing cores. However, we were unable to achieve more than ∼ 2.7 parallel speedup with saturation reached at about 10 cores. Conclusions The results suggest that csdR is a useful tool for differential co-expression analysis and is able to generate robust results within a workday on datasets of realistic sizes when run on a workstation or compute server.
Biotechnology and BioengineeringVolume 118, Issue 5 p. 1757-1761 ISSUE INFORMATIONFree Access Biotechnology and Bioengineering: Volume 118, Number 5, May 2021 First published: 14 April 2021 https://doi.org/10.1002/bit.27401AboutPDF ToolsRequest permissionExport citationAdd to favoritesTrack citation ShareShare Give accessShare full text accessShare full-text accessPlease review our Terms and Conditions of Use and check box below to share full-text version of article.I have read and accept the Wiley Online Library Terms and Conditions of UseShareable LinkUse the link below to share a full-text version of this article with your friends and colleagues. Learn more.Copy URL Share a linkShare onEmailFacebookTwitterLinked InRedditWechat Volume118, Issue5May 2021Pages 1757-1761 RelatedInformation