Background: Knowledge about the origin of SARS-CoV-2 is necessary for both a biological and epidemiological understanding of the COVID-19 pandemic. Evidence suggests that a proximal evolutionary ancestor of SARS-CoV-2 belongs to the bat coronavirus family. However, as further evidence for a direct zoonosis remains limited, alternative modes of SARS-CoV-2 biogenesis should be considered. Results: Here we show that the genomes from SARS-CoV-2 and from SARS-CoV-1 are differentially enriched with short chromosomal sequences from the yeast S. cerevisiae at focal positions that are known to be critical for host cell invasion, virus replication, and host immune response. For SARS-CoV-1, we identify two sites: one at the start of the RNA dependent RNA polymerase gene, and the other at the start of the spike protein’s receptor binding domain; for SARS-CoV-2, one at the start of the viral replicase domain, and the other toward the end of the spike gene past its domain junction. At this junction, we detect a highly specific stretch of yeast DNA origin covering the critical furin cleavage site insert PRRA, which has not been seen in other lineage b betacoronaviruses. As yeast is not a natural host for this virus family, we propose a passage model for viral constructs in yeast cells based on co-transformation of virus DNA plasmids carrying yeast selectable genetic markers followed by intra-chromosomal homologous recombination through gene conversion. Highly differential sequence homology data across yeast chromosomes congruent with chromosomes harboring specific auxotrophic markers further support this passage model. Conclusions: These results provide evidence that among SARS-like coronaviruses only the genomes of SARS-CoV-1 and SARS-CoV-2 contain information that points to a synthetic passage in genetically modified yeast cells. Our data specifically allow the identification of the yeast S. cerevisiae as a potential recombination donor for the critical furin cleavage site in SARS-CoV-2.
Background: Knowledge about the origin of SARS-CoV-2 is necessary for both a biological and epidemiological understanding of the COVID-19 pandemic. Evidence suggests that a proximal evolutionary ancestor of SARS-CoV-2 belongs to the bat coronavirus family. However, as further evidence for a direct zoonosis remains limited, alternative modes of SARS-CoV-2 biogenesis should be also considered. Results: Here we show that the genomes from SARS-CoV-2 and from SARS-CoV-1 are differentially enriched with short chromosomal sequences from the yeast S. cerevisiae at focal positions that are known to be critical for virus replication, host cell invasion, and host immune response. Specifically, for SARS-CoV-2, we identify two sites: one at the start of the viral replicase domain, and the other at the end of the spike gene past its critical domain junction; for SARS-CoV-1, one at the start of the RNA dependent RNA polymerase gene, and the other at the start of the spike protein’s receptor binding domain. As yeast is not a natural host for this virus family, we propose a directed passage model for viral constructs, including virus replicase, in yeast cells based on co-transformation of virus DNA plasmids carrying yeast selectable genetic markers followed by intra-chromosomal homologous recombination through gene conversion. Highly differential sequence homology data across yeast chromosomes congruent with chromosomes harboring specific auxotrophic markers further support this passage model. Model and data together allow us to infer a hypothetical tripartite genome assembly scheme for the synthetic biogenesis of SARS-CoV-2 and SARS-CoV-1. Conclusions: These results provide evidence that the genome sequences of SARS-CoV-1, SARS-CoV-2, but not that of RaTG13, BANAL-20-52 and all other closest SARS coronavirus family members identified, are carriers of distinct homology signals that might point to large-scale genomic editing during a passage of directed replication and chromosomal integration inside genetically modified yeast cells.
Recent independent results by Zhang et al. (Reverse-transcribed SARS-CoV-2 RNA can integrate into the genome of cultured human cells and can be expressed in patient-derived tissues. Proc. Natl. Acad. Sci. U. S. A. 118, e2105968118, 2021) and by Briggs et al. (Assessment of potential SARS-CoV-2 virus integration into human genome reveals no significant impact on RT-qPCR COVID-19 testing. Proc. Natl. Acad. Sci. U. S. A. 118, e2113065118, 2021) suggest that LINE1-mediated retrotransposition and integration of the SARS-CoV-2 genome, by means of recombinant complementary DNA, into the host genomes of SARS-CoV-2 infected humans are detectable with standard PCR diagnostic methods. On this background, a new and far-reaching patent for a SARS-CoV-2 reverse genetic system (published 30 September 2021 as US patent US2021024532W and World Intellectual Property Organization WO2021195596A2, https://worldwide.espacenet.com/patent/search?q=pn%3DWO2021195596A2) with legal claims toward artificial SARS-CoV-2 complementary DNA and its sequence variants in host cells is briefly discussed. It is argued that this might pose historically a unique situation, where genomes of living persons may have become both the targets of recombination with artificial DNA and the objects of emerging patent claims. This situation and its potential implications need an open and transparent public discussion.
Both vaccine candidates offer protection to older adults
Recently released interim numbers from advanced vaccine candidate clinical trials suggest that a COVID-19 vaccine effectiveness (VE) of >90% is achievable. However, SARS-CoV-2 transmission dynamics are highly heterogeneous and exhibit localized bursts of transmission, which may lead to sharp localized peaks in the number of new cases, often followed by longer periods of low incidence. Here we show that, for interim estimates of VE, these characteristic bursts in SARS-CoV-2 infection may introduce a strong positive bias in VE. Specifically, we generate null models of vaccine effectiveness, i.e., random models with bursts that over longer periods converge to zero VE but that for interim periods frequently produce apparent VE near 100%. As an example, by following the relevant clinical trial protocol, we can reproduce recently reported interim outcomes from an ongoing phase 3 clinical trial of an RNA-based vaccine candidate. Thus, to avoid potential random biases in VE, it is suggested that interim estimates on COVID-19 VE should control for the intrinsic inhomogeneity in both SARS-CoV-2 infection dynamics and reported cases.
From December 2019 to early March 2020, the local outbreak of novel corona virus disease (COVID-19) in central China’s Hubei region has grown into a worldwide pandemic. This rapid and catastrophic escalation makes the search for and understanding of the underlying mechanisms of infection and disease, caused by severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), as well as of their associated risk factors an urgent priority. In particular, strong variations in COVID-19 infection rates as seen internationally require a better understanding. Here we show that reported influenza vaccination coverage rates for 29 OECD countries are associated significantly with recently observed SARS-CoV-2 infection rates in these countries. This early observation, which merits further investigation, suggests that during the current coronavirus outbreak an influenza vaccination background might be a relevant factor for SARS-CoV-2 infection. The observed phenomenon is discussed in the context of vaccine associated virus interference and antibody-dependent enhancement of viral infectivity.
In the human malaria parasite Plasmodium falciparum, membrane glutathione S-transferases (GST) have recently emerged as potential cellular detoxifying units and as drug target candidates with the artemisinin (ART) class of antimalarials inhibiting their activity at single-digit nanomolar potency when activated by iron sources such as cytotoxic hematin. Here we put forward the hypothesis that the membrane GST Plasmodium falciparum exported protein 1 (PfEXP1, PF3D7_1121600) might be directly involved in the mode of action of the unrelated antimalarial 4-aminoquinoline drug chloroquine (CQ). Along this line we report potent biochemical inhibition of membrane glutathione S-transferase activity in recombinant PfEXP1 through CQ at half maximal inhibitory CQ concentrations of 9.02 nM and 19.33 nM when using hematin and the iron deficient 1-chloro-2,4-dinitrobenzene (CDNB) as substrate, respectively. Thus, in contrast to ART, CQ may not require activation through an iron source such as hematin for a potent inhibition of membrane GST activity. Arguably, these data represent the first instance of low nanomolar inhibition of an essential Plasmodium falciparum enzyme through a 4-aminoquinoline and might encourage further investigation of PfEXP1 as a potential CQ target candidate.
We present KnIT, the Knowledge Integration Toolkit, a system for accelerating scientific discovery and predicting previously unknown protein protein interactions. Such predictions enrich biological research and are pertinent to drug discovery and the understanding of disease. Unlike a prior study, KnIT is now fully automated and demonstrably scalable. It extracts information from the scientific literature, automatically identifying direct and indirect references to protein interactions, which is knowledge that can be represented in network form. It then reasons over this network with techniques such as matrix factorization and graph diffusion to predict new, previously unknown interactions. The accuracy and scope of KnIT's knowledge extractions are validated using comparisons to structured, manually curated data sources as well as by performing retrospective studies that predict subsequent literature discoveries using literature available prior to a given date. The KnIT methodology is a step towards automated hypothesis generation from text, with potential application to other scientific domains.
A central problem in biology is to identify gene function. One approach is to infer function in large supergenomic networks of interactions and ancestral relationships among genes; however, their analysis can be computationally prohibitive. We show here that these biological networks are compressible. They can be shrunk dramatically by eliminating redundant evolutionary relationships, and this process is efficient because in these networks the number of compressible elements rises linearly rather than exponentially as in other complex networks. Compression enables global network analysis to computationally harness hundreds of interconnected genomes and to produce functional predictions. As a demonstration, we show that the essential, but functionally uncharacterized Plasmodium falciparum antigen EXP1 is a membrane glutathione S-transferase. EXP1 efficiently degrades cytotoxic hematin, is potently inhibited by artesunate, and is associated with artesunate metabolism and susceptibility in drug-pressured malaria parasites. These data implicate EXP1 in the mode of action of a frontline antimalarial drug.
Keeping up with the ever-expanding flow of data and publications is untenable and poses a fundamental bottleneck to scientific progress. Current search technologies typically find many relevant documents, but they do not extract and organize the information content of these documents or suggest new scientific hypotheses based on this organized content. We present an initial case study on KnIT, a prototype system that mines the information contained in the scientific literature, represents it explicitly in a queriable network, and then further reasons upon these data to generate novel and experimentally testable hypotheses. KnIT combines entity detection with neighbor-text feature analysis and with graph-based diffusion of information to identify potential new properties of entities that are strongly implied by existing relationships. We discuss a successful application of our approach that mines the published literature to identify new protein kinases that phosphorylate the protein tumor suppressor p53. Retrospective analysis demonstrates the accuracy of this approach and ongoing laboratory experiments suggest that kinases identified by our system may indeed phosphorylate p53. These results establish proof of principle for automated hypothesis generation and discovery based on text mining of the scientific literature.
Membrane glutathione S-transferases from the class of membrane-associated proteins in eicosanoid and glutathione metabolism (MAPEG) form a superfamily of detoxification enzymes that catalyze the conjugation of reduced glutathione (GSH) to a broad spectrum of xenobiotics and hydrophobic electrophiles. Evolutionarily unrelated to the cytosolic glutathione S-transferases, they are found across bacterial and eukaryotic domains, for example in mammals, plants, fungi and bacteria in which significant levels of glutathione are maintained. Species of genus Plasmodium, the unicellular protozoa that are commonly known as malaria parasites, do actively support glutathione homeostasis and maintain its metabolism throughout their complex parasitic life cycle. In humans and in other mammals, the asexual intraerythrocytic stage of malaria, when the parasite feeds on hemoglobin, grows and eventually asexually replicates inside infected red blood cells (RBCs), is directly associated with host disease symptoms and during this critical stage GSH protects the host RBC and the parasite against oxidative stress from parasite-induced hemoglobin catabolism. In line with these observations, several GSH-dependent Plasmodium enzymes have been characterized including glutathione reductases, thioredoxins, glyoxalases, glutaredoxins and glutathione S-transferases (GSTs); furthermore, GSH itself have been found to associate spontaneously and to degrade free heme and its hydroxide, hematin, which are the main cytotoxic byproducts of hemoglobin catabolism. However, despite the apparent importance of glutathione metabolism for the parasite, no membrane associated glutathione S-transferases of genus Plasmodium have been previously described. We recently reported the first examples of MAPEG members among Plasmodium spp.
Automated annotation of protein function is challenging. As the number of sequenced genomes rapidly grows, the overwhelming majority of protein products can only be annotated computationally. If computational predictions are to be relied upon, it is crucial that the accuracy of these methods be high. Here we report the results from the first large-scale community-based critical assessment of protein function annotation (CAFA) experiment. Fifty-four methods representing the state of the art for protein function prediction were evaluated on a target set of 866 proteins from 11 organisms. Two findings stand out: (i) today's best protein function prediction algorithms substantially outperform widely used first-generation methods, with large gains on all types of targets; and (ii) although the top methods perform well enough to guide experiments, there is considerable need for improvement of currently available tools.
Background Annotating protein function with both high accuracy and sensitivity remains a major challenge in structural genomics. One proven computational strategy has been to group a few key functional amino acids into templates and search for these templates in other protein structures, so as to transfer function when a match is found. To this end, we previously developed Evolutionary Trace Annotation (ETA) and showed that diffusing known annotations over a network of template matches on a structural genomic scale improved predictions of function. In order to further increase sensitivity, we now let each protein contribute multiple templates rather than just one, and also let the template size vary. Results Retrospective benchmarks in 605 Structural Genomics enzymes showed that multiple templates increased sensitivity by up to 14% when combined with single template predictions even as they maintained the accuracy over 91%. Diffusing function globally on networks of single and multiple template matches marginally increased the area under the ROC curve over 0.97, but in a subset of proteins that could not be annotated by ETA, the network approach recovered annotations for the most confident 20-23 of 91 cases with 100% accuracy. Conclusions We improve the accuracy and sensitivity of predictions by using multiple templates per protein structure when constructing networks of ETA matches and diffusing annotations.
Mechanisms of DNA repair and mutagenesis are defined on the basis of relatively few proteins acting on DNA, yet the identities and functions of all proteins required are unknown. Here, we identify the network that underlies mutagenic repair of DNA breaks in stressed Escherichia coli and define functions for much of it. Using a comprehensive screen, we identified a network of ≥93 genes that function in mutation. Most operate upstream of activation of three required stress responses (RpoS, RpoE, and SOS, key network hubs), apparently sensing stress. The results reveal how a network integrates mutagenic repair into the biology of the cell, show specific pathways of environmental sensing, demonstrate the centrality of stress responses, and imply that these responses are attractive as potential drug targets for blocking the evolution of pathogens.
Genomic centers discover increasingly many protein sequences and structures, but not necessarily their full biological functions. Thus, currently, less than one percent of proteins have experimentally verified biochemical activities. To fill this gap, function prediction algorithms apply metrics of similarity between proteins on the premise that those sufficiently alike in sequence, or structure, will perform identical functions. Although high sensitivity is elusive, network analyses that integrate these metrics together hold the promise of rapid gains in function prediction specificity.
In many graph-based semi-supervised learning algorithms, edge weights are assumed to be fixed and determined by the data points' (often symmetric) relationships in input space, without considering directionality. However, relationships may be more informative in one direction (e.g. from labelled to unlabelled) than in the reverse direction, and some relationships (e.g. strong weights between oppositely labelled points) are unhelpful in either direction. Undesirable edges may reduce the amount of influence an informative point can propagate to its neighbours - the point and its outgoing edges have been ''blunted.'' We present an approach to ''sharpening'' in which weights are adjusted to meet an optimization criterion wherever they are directed towards labelled points. This principle can be applied to a wide variety of algorithms. In this paper, we present one solution satisfying the principle, in order to show that it can improve performance on a number of publicly available bench-mark data sets. When tested on a real-world problem, protein function classification with four vastly different molecular similarity graphs, sharpening improved ROC scores by 16% on average, at negligible computational cost.
High-throughput Structural Genomics yields many new protein structures without known molecular function. This study aims to uncover these missing annotations by globally comparing select functional residues across the structural proteome. First, Evolutionary Trace Annotation, or ETA, identifies which proteins have local evolutionary and structural features in common; next, these proteins are linked together into a proteomic network of ETA similarities; then, starting from proteins with known functions, competing functional labels diffuse link-by-link over the entire network. Every node is thus assigned a likelihood z-score for every function, and the most significant one at each node wins and defines its annotation. In high-throughput controls, this competitive diffusion process recovered enzyme activity annotations with 99% and 97% accuracy at half-coverage for the third and fourth Enzyme Commission (EC) levels, respectively. This corresponds to false positive rates 4-fold lower than nearest-neighbor and 5-fold lower than sequence-based annotations. In practice, experimental validation of the predicted carboxylesterase activity in a protein from Staphylococcus aureus illustrated the effectiveness of this approach in the context of an increasingly drug-resistant microbe. This study further links molecular function to a small number of evolutionarily important residues recognizable by Evolutionary Tracing and it points to the specificity and sensitivity of functional annotation by competitive global network diffusion. A web server is at http://mammoth.bcm.tmc.edu/networks.
Recurrent international financial crises inflict significant damage to societies and stress the need for mechanisms or strategies to control risk and tamper market uncertainties. Unfortunately, the complex network of market interactions often confounds rational approaches to optimize financial risks. Here we show that investors can overcome this complexity and globally minimize risk in portfolio models for any given expected return, provided the margin requirement remains below a critical, empirically measurable value. In practice, for markets with centrally regulated margin requirements, a rational stabilization strategy would be keeping margins small enough. This result follows from ground states of the random field spin glass Ising model that can be calculated exactly through convex optimization when relative spin coupling is limited by the norm of the network's Laplacian matrix. In that regime, this novel approach is robust to noise in empirical data and may be also broadly relevant to complex networks with frustrated interactions that are studied throughout scientific fields.