Selecting an optimal antigen is a crucial step in vaccine development, significantly influencing both the vaccine’s effectiveness and the breadth of protection it provides. High antigen sequence variability, as seen in pathogens like rhinovirus, HIV, influenza virus, complicates the design of a single cross-protective antigen. Consequently, vaccination with a single antigen molecule often confers protection against only a single variant. In this study, machine learning methods were applied to the design of factor H binding protein (fHbp), an antigen from the bacterial pathogen Neisseria meningitidis. The vast number of potential antigen mutants presents a significant challenge for improving fHbp antigenicity. Moreover, limited data on antigen-antibody binding in public databases constrains the training of machine learning models. To address these challenges, we used computational models to predict fHbp properties and machine learning was applied to select both the most promising and informative mutants using a Gaussian process (GP) model. These mutants were experimentally evaluated to both confirm promising leads and refine the machine learning model for future iterations. In our current model, mutants were designed that enabled the transfer of fHbp v1.1 specific conformational epitopes onto fHbp v3.28, while maintaining binding to overlapping cross-reactive epitopes. The top mutant identified underwent biophysical and x-ray crystallographic characterization to confirm that the overall structure of fHbp was maintained throughout this epitope engineering experiment. The integrated strategy presented here could form the basis of a next-generation, iterative antigen design platform, potentially accelerating the development of new broadly protective vaccines.
ABSTRACTThe COVID-19 pandemic underscored the promise of monoclonal antibody-based prophylactic and therapeutic drugs1–3, but also revealed how quickly viral escape can curtail effective options4, 5. With the emergence of the SARS-CoV-2 Omicron variant in late 2021, many clinically used antibody drug products lost potency, including EvusheldTMand its constituent, cilgavimab4, 6. Cilgavimab, like its progenitor COV2-2130, is a class 3 antibody that is compatible with other antibodies in combination4and is challenging to replace with existing approaches. Rapidly modifying such high-value antibodies with a known clinical profile to restore efficacy against emerging variants is a compelling mitigation strategy. We sought to redesign COV2-2130 to rescue in vivo efficacy against Omicron BA.1 and BA.1.1 strains while maintaining efficacy against the contemporaneously dominant Delta variant. Here we show that our computationally redesigned antibody, 2130-1-0114-112, achieves this objective, simultaneously increases neutralization potency against Delta and many variants of concern that subsequently emerged, and provides protectionin vivoagainst the strains tested, WA1/2020, BA.1.1, and BA.5. Deep mutational scanning of tens of thousands pseudovirus variants reveals 2130-1-0114-112 improves broad potency without incurring additional escape liabilities. Our results suggest that computational approaches can optimize an antibody to target multiple escape variants, while simultaneously enriching potency. Because our approach is computationally driven, not requiring experimental iterations or pre-existing binding data, it could enable rapid response strategies to address escape variants or pre-emptively mitigate escape vulnerabilities.
Viral populations in natural infections can have a high degree of sequence diversity, which can directly impact immune escape. However, antibody potency is often tested in vitro with a relatively clonal viral populations, such as laboratory virus or pseudotyped virus stocks, which may not accurately represent the genetic diversity of circulating viral genotypes. This can affect the validity of viral phenotype assays, such as antibody neutralization assays. To address this issue, we tested whether recombinant virus carrying SARS-CoV-2 spike (VSV-SARS-CoV-2-S) stocks could be made more genetically diverse by passage, and if a stock passaged under selective pressure was more capable of escaping monoclonal antibody (mAb) neutralization than unpassaged stock or than viral stock passaged without selective pressures. We passaged VSV-SARS-CoV-2-S four times concurrently in three cell lines and then six times with or without polyclonal antiserum selection pressure. All three of the monoclonal antibodies tested neutralized the viral population present in the unpassaged stock. The viral inoculum derived from serial passage without antiserum selection pressure was neutralized by two of the three mAbs. However, the viral inoculum derived from serial passage under antiserum selection pressure escaped neutralization by all three mAbs. Deep sequencing revealed the rapid acquisition of multiple mutations associated with antibody escape in the VSV-SARS-CoV-2-S that had been passaged in the presence of antiserum, including key mutations present in currently circulating Omicron subvariants. These data indicate that viral stock that was generated under polyclonal antiserum selection pressure better reflects the natural environment of the circulating virus and may yield more biologically relevant outcomes in phenotypic assays. Thus, mAb assessment assays that utilize a more genetically diverse, biologically relevant, virus stock may yield data that are relevant for prediction of mAb efficacy and for enhancing biosurveillance.
Minimizing the human and economic costs of the COVID-19 pandemic and future pandemics requires the ability to develop and deploy effective treatments for novel pathogens as soon as possible after they emerge. To this end, we introduce a new computational pipeline for the rapid identification and characterization of binding sites in viral proteins along with the key chemical features, which we call chemotypes, of the compounds predicted to interact with those same sites. The composition of source organisms for the structural models associated with an individual binding site is used to assess the site's degree of structural conservation across different species, including other viruses and humans. We propose a search strategy for novel therapeutics that involves the selection of molecules preferentially containing the most structurally rich chemotypes identified by our algorithm. While we demonstrate the pipeline on SARS-CoV-2, it is generalizable to any new virus, as long as either experimentally solved structures for its proteins are available or sufficiently accurate predicted structures can be constructed.
Protein-ligand interactions are essential to drug discovery and drug development efforts. Desirable on-target or multitarget interactions are the first step in finding an effective therapeutic, while undesirable off-target interactions are the first step in assessing safety. In this work, we introduce a novel ligand-based featurization and mapping of human protein pockets to identify closely related protein targets and to project novel drugs into a hybrid protein-ligand feature space to identify their likely protein interactions. Using structure-based template matches from PDB, protein pockets are featured by the ligands that bind to their best co-complex template matches. The simplicity and interpretability of this approach provide a granular characterization of the human proteome at the protein-pocket level instead of the traditional protein-level characterization by family, function, or pathway. We demonstrate the power of this featurization method by clustering a subset of the human proteome and evaluating the predicted cluster associations of over 7000 compounds.
Alchemical free energy perturbation (FEP) is a rigorous and powerful technique to calculate the free energy difference between distinct chemical systems. Here we report our implementation of automated large-scale FEP calculations, using the Amber software package, to facilitate antibody design and evaluation. In combination with Hamiltonian replica exchange, our FEP simulations aim to predict the effect of mutations on both the binding affinity and the structural stability. Importantly, we incorporate multiple strategies to faithfully estimate the statistical uncertainties in the FEP results. As a case study, we apply our protocols to systematically evaluate variants of the m396 antibody for their conformational stability and their binding affinity to the spike proteins of SARS-CoV-1 and SARS-CoV-2. By properly adjusting relevant parameters, the particle collapse problems in the FEP simulations are avoided. Furthermore, large statistical errors in a small fraction of the FEP calculations are effectively reduced by extending the sampling, such that acceptable statistical uncertainties are achieved for the vast majority of the cases with a modest total computational cost. Finally, our predicted conformational stability for the m396 variants is qualitatively consistent with the experimentally measured melting temperatures. Our work thus demonstrates the applicability of FEP in computational antibody design.
We present a structure-based method for finding and evaluating structural similarities in protein regions relevant to ligand binding. PDBspheres comprises an exhaustive library of protein structure regions ('spheres') adjacent to complexed ligands derived from the Protein Data Bank (PDB), along with methods to find and evaluate structural matches between a protein of interest and spheres in the library. PDBspheres uses the LGA (Local-Global Alignment) structure alignment algorithm as the main engine for detecting structural similarities between the protein of interest and template spheres from the library, which currently contains >2 million spheres. To assess confidence in structural matches, an all-atom-based similarity metric takes side chain placement into account. Here, we describe the PDBspheres method, demonstrate its ability to detect and characterize binding sites in protein structures, show how PDBspheres-a strictly structure-based method-performs on a curated dataset of 2528 ligand-bound and ligand-free crystal structures, and use PDBspheres to cluster pockets and assess structural similarities among protein binding sites of 4876 structures in the 'refined set' of the PDBbind 2019 dataset.
Minimizing the human and economic costs of the COVID-19 pandemic and of future pandemics requires the ability to develop and deploy effective treatments for novel pathogens as soon as possible after they emerge. To this end, we introduce a unique, computational pipeline for the rapid identification and characterization of binding sites in the proteins of novel viruses as well as the core chemical components with which these sites interact. We combine molecular-level structural modeling of proteins with clustering and cheminformatic techniques in a computationally efficient manner. Similarities between our results, experimental data, and other computational studies provide support for the effectiveness of our predictive framework. While we present here a demonstration of our tool on SARS-CoV-2, our process is generalizable and can be applied to any new virus, as long as either experimentally solved structures for its proteins are available or sufficiently accurate homology models can be constructed.
For over a decade, drug-induced liver injury (DILI) has posed significant drawbacks in the synthesis and development of drugs and remains a consequential concern. With finite success within the existing preclinical models, DILI is one of the main causes of drug withdrawal or termination from the market. Particularly, this withdrawal occurs during the late stages of drug development (Kullak-Ublick, 2017). Since DILI is difficult to diagnose and treat, it has become an obstacle in the drug production market that in turn affects clinicians, pharmaceutical companies, and consumers. We propose a method for learning features of DILI-positive drugs based on the graphical relationships and patterns they possess within a network of biological databases. We also train various statistical and machine learning models on these learned features in order to classify the drugs as DILI-positive or negative. Our methods include Random Forest, Neural networks, and logistic regression classification. We utilize labeled DILI-positive and DILI-negative datasets, which were developed by the FDA and the National center for toxicological research, as well as additional literature datasets (Thakkar, 2020) in order to validate our results and assess our featurization and model accuracy. Keywords: liver toxicity, hepatoxic drug analysis, drug classification, FDA clinical trials, graph databases, data processing, graph embeddings, classification models, machine-learning featurization, model comparison.
A rapid response is necessary to contain emergent biological outbreaks before they can become pandemics. The novel coronavirus (SARS-CoV-2) that causes COVID-19 was first reported in December of 2019 in Wuhan, China and reached most corners of the globe in less than two months. In just over a year since the initial infections, COVID-19 infected almost 100 million people worldwide. Although similar to SARS-CoV and MERS-CoV, SARS-CoV-2 has resisted treatments that are effective against other coronaviruses. Crystal structures of two SARS-CoV-2 proteins, spike protein and main protease, have been reported and can serve as targets for studies in neutralizing this threat. We have employed molecular docking, molecular dynamics simulations, and machine learning to identify from a library of 26 million molecules possible candidate compounds that may attenuate or neutralize the effects of this virus. The viability of selected candidate compounds against SARS-CoV-2 was determined experimentally by biolayer interferometry and FRET-based activity protein assays along with virus-based assays. In the pseudovirus assay, imatinib and lapatinib had IC50 values below 10 μM, while candesartan cilexetil had an IC50 value of approximately 67 µM against Mpro in a FRET-based activity assay. Comparatively, candesartan cilexetil had the highest selectivity index of all compounds tested as its half-maximal cytotoxicity concentration 50 (CC50) value was the only one greater than the limit of the assay (>100 μM).
Predicting accurate protein-ligand binding affinities is an important task in drug discovery but remains a challenge even with computationally expensive biophysics-based energy scoring methods and state-of-the-art deep learning approaches. Despite the recent advances in the application of deep convolutional and graph neural network-based approaches, it remains unclear what the relative advantages of each approach are and how they compare with physics-based methodologies that have found more mainstream success in virtual screening pipelines. We present fusion models that combine features and inference from complementary representations to improve binding affinity prediction. This, to our knowledge, is the first comprehensive study that uses a common series of evaluations to directly compare the performance of three-dimensional (3D)-convolutional neural networks (3D-CNNs), spatial graph neural networks (SG-CNNs), and their fusion. We use temporal and structure-based splits to assess performance on novel protein targets. To test the practical applicability of our models, we examine their performance in cases that assume that the crystal structure is not available. In these cases, binding free energies are predicted using docking pose coordinates as the inputs to each model. In addition, we compare these deep learning approaches to predictions based on docking scores and molecular mechanic/generalized Born surface area (MM/GBSA) calculations. Our results show that the fusion models make more accurate predictions than their constituent neural network models as well as docking scoring and MM/GBSA rescoring, with the benefit of greater computational efficiency than the MM/GBSA method. Finally, we provide the code to reproduce our results and the parameter files of the trained models used in this work. The software is available as open source at https://github.com/llnl/fast. Model parameter files are available at ftp://gdo-bioinformatics.ucllnl.org/fast/pdbbind2016_model_checkpoints/.
SummaryRapid assessment of whether a pandemic pathogen may have increased transmissibility or be capable of evading existing vaccines and therapeutics is critical to mounting an effective public health response. Over the period of seven days, we utilized rapid computational prediction methods to evaluate potential public health implications of the emerging SARS-CoV-2 Omicron variant. Specifically, we modeled the structure of the Omicron variant, examined its interface with human angiotensin converting enzyme 2 (ACE-2) and evaluated the change in binding affinity between Omicron, ACE-2 and publicly known neutralizing antibodies. We also compared the Omicron variant to known Variants of Concern (VoC). Seven of the 15 Omicron mutations occurring in the spike protein receptor binding domain (RBD) occur at the ACE-2 cell receptor interface, and therefore may play a critical role in enhancing binding to ACE-2. Our estimates of Omicron RBD-ACE-2 binding affinities indicate that at least two of RBD mutations, Q493R and N501Y, contribute to enhanced ACE-2 binding, nearly doubling delta-delta-G (ddG) free energies calculated for other VoC’s. Binding affinity estimates also were calculated for 55 known neutralizing SARS-CoV-2 antibodies. Analysis of the results showed that Omicron substantially degrades binding for more than half of these neutralizing SARS-CoV-2 antibodies, and for roughly 10 times as many of the antibodies than the currently dominant Delta variant. This early study lends support to use of rapid computational risk assessments to inform public health decision-making while awaiting detailed experimental characterization and confirmation.
Structure-based Deep Fusion models were recently shown to outperform several physics- and machine learning-based protein-ligand binding affinity prediction methods. As part of a multi-institutional COVID-19 pandemic response, over 500 million small molecules were computationally screened against four protein structures from the novel coronavirus (SARS-CoV-2), which causes COVID-19. Three enhancements to Deep Fusion were made in order to evaluate more than 5 billion docked poses on SARS-CoV-2 protein targets. First, the Deep Fusion concept was refined by formulating the architecture as one, coherently backpropagated model (Coherent Fusion) to improve binding-affinity prediction accuracy. Secondly, the model was trained using a distributed, genetic hyper-parameter optimization. Finally, a scalable, high-throughput screening capability was developed to maximize the number of ligands evaluated and expedite the path to experimental evaluation. In this work, we present both the methods developed for machine learning-based high-throughput screening and results from using our computational pipeline to find SARS-CoV-2 inhibitors.
The 2014–2016 Zika virus (ZIKV) epidemic in the Americas resulted in large deposits of next-generation sequencing data from clinical samples. This resource was mined to identify emerging mutations and trends in mutations as the outbreak progressed over time. Information on transmission dynamics, prevalence, and persistence of intra-host mutants, and the position of a mutation on a protein were then used to prioritize 544 reported mutations based on their ability to impact ZIKV phenotype. Using this criteria, six mutants (representing naturally occurring mutations) were generated as synthetic infectious clones using a 2015 Puerto Rican epidemic strain PRVABC59 as the parental backbone. The phenotypes of these naturally occurring variants were examined using both cell culture and murine model systems. Mutants had distinct phenotypes, including changes in replication rate, embryo death, and decreased head size. In particular, a NS2B mutant previously detected during in vivo studies in rhesus macaques was found to cause lethal infections in adult mice, abortions in pregnant females, and increased viral genome copies in both brain tissue and blood of female mice. Additionally, mutants with changes in the region of NS3 that interfaces with NS5 during replication displayed reduced replication in the blood of adult mice. This analytical pathway, integrating both bioinformatic and wet lab experiments, provides a foundation for understanding how naturally occurring single mutations affect disease outcome and can be used to predict the of severity of future ZIKV outbreaks. To determine if naturally occurring individual mutations in the Zika virus epidemic genotype affect viral virulence or replication rate in vitro or in vivo, we generated an infectious clone representing the epidemic genotype of stain Puerto Rico, 2015. Using this clone, six mutants were created by changing nucleotides in the genome to cause one to two amino acid substitutions in the encoded proteins. The six mutants we generated represent mutations that differentiated the early epidemic genotype from genotypes that were either ancestral or that occurred later in the epidemic. We assayed each mutant for changes in growth rate, and for virulence in adult mice and pregnant mice. Three of the mutants caused catastrophic embryo effects including increased embryonic death or significant decrease in head diameter. Three other mutants that had mutations in a genome region associated with replication resulted in changes in in vitro and in vivo replication rates. These results illustrate the potential impact of individual mutations in viral phenotype.
Summary Rapidly responding to novel pathogens, such as SARS-CoV-2, represents an extremely challenging and complex endeavor. Numerous promising therapeutic and vaccine research efforts to mitigate the catastrophic effects of COVID-19 pandemic are underway, yet an efficacious countermeasure is still not available. To support these global research efforts, we have used a novel computational pipeline combining machine learning, bioinformatics, and supercomputing to predict antibody structures capable of targeting the SARS-CoV-2 receptor binding domain (RBD). In 22 days, using just the SARS-CoV-2 sequence and previously published neutralizing antibody structures for SARS-CoV-1, we generated 20 initial antibody sequences predicted to target the SARS-CoV-2 RBD. As a first step in this process, we predicted (and publicly released) structures of the SARS-CoV-2 spike protein using homology-based structural modeling. The predicted structures proved to be accurate within the targeted RBD region when compared to experimentally derived structures published weeks later. Next we used our in silico design platform to iteratively propose mutations to SARS-CoV-1 neutralizing antibodies (known not to bind SARS-Cov-2) to enable and optimize binding within the RBD of SARS-CoV-2. Starting from a calculated baseline free energy of −48.1 kcal/mol (± 8.3), our 20 selected first round antibody structures are predicted to have improved interaction with the SARS-CoV-2 RBD with free energies as low as −82.0 kcal/mole. The baseline SARS-CoV-1 antibody in complex with the SARS-CoV-1 RBD has a calculated interaction energy of −52.2 kcal/mole and neutralizes the virus by preventing it from binding and entering the human ACE2 receptor. These results suggest that our predicted antibody mutants may bind the SARS-CoV-2 RBD and potentially neutralize the virus. Additionally, our selected antibody mutants score well according to multiple antibody developability metrics. These antibody designs are being expressed and experimentally tested for binding to COVID-19 viral proteins, which will provide invaluable feedback to further improve the machine learning–driven designs. This technical report is a high-level description of that effort; the Supplementary Materials includes the homology-based structural models we developed and 178,856 in silico free energy calculations for 89,263 mutant antibodies derived from known SARS-CoV-1 neutralizing antibodies.
The question of how Zika virus (ZIKV) changed from a seemingly mild virus to a human pathogen capable of microcephaly and sexual transmission remains unanswered. The unexpected emergence of ZIKV's pathogenicity and capacity for sexual transmission may be due to genetic changes, and future changes in phenotype may continue to occur as the virus expands its geographic range. Alternatively, the sheer size of the 2015-16 epidemic may have brought attention to a pre-existing virulent ZIKV phenotype in a highly susceptible population. Thus, it is important to identify patterns of genetic change that may yield a better understanding of ZIKV emergence and evolution. However, because ZIKV has an RNA genome and a polymerase incapable of proofreading, it undergoes rapid mutation which makes it difficult to identify combinations of mutations associated with viral emergence. As next generation sequencing technology has allowed whole genome consensus and variant sequence data to be generated for numerous virus samples, the task of analyzing these genomes for patterns of mutation has become more complex. However, understanding which combinations of mutations spread widely and become established in new geographic regions versus those that disappear relatively quickly is essential for defining the trajectory of an ongoing epidemic. In this study, multiscale analysis of the wealth of genomic data generated over the course of the epidemic combined with in vivo laboratory data allowed trends in mutations and outbreak trajectory to be assessed. Mutations were detected throughout the genome via deep sequencing, and many variants appeared in multiple samples and in some cases become consensus. Similarly, amino acids that were previously consensus in pre-outbreak samples were detected as low frequency variants in epidemic strains. Protein structural models indicate that most of the mutations associated with the epidemic transmission occur on the exposed surface of viral proteins. At the macroscale level, consensus data was organized into large and interactive databases to allow the spread of individual mutations and combinations of mutations to be visualized and assessed for temporal and geographical patterns. Thus, the use of multiscale modeling for identifying mutations or combinations of mutations that impact epidemic transmission and phenotypic impact can aid the formation of hypotheses which can then be tested using reverse genetics.
Francisella tularensis is classified as a Class A bioterrorism agent by the U.S. government due to its high virulence and the ease with which it can be spread as an aerosol. It is a facultative intracellular pathogen and the causative agent of tularemia. Ciprofloxacin (Cipro) is a broad spectrum antibiotic effective against Gram-positive and Gram-negative bacteria. Increased Cipro resistance in pathogenic microbes is of serious concern when considering options for medical treatment of bacterial infections. Identification of genes and loci that are associated with Ciprofloxacin resistance will help advance the understanding of resistance mechanisms and may, in the future, provide better treatment options for patients. It may also provide information for development of assays that can rapidly identify Cipro-resistant isolates of this pathogen. In this study, we selected a large number of F. tularensis live vaccine strain (LVS) isolates that survived in progressively higher Ciprofloxacin concentrations, screened the isolates using a whole genome F. tularensis LVS tiling microarray and Illumina sequencing, and identified both known and novel mutations associated with resistance. Genes containing mutations encode DNA gyrase subunit A, a hypothetical protein, an asparagine synthase, a sugar transamine/perosamine synthetase and others. Structural modeling performed on these proteins provides insights into the potential function of these proteins and how they might contribute to Cipro resistance mechanisms.
ABSTRACT Although it is becoming clear that many microbial primary producers can also play a role as organic consumers, we know very little about the metabolic regulation of photoautotroph organic matter consumption. Cyanobacteria in phototrophic biofilms can reuse extracellular organic carbon, but the metabolic drivers of extracellular processes are surprisingly complex. We investigated the metabolic foundations of organic matter reuse by comparing exoproteome composition and incorporation of 13C-labeled and 15N-labeled cyanobacterial extracellular organic matter (EOM) in a unicyanobacterial biofilm incubated using different light regimes. In the light and the dark, cyanobacterial direct organic C assimilation accounted for 32% and 43%, respectively, of all organic C assimilation in the community. Under photosynthesis conditions, we measured increased excretion of extracellular polymeric substances (EPS) and proteins involved in micronutrient transport, suggesting that requirements for micronutrients may drive EOM assimilation during daylight hours. This interpretation was supported by photosynthesis inhibition experiments, in which cyanobacteria incorporated N-rich EOM-derived material. In contrast, under dark, C-starved conditions, cyanobacteria incorporated C-rich EOM-derived organic matter, decreased excretion of EPS, and showed an increased abundance of degradative exoproteins, demonstrating the use of the extracellular domain for C storage. Sequence-structure modeling of one of these exoproteins predicted a specific hydrolytic activity that was subsequently detected, confirming increased EOM degradation in the dark. Associated heterotrophic bacteria increased in abundance and upregulated transport proteins under dark relative to light conditions. Taken together, our results indicate that biofilm cyanobacteria are successful competitors for organic C and N and that cyanobacterial nutrient and energy requirements control the use of EOM. IMPORTANCE Cyanobacteria are globally distributed primary producers, and the fate of their fixed C influences microbial biogeochemical cycling. This fate is complicated by cyanobacterial degradation and assimilation of organic matter, but because cyanobacteria are assumed to be poor competitors for organic matter consumption, regulation of this process is not well tested. In mats and biofilms, this is especially relevant because cyanobacteria produce an extensive organic extracellular matrix, providing the community with a rich source of nutrients. Light is a well-known regulator of cyanobacterial metabolism, so we characterized the effects of light availability on the incorporation of organic matter. Using stable isotope tracing at the single-cell level, we quantified photoautotroph assimilation under different metabolic conditions and integrated the results with proteomics to elucidate metabolic status. We found that cyanobacteria effectively compete for organic matter in the light and the dark and that nutrient requirements and community interactions contribute to cycling of extracellular organic matter.
In vivo serial passage of non-pathogenic viruses has been shown to lead to increased viral virulence, and although the precise mechanism(s) are not clear, it is known that both host and viral factors are associated with increased pathogenicity. Under- or overnutrition leads to a decreased or dysregulated immune response and can increase viral mutant spectrum diversity and virulence. The objective of this study was to identify the role of viral mutant spectra dynamics and host immunocompetence in the development of pathogenicity during in vivo passage. Because the nutritional status of the host has been shown to affect the development of viral virulence, the diet of animal model reflected two extremes of diets which exist in the global population, malnutrition and obesity. Sendai virus was serially passaged in groups of mice with differing nutritional status followed by transmission of the passaged virus to a second host species, guinea pigs. Viral population dynamics were characterized using deep sequence analysis and computational modeling. Histopathology, viral titer and cytokine assays were used to characterize viral virulence. Viral virulence increased with passage and the virulent phenotype persisted upon passage to a second host species. Additionally, nutritional status of mice during passage influenced the phenotype. Sequencing revealed the presence of several non-synonymous changes in the consensus sequence associated with passage, a majority of which occurred in the hemagglutinin-neuraminidase and polymerase genes, as well as the presence of persistent high frequency variants in the viral population. In particular, an N1124D change in the consensus sequences of the polymerase gene was detected by passage 10 in a majority of the animals. In vivo comparison of an 1124D plaque isolate to a clone with 1124N genotype indicated that 1124D was associated with increased virulence.
Hepatitis C Virus (HCV) infects 200 million individuals worldwide. Although several FDA approved drugs targeting the HCV serine protease and polymerase have shown promising results, there is a need for better drugs that are effective in treating a broader range of HCV genotypes and subtypes without being used in combination with interferon and/or ribavirin. Recently, two crystal structures of the core of the HCV E2 protein (E2c) have been determined, providing structural information that can now be used to target the E2 protein and develop drugs that disrupt the early stages of HCV infection by blocking E2's interaction with different host factors. Using the E2c structure as a template, we have created a structural model of the E2 protein core (residues 421-645) that contains the three amino acid segments that are not present in either structure. Computational docking of a diverse library of 1,715 small molecules to this model led to the identification of a set of 34 ligands predicted to bind near conserved amino acid residues involved in the HCV E2: CD81 interaction. Surface plasmon resonance detection was used to screen the ligand set for binding to recombinant E2 protein, and the best binders were subsequently tested to identify compounds that inhibit the infection of Huh-7 cells by HCV. One compound, 281816, blocked E2 binding to CD81 and inhibited HCV infection in a genotype-independent manner with IC50's ranging from 2.2 µM to 4.6 µM. 281816 blocked the early and late steps of cell-free HCV entry and also abrogated the cell-to-cell transmission of HCV. Collectively the results obtained with this new structural model of E2c suggest the development of small molecule inhibitors such as 281816 that target E2 and disrupt its interaction with CD81 may provide a new paradigm for HCV treatment.