BACKGROUND:The HVTN108 trial evaluated the safety and immunogenicity of a DNA prime, adjuvanted protein boost HIV vaccine in the United States and South Africa. The underlying factors influencing individual variation in vaccine responsiveness are unknown. In this study, we defined the IgG Fc and Fc receptor (FcR) genotypes in the HVTN108 cohort to test our hypothesis that IgG and FcR genetic variation can affect vaccine-elicited functional antibody responses. METHODS:IgG Fc and FcR alleles were determined by targeted PCR amplification and next-generation sequencing. Vaccine-elicited functional antibody responses, including binding antibody multiplex assay (BAMA), antibody-dependent cellular cytotoxicity (ADCC), and antibody-dependent cellular phagocytosis (ADCP) activity, were measured using standardized and qualified methods. Relationships between alleles and antibody responses were identified by linear regression controlling for treatment group and region. RESULTS:The distribution of many polymorphisms significantly differed between the United States and South Africa. Within the subset of the cohort tested for functional antibody responses (IgG, n = 41; FcR, n = 55), IgG genotypes such as IGHG1*12 ( P = 0.012), IGHG3*11 ( P = 0.033), IGHG2*02 ( P = 0.038), IGHG4*07 ( P = 0.076), and others were associated with ADCC antibody responses when corrected for vaccine group and regional effects. In the same way, we identified that the FCER1A rs2427827 mutation had a significant association with lower peak ADCC activity and the FCER2 rs2228137 mutation was associated with lower antibody binding to Con6 gp120 protein. CONCLUSIONS:Genetic variation in both antibodies and FcRs associated with levels of HIV-vaccine-elicited functional antibodies. Significant regional differences in distribution of this variation support the need for vaccine testing in diverse populations.
HLA-DR genes are associated with the progression from stage 1 and stage 2 to onset of stage 3 type 1 diabetes (T1D), after accounting HLA-DQ genes with which they are in high linkage disequilibrium. Based on an integrated cohort of participants from 2 completed clinical trials, this investigation finds that, sharing a haplotype with the DRB1*03:01 (DR3) allele, DRB3*01:01:02 and *02:02:01 have respectively negative and positive associations with the progression. Furthermore, we uncovered 2 residues (β11, β26, participating in pockets 6 and 4, respectively) on the DRB3 molecule responsible for the progression among DR3 carriers; motif RY and LF respectively delay and promote the progression (hazard ratio [HR] = 0.73 and 2.38, P = 0.039 and 0.017, respectively). Two anchoring pockets 6 and 4 probably bind differential autoantigenic epitopes. We further investigated the progression association with the motifs RY and LF among carriers of DR3 and found that carriers of the motif LF have significantly faster progression than carriers of RY (HR = 1.48, P = 0.019 in unadjusted analysis; HR = 1.39, P = 0.047 in adjusted analysis), results of which provide an impetus to examine the possible role of specific DRB3-binding peptides in the progression to T1D.
The aim of this work was to explore associations between type 1 diabetes progression from stages 1 or 2 to stage 3 and interacting ligand–receptor complexes of HLA class I (HLA-I) and KIR gene products. Applying next-generation sequencing technology to genotype HLA-I genes (HLA-A, -B, -C) and KIR genes (KIR2DL1, KIR2DL2, KIR2DL3, KIR2DL4, KIR2DL5, KIR2DS1, KIR2DS2, KIR2DS3, KIR2DS4, KIR2DS5, KIR3DL1, KIR3DL3, KIR3DS1, KIR2DP1, KIR3DP1) from 1215 participants in the Diabetes Prevention Trial-Type 1 (DPT-1) and the Diabetes Prevention Trial (TN07), we systematically explored associations of HLA-I–KIR ligand–receptor interactions (LRIs) with disease progression via a Cox regression model. We investigated the structural properties of identified LRI complexes. KIR and HLA-I genes had no or sporadic associations with disease progression. Out of all possible LRIs, nine HLA-A Ligands and 14 HLA-B ligands with corresponding receptors had modest associations with progression (p<0.05). As an example, carriers of A*03:01-KIR2DS4 had slower progression (HR 0.36, p=3.06 × 10−2), as did B*07:02-KIR2DL3 carriers (HR 0.26, p=7.76 × 10−3). Structural investigations of KIR–HLA-I complexes via homology modelling based on already-solved respective complex structures suggested that the respective electrostatic and van der Waals interactions encoded in the protein sequences result in strong biophysical LRIs, which could alter the progression of type 1 diabetes. These results reveal that LRIs of KIR–HLA-I gene products, rather than individual genes, contribute to type 1 diabetes progression, and such interactions are likely to be stabilised by electrostatic and van der Waals forces. As the KIR–HLA-I interactions involve part of the C-terminus of the antigen-binding groove of HLA-I, but may be affected by the respective bound peptide, this suggests a new mechanism for type 1 diabetes pathogenesis. Clinical data on participants in DPT-1 and TN07 can be obtained from the NIDDK-Central Repository ( https://repository.niddk.nih.gov/home ) following the formal approval process.
Chimerism analysis by Next-Generation Sequencing (NGS) is an emerging method for engraftment monitoring post-allogeneic hematopoietic cell transplantation. A high-sensitivity methodology is required for the detection of microchimerism (< 1 % chimerism) which may have clinical utility in early relapse detection, allograft monitoring in organ transplantation and other allogeneic cellular therapies (such as microtransplantations). As more clinical laboratories adopt this methodology, a thorough assessment of performance is needed. This study evaluated one such NGS-based assay that utilizes both SNPs and InDels as genetic markers. An assessment of accuracy, linearity, sensitivity and reproducibility was performed. Analytical sensitivity was 0.2 % donor for single donor and 0.5 % donors for double donors. The assay showed a high degree of reproducibility over a full range of chimerism. Comparison to STR-PCR showed high concordance; yet below 5 % chimerism was consistently detected by NGS, but not by STR-PCR. Comparison to qPCR showed high concordance, but with lower correlation in the mid-range (40-60 % chimerism). Overall, the assay showed consistent performance with high sensitivity and accuracy compared to STR-PCR and qPCR across a full range of chimerism in the setting of single and multi-donor transplantations. In addition, criteria for quality metrics were established for sequencing performance and data analysis. We conclude with a discussion of considerations for clinical laboratory validation of NGS-based chimerism assay and analysis software.
Objective To explore if oral insulin could delay onset of stage 3 type 1 diabetes (T1D) among stage 1/2 patients who carry HLA DR4-DQ8 and/or have elevated levels of IA-2 autoantibodies (IA-2A). Research and methods Next generation targeted sequencing technology was used to genotype eight HLA class II genes (DQA1, DQB1, DRB1, DRB3, DRB4, DRB5, DPA1, and DPB1) in 546 participants in the TrialNet oral insulin preventative trial (TN07). Baseline levels of autoantibodies against insulin (IAA), GAD65 (GADA), and IA-2A were determined prior to treatment assignment. Available clinical and demographic covariables from TN07 were used in this post-hoc analysis with the Cox regression model to quantify the preventative efficacy of oral insulin . Results 1) Oral insulin was found to reduce the frequency of T1D onset among participants with elevated IA-2A levels (HR=0.62, p=0.012), but had no preventive effect among those with low IA-2A levels (HR=1.03, p=0.91). 2) High IA-2A levels were found to be positively associated with the HLA DR4-DQ8 (OR=1.63, p=6.37*10-6) haplotype and negatively associated with the HLA DR7-containing DRB1*07:01-DRB4*01:01-DQA1*02:01-DQB1*02:02 (OR=0.49, p=0.037) extended haplotype. 3) Among DR4-DQ8 carriers, oral insulin delayed the progression towards stage 3 T1D onset (HR=0.59, p=0.027), especially if participants also had high IA-2A levels (HR=0.50, p=0.028). Conclusions These results suggest the presence of a T1D endotype characterized by HLA DR4-DQ8 and/or elevated IA-2A levels, and, for those stage 1/2 patients with such an endotype, oral insulin is found to delaythe clinical T1D onset.
The aim of this work was to explore molecular amino acids (AAs) and related structures of HLA-DQA1-DQB1 that underlie its contribution to the progression from stages 1 or 2 to stage 3 type 1 diabetes. Using high-resolution DQA1 and DQB1 genotypes from 1216 participants in the Diabetes Prevention Trial-Type 1 and the Diabetes Prevention Trial, we applied hierarchically organised haplotype association analysis (HOH) to decipher which AAs contributed to the associations of DQ with disease and their structural properties. HOH relied on the Cox regression to quantify the association of DQ with time-to-onset of type 1 diabetes. By numerating all possible DQ heterodimers of α- and β-chains, we showed that the heterodimerisation increases genetic diversity at the cellular level from 43 empirically observed haplotypes to 186 possible heterodimers. Heterodimerisation turned several neutral haplotypes (DQ2.2, DQ2.3 and DQ4.4) to risk haplotypes (DQ2.2/2.3-DQ4.4 and DQ4.4-DQ2.2). HOH uncovered eight AAs on the α-chain (−16α, −13α, −6α, α22, α23, α44, α72, α157) and six AAs on the β-chain (−18β, β9, β13, β26, β57, β135) that contributed to the association of DQ with progression of type 1 diabetes. The specific AAs concerned the signal peptide (minus sign, possible linkage to expression levels), pockets 1, 4 and 9 in the antigen-binding groove of the α1β1 domain, and the putative homodimerisation of the αβ heterodimers. These results unveil the contribution made by DQ to type 1 diabetes progression at individual residues and related protein structures, shedding light on its immunological mechanisms and providing new leads for developing treatment strategies. Clinical trial data and biospecimen samples are available through the National Institute of Diabetes and Digestive and Kidney Diseases Central Repository portal ( https://repository.niddk.nih.gov/studies ).
Plasmodium falciparum reticulocyte-binding protein homolog 5 (RH5) is the most advanced blood-stage malaria vaccine candidate and is being evaluated for efficacy in endemic regions, emphasizing the need to study the underlying antibody response to RH5 during natural infection, which could augment or counteract responses to vaccination. Here, we found that RH5-reactive B cells were rare, and circulating immunoglobulin G (IgG) responses to RH5 were short-lived in malaria-exposed Malian individuals, despite repeated infections over multiple years. RH5-specific monoclonal antibodies isolated from eight malaria-exposed individuals mostly targeted non-neutralizing epitopes, in contrast to antibodies isolated from five RH5-vaccinated, malaria-naive UK individuals. However, MAD8-151 and MAD8-502, isolated from two malaria-exposed Malian individuals, were among the most potent neutralizers out of 186 antibodies from both cohorts and targeted the same epitopes as the most potent vaccine-induced antibodies. These results suggest that natural malaria infection may boost RH5-vaccine-induced responses and provide a clear strategy for the development of next-generation RH5 vaccines.
Objective To explore associations of HLA class II genes (HLAII) with the progression of islet autoimmunity from asymptomatic to symptomatic type 1 diabetes (T1D). Research design and methods Next-generation targeted sequencing was used to genotype eight HLAII genes (DQA1, DQB1, DRB1, DRB3, DRB4, DRB5, DPA1, DPB1) in 1,216 participants from the Diabetes Prevention Trial-1 (DPT-1) and Randomized Diabetes Prevention Trial with Oral Insulin sponsored by TrialNet (TN07). By the linkage-disequilibrium, DQA1 and DQB1 are haplotyped to form DQ haplotypes; DP and DR haplotypes are similarly constructed. Together with available clinical covariables, we applied the Cox regression model to assess HLAII immunogenic associations with the disease progression. Results 1) The current investigation updated the previously reported genetic associations of DQA1*03:01-DQB1*03:02 (HR=1.25, p=3.50*10-3 ) and DQA1*03:03-DQB1*03:01 (HR=0.56, p=1.16*10-3), and also uncovered a risk association with DQA1*05:01-DQB1*02:01 (HR=1.19, p=0.041). 2) After adjusting for DQ, DPA1*02:01-DPB1*11:01 and DPA1*01:03-DPB1*03:01 were found to have opposite associations with progression (HR=1.98 and 0.70, p=0.021 and 6.16*10-3, respectively). 3) DRB1*03:01-DRB3*01:01 and DRB1*03:01-DRB3*02:02, sharing the DRB1*03:01, had opposite associations (HR=0.73 and 1.44, p=0.04 and 0.019, respectively), indicating a role of DRB3. Meanwhile, DRB1*12:01-DRB3*02:02 and DRB1*01:03 alone were found to associate with progression (HR=2.6 and 2.32, p=0.018 and 0.039, respectively). 4) through enumerating all heterodimers, it was found that both DQ and DP could exhibit associations with disease progression. Conclusions These results suggest that HLAII polymorphisms influence progression from islet autoimmunity to T1D among at-risk subjects with islet autoantibodies.
Importance Earlier detection of emerging novel SARS-COV-2 variants is important for public health surveillance of potential viral threats and for earlier prevention research. Artificial intelligence may facilitate early detection of SARS-CoV2 emerging novel variants based on variant-specific mutation haplotypes and, in turn, be associated with enhanced implementation of risk-stratified public health prevention strategies. Objective To develop a haplotype-based artificial intelligence (HAI) model for identifying novel variants, including mixture variants (MVs) of known variants and new variants with novel mutations. Design, Setting, and Participants This cross-sectional study used serially observed viral genomic sequences globally (prior to March 14, 2022) to train and validate the HAI model and used it to identify variants arising from a prospective set of viruses from March 15 to May 18, 2022. Main Outcomes and Measures Viral sequences, collection dates, and locations were subjected to statistical learning analysis to estimate variant-specific core mutations and haplotype frequencies, which were then used to construct an HAI model to identify novel variants. Results Through training on more than 5 million viral sequences, an HAI model was built, and its identification performance was validated on an independent validation set of more than 5 million viruses. Its identification performance was assessed on a prospective set of 344 901 viruses. In addition to achieving an accuracy of 92.8% (95% CI within 0.1%), the HAI model identified 4 Omicron MVs (Omicron-Alpha, Omicron-Delta, Omicron-Epsilon, and Omicron-Zeta), 2 Delta MVs (Delta-Kappa and Delta-Zeta), and 1 Alpha-Epsilon MV, among which Omicron-Epsilon MVs were most frequent (609/657 MVs [92.7%]). Furthermore, the HAI model found that 1699 Omicron viruses had unidentifiable variants given that these variants acquired novel mutations. Lastly, 524 variant-unassigned and variant-unidentifiable viruses carried 16 novel mutations, 8 of which were increasing in prevalence percentages as of May 2022. Conclusions and Relevance In this cross-sectional study, an HAI model found SARS-COV-2 viruses with MV or novel mutations in the global population, which may require closer examination and monitoring. These results suggest that HAI may complement phylogenic variant assignment, providing additional insights into emerging novel variants in the population.
BACKGROUNDMosaic and consensus HIV-1 immunogens provide two distinct approaches to elicit greater breadth of coverage against globally circulating HIV-1 and have shown improved immunologic breadth in nonhuman primate models.METHODSThis double-blind randomized trial enrolled 105 healthy HIV-uninfected adults who received 3 doses of either a trivalent global mosaic, a group M consensus (CON-S), or a natural clade B (Nat-B) gp160 env DNA vaccine followed by 2 doses of a heterologous modified vaccinia Ankara-vectored HIV-1 vaccine or placebo. We performed prespecified blinded immunogenicity analyses at day 70 and day 238 after the first immunization. T cell responses to vaccine antigens and 5 heterologous Env variants were fully mapped.RESULTSEnv-specific CD4+ T cell responses were induced in 71% of the mosaic vaccine recipients versus 48% of the CON-S recipients and 48% of the natural Env recipients. The mean number of T cell epitopes recognized was 2.5 (95% CI, 1.2-4.2) for mosaic recipients, 1.6 (95% CI, 0.82-2.6) for CON-S recipients, and 1.1 (95% CI, 0.62-1.71) for Nat-B recipients. Mean breadth was significantly greater in the mosaic group than in the Nat-B group using overall (P = 0.014), prime-matched (P = 0.002), heterologous (P = 0.046), and boost-matched (P = 0.009) measures. Overall T cell breadth was largely due to Env-specific CD4+ T cell responses.CONCLUSIONPriming with a mosaic antigen significantly increased the number of epitopes recognized by Env-specific T cells and enabled more, albeit still limited, cross-recognition of heterologous variants. Mosaic and consensus immunogens are promising approaches to address global diversity of HIV-1.TRIAL REGISTRATIONClinicalTrials.gov NCT02296541.FUNDINGUS NIH grants UM1 AI068614, UM1 AI068635, UM1 AI068618, UM1 AI069412, UL1 RR025758, P30 AI064518, UM1 AI100645, and UM1 AI144371, and Bill & Melinda Gates Foundation grant OPP52282.
Plasmodium falciparum RH5 is the most advanced blood-stage malaria vaccine candidate and is under evaluation for efficacy in endemic regions, emphasizing the need to study the underlying antibody response to RH5 during natural infection. Here, we found that RH5-reactive B cells were rare in malaria-exposed individuals despite repeated infections over multiple years. RH5-specific monoclonal antibodies isolated from these individuals were extensively mutated but mostly targeted non-neutralizing epitopes, in contrast to antibodies from RH5-vaccinated, malaria-naive individuals. However, infection-derived MAD8-151 and MAD8-502 were among the most potent neutralizers out of 186 antibodies isolated from both cohorts and target the same epitopes as the most effective vaccine-induced antibodies. Binding to basigin receptor-proximal epitopes was the primary factor governing the potency of RH5-specific antibodies from both natural infection and vaccination, followed by the strength of binding. These results indicate a clear strategy for the development of next-generation RH5 vaccines for use in malaria-endemic regions.
The engineered outer domain germline targeting version 8 (eOD-GT8) 60-mer nanoparticle was designed to prime VRC01-class HIV-specific B cells that would need to be matured, through additional heterologous immunizations, into B cells that are able to produce broadly neutralizing antibodies. CD4 T cell help will be critical for the development of such high-affinity neutralizing antibody responses. Thus, we assessed the induction and epitope specificities of the vaccine-specific T cells from the IAVI G001 phase 1 clinical trial that tested immunization with eOD-GT8 60-mer adjuvanted with AS01 B . Robust polyfunctional CD4 T cells specific for eOD-GT8 and the lumazine synthase (LumSyn) component of eOD-GT8 60-mer were induced after two vaccinations with either the 20- or 100-microgram dose. Antigen-specific CD4 T helper responses to eOD-GT8 and LumSyn were observed in 84 and 93% of vaccine recipients, respectively. CD4 helper T cell epitope “hotspots” preferentially targeted across participants were identified within both the eOD-GT8 and LumSyn proteins. CD4 T cell responses specific to one of these three LumSyn epitope hotspots were observed in 85% of vaccine recipients. Last, we found that induction of vaccine-specific peripheral CD4 T cells correlated with expansion of eOD-GT8–specific memory B cells. Our findings demonstrate strong human CD4 T cell responses to an HIV vaccine candidate priming immunogen and identify immunodominant CD4 T cell epitopes that might improve human immune responses either to heterologous boost immunogens after this prime vaccination or to other human vaccine immunogens.
Humanity has faced three recent outbreaks of novel betacoronaviruses, emphasizing the need to develop approaches that broadly target coronaviruses. Here, we identify 55 monoclonal antibodies from COVID-19 convalescent donors that bind diverse betacoronavirus spike proteins. Most antibodies targeted an S2 epitope that included the K814 residue and were non-neutralizing. However, 11 antibodies targeting the stem helix neutralized betacoronaviruses from different lineages. Eight antibodies in this group, including the six broadest and most potent neutralizers, were encoded by IGHV1-46 and IGKV3-20. Crystal structures of three antibodies of this class at 1.5-1.75-Å resolution revealed a conserved mode of binding. COV89-22 neutralized SARS-CoV-2 variants of concern including Omicron BA.4/5 and limited disease in Syrian hamsters. Collectively, these findings identify a class of IGHV1-46/IGKV3-20 antibodies that broadly neutralize betacoronaviruses by targeting the stem helix but indicate these antibodies constitute a small fraction of the broadly reactive antibody response to betacoronaviruses after SARS-CoV-2 infection.
IMPORTANCE With timely collection of SARS-CoV-2 viral genome sequences, it is important to apply efficient data analytics to detect emerging variants at the earliest time. OBJECTIVE To evaluate the application of a statistical learning strategy (SLS) to improve early detection of novel SARS-CoV-2 variants using viral sequence data from global surveillance. DESIGN, SETTING, AND PARTICIPANTS This case series applied an SLS to viral genomic sequence data collected from 63 686 individuals in Africa and 531 827 individuals in the United States with SARS-CoV-2. Data were collected from January 1, 2020, to December 28, 2021. MAIN OUTCOMES AND MEASURES The outcome was an indicator of Omicron variant derived from viral sequences. Centering on a temporally collected outcome, the SLS used the generalized additive model to estimate locally averaged Omicron caseload percentages (OCPs) over time to characterize Omicron expansion and to estimate when OCP exceeded 10%, 25%, 50%, and 75% of the caseload. Additionally, an unsupervised learning technique was applied to visualize Omicron expansions, and temporal and spatial distributions of Omicron cases were investigated. RESULTS In total, there were 2698 cases of Omicron in Africa and 12 141 in the United States. The SLS found that Omicron was detectable in South Africa as early as December 31, 2020. With 10% OCP as a threshold, it may have been possible to declare Omicron a variant of concern as early as November 4, 2021, in South Africa. In the United States, the application of SLS suggested that the first case was detectable on November 21, 2021. CONCLUSIONS AND RELEVANCE The application of SLS demonstrates how the Omicron variant may have emerged and expanded in Africa and the United States. Earlier detection could help the global effort in disease prevention and control. To optimize early detection, efficient data analytics, such as SLS, could assist in the rapid identification of new variants as soon as they emerge, with or without lineages designated, using viral sequence data from global surveillance.
Rhesus cytomegalovirus (RhCMV)-based vaccination against Simian Immunodeficiency virus (SIV) elicits MHC-E-restricted CD8+ T cells that stringently control SIV infection in ~55% of vaccinated rhesus macaques (RM). However, it is unclear how accurately the RM model reflects HLA-E immunobiology in humans. Using long-read sequencing, we identified 16 Mamu-E isoforms and all Mamu-E splicing junctions were detected among HLA-E isoforms in humans. We also obtained the complete Mamu-E genomic sequences covering the full coding regions of 59 RM from a RhCMV/SIV vaccine study. The Mamu-E gene was duplicated in 32 (54%) of 59 RM. Among four groups of Mamu-E alleles: three ~5% divergent full-length allele groups (G1, G2, G2_LTR) and a fourth monomorphic group (G3) with a deletion encompassing the canonical Mamu-E exon 6, the presence of G2_LTR alleles was significantly (p = 0.02) associated with the lack of RhCMV/SIV vaccine protection. These genomic resources will facilitate additional MHC-E targeted translational research.
SARS-CoV-2 provokes a brisk T cell response. Peptide-based studies exclude antigen processing and presentation biology and may influence T cell detection studies. To focus on responses to whole virus and complex antigens, we used intact SARS-CoV-2 and full-length proteins with DC to activate CD8 and CD4 T cells from convalescent persons. T cell receptor (TCR) sequencing showed partial repertoire preservation after expansion. Resultant CD8 T cells recognize SARS-CoV-2-infected respiratory cells, and CD4 T cells detect inactivated whole viral antigen. Specificity scans with proteome-covering protein/peptide arrays show that CD8 T cells are oligospecific per subject and that CD4 T cell breadth is higher. Some CD4 T cell lines enriched using SARS-CoV-2 cross-recognize whole seasonal coronavirus (sCoV) antigens, with protein, peptide, and HLA restriction validation. Conversely, recognition of some epitopes is eliminated for SARS-CoV-2 variants, including spike (S) epitopes in the alpha, beta, gamma, and delta variant lineages.
Extensive mutations in the Omicron spike protein appear to accelerate the transmission of SARS-CoV-2, and rapid infections increase the odds that additional mutants will emerge. To build an investigative framework, we have applied an unsupervised machine learning approach to 4296 Omicron viral genomes collected and deposited to GISAID as of December 14, 2021, and have identified a core haplotype of 28 polymutants (A67V, T95I, G339D, R346K, S371L, S373P, S375F, K417N, N440K, G446S, S477N, T478K, E484A, Q493R, G496S, Q498R, N501Y, Y505H, T547K, D614G, H655Y, N679K, P681H, N764K, K796Y, N856K, Q954H, N69K, L981F) in the spike protein and a separate core haplotype of 17 polymutants in non-spike genes: (K38, A1892) in nsp3, T492 in nsp4, (P132, V247, T280, S284) in 3C-like proteinase, I189 in nsp6, P323 in RNA-dependent RNA polymerase, I42 in Exonuclease, T9 in envelope protein, (D3, Q19, A63) in membrane glycoprotein, and (P13, R203, G204) in nucleocapsid phosphoprotein. Using these core haplotypes as reference, we have identified four newly emerging polymutants (R346, A701, I1081, N1192) in the spike protein (p-value=9.37*10−4, 1.0*10−15, 4.76*10−7 and 1.56*10−4, respectively), and five additional polymutants in non-spike genes (D343G in nucleocapsid phosphoprotein, V1069I in nsp3, V94A in nsp4, F694Y in the RNA-dependent RNA polymerase and L106L/F of ORF3a) that exhibit significant increasing trajectories (all p-values < 1.0*10−15). In the absence of relevant clinical data for these newly emerging mutations, it is important to monitor them closely. Two emerging mutations may be of particular concern: the N1192S mutation in spike protein locates in an extremely highly conserved region of all human coronaviruses that is integral to the viral fusion process, and the F694Y mutation in the RNA polymerase may induce conformational changes that could impact Remdesivir binding.
SARS-CoV-2 is spreading worldwide with continuously evolving variants, some of which occur in the Spike protein and appear to increase viral transmissibility. However, variants that cause severe COVID-19 or lead to other breakthroughs have not been well characterized. To discover such viral variants, we assembled a cohort of 683 COVID-19 patients; 388 inpatients (“cases”) and 295 outpatients (“controls”) from April to August 2020 using electronically captured COVID test request forms and sequenced their viral genomes. To improve the analytical power, we accessed 7137 viral sequences in Washington State to filter out viral single nucleotide variants (SNVs) that did not have significant expansions over the collection period. Applying this filter led to the identification of 53 SNVs that were statistically significant, of which 13 SNVs each had 3 or more variant copies in the discovery cohort. Correlating these selected SNVs with case/control status, eight SNVs were found to significantly associate with inpatient status (q-values < 0.01). Using temporal synchrony, we identified a four SNV-haplotype (t19839-g28881-g28882-g28883) that was significantly associated with case/control status (Fisher’s exact p = 2.84 × 10 –11 ). This haplotype appeared in April 2020, peaked in June, and persisted into January 2021. The association was replicated (OR = 5.46, p -value = 4.71 × 10 −12 ) in an independent cohort of 964 COVID-19 patients (June 1, 2020 to March 31, 2021). The haplotype included a synonymous change N73N in endoRNase, and three non-synonymous changes coding residues R203K, R203S and G204R in the nucleocapsid protein. This discovery points to the potential functional role of the nucleocapsid protein in triggering “cytokine storms” and severe COVID-19 that led to hospitalization. The study further emphasizes a need for tracking and analyzing viral sequences in correlations with clinical status.