Abstract Therapeutic antibodies are challenged by rapidly evolving pathogens that exploit glycosylation to shield epitopes. SARS-CoV-2 JN.1 exemplifies this, escaping antibodies through the N354-linked glycan. However, targeting glycosylated epitopes remains vacant, as scarce and heterogeneous glycan structures render existing approaches ineffective. Here, we introduce the Antibody Evolution Nexus with Causal-Driven Simulation (AENCS), integrating molecular simulation with causal inference. Applying AENCS to restore S309 efficacy against JN.1, we identified ACC01, exhibiting ∼24-fold improved neutralization. With limited prior knowledge of the N354 glycosylation site, ACC01 stabilized this glycan conformation, facilitating the determination of its cryo-EM structure. Causal dissection revealed how this glycan shield is functionally inverted into a binding anchor through multi-layered interactions. This mechanistic conversion, combined with the conservation of N354 glycosylation, enabled ACC01 to maintain potent activity against the latest variant NB.1.8.1. Collectively, AENCS demonstrates causal-driven antibody engineering can illuminate cryptic glycosylated epitopes, providing viable paradigms for exploring this vacant frontier.
IntroductionNorovirus is a key pathogen of acute gastroenteritis and poses a significant burden on both the economy and public health. This study focuses on continuous monitoring of norovirus in Shenzhen, China, from 2016 to 2022, aiming to analyze the epidemic characteristics and genetic diversity of norovirus in the context of global sequence data.MethodsThe study was based on data collected from local sentinel hospitals. It involved analyzing the demographic, spatial, and temporal distribution of norovirus infections. Phylogenetic analysis was conducted, and genotype dynamics were compared across geographic levels. Mutations affecting protein stability were evaluated, and recombination analysis was performed to identify critical breakpoints and fragments for norovirus.ResultsThe study found that norovirus primarily infected infants under 3 years old, with epidemics occurring in winter and concentrated in developed districts. Phylogenetic analysis revealed both similarities and differences in the evolutionary patterns of various genotypes at different geographical levels. Mutations in the VP1 protein, based on the protein structure of GII.4_Sydney[P31], provided insights into the evolutionary trends of key genotypes. Additionally, recombination analysis identified important breakpoints and fragments for norovirus.DiscussionThe findings offer valuable insights to evolution and transmission of norovirus. These results can serve as a reference for future research, and they may aid in vaccine development efforts aimed at controlling norovirus outbreaks.
BACKGROUND:Rotavirus A group (RVA) is a leading cause of viral diarrhea, posing a substantial economic and public health burden. Compared to other enteric viruses, RVA possesses diverse genetic mechanisms, making it more challenging to control and prevent. Moreover, surveillance and evolutionary studies on RVA remain limited in Southern China. METHODS:We collected diarrheal stool samples from sentinel hospitals in Shenzhen and Zhuhai between 2020 and 2023. RVA-positive samples were identified via RT-PCR, followed by RNA extraction, sequencing, and genome assembly, yielding 57 RVA strains, comprising 604 sequences. Genotype trends were analyzed statistically. For phylogenetic analysis, global sequences were curated by CD-HIT, aligned with contributed sequences by MAFFT, and analyzed using IQ-TREE. Recombination and reassortment events were detected via RDP4. RESULTS:We analyzed the temporal distribution and genetic diversity of 57 newly sequenced strains from Shenzhen and Zhuhai in the context of global sequences. Our findings reveal that the prevalent genotypes of RVA in China have undergone changes over time with the decreasing of G9P[8] and the rising of G8P[8]. Phylogenetic analysis focusing on the VP7 and VP4 genes revealed distinct evolutionary patterns among different genotypes across temporal and geographical dimensions. Additionally, we discovered one reassortment event in the VP7 gene and two recombination events in the NSP1 and NSP5/6 gene. CONCLUSIONS:we observed significant variability and complexity in the evolutionary characteristics of RVA in Shenzhen and Zhuhai. These insights enhance our understanding of global evolution and transmission of RVA and provide guidance for future research and vaccine development.
BACKGROUND:Coinfections involving multiple diarrheal viruses have gained increasing recognition as a significant cause of acute gastroenteritis in recent years. Understanding the genetic diversity and evolutionary relationships of these viruses is crucial for effective outbreak identification and tracking. OBJECTIVE:To report two cases of HAdV and SaV coinfections and elucidate the genetic diversity and evolutionary patterns of these viruses through whole-genome sequencing (WGS) and phylogenetic analysis. METHODS:A total of 873 diarrheal stool samples were collected from sentinel hospitals in Shenzhen, China, in 2021. The collected stool samples were identified using RT-PCR and positive samples were subjected to WGS on the NovaSeq platform. phylogenetic trees were constructed using MEGA to analyze genetic relationships. RESULTS:The sequencing results showed that both samples were human adenovirus type 41, which clustered in two distinct evolutionary clades. Additionally, we also retrieved the complete genome of sapovirus (GI.1 genotype) from the same sample. Phylogenetic analysis revealed that they were similar to previously reported strains, belonging to the clade predominating in China. CONCLUSIONS:This study reveals the genetic diversity of epidemic strains involved in coinfections of human adenovirus and sapovirus. The findings establish a groundwork for the identification and traces of acute gastroenteritis outbreaks.
Influenza A viruses (IAVs) have historically posed significant public health threats, causing severe pandemics. Viral host specificity is typically constrained by host barriers, limiting the range of species that can be infected. However, these barriers are not absolute, and occasionally, cross-species transmission occurs, leading to human outbreaks. Early identification of changes in IAV host specificity is, therefore, critical. Despite advancements, identifying host susceptibility from genomic sequences during outbreaks remains challenging. Timely predictions are critical for effective real-time outbreak management and risk mitigation during the early stages of an epidemic. To address this, we proposed Flu-level Convolutional Neural Networks (Flu-CNN), a model designed to analyze genomic segments and identify IAV host specificity, with a particular focus on avian influenza viruses that could potentially infect humans. Extensive evaluations on large-scale genomic datasets containing 911,098 sequences show that Flu-CNN achieves an impressive 99% accuracy in determining host specificity from a single genomic segment, even for high-risk subtypes like H5N1, H7N9, and H9N2, which have a limited number of viral strains. Given its high level of accuracy, the model was applied to identify key mutations and assess the zoonotic potential of these strains. Furthermore, our study presents a pioneering approach for predicting IAV host specificity, offering novel insights into the evolutionary trajectory of these viruses. The model's significance extends beyond evolutionary analysis, playing a pivotal role in outbreak surveillance and contributing to efforts aimed at preventing the viral spread on a global scale.
IntroductionAntibiotic resistance is emerging as a critical global public health threat. The precise prediction of bacterial antibiotic resistance genes (ARGs) and phenotypes is essential to understand resistance mechanisms and guide clinical antibiotic use. Although high-throughput DNA sequencing provides a foundation for identification, current methods lack precision and often require manual intervention.MethodsWe developed a novel deep learning model for ARG prediction by integrating bacterial protein sequences using two protein language models, ProtBert-BFD and ESM-1b. The model further employs data augmentation techniques and Long Short-Term Memory (LSTM) networks to enhance feature extraction and classification performance.ResultsThe proposed model demonstrated superior performance compared to existing methods, achieving higher accuracy, precision, recall, and F1-score. It significantly reduced both false negative and false positive predictions in identifying ARGs, providing a robust computational tool for reliable gene-level resistance detection. Moreover, the model was successfully applied to predict bacterial resistance phenotypes, demonstrating its potential for clinical applicability.DiscussionThis study presents an accurate and automated approach for predicting antibiotic resistance genes and phenotypes, reducing the need for manual verification. The model offers a powerful technical tool that can support clinical decision-making and guide antibiotic use, thereby addressing an urgent need in the fight against antimicrobial resistance.
OBJECTIVE:This study aims to systematically investigate the molecular epidemiology and genomic characteristics of Enterobacter cloacae complex (ECC) strains globally harbouring blaKPC and mcr, as well as the co-existence of drug resistance genes. The goal is to provide insights and recommendations for monitoring clinical drug-resistant strains and super-resistant plasmids. METHODS:This study analysed 281 ECC isolates harbouring both blaKPC and mcr, obtained from NCBI GenBank database (2003-2024) with whole-genome sequencing data. We constructed a phylogenetic tree of ECC strains for phylogenetic analysis. Resistance genes were identified using the ABRicate and CARD databases, and their distribution was examined. Plasmid replicon types for blaKPC and mcr were analysed with PlasmidFinder, including upstream and downstream genetic environments of these genes. RESULTS:The predominant genotype combinations harbouring both blaKPC and mcr are blaKPC-3 and mcr-9.1, followed by blaKPC-4 and mcr-9.1, with the dominant strain being E. hormaechei ssp. xiangfangensis. The IncHI2(2A) plasmid co-harbouring blaKPC and mcr-9 was mainly detected in genomes from the United States. Phylogenetic analysis indicated that the blaKPC-mcr-9-IncHI2(2A) plasmids in ECC strains have a high genetic similarity to mcr-9-IncHI2(2A), suggesting that the latter may have acquired the additional highly transferable blaKPC. CONCLUSIONS:ECC strains have become an important reservoir cluster for blaKPC and mcr-9, and the IncHI2(2A) plasmid is a potential vector for the horizontal co-transmission of carbapenem and colistin resistance genes. Effective monitoring should be implemented to assess the prevalence of co-harbouring blaKPC and mcr-9 in individual ECC isolates and even in single plasmid.
Diarrhea is one of the major public health issues worldwide. Although the infections of individual enteric virus have been extensively studied, elucidation of the coinfection involving multiple viruses is still limited. In this study, we identified the coinfection of human adenovirus (HAdV) and human astrovirus (HAstV) in a child with acute gastroenteritis, analyzed their genotypes and molecular evolution characteristics. The sample was collected and identified using RT-PCR and subjected to whole-genome sequencing on the NovaSeq (Illumina) platform. Obtained sequences were assembled into the complete genome of HAdV and the ORF1 of HAstV. We conducted phylogenetic analysis using IQ-TREE software and conducted recombination analysis with the Recombination Detection Program. The sequenced HAdV was confirmed to be genotype 41, and was genetically close to some European strains. Phylogenetic analysis revealed that the HAstV was genetically close to both HAstV-2 and HAstV-4 and was different from the genotype prevalent in Shenzhen before. The recombination analysis confirmed that the sequenced HAstV strain is a recombinant of HAstV-2 and HAstV-4. Our analysis has shown that the strains in this coinfection are both uncommon variants in this geographical region, instead of dominant subtypes that have prevailed for years. This study presents a coinfection of HAdV and HAstV and conducts an evolutionary analysis on involved viruses, which reveals the genetic diversity of epidemic strains in Southern China and offers valuable insights into vaccine and medical research.
Modeling and predicting mutations are critical for COVID-19 and similar pandemic preparedness. However, existing predictive models have yet to integrate the regularity and randomness of viral mutations with minimal data requirements. Here, we develop a non-demanding language model utilizing both regularity and randomness to predict candidate SARS-CoV-2 variants and mutations that might prevail. We constructed the “grammatical frameworks” of the available S1 sequences for dimension reduction and semantic representation to grasp the model’s latent regularity. The mutational profile, defined as the frequency of mutations, was introduced into the model to incorporate randomness. With this model, we successfully identified and validated several variants with significantly enhanced viral infectivity and immune evasion by wet-lab experiments. By inputting the sequence data from three different time points, we detected circulating strains or vital mutations for XBB.1.16, EG.5, JN.1, and BA.2.86 strains before their emergence. In addition, our results also predicted the previously unknown variants that may cause future epidemics. With both the data validation and experiment evidence, our study represents a fast-responding, concise, and promising language model, potentially generalizable to other viral pathogens, to forecast viral evolution and detect crucial hot mutation spots, thus warning the emerging variants that might raise public health concern.
Continuous gain and loss of genes are the primary driving forces of bacterial evolution and environmental adaptation. Studying bacterial evolution in terms of protein domain, which is the fundamental function and evolutionary unit of proteins, can provide a more comprehensive understanding of bacterial differentiation and phenotypic adaptation processes. Therefore, we proposed a phylogenetic tree-based method for detecting genetic gain and loss events in terms of protein domains. Specifically, the method focuses on a single domain to trace its evolution process or on multiple domains to investigate their co-evolution principles. This novel method was validated using 122 Shigella isolates. We found that the loss of a significant number of domains was likely the main driving force behind the evolution of Shigella, which could reduce energy expenditure and preserve only the most essential functions. Additionally, we observed that simultaneously gained and lost domains were often functionally related, which can facilitate and accelerate phenotypic evolutionary adaptation to the environment. All results obtained using our method agree with those of previous studies, which validates our proposed method.
Background Lassa fever is a hemorrhagic disease caused by Lassa virus (LASV), which has been classified by the World Health Organization as one of the top infectious diseases requiring prioritized research. Previous studies have provided insights into the classification and geographic characteristics of LASV lineages. However, the factor of the distribution and evolution characteristics and phylodynamics of the virus was still limited. Methods To enhance comprehensive understanding of LASV, we employed phylogenetic analysis, reassortment and recombination detection, and variation evaluation utilizing publicly available viral genome sequences. Results The results showed the estimated the root of time of the most recent common ancestor (TMRCA) for large (L) segment was approximately 634 (95% HPD: [385879]), whereas the TMRCA for small (S) segment was around 1224 (95% HPD: [10301401]). LASV primarily spread from east to west in West Africa through two routes, and in route 2, the virus independently spread to surrounding countries through Liberia, resulting in a wider spread of LASV. From 1969 to 2018, the effective population size experienced two significant increased, indicating the enhanced genetic diversity of LASV. We also found the evolution rate of L segment was faster than S segment, further results showed zinc-binding protein had the fastest evolution rate. Reassortment events were detected in multiple lineages including sub-lineage IIg, while recombination events were observed within lineage V. Significant amino acid changes in the glycoprotein precursor of LASV were identified, demonstrating sequence diversity among lineages in LASV. Conclusion This study comprehensively elucidated the transmission and evolution of LASV in West Africa, providing detailed insights into reassortment events, recombination events, and amino acid variations.
Driven by various mutations on the viral Spike protein, diverse variants of SARS-CoV-2 have emerged and prevailed repeatedly, significantly prolonging the pandemic. This phenomenon necessitates the identification of key Spike mutations for fitness enhancement. To address the need, this manuscript formulates a well-defined framework of causal inference methods for evaluating and identifying key Spike mutations to the viral fitness of SARS-CoV-2. In the context of large-scale genomes of SARS-CoV-2, it estimates the statistical contribution of mutations to viral fitness across lineages and therefore identifies important mutations. Further, identified key mutations are validated by computational methods to possess functional effects, including Spike stability, receptor-binding affinity, and potential for immune escape. Based on the effect score of each mutation, individual key fitness-enhancing mutations such as D614G and T478K are identified and studied. From individual mutations to protein domains, this paper recognizes key protein regions on the Spike protein, including the receptor-binding domain and the N-terminal domain. This research even makes further efforts to investigate viral fitness via mutational effect scores, allowing us to compute the fitness score of different SARS-CoV-2 strains and predict their transmission capacity based solely on their viral sequence. This prediction of viral fitness has been validated using BA.2.12.1, which is not used for regression training but well fits the prediction. To the best of our knowledge, this is the first research to apply causal inference models to mutational analysis on large-scale genomes of SARS-CoV-2. Our findings produce innovative and systematic insights into SARS-CoV-2 and promotes functional studies of its key mutations, serving as reliable guidance about mutations of interest.
Throughout history, Influenza A viruses (IAVs) have caused significant harm and catastrophic pandemics. The presence of host barriers results in viral host tropism, where infected hosts are subject to strict restrictions due to the hindered spread of viruses across hosts. Therefore, the identification of host tropism of IAVs, particularly in humans, is crucial to preventing the cross-host transmission of avian viruses and their outbreaks in humans. Nevertheless, efficiently and effectively identifying host tropism, especially for early host susceptibility warnings based on viral genome sequences during outbreak onset, remains challenging. To address this challenge, we propose Flu-CNN, a deep neural network model based on classical character-level convolutional networks. By analyzing the genomic segments of IAVs, Flu-CNN can accurately identify the host tropism, with a particular focus on avian influenza viruses that may infect humans. According to our experimental evaluations, Flu-CNN achieved an accuracy of 99% in identifying virus hosts via only a single genomic segment, even for subtypes with a relatively small number of viral strains such as H5N1, H7N9, and H9N2. The superiority of Flu-CNN demonstrates its effectiveness in screening for critical amino acid mutations, which is important to host adaptation, and zoonotic risk prediction of viral strains. Flu-CNN is a valuable tool for identifying evolutionary characterization, monitoring potential outbreaks, and preventing epidemical spreads of IAVs, which contribute to the effective surveillance of influenza A viruses.
The outbreak and spread of COVID-19 remind us again of the devastating attack that human-to-human transmitted respiratory infectious diseases (H-HRIDs) bring to global economics and public health. Predicting the trend of H-HRIDs yields important suggestions for making response strategies. Existing methods predict the future trend of H-HRIDs based on the prophase epidemiological data. However, it is also crucial to predict the potential outbreak trend of H-HRIDs in regions without prophase epidemiological data based on those with epidemiological data available. This prediction can provide constructive pharmaceutical strategies and non-pharmaceutical interventions such as pre-allocating the limited preventions to regions with different potential trends, especially at the early stage of an H-HRID. In this paper, transform the problem of predicting the early-stage trend of an H-HRID in new regions without prophase epidemiological data into the problem of learning regions’ geospatial features via a Contrastive learning-based Hierarchical Graph Convolutional Neural Network (CHGCN). Specifically, CHGCN first generates training data from regions with epidemiological data available via a Relaxed Graph Edit Distance (RGED) algorithm. CHGCN then uses a constructive learning schema to learn region embeddings, where regions with similar geospatial features and H-HRID trends exhibit similar embeddings. Once trained, CHGCN searches for the region with epidemiological data available that has the highest embedding similarity to the given query region without prophase epidemiological data, so as to achieve early-stage H-HRID trend prediction for the query region. Experimental results on COVID-19 demonstrate that the proposed CHGCN achieves state-of-the-art performance in predicting the early-stage trend of H-HRIDs, compared with the baselines.
Please note: The publisher is not responsible for the content or functionality of any supporting information supplied by the authors. Any queries (other than missing content) should be directed to the corresponding author for the article.
Human astrovirus (HAstV) is a single-stranded, positive-sense RNA virus and is the leading cause of viral gastroenteritis. However, despite its prevalence, astroviruses still remain one of the least studied enteroviruses. In this study, we sequenced 11 classical astrovirus strains from clinical samples collected in Shenzhen, China from 2016 to 2019, analyzed their genetic characteristics, and deposited them into GenBank. We conducted phylogenetic analysis using IQ-TREE software, with references to astrovirus sequences worldwide. The phylogeographic analysis was performed using the Bayesian Evolutionary Analysis Sampling Trees program, through Bayesian Markov Chain Monte Carlo sampling. We also conducted recombination analysis with the Recombination Detection Program. The newly sequenced strains were categorized as HAstV genotype 1, which is the predominant genotype in Shenzhen. Phylogeographic reconstruction indicated that HAstV-1 may have migrated from the United States to China, followed by frequent transmission between China and Japan. The recombination analysis revealed recombination events within and across genotypes, and identified a recombination-prone region that produced relatively uniform recombination breakpoints and fragment lengths. The genetic analysis of HAstV strains in Shenzhen addresses the current lack of astrovirus data in the region of Shenzhen and provides key insights to the evolution and transmission of astroviruses worldwide. These findings highlight the importance of improving surveillance of astroviruses.
目的 探究甲型流感病毒宿主分布规律,量化评估其多宿主传播趋势.方法 从流感病毒公共数据库中下载流感病毒全基因组数据,对编码序列进行特征提取,采用自然聚类结果对流感病毒进行宿主分布与进化态势分析,建立宿主谱熵值计算方法量化评估流感病毒多宿主传播扩散趋势,并使用机器学习方法对分析结果进行验证.结果 H3N2型流感病毒多宿主分布界限清晰,具有最低宿主谱熵值;H9N2型流感病毒多宿主分布混乱,不同宿主的毒株序列相似程度高,具有最高宿主谱熵值.不同类型甲型流感病毒的宿主谱熵值大小与多宿主传播扩散趋势紧密相关.结论 受传播环境与选择压力的影响,不同流感病毒的进化方向与多宿主传播趋势呈现出较为明显的差异,较为适应人类宿主的H3N2流感病毒的多宿主传播趋势更有序,而频繁发生跨宿主传播的H7N9、H9N2等禽流感病毒的多宿主传播趋势更混乱,进化方向也更多样.宿主谱熵值是衡量流感病毒多宿主传播趋势与进化方向有序性的高效计算方法,利用该方法有助于了解流感病毒的进化与多宿主分布,评估传播风险,可为流感病毒的监测和预警提供新的见解.
Klebsiella pneumoniae is a common human commensal and opportunistic pathogen. In recent years, the clinical isolation and resistance rates of K. pneumoniae have shown a yearly increase, leading to a special interest in mobile genetic elements. Prophages are a representative class of mobile genetic elements that can carry host-friendly genes, transfer horizontally between strains, and coevolve with the host's genome. In this study, we identified 15,946 prophages from the genomes of 1437 fully assembled K. pneumoniae deposited in the NCBI database, with 9755 prophages on chromosomes and 6191 prophages on plasmids. We found prophages to be notably diverse and widely disseminated in the K. pneumoniae genomes. The K. pneumoniae prophages encoded multiple putative virulence factors and antibiotic resistance genes. The comparison of strain types with prophage types suggests that the two may be related. The differences in GC content between the same type of prophages and the genomic region in which they were located indicates the alien properties of the prophages. The overall distribution of GC content suggests that prophages integrated on chromosomes and plasmids may have different evolutionary characteristics. These results suggest a high prevalence of prophages in the K. pneumoniae genome and highlight the effect of prophages on strain characterization.
目的 分析新型冠状病毒(新冠病毒)各个编码序列可能的重组来源,为理解新冠病毒的重组进化规律提供新的见解.方法 选取公共数据库中与新冠病毒序列相关的典型β属冠状病毒全基因组序列,采用特征提取和自组织映射神经网络的方法对冠状病毒各个编码序列进行聚类分析,进而实现对新冠病毒进行重组来源分析.结果 Wuhan?Hu?1的ORF1a基因和ORF1b基因序列与Yunnan RmYN02序列特征相似度较高,M和N编码序列与Yunnan RaTG13序列特征更为接近,而S基因序列则与Yunnan RaTG13和穿山甲冠状病毒中相关序列特征相似.结论 重组是冠状病毒进化的重要动力,对冠状病毒重组进化规律的分析有助于了解新冠病毒的起源.从序列特征的角度分析了新冠病毒5个长编码序列可能的重组起源,但仍存在一些短编码序列的重组起源并不清楚,亟需更多关于新冠病毒溯源的采样研究.