Fully understanding traditional Chinese medicines (TCMs) is still challenging because of the extreme complexity of their chemical components and mechanisms of action. The TCM Plant Genome Project aimed to obtain genetic information, determine gene functions, discover regulatory networks of herbal species, and elucidate the molecular mechanisms involved in the disease prevention and treatment, thereby accelerating the modernization of TCMs. A comprehensive database that contains TCM‐related information will provide a vital resource. Here, we present an integrative genome database of TCM plants (IGTCM) that contains 14,711,220 records of 83 annotated TCM‐related herb genomes, including 3,610,350 genes, 3,534,314 proteins and corresponding coding sequences, and 4,032,242 RNAs, as well as 1033 non‐redundant component records for 68 herbs, downloaded and integrated from the GenBank and RefSeq databases. For minimal interconnectivity, each gene, protein, and component was annotated using the eggNOG‐mapper tool and Kyoto Encyclopedia of Genes and Genomes database to acquire pathway information and enzyme classifications. These features can be linked across several species and different components. The IGTCM database also provides visualization and sequence similarity search tools for data analyses. These annotated herb genome sequences in IGTCM database are a necessary resource for systematically exploring genes related to the biosynthesis of compounds that have significant medicinal activities and excellent agronomic traits that can be used to improve TCM‐related varieties through molecular breeding. It also provides valuable data and tools for future research on drug discovery and the protection and rational use of TCM plant resources. The IGTCM database is freely available at http://yeyn.group:96/.
Essential ncRNA is a type of ncRNA which is indispensable for the sur-vival of organisms. Although essential ncRNAs cannot encode proteins, they are as important as essential coding genes in biology. They have got wide variety of applications such as antimicrobial target discovery, minimal genome construction and evolution analysis. At present, the number of species required for the deter-mination of essential ncRNAs in the whole genome scale is still very few due to the traditional methods are time-consuming, laborious and costly. In addition, tra-ditional experimental methods are limited by the organisms as less than 1% of bacteria can be cultured in the laboratory. Therefore, it is important and necessary to develop theories and methods for the recognition of essential non-coding RNA. In this paper, we present a novel method for predicting essential ncRNA by using both compositional and derivative features calculated by information theory of ncRNA sequences. The method was developed with Support Vector Machine (SVM). The accuracy of the method was evaluated through cross-species cross -vali-dation and found to be between 0.69 and 0.81. It shows that the features we selected have good performance for the prediction of essential ncRNA using SVM. Thus, the method can be applied for discovering essential ncRNAs in bacteria.
This study systematically analyses the mechanism of Spike(S) protein and bioinformatics research progress by mining domestic and foreign literatures. This paper summarizes the current research on use of bioinformatics tools by scientists at home and abroad to develop drugs and vaccines research of Spike protein. In addition, the existing Spike protein data resources are systematically organized. Related bioinformatics tools and databases are summarized, and the research prospects of Spike protein in this field are proposed. This review aims to provide systematic Spike protein research references for biologists,virologists and other researchers who study Spike proteins. It promotes the research of Spike protein mechanism and the development of effective drugs and vaccines.
细菌非编码RNA(non-coding RNA, ncRNA)是近年来在细菌基因组内新发现的一类基因表达调控因子,与必需基因概念类似,有一部分ncRNA是生物体生存所必不可少的,称之为"必需非编码RNA".因此,细菌的必需ncRNA可以作为药物开发的潜在靶标,以降低致病菌的耐药性.同时,必需ncRNA也成为最小基因组研究的重要对象之一.目前已经通过湿实验系统地确定了10余种细菌的必需ncRNA,然而还没有一个专门的必需ncRNA数据库,导致对必需ncRNA的研究远远跟不上科学研究和药物设计的需要.因此,该研究构建了一个专门的细菌必需ncRNA数据库DBEncRNA,以帮助研究人员开发高效的必需ncRNA计算机识别方法,用于进一步研究抗菌药物靶标发现和最小基因组.DBEncRNA数据库可以通过http://yeyn.group:86/免费访问使用.
Abstract Spike protein is a class I virus fusion glycoprotein that plays an important role in cell infection by mediating receptor binding and membrane fusion during virus invasion. The fusion mechanism between virus and host cell mediated by spike protein is complex, and alterations via mutation can affect viral pathogenicity and have been linked with virus pandemics. In-depth studies of the biology and toxicology of viruses that pose a potential public health threat have promoted the use of bioinformatic approaches to explore the structure, function, and evolution of spike proteins. Therefore, data related to spike proteins have increased exponentially. To facilitate further biomedical research, these data require integration. Here, we developed a database of virus spike proteins (DVSP), which provides a free data and bioinformatics service for the scientific community. The DVSP contains 35,579 spike protein and 35,175 nucleotide sequences collected from viruses. Among these, 16,027 spike proteins and 15,623 pre-translational nucleotide sequences are derived from severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), the causative agent of the ongoing corona virus disease 2019 (COVID-19) pandemic. Overall, the DVSP provides 19,552 virus classifications based on integration and virus taxonomy, annotation information from the GenBank database, and sequence information. The sequences, virus taxonomy, and host information contained in the DVSP provide an important resource for future pathogenicity, evolutionary, and drug development studies. It serves as a valuable and comprehensive knowledge source for exploring the distribution, structure, and phylogeny of spike proteins in viruses, thereby promoting the formulation of new hypotheses and experimental designs for spike protein studies. The DVSP is a centralized reference and resource for biological information about the spike proteins of viruses (including coronavirus), which is available at http://yeyn.group:90/.
Studies have confirmed that the growth, development and maintenance structure and function of all living organisms are determined by their genes, but the speed of gene translation products is affected by various internal and external factors. Therefore, studying the gene translation speed is extremely important for production and improving the speed through a variety of ways has always been an important subject. In recent years, the found of codon usage bias, tRNA recycling model and SD sequence-like structure provide another way to optimize the translation efficiency. These methods are almost based on codon usage theory to optimize gene sequence to speed up or slow down the translation speed so as to regulate the cell expression and metabolism. This paper is combined the three theories above, aimed at E.coli to design some reasonable algorithm to optimize gene sequence and thus using codon adaptation index (CAI) to evaluate the optimized effect.
The shape of helicobacter pylori (HP) is usually helical or S-shaped, which makes it to have the feature of highly divergent. Besides, its distribution also is different greatly in space and time, which is the factor to lead to multiple infection. In addition, in the same infected person, multiple infection also appeared in different regions of strain gene infection origin. Therefore, it is of great significance to study the genomic diversity and evolutionary analysis of Helicobacter pylori. In this paper, we are to analyze the homology of multiple sequences using ClustalX1.83 and meanwhile to construct the homologous genome evolutionary tree by MEGA7.0 to analyze the its function. The purpose of this study is to compare the genomic evolutionary differences of Helicobacter pylori among populations in different regions, and to discuss the significance of its infection prevalence and genomic variation in geographical changes, so as to conduct further discussion on the molecular level research mechanism of Helicobacter pylori.
Essential genes are indispensable for biological survival. Thus it is of great significance to identify and study essential genes. A machine learning method, K-Nearest Neighbor, is used for development of predicting essential bacterial genes. The homologous features, including sequence homology and functional homology, of the bacterial genomes are extracted for determining essential genes. Based on the features, we use K-Nearest Neighbor algorithm for determining of gene function. And we tune the minimum matching parameter (K) in the essential gene predicted model for building an optimal model of the Escherichia coli specificity model. The corresponding optimal parameter (K) is then extended to other bacterial essential genes predicting models. After cross validation, the highest accuracy is 0.89 while K between 5 and 7. Therefore, the features we extracted can increase the accuracy of the bacterial essential gene prediction. In the premise, we found that the prediction accuracy of the prediction model based on K-Nearest Neighbor was not significantly different in different evolutionary distances between organisms in the database and the investigated species. That means the machine learning model can be extended to more distant species. It wills have a better predictive performance for predicting essential genes of distant species than the usual sequence-based methods.
Spike蛋白是I类融合糖蛋白,主要发挥糖蛋白的作用。尤其广泛存在于病毒中,在病毒入侵细胞时介导受体结合和膜融合,从而影响细胞的感染倾向。随着新型冠状病毒肺炎(Corona Virus Disease 2019, COVID-19)疫情爆发,生物学和毒理学家们利用生物信息学手段对Spike蛋白进行测序、注释等方法探索该蛋白的结构、功能和进化关系,发现S蛋白是病毒有效治疗、疫苗研发以及临床诊断的重要靶点。本研究通过挖掘国内外文献,系统地分析了S蛋白的作用机制和生物信息学方面的研究进展。总结分析了现有国内外与Spike蛋白相关的生物信息学工具及数据库,最后对病毒Spike蛋白在该领域的研究提出了展望。