Abstract Background Accurately identifying essential genes in bacteria is critical for understanding microbial biology and developing novel antibiotics. However, the heterogeneity of biological data poses a challenge for reliable prediction. This study aims to enhance prediction accuracy by integrating diverse biological features through a multi-feature fusion framework. Results This study combined sequence data, gene annotations, protein–protein interaction networks, and subcellular localization information to construct a convolutional neural network (CNN)-based model, CNN4Essential. Feature importance was assessed using a random forest algorithm, and dimensionality reduction was performed with truncated singular value decomposition. The model was evaluated through intra-species prediction (INSP) and leave-a-species-out prediction (LASP) across 22 prokaryotic species. CNN4Essential achieved an average AUC of 0.884 in INSP, 0.726 in LASP, and 0.851 across all species, outperforming existing methods. Furthermore, predictions for Haemophilus influenzae were compared with known drug-target genes from DrugBank. A positive correlation between prediction scores and target gene matching rates was observed. Conclusions The integration of multi-source features with a deep learning model significantly improves bacterial essential gene prediction. CNN4Essential not only surpasses single-feature and shallow models in performance but also holds promise for identifying potential drug targets.
Urinary calculi is a common urological disease. This study aimed to build a prediction model for the risk factors of urinary calculi based on the nomogram in machine learning, providing a basis for the prevention and treatment of this disorder. The clinical data of 205 non-urinary calculi volunteers and 105 urinary calculi patients were finally included. The differences in basic clinical data between the control group (non-suffering) and the observation group (affected) were analyzed. Univariate and multivariate Logistic regression analyses were used to determine the risk and predictive factors of urinary uric acid calculi, and a prediction model for the risk factors of uric acid calculi was thus formulated. The calibration curve and ROC curve were used to evaluate the calibration and performance of the prediction model. The results showed that age, red blood cells, white blood cells, urate crystals, protein, ketone bodies, and platelets are the risk factors for urinary calculi. The area under the ROC curve of the prediction model is 1, indicating an ideal model. The calibration curve suggested that the predicted values of the independent predictive factors for urinary calculi are approximately consistent with the actual values. In conclusion, the prediction model constructed by age, red blood cells, white blood cells, urate crystals, protein, ketone bodies, and platelets has a high predictive value, which is beneficial for improving the early diagnosis ability of urinary calculi and achieving individualized prevention. Future studies can further explore the application of this model in clinical practice.
Accurately reconstructing Gene Regulatory Networks (GRNs) is crucial for understanding gene functions and disease mechanisms. Single-cell RNA sequencing (scRNA-seq) technology provides vast data for computational GRN reconstruction. Since GRNs are ideally modeled as signed directed graphs to capture activation/inhibition relationships, the most intuitive and reasonable approach is to design feature extractors based on the topological structure of GRNs to extract structural features, then combine them with biological characteristics for research. However, traditional spectral graph convolution struggles with this representation. Thus, we propose MSGRNLink, a novel framework that explicitly models GRNs as signed directed graphs and employs magnetic signed Laplacian convolution. Experiments across simulated and real datasets demonstrate that MSGRNLink outperforms all baseline models in AUROC. Parameter sensitivity analysis and ablation studies confirmed its robustness and the importance of each module. In a bladder cancer case study, MSGRNLink predicted more known edges and edge signs than benchmark models, further validating its biological relevance.
With the rapid advancements in sequencing technology, an increasing number of genome sequences for medicinal plants have become available. In 2021, TCMPG was introduced as a comprehensive database dedicated to traditional Chinese medicine (TCM) plant genomes, capturing significant attention within the scientific community. To offer invaluable resources to researchers, this database received an upgrade to its latest version, named TCMPG 2.0, accessible at http://cbcb.cdutcm.edu.cn/TCMPG2/. TCMPG 2.0 surpasses its predecessor by adding 114 medicinal plants, 129 genomes, and 159 related herbs. Recognizing the critical role of ingredients in assessing the pharmacological effects of Chinese medicines, TCMPG 2.0 has also amassed data on 13,868 herbal ingredients. Additionally, four practical analytical tools, including Heatmap, Primer3, PlantiSMASH, and CRISPRCasFinder, have been integrated into the toolbox. These valuable enhancements significantly enhance the database's utility for researchers in the field.
Fully understanding traditional Chinese medicines (TCMs) is still challenging because of the extreme complexity of their chemical components and mechanisms of action. The TCM Plant Genome Project aimed to obtain genetic information, determine gene functions, discover regulatory networks of herbal species, and elucidate the molecular mechanisms involved in the disease prevention and treatment, thereby accelerating the modernization of TCMs. A comprehensive database that contains TCM‐related information will provide a vital resource. Here, we present an integrative genome database of TCM plants (IGTCM) that contains 14,711,220 records of 83 annotated TCM‐related herb genomes, including 3,610,350 genes, 3,534,314 proteins and corresponding coding sequences, and 4,032,242 RNAs, as well as 1033 non‐redundant component records for 68 herbs, downloaded and integrated from the GenBank and RefSeq databases. For minimal interconnectivity, each gene, protein, and component was annotated using the eggNOG‐mapper tool and Kyoto Encyclopedia of Genes and Genomes database to acquire pathway information and enzyme classifications. These features can be linked across several species and different components. The IGTCM database also provides visualization and sequence similarity search tools for data analyses. These annotated herb genome sequences in IGTCM database are a necessary resource for systematically exploring genes related to the biosynthesis of compounds that have significant medicinal activities and excellent agronomic traits that can be used to improve TCM‐related varieties through molecular breeding. It also provides valuable data and tools for future research on drug discovery and the protection and rational use of TCM plant resources. The IGTCM database is freely available at http://yeyn.group:96/.
Essential ncRNA is a type of ncRNA which is indispensable for the sur-vival of organisms. Although essential ncRNAs cannot encode proteins, they are as important as essential coding genes in biology. They have got wide variety of applications such as antimicrobial target discovery, minimal genome construction and evolution analysis. At present, the number of species required for the deter-mination of essential ncRNAs in the whole genome scale is still very few due to the traditional methods are time-consuming, laborious and costly. In addition, tra-ditional experimental methods are limited by the organisms as less than 1% of bacteria can be cultured in the laboratory. Therefore, it is important and necessary to develop theories and methods for the recognition of essential non-coding RNA. In this paper, we present a novel method for predicting essential ncRNA by using both compositional and derivative features calculated by information theory of ncRNA sequences. The method was developed with Support Vector Machine (SVM). The accuracy of the method was evaluated through cross-species cross -vali-dation and found to be between 0.69 and 0.81. It shows that the features we selected have good performance for the prediction of essential ncRNA using SVM. Thus, the method can be applied for discovering essential ncRNAs in bacteria.
Dendritic cells (DCs) can mediate immune responses or immune tolerance depending on their immunophenotype and functional status. Remodeling of DCs’ immune functions can develop proper therapeutic regimens for different immune-mediated diseases. In the immunopathology of autoimmune diseases (ADs), activated DCs notably promote effector T-cell polarization and exacerbate the disease. Recent evidence indicates that metformin can attenuate the clinical symptoms of ADs due to its anti-inflammatory properties. Whether and how the therapeutic effects of metformin on ADs are associated with DCs remain unknown. In this study, metformin was added to a culture system of LPS-induced DC maturation. The results revealed that metformin shifted DC into a tolerant phenotype, resulting in reduced surface expression of MHC-II, costimulatory molecules and CCR7, decreased levels of proinflammatory cytokines (TNF-α and IFN-γ), increased level of IL-10, upregulated immunomodulatory molecules (ICOSL and PD-L) and an enhanced capacity to promote regulatory T-cell (T reg ) differentiation. Further results demonstrated that the anti-inflammatory effects of metformin in vivo were closely related to remodeling the immunophenotype of DCs. Mechanistically, metformin could mediate the metabolic reprogramming of DCs through FoxO3a signaling pathways, including disturbing the balance of fatty acid synthesis (FAS) and fatty acid oxidation (FAO), increasing glycolysis but inhibiting the tricarboxylic acid cycle (TAC) and pentose phosphate pathway (PPP), which resulted in the accumulation of fatty acids (FAs) and lactic acid, as well as low anabolism in DCs. Our findings indicated that metformin could induce tolerance in DCs by reprogramming their metabolic patterns and play anti-inflammatory roles in vitro and in vivo.
Recently,Worobey et al.(2022)pub-lished a report entitled 'The Huanan Seafood Wholesale Market in Wuhan was the early epicenter of the COVID-19pan-demic'that succinctly summarizes their study[1].A pre-print version of this study had earlier elicited a series of high-profile media coverages[2,3].All these reports deliver a social-political message that the Huanan market is the epicenter of COVID-19.
Transcription factors (TFs) play a part in gene expression. TFs can form complex gene expression regulation system by combining with DNA. Thereby, identifying the binding regions has become an indispensable step for understanding the regulatory mechanism of gene expression. Due to the great achievements of applying deep learning (DL) to computer vision and language processing in recent years, many scholars are inspired to use these methods to predict TF binding sites (TFBSs), achieving extraordinary results. However, these methods mainly focus on whether DNA sequences include TFBSs. In this paper, we propose a fully convolutional network (FCN) coupled with refinement residual block (RRB) and global average pooling layer (GAPL), namely FCNARRB. Our model could classify binding sequences at nucleotide level by outputting dense label for input data. Experimental results on human ChIP-seq datasets show that the RRB and GAPL structures are very useful for improving model performance. Adding GAPL improves the performance by 9.32% and 7.61% in terms of IoU (Intersection of Union) and PRAUC (Area Under Curve of Precision and Recall), and adding RRB improves the performance by 7.40% and 4.64%, respectively. In addition, we find that conservation information can help locate TFBSs.
The performance of implanted biomaterials is largely determined by their interaction with the host immune system. As a fibrous-like 3D network, fibrin matrix formed at the interfaces of tissue and material, whose effects on dendritic cells (DCs) remain unknown. Here, a bone plates implantation model was developed to evaluate the fibrin matrix deposition and DCs recruitment in vivo. The DCs responses to fibrin matrix were further analyzed by a 2D and 3D fibrin matrix model in vitro. In vivo results indicated that large amount of fibrin matrix deposited on the interface between the tissue and bone plates, where DCs were recruited. Subsequent in vitro testing denoted that DCs underwent significant shape deformation and cytoskeleton reorganization, as well as mechanical property alteration. Furthermore, the immune function of imDCs and mDCs were negatively and positively regulated, respectively. The underlying mechano-immunology coupling mechanisms involved RhoA and CDC42 signaling pathways. These results suggested that fibrin plays a key role in regulating DCs immunological behaviors, providing a valuable immunomodulatory strategy for tissue healing, regeneration and implantation.
This study systematically analyses the mechanism of Spike(S) protein and bioinformatics research progress by mining domestic and foreign literatures. This paper summarizes the current research on use of bioinformatics tools by scientists at home and abroad to develop drugs and vaccines research of Spike protein. In addition, the existing Spike protein data resources are systematically organized. Related bioinformatics tools and databases are summarized, and the research prospects of Spike protein in this field are proposed. This review aims to provide systematic Spike protein research references for biologists,virologists and other researchers who study Spike proteins. It promotes the research of Spike protein mechanism and the development of effective drugs and vaccines.
目的 探讨基于双线性卷积神经网络(BRNV)模型的阿尔茨海默病(AD)自动诊断.方法 选取AD神经成像倡议(ADNI)数据库中的AD(n=93)、轻度认知功能障碍(MCI,n=76)及正常认知(NC,n=100)受试者的核磁共振图像(MRI)作为数据集,预处理后按照8:2的比例分为训练集和验证集,同时另取ADNI中不同于以上已经选取的受试者数据150例作为测试集(AD、MCI及NC受试者各50例),将每名受试者经过预处理后的三维MRI数据转换为矢状面、冠状面及横断面2 D切片123张,最终获得AD组MCI组及NC组受试者训练集(n=9225、7503、9840)、验证集(n=2214、1845、2460)及测试集(n=6150、6150、6150)2 D切片;设计BRNV模型对预处理后的MRI数据进行分类预测,采用迁移学习方法为模型寻找最优的初始网络参数权重,通过BRNV模型对AD的分类效果绘制受试者工作特征曲线(ROC),以ROC曲线下面积(AUC)、准确率、特异性及敏感性评估模型的诊断价值.结果 BRNV模型对AD受试者诊断分类,准确率、AUC、特异性及敏感性分别达85.4%、91.3%、86.1%及84.2%;BRNV模型对MCI受试者诊断分类,准确率、AUC、特异性及敏感性分别达73.3%、75.1%、72.0%及74.0%.结论 BRNV模型在AD自动诊断中具有较高的准确率,有助于针对AD的计算机辅助诊断系统开发,并帮助医生提高AD的诊断效率.
细菌非编码RNA(non-coding RNA, ncRNA)是近年来在细菌基因组内新发现的一类基因表达调控因子,与必需基因概念类似,有一部分ncRNA是生物体生存所必不可少的,称之为"必需非编码RNA".因此,细菌的必需ncRNA可以作为药物开发的潜在靶标,以降低致病菌的耐药性.同时,必需ncRNA也成为最小基因组研究的重要对象之一.目前已经通过湿实验系统地确定了10余种细菌的必需ncRNA,然而还没有一个专门的必需ncRNA数据库,导致对必需ncRNA的研究远远跟不上科学研究和药物设计的需要.因此,该研究构建了一个专门的细菌必需ncRNA数据库DBEncRNA,以帮助研究人员开发高效的必需ncRNA计算机识别方法,用于进一步研究抗菌药物靶标发现和最小基因组.DBEncRNA数据库可以通过http://yeyn.group:86/免费访问使用.
PurposeAnaplastic thyroid carcinoma (ATC) and primary squamous cell carcinoma of the thyroid (PSCCTh) have similar histological findings and are currently treated using the same approaches; however, the characteristics and prognosis of these cancers are poorly researched. The objective of this study was to determine the differences in characteristics between ATC and PSCCTh and establish prognostic models.Patients and MethodsAll variables of patients with ATC and PSCCTh, diagnosed from 2004–2015, were retrieved from the Surveillance, Epidemiology, and End Results Program (SEER) database. Percentage differences for categorical data were compared using the Chi-square test. Kaplan-Meier curves, log-rank test, and Cox-regression for survival analysis, and C-index value was used to evaluate the performance of the prognostic models.ResultsAfter application of the inclusion and exclusion criteria, a total of 1164 ATC and 124 PSCCTh patients, diagnosed from 2004 to 2015, were included in the study. There were no differences in sex, ethnicity, age, marital status, or percentage of proximal metastases between the two cancers; however, radiotherapy, chemotherapy, incidence of surgical treatment, and presence of multiple primary tumors were higher in patients with ATC than those with PSCCTh. Further cancer-specific survival (CSS) of patients with PSCCTh was better than that of patients with ATC. Prognostic factors were not identical for the two cancers. Multivariate Cox model analysis indicated that age, sex, radiotherapy, chemotherapy, surgery, multiple primary tumors, marital status, and distant metastasis status are independent prognostic factors for CSS in patients with ATC, while for patients with PSCCTh, the corresponding factors are age, radiotherapy, multiple primary tumors, and surgery. The C-index values of the two models were both > 0.8, indicating that the models exhibited good discriminative ability.ConclusionPrognostic factors influencing CSS were not identical in patients with ATC and PSCCTh. These findings indicate that different clinical treatment and management plans are required for patients with these two types of thyroid cancer.
Background This study aimed to establish and validate an accurate prognostic model, based on demographic and clinical parameters, for predicting the cancer-specific survival (CSS) of patients with poorly differentiated thyroid carcinoma (PDTC). Materials and methods Patients diagnosed with PDTC between 2004 to 2015 were obtained from the Surveillance, Epidemiology, and End Results (SEER) database. Randomly split the data into training and validation sets. Kaplan–Meier analysis with the log-rank test was performed to compare the survival distribution among cases. Univariate and multivariate Cox proportional hazards regression analyses were used to identify independent prognostic factors, which were subsequently utilized to construct a nomogram for predicting the 5- and 10-year cancer-specific survival of patients with PDTC. The discriminative ability and calibration of the nomogram model were assessed using the concordance index and calibration plots, respectively. In addition, we performed a decision curve analysis to assess the clinical value of the nomogram. Simultaneously, we compared the predictive performance of the nomogram model against that of the American Joint Committee on Cancer (AJCC) T-, N-, M-stage. Results A total of 970 eligible patients were randomly assigned to either a training cohort (n = 679) or a validation cohort (n = 291). The Kaplan–Meier analysis revealed that there were no significant differences in cumulative survival based on the race, radiation, and marital status of patients. The stepwise Cox regression model showed that the model was optimal when the following five variables were included: age, tumor size, T-, N-, and M-stage. A nomogram was developed as a graphical representation of the model and exhibited good calibration and discriminative ability in the study. Compared to the T-, N-, and M-stage, the C-index of nomogram (training group: 0.807, validation group: 0.802), the areas under the receiver operating characteristic curve of the training set (5-year AUC: 0.843, 10-year AUC:0.834) and the validation set (5-year AUC:0.878, 10-year AUC:0.811), and the calibration plots of this model all exhibited better performance. At last, compared with T-, N-, and M-stage, the decision curve analysis indicated that the nomogram had excellent clinical net benefit. Conclusions The nomogram developed by us can accurately predict the CSS of PDTC patients. It can help clinicians determine appropriate treatment strategies for poorly differentiated thyroid carcinoma patients.
Abstract Spike protein is a class I virus fusion glycoprotein that plays an important role in cell infection by mediating receptor binding and membrane fusion during virus invasion. The fusion mechanism between virus and host cell mediated by spike protein is complex, and alterations via mutation can affect viral pathogenicity and have been linked with virus pandemics. In-depth studies of the biology and toxicology of viruses that pose a potential public health threat have promoted the use of bioinformatic approaches to explore the structure, function, and evolution of spike proteins. Therefore, data related to spike proteins have increased exponentially. To facilitate further biomedical research, these data require integration. Here, we developed a database of virus spike proteins (DVSP), which provides a free data and bioinformatics service for the scientific community. The DVSP contains 35,579 spike protein and 35,175 nucleotide sequences collected from viruses. Among these, 16,027 spike proteins and 15,623 pre-translational nucleotide sequences are derived from severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), the causative agent of the ongoing corona virus disease 2019 (COVID-19) pandemic. Overall, the DVSP provides 19,552 virus classifications based on integration and virus taxonomy, annotation information from the GenBank database, and sequence information. The sequences, virus taxonomy, and host information contained in the DVSP provide an important resource for future pathogenicity, evolutionary, and drug development studies. It serves as a valuable and comprehensive knowledge source for exploring the distribution, structure, and phylogeny of spike proteins in viruses, thereby promoting the formulation of new hypotheses and experimental designs for spike protein studies. The DVSP is a centralized reference and resource for biological information about the spike proteins of viruses (including coronavirus), which is available at http://yeyn.group:90/.
编程类课程属于实践性课程,传统的实验课程设置是以基础实验为主,导致每个学生的实验成绩区分度不高,学生缺乏独立思考意识,进而打击了学生的积极性并抑制了学生的个性发展.通过比较传统的基础实验和综合性实验设置对学生学习成效的影响,发现综合性实验能有效提高学生的学习成效,充分发挥学生的主观能动性和独立操作能力,以培养学生的综合设计能力和创新意识,使学生具备可视化开发环境下的编程设计能力、良好的程序设计素养与规范的程序设计方法,从而能独立开发出具有实际意义的程序,毕业时能更好地适应市场的需求.
The ring-like In(Ga)As quantum dot pairs (QDPs) is prepared on GaAs ring-ring-disk (R-R-D) nanostructures template by droplet epitaxy (DE) method. The surface morphology of GaAs samples are characterized with Scanning Tunneling Microscope (STM), and studying local controlled growth mechanism of quantum dots (QDs) on GaAs R-R-D nanostructures. STM images show that In(Ga)As QDPs self-assembled a kind of complicated structure with one quantum dot (QD) pair in the center of nanostructures, other QDPs appearing both double rings and disk areas of original GaAs R-R-D nanostructures distributed like a exquisite ring. These In(Ga)As QDPs are uniformly arranged along the [1 1 ¯ 0] direction of GaAs(001) surface. Combining Stranski-Krastanov (S-K) growth mode and the tensile effect of strain field in the [1 1 ¯ 0] direction on sample surface, the nucleation position and distribution shape of ring-like In(Ga)As QDPs is locally controllable. It is found that high uniformly QDPs can be locally controlled by the original GaAs R-R-D nanostructures templates. These results can provide certain guiding significance for the controllable growth of QD by DE method.