Burn-hemorrhagic shock combined injury, a severe condition causing complex stress responses and metabolic disturbances that significantly affect clinical outcomes in both military and civilian settings, was modeled in swine to investigate the associated metabolomic and proteomic changes and identify potential biomarkers for disease prognosis. Eight clean-grade adult male Landrace pigs (4–5 months, average weight 60–70 kg) were used to model burn-hemorrhagic shock combined injury. Serum samples collected at 0 h and 2 h post-injury were analyzed using metabolomic and proteomic measurements. The metabolomic and proteomic data were processed through partial least squares–discriminant analysis (PLS–DA) and the KEGG enrichment etc. Furthermore, the integrate analysis of the metabolomic and proteomic data was generalized by canonical correlation discriminant analysis, and the correlation between metabolites and mortality of the swine model was predicted using a multiple linear regression model by Pearson analysis. PLS–DA revealed a global shift in each of the metabolomic and proteomic profiles following injury. The levels of 87 signature metabolites including various types of amino acids, fatty acids and acyl-carnitines of different lengths, and many metabolites in the gluconeogenesis, glycolysis, and tricarboxylic acid (TCA) cycle are generally increased (P < 0.05) after injury and can be used as biomarkers. Pathways related to amino acids metabolism and TCA cycle were significantly enriched (P < 0.01). In proteome analysis, we found dramatically altered (P < 0.05) levels of matrix and red blood cell-related proteins, such as type I collagen and hemoglobin. Most importantly, we found that the markedly elevated (P < 0.01) succinic acid, glutaric acid, and malic acid are closely associated (r = 0.863, 0.861, and 0.821, respectively) with injury severity by Pearson analysis, and can predict mortality using a multiple linear regression model. The study provides compelling observations that burn-shock swine model undergoes dramatic changes in the acute phase and present a valuable panel for clinical use of prognosis.
This paper presents Py2ONTO-Edit, an ontology editing tool that integrates the low-level functionality of Owlready2 to simplify the extraction and translation of ontology terms. It offers two extraction methods: 1. Global extraction method. 2. Selective-depth extraction method. Another key feature is the translation of ontology terms using multiple translation packages to add non-English labels (e.g., Chinese, French, German) to the ontology. This paper presents two main contributions: 1. Implementation of flexible features for term extraction. 2. Enabling of multilingual translation of ontology terms. Py2ONTO-Edit is an easy-to-use Python tool for developers focused on ontology term reuse and translation.
With the rise of data-intensive research, data literacy has become a critical capability for improving scientific data quality and achieving artificial intelligence (AI) readiness. In the biomedical domain, data are characterized by high complexity and privacy sensitivity, calling for robust and systematic data management skills. This paper reviews current trends in scientific data governance and the evolving policy landscape, highlighting persistent challenges such as inconsistent standards, semantic misalignment, and limited awareness of compliance. These issues are largely rooted in the lack of structured training and practical support for researchers. In response, this study builds on existing data literacy frameworks and integrates the specific demands of biomedical research to propose a comprehensive, lifecycle-oriented data literacy competency model with an emphasis on ethics and regulatory awareness. Furthermore, it outlines a tiered training strategy tailored to different research stages-undergraduate, graduate, and professional, offering theoretical foundations and practical pathways for universities and research institutions to advance data literacy education.
Background Several clinical cases and experiments have demonstrated the effectiveness of traditional Chinese medicine (TCM) formulas in treating and preventing diseases. These formulas contain critical information about their ingredients, efficacy, and indications. Classifying TCM formulas based on this information can effectively standardize TCM formulas management, support clinical and research applications, and promote the modernization and scientific use of TCM. To further advance this task, TCM formulas can be classified using various approaches, including manual classification, machine learning, and deep learning. Additionally, large language models (LLMs) are gaining prominence in the biomedical field. Integrating LLMs into TCM research could significantly enhance and accelerate the discovery of TCM knowledge by leveraging their advanced linguistic understanding and contextual reasoning capabilities. Objective The objective of this study is to evaluate the performance of different LLMs in the TCM formula classification task. Additionally, by employing ensemble learning with multiple fine-tuned LLMs, this study aims to enhance classification accuracy. Methods The data for the TCM formula were manually refined and cleaned. We selected 10 LLMs that support Chinese for fine-tuning. We then employed an ensemble learning approach that combined the predictions of multiple models using both hard and weighted voting, with weights determined by the average accuracy of each model. Finally, we selected the top 5 most effective models from each series of LLMs for weighted voting (top 5) and the top 3 most accurate models of 10 for weighted voting (top 3). Results A total of 2441 TCM formulas were curated manually from multiple sources, including the Coding Rules for Chinese Medicinal Formulas and Their Codes, the Chinese National Medical Insurance Catalog for proprietary Chinese medicines, textbooks of TCM formulas, and TCM literature. The dataset was divided into a training set of 1999 TCM formulas and test set of 442 TCM formulas. The testing results showed that Qwen-14B achieved the highest accuracy of 75.32% among the single models. The accuracy rates for hard voting, weighted voting, weighted voting (top 5), and weighted voting (top 3) were 75.79%, 76.47%, 75.57%, and 77.15%, respectively. Conclusions This study aims to explore the effectiveness of LLMs in the TCM formula classification task. To this end, we propose an ensemble learning method that integrates multiple fine-tuned LLMs through a voting mechanism. This method not only improves classification accuracy but also enhances the existing classification system for classifying the efficacy of TCM formula.
This paper presents a large publicly available benchmark dataset (TCMEval-SDT) for the thought process involved in syndrome differentiation in traditional Chinese medicine (TCM). The dataset consists of 300 TCM syndrome diagnosis cases sourced from the internet, classical Chinese medical texts, and medical records from hospitals, with metadata adhering to the Findable, Accessible, Interoperable, and Reusable (FAIR) principles. Each case has been annotated and curated by TCM experts and includes medical record ID, clinical data, explanatory summary, TCM syndrome, clinical information, and TCM pathogenesis, to support algorithms or models in emulating the diagnostic process of TCM clinicians. To provide a comprehensive description of the TCM syndrome diagnosis process, we summarize the diagnosis into four steps: (1) clinical information extraction, (2) TCM pathogenesis reasoning, (3) TCM syndrome reasoning, and (4) explanatory summary. We have also established validation criteria to evaluate their ability in TCM clinical diagnosis using this dataset. To facilitate research and evaluation in syndrome diagnosis of TCM, the TCMEval-SDT dataset is made publicly available under the CC-BY 4.0 license.
BACKGROUND:Ovarian cancer (OC) is a major global cause of death among gynecological cancers, with a high mortality rate. Early diagnosis, distinguishing between benign conditions and early malignant OC forms, is vital for successful treatment. This research investigates serum metabolites to find diagnostic biomarkers for early OC identification. METHODS:Metabolomic profiles derived from the serum of 60 patients with benign conditions and 60 patients with malignant OC were examined using ultra-performance liquid chromatography coupled with tandem mass spectrometry (UPLC-MS/MS). Comparative analysis revealed differential metabolites linked to OC, aiding biomarker identification for early-diagnosis of OC via machine learning features. The predictive ability of these biomarkers was evaluated against the traditional biomarker, cancer antigen 125 (CA125). RESULTS:84 differential metabolites were identified, including 2-Thiothiazolidine-4-carboxylic acid (TTCA), Methionyl-Cysteine, and Citrulline that could serve as potential biomarkers to identify benign conditions and malignant OC. In the diagnosis of early-stage OC, the area under the curve (AUC) for Citrulline was 0.847 (95 % Confidence Interval (CI): 0.719-0.974), compared to 0.770 (95 % CI: 0.596-0.944) for TTCA, and 0.754 for Methionine-Cysteine (95 % CI: 0.589-0.919). These metabolites demonstrate a superior diagnostic capability relative to CA125, which has an AUC of 0.689 (95 % CI: 0.448-0.931). Among these biomarkers, Citrulline stands out as the most promising. Additionally, in the diagnosis of benign conditions and malignant OC, using logistic regression to combine potential biomarkers with CA125 has an AUC of 0.987 (95 % CI: 0.9708-1) has been proven to be more effective than relying solely on the traditional biomarker CA125 with an AUC of 0.933 (95 % CI: 0.870-0.996). Furthermore, among all the differential metabolites, lipid metabolites dominate, significantly impacting glycerophospholipid metabolism pathway. CONCLUSION:The discovered serum metabolite biomarkers demonstrate excellent diagnostic performance for distinguishing between benign conditions and malignant OC and for early diagnosis of malignant OC.
Gene expression levels serve as valuable markers for assessing prognosis in cancer patients. To understand the mechanisms underlying prognosis and explore potential therapeutics across diverse cancers, we developed CancerPro (https:/medcode.link/cancerpro). This knowledge network platform integrates comprehensive biomedical data on genes, drugs, diseases and pathways, along with their interactions. By integrating ontology and knowledge graph technologies, CancerPro offers a user-friendly interface for analyzing pan-cancer prognostic markers and exploring genes or drugs of interest. CancerPro implements three core functions: gene set enrichment analysis based on multiple annotations; in-depth drug analysis; and in-depth gene list analysis. Using CancerPro, we categorized genes and cancers into distinct groups and utilized network analysis to identify key biological pathways associated with unfavorable prognostic genes. The platform further pinpoints potential drug targets and explores potential links between prognostic markers and patient characteristics such as glutathione levels and obesity. For renal and prostate cancer, CancerPro identified risk genes linked to immune deficiency pathways and alternative splicing abnormalities. This research highlights CancerPro's potential as a valuable tool for researchers to explore pan-cancer prognostic markers and uncover novel therapeutic avenues. Its flexible tools support a wide range of biological investigations, making it a versatile asset in cancer research and beyond.
The significance of small molecule metabolites as biomarkers for disease diagnosis and prognosis is growing increasingly evident, necessitating the development of highly sensitive qualitative and quantitative methods. Herein, multi-chemoselective probes are synthesized and applied for profiling metabolites, including carboxyl, phosphate, hydroxyl, amino, thiol, and carbonyl compounds. This approach seamlessly integrates magnetic solid-phase materials, orthogonal cleavage sites, isotopic tags, and selective coupling sites, minimizes matrix interference, and enhances quantitative accuracy. Meanwhile, a homemade program, High-Resolution Isotope-Assisted Identification and Quantitative (HRIAIQuant) is developed to process the data, which adeptly filters through 33,874 ion pairs present in human serum, leading to the identification of 701 known metabolites and a remarkable 1,062 potential novel ones. This method is successfully applied to analyze metabolites in multiple brain regions of SAMP8 and SAMR1 models, offering a novel tool for Alzheimer's disease research.
In recent years, the morbidity and mortality of lung cancer have been rising continuously, and it has become one of the most dangerous malignant tumors that threaten human life and health. The incidence of non-small cell lung cancer(NSCLC) accounts for more than 80% of the total incidence of lung cancer. Due to its complicated diagnostic process and high diagnostic cost, the effective diagnosis and treatment of NSCLC have become a great challenge for doctors. It has found that tumor mutation burden(TMB) is positively correlated with the efficacy of NSCLC immunotherapy, and TMB value has a certain predictive effect on the efficacy of targeted therapy and chemotherapy. Based on the above findings, a deep learning model(RcaNet) was proposed. In this model, a residual network(ResNet) was taken as the backbone network, and multi-dimensional feature attention and multi-scale information fusion were added in the network, enhancing the ability of the network in paying attention to and extract the deep features of lung cancer pathological sections. Experiments were performed with RcaNet and the mainstream deep learning models on the TCGA public data set with experimental training samples of 925 954. The results showed that the average area under the curve(AUC) of the RcaNet model is 0.883 0, which is 6.8% higher than that of CAIM model, 4.2% higher than that of ResNeSt model, and 5.3% higher than that of ResNet model. Our proposed method has guiding significance and application value for the diagnosis and treatment of NSCLC.
The past twenty years have seen the increasingly important role of ontology in traditional Chinese medicine (TCM). However, the development of TCM ontology faces many challenges. Since the epistemologies dramatically differ between TCM and contemporary biomedicine, it is hard to apply the existing top-level ontology mechanically. "Data silos" are widely present in the currently available terminology standards, term sets, and ontologies. The formal representation of ontology needs to be further improved in TCM. Therefore, we propose a unified basic semantic framework of TCM based on in-depth theoretical research on the existing top-level ontology and a re-study of important concepts in TCM. Under such a framework, ontologies in TCM sub-domains should be built collaboratively and be represented formally in a common format. Besides, extensive cooperation should be encouraged by establishing ontology research communities to promote ontology peer review and reuse.
Background The current COVID-19 pandemic and the previous SARS/MERS outbreaks of 2003 and 2012 have resulted in a series of major global public health crises. We argue that in the interest of developing effective and safe vaccines and drugs and to better understand coronaviruses and associated disease mechenisms it is necessary to integrate the large and exponentially growing body of heterogeneous coronavirus data. Ontologies play an important role in standard-based knowledge and data representation, integration, sharing, and analysis. Accordingly, we initiated the development of the community-based Coronavirus Infectious Disease Ontology (CIDO) in early 2020. Results As an Open Biomedical Ontology (OBO) library ontology, CIDO is open source and interoperable with other existing OBO ontologies. CIDO is aligned with the Basic Formal Ontology and Viral Infectious Disease Ontology. CIDO has imported terms from over 30 OBO ontologies. For example, CIDO imports all SARS-CoV-2 protein terms from the Protein Ontology, COVID-19-related phenotype terms from the Human Phenotype Ontology, and over 100 COVID-19 terms for vaccines (both authorized and in clinical trial) from the Vaccine Ontology. CIDO systematically represents variants of SARS-CoV-2 viruses and over 300 amino acid substitutions therein, along with over 300 diagnostic kits and methods. CIDO also describes hundreds of host-coronavirus protein-protein interactions (PPIs) and the drugs that target proteins in these PPIs. CIDO has been used to model COVID-19 related phenomena in areas such as epidemiology. The scope of CIDO was evaluated by visual analysis supported by a summarization network method. CIDO has been used in various applications such as term standardization, inference, natural language processing (NLP) and clinical data integration. We have applied the amino acid variant knowledge present in CIDO to analyze differences between SARS-CoV-2 Delta and Omicron variants. CIDO's integrative host-coronavirus PPIs and drug-target knowledge has also been used to support drug repurposing for COVID-19 treatment. Conclusion CIDO represents entities and relations in the domain of coronavirus diseases with a special focus on COVID-19. It supports shared knowledge representation, data and metadata standardization and integration, and has been used in a range of applications.
目的:开发一个简单易用的本体构建工具Py2ONTO,探索人工方式和自动化方式相结合的本体构建方案.方法:通过将Owlready常用的底层函数封装为功能函数,在此基础上设计一个基于CSV模板的自动构建本体的流程,并在代码中提供功能函数的调用,通过BioPortal、MedPortal等本体数据库实现术语映射功能.结果:开发了基于Python的本体建设工具Py2ONTO,该工具可以通过命令行方式及Python脚本调用功能函数方式构建本体.通过该工具构建了两个本体,测试结果证明本文构建的本体满足其术语覆盖度和可推理性评估要求.结论:为本体构建提供了一种较为简单且可靠的选择,有助于科研人员构建高质量的本体.
目的:基于FDA公共数据开放项目(openFDA)中诺西那生不良事件的数据,分析研究诺西那生的不良反应,为临床安全用药提供参考.方法:收集2016年1月1日至2020年12月31日美国FDA不良事件报告系统(FAERS)中诺西那生相关的不良事件(AE)报告,提取报告数排名前50位不良事件报告,采用报告比值比法(ROR)和比例报告比值法(PRR)挖掘诺西那生不良反应风险信号.结果:设定时段内FAERS共收到以诺西那生为首要怀疑药物的不良事件报告3497例,对报告数排名前50位不良事件进行药物不良反应风险信号分析,检测出35个不良反应风险信号,其中30个为说明书中未收录的不良事件.不良事件报告数前5位的不良事件依次为头痛、腰穿后综合征、发热、感染性肺炎、呕吐.不良反应风险信号强度排名前5位的依次为腰穿后综合征、脑脊液漏、鼻病毒感染、尿蛋白检出、呼吸道合胞病毒感染.结论:利用诺西那生不良事件的真实世界数据深入分析其潜在的不良反应,提示临床予以关注及进一步进行安全性评价.其中尿蛋白检出不良反应风险信号强度高,建议在每次给药前进行尿蛋白检测.
肿瘤突变负荷(TMB)与非小细胞肺癌(NSCLC)的免疫治疗疗效呈正相关,并且在近期的相关研究中,肿瘤突变负荷对靶向治疗及化疗的疗效也有一定的预测作用.因此,提出一种融合注意力机制的Inception深度学习模型(CAIM),用于对TCGA数据库中的非小型细胞肺癌中的肺腺癌的病理切片进行识别.首先,对数据样本进行切分,裁剪成小切片;然后送入到深度学习模型中,通过卷积学习图像特征,再与注意力机制结合进一步加强特征的提取;最后通过对小切片预测信息的整合,自动判别肺腺癌病理切片TMB值的高低.数据集由337张肺腺癌病理组织切片组成,其中高TMB值的数据271张,低TMB值的数据实验66张.结果 表明,所提方法的性能平均曲线下面积(AUC)为0.82,明显高于图像分类方法残差网络(ResNet)的AUC值0.66.研究结果对临床实践中肿瘤突变负荷的检测和辅助诊断具有重要意义.
Research on drugs against SARS-CoV-2 (cause of COVID-19) has been one of the major world concerns at present. There have been abundant research data and findings in this field. The interference of drugs on gene expression in cell lines, drug-target, protein-virus receptor networks, and immune cell infiltration of the host may provide useful information for anti-SARS-CoV-2 drug research. To simplify the complex bioinformatics analysis and facilitate the evaluation of the latest research data, we developed OmiczViz ( http://medcode.link/omicsviz ), a web tool that has integrated drug-cell line interference data, virus-host protein-protein interactions, and drug-target interactions. To demonstrate the usages of OmiczViz, we analyzed the gene expression data from cell lines treated with chloroquine and ruxolitinib, the drug-target protein networks of 48 anti-coronavirus drugs and drugs bound with ACE2, and the profiles of immune cell infiltration between different COVID-19 patient groups. Our research shows that chloroquine had a regulatory role of the immune response in renal cell line but not in lung cell line. The anti-coronavirus drug-target network analysis suggested that antihistamine of promethaziney and dietary supplement of Zinc might be beneficial when used jointly with antiviral drugs. The immune infiltration analysis indicated that both the COVID-19 patients admitted to the ICU and the elderly with infection showed immune exhaustion status, yet with different molecular mechanisms. The interactive graphic interface of OmiczViz also makes it easier to analyze newly discovered and user-uploaded data, leading to an in-depth understanding of existing findings and an expansion of existing knowledge of SARS-CoV-2. Collectively, OmicsViz is web program that promotes the research on medical agents against SARS-CoV-2 and supports the evaluation of the latest research findings.
The construction of a professional learning management system(LMS)platform is one of the important tasks in modern medical education reform. Peking Union Medical College (PUMC) had built the LMS based on open-source software modular object-oriented dynamic learning environment(Moodle)that supports the implementation of blended teaching methods. First, a teaching model for performing blended teaching effectively based on Moodle was built. Under the guidance of such model, the autonomous learning environment was established, which provided orderly, abundant and various learning resources for 46 medicine related courses in PUMC. In such a virtual environment, the students' self-directed learning could guided, supervised and provided feedback. The teachers could combine the virtual learning environment with classroom teaching closely, and develop more innovative and exploratory teaching activities. Furthermore, the data from whole learning process could be accumulated by Moodle easily to evaluate the teaching quality. Through Moodle, modern teaching methods could be applied in the medical education in PUMC effectively and efficiently.
近年来,互联网与物联网等技术的快速发展为探寻精准医学与人工智能之间的关联应用提供了战略机遇。然而,如何组织、集成和共享复杂异构的生物医学数据业已成为阻碍该领域发展的严重技术挑战。本体因其能够提供知识与元数据层面的语义基础而推动了生物医学人工智能的发展,而基于互操作性的本体对异构知识与数据的整合和分析发挥着关键性的作用。将人工智能与精准医疗的有机集合称之为智能精准医疗,并提出一个"河马假设",用以阐明互操作性本体与智能精准医疗之间的正相关关系及如何发挥协同增效作用;同时,提出和展示了使用可扩展本体开发的原理和工具来实现本体的互操作性,进而支持智能精准医疗的应用与发展。此外,还对国内外互操作性本体的研究现状、本体中国的成立与发展以及医学伦理在智能精准医疗中的重要性等进行了综述和深入探讨。
Current COVID-19 pandemic and previous SARS/MERS outbreaks have caused a series of major crises to global public health We must integrate the large and exponentially growing amount of heterogeneous coronavirus data to better understand coronaviruses and associated disease mechanisms, in the interest of developing effective and safe vaccines and drugs Ontologies have emerged to play an important role in standard knowledge and data representation, integration, sharing, and analysis We have initiated the development of the community-based Coronavirus Infectious Disease Ontology (CIDO) As an Open Biomedical Ontology (OBO) library ontology, CIDO is an open source and interoperable with other existing OBO ontologies In this article, the general architecture and the design patterns of the CIDO are introduced, CIDO representation of coronaviruses, phenotypes, anti-coronavirus drugs and medical devices (e g ventilators) are illustrated, and an application of CIDO implemented to identify repurposable drug candidates for effective and safe COVID-19 treatment is presented Copyright © 2020 for this paper by its authors
Background: Tumor mutation burden (TMB) is associated with the formation of tumors and the outcomes of immunotherapy, especially in non-small cell lung cancers (NSCLC). The aim of this study was to study immune features of NSCLC with different mutation burdens. We performed a comprehensive comparison of immune features, including driver mutations, up- or down-regulated expression of key biological pathways, tumor cell evolution, tumor cell stemness, etiology, and immune cell infiltration, between ultra-low (UL) and ultra-high (UH) TMB NSCLC. Results: Generally, we found that UL-TMB tumors contained a higher proportion of driver mutations and maintained an “immune-desert” tumor microenvironment with reduced effector cell infiltration and increased suppressor cell infiltration. In UH-TMB tumors, the expression levels of cell cycles, DNA replication and mismatch repair pathways were up-regulated. We discovered that UL-TMB tumors had different up-regulated pathways between a denocarcinoma (LUAD) and squamous cell carcinoma (LUSC). Basically, the expression levels of primary bile acid biosynthesis, arachidonic acid metabolism, and drug metabolism pathways were up-regulated in LUAD; however, the expression level of CaM pathway and the PLC-gamma1 signaling pathway were up-regulated in LUSC. Most mutations in UH-TMB tumors occurred in the early stages, while UL-TMB tumors showed the opposite trend. UH-TMB tumors had a relatively high stemness score. Despite the mutations and neo-antigens increasing in UH-TMB tumors, the overall antigen presentation capacity was weak. Conclusions: These results support rational design of therapeutic strategies for NSCLC with different mutation burdens.
The Coronavirus Infectious Disease Ontology (CIDO) is a community-based ontology that supports coronavirus disease knowledge and data standardization, integration, sharing, and analysis.