Current methods for breast cancer diagnosis often face limitations in integrating prior knowledge, fusing multimodal data, and providing interpretability. To this end, we propose a novel knowledge-augmented multimodal learning framework. Based on a clinically guided breast cancer knowledge graph, our approach enhances patient clinical representations using Graph Attention Networks (GAT) and extracts pathological image features through weak supervision learning. A bidirectional cross-attention fusion mechanism enables interactive alignment and integration of multimodal heterogeneous information in the feature space. Extensive experiments on a real-world PathologicalEMR dataset demonstrate that our method achieves outstanding performance, with an AUC of 0.9963 in distinguishing benign from malignant cases. Further ablation studies validate the contributions of various multimodal feature combinations and knowledge enhancement strategies, underscoring the framework's effectiveness and interpretability. The proposed solution significantly advances the capability and trustworthiness of multi-source information integration for complex medical diagnostics, offering a promising tool for clinical decision support.
Background: Drug repositioning is a pivotal strategy in pharmaceutical research, offering accelerated and cost-effective therapeutic discovery. However, biomedical information relevant to drug repositioning is often complex, dispersed, and underutilized due to limitations in traditional extraction methods, such as reliance on annotated data and poor generalizability. Large language models (LLMs) show promise but face challenges such as hallucinations and interpretability issues. Objective: This study proposed long chain-of-thought for drug repositioning knowledge extraction (LCoDR-KE), a lightweight and domain-specific framework to enhance LLMs' accuracy and adaptability in extracting structured biomedical knowledge for drug repositioning. Methods: A domain-specific schema defined 11 entities (eg, drug, disease) and 18 relationships (eg, treats, is biomarker of). Following the established schema architecture, we constructed automatic annotation based on 10,000 PubMed abstracts via chain-of-thought prompt engineering. A total of 1000 expert-validated abstracts were curated into a drug repositioning corpus, a high-quality specialized corpus, while the remaining entries were allocated for model training purposes. Then, the proposed LCoDR-KE framework combined supervised fine-tuning of the Qwen2.5-7B-Instruct model with reinforcement learning and dual-reward mechanisms. Performance was evaluated against state-of-the-art models (eg, conditional random fields, Bidirectional Encoder Representations From Transformers, BioBERT, Qwen2.5, DeepSeek-R1, OpenBioLLM-70B, and model variants) using precision, recall, and F1-score. In addition, the convergence of the training method was assessed by analyzing performance progression across iteration steps. Results: LCoDR-KE achieved an entity F1 of 81.46% (eg, drug 95.83%, disease 90.52%) and triplet F1 of 69.04%, outperforming traditional models and rivaling larger LLMs (DeepSeek-R1: entity F1=84.64%, triplet F1=69.02%). Ablation studies confirmed the contributions of supervised fine-tuning (8.61% and 20.70% F1 drop if removed) and reinforcement learning (6.09% and 14.09% F1 drop if removed). The training process demonstrated stable convergence, validated through iterative performance monitoring. Qualitative analysis of the model's chain-of-thought outputs showed that LCoDR-KE performed structured and schema-aware reasoning by validating entity types, rejecting incompatible relations, enforcing constraints, and generating compliant JSON. Error analysis revealed 4 main types of mistakes and challenges for further improvement. Conclusions: LCoDR-KE enhances LLMs' domain-specific adaptability for drug repositioning by offering an open-source drug repositioning corpus and a long chain-of-thought framework based on a lightweight LLM model. This framework supports drug discovery and knowledge reasoning while providing scalable, interpretable solutions applicable to broader biomedical knowledge extraction tasks.
With the proliferation of online medical communities, vast amounts of valuable medical and health question-and-answer data have emerged, paving new avenues for the extraction of public health information and the enhancement of question-and-answer system efficiency. To classify public health-related questions more accurately, this paper introduces a novel model ensemble approach that skillfully integrates the advantages of BERT and two large language models, aiming to achieve precise categorization of public health questions by aggregating predictions from each model. To validate the effectiveness of this method, we conducted detailed experiments using the Chinese Medical Intent Dataset. The proposed method was then compared with various existing methods, showing a significant improvement in accuracy, demonstrating the outstanding efficacy of combining model integration with a voting mechanism in the classification task of public health questions.
Studying patients’ physician selection behavior through multimodal data is a key focus in online healthcare community. However, there is limited research on multimodal data fusion in this field. This paper aims to analyze how different types of data influence patients’ online physician selection from the perspective of multimodal data fusion, providing insights into patients’ decision-making in the era of big data. Based on multimodal data from the Haodf platform, this study includes a sample of 13,166 records. Two predictive models were developed: a feature-engineered neural network model and an end-to-end deep learning model. The former combines selected features extracted from images and text with structured data to predict physician selection behavior. The latter uses raw data, employing ResNet50 and BERT for image and text feature extraction, then integrates structured data for prediction. The analysis confirmed that both models can effectively predict patients’ physician selection behavior and offer a certain degree of interpretability. The feature-engineered model achieved a test set MSE of 0.4124, MAE of 0.4579, and an R2 of 0.5304, while the end-to-end model achieved a test set MSE of 0.3991, MAE of 0.4327, and an R2 of 0.6321. Overall, the end-to-end model outperformed the feature-engineered model. Both the feature-engineered and the end-to-end models effectively explore patients’ physician selection behavior in online healthcare community. The results demonstrate that the multimodal information—including doctor-generated content, patient-generated content, and system-generated content—significantly influences patient decision-making. Additionally, different features from images and text influence patients’ choices in distinct ways.
With the booming development of online medical communities, a large amount of valuable medical and health question-and-answer data emerges, providing the possibility for the extraction of public health information and the improvement of question and answer system efficiency. To classify public health questions more effectively, this paper proposes a model ensemble approach that utilizes BERT and two large language models to accurately predict the categories of public health questions. It integrates the prediction results from each model through a voting mechanism to derive the final classification outcome. To verify the effectiveness of this method, this paper conducts experiments based on the Chinese Medical Intent Dataset (CMID) and compares it with various other methods. The experimental results demonstrate that the method proposed in this paper has achieved significant improvement in accuracy, effectively proving the effectiveness of the approach combining model integration and voting mechanism in the classification task of public health questions.
Exploring the potential efficacy of a drug is a valid approach for drug discovery with shorter development times and lower costs. Recently, several computational drug repositioning methods have been introduced to learn multi-features for potential association prediction. A drug repositioning knowledge graph of drugs, diseases, targets, genes and side effects was introduced in our study to impose an explicit structure to integrate heterogeneous biomedical data. We revealed drug and disease embeddings from the constructed knowledge graph via a two-layer graph convolutional network with an attention mechanism. Finally, KGCN-DDA achieved superior performance in drug-disease association prediction with an AUC value of 0.8818 and an AUPR value of 0.5916, a relative improvement of 31.67
Objective To develop a traceable cancer hallmark ontology with terminology including gene mutation,cancer hallmark,and cell line for knowledge integration,standardization,correlation,and discovery.Methods The Ontology Development 101 and the current ontology development methods were employed to determine the content coverage,structural layers,reusable terms,and new terms of the cancer hallmark ontology.Taking colorectal cancer as a study case,we extracted the knowledge related with colorectal cancer hallmarks using text mining and text classification technology from PubMed,and then formalized the extracted knowledge into the cancer hallmark ontology.Moreover,we made use of existing cancer hallmark evidence in Catalogue of Somatic Mutations in Cancer and further semantic retrieval to discover new knowledge.Results The established cancer hallmark ontology comprised 9910 classes and 6138 instances,which realized the semantic representation of 2310 article abstracts about colorectal cancer and 26 pieces of evidence about genes and their cancer hallmarks.Compared with the Catalogue of Somatic Mutations in Cancer,new evidence for more genes associated with colorectal cancer hallmarks was found based on cancer hallmark ontology.Conclusion This study is of great significance to the research on the cancer pathogenesis at the molecular level,the revealing of specific roles of genes and mutations in the occurrence of cancer,and the rapid knowledge discovery of cancer hallmarks.
构建基于药物多特征融合的药物疾病关联预测模型,为药物知识发现提供新思路.借助药物的化学结构、药物-副作用关联、药物-靶标关联的3个特征,构建融合的药物综合相似度及基于MeSH的疾病语义相似度特征表示方法.利用图卷积神经网络模型抽取药物-疾病图数据特征信息,构建基于多特征融合的药物疾病关联预测模型(MFFGCN),进而实现未知的药物疾病关联发现.利用269种药物、598种疾病及其之间的18 416种关联关系,对药物疾病存在的未知关联进行预测,借助AUC、AUPR、准确率、灵敏度、召回率、F1等多个评价指标进行评价.结果表明,多特征融合的药物疾病关联预测方法的AUC指标为0.866 2,较单一特征的平均预测指标最大相对提升为2.48%,较4种代表性基线方法的指标最大相对提升为1.67%;AUPR指标为0.341 2,较单一特征预测结果最大相对提升为1.67%,较4种代表性基线方法提升27.49%.对预测结果中药物-疾病预测关联得分中排名前10的组合及阿霉素为例的单一药物预测组合进行文献研究验证、临床治疗验证,同样证明MFFGCN在未知的药物疾病关联预测上表现良好,能有效地发现药物的新适应症,为药物重定位提供方法借鉴和理论依据.
Medical procedure entity normalization is an important task to realize medical information sharing at the semantic level; it faces main challenges such as variety and similarity in real-world practice. Although deep learning-based methods have been successfully applied to biomedical entity normalization, they often depend on traditional context-independent word embeddings, and there is minimal research on medical entity recognition in Chinese Regarding the entity normalization task as a sentence pair classification task, we applied a three-step framework to normalize Chinese medical procedure terms, and it consists of dataset construction, candidate concept generation and candidate concept ranking. For dataset construction, external knowledge base and easy data augmentation skills were used to increase the diversity of training samples. For candidate concept generation, we implemented the BM25 retrieval method based on integrating synonym knowledge of SNOMED CT and train data. For candidate concept ranking, we designed a stacking-BERT model, including the original BERT-based and Siamese-BERT ranking models, to capture the semantic information and choose the optimal mapping pairs by the stacking mechanism. In the training process, we also added the tricks of adversarial training to improve the learning ability of the model on small-scale training data. Based on the clinical entity normalization task dataset of the 5th China Health Information Processing Conference, our stacking-BERT model achieved an accuracy of 93.1%, which outperformed the single BERT models and other traditional deep learning models. In conclusion, this paper presents an effective method for Chinese medical procedure entity normalization and validation of different BERT-based models. In addition, we found that the tricks of adversarial training and data augmentation can effectively improve the effect of the deep learning model for small samples, which might provide some useful ideas for future research.
Transcriptome-wide association studies (TWASs), as a practical and prevalent approach for detecting the associations between genetically regulated genes and traits, are now leading to a better understanding of the complex mechanisms of genetic variants in regulating various diseases and traits. Despite the ever-increasing TWAS outputs, there is still a lack of databases curating massive public TWAS information and knowledge. To fill this gap, here we present TWAS Atlas (https://ngdc.cncb.ac.cn/twas/), an integrated knowledgebase of TWAS findings manually curated from extensive literature. In the current implementation, TWAS Atlas collects 401,266 high-quality human gene-trait associations from 200 publications, covering 22,247 genes and 257 traits across 135 tissue types. In particular, an interactive knowledge graph of the collected gene-trait associations is constructed together with single nucleotide polymorphism (SNP)-gene associations to build up comprehensive regulatory networks at multi-omics levels. In addition, TWAS Atlas, as a user-friendly web interface, efficiently enables users to browse, search and download all association information, relevant research metadata and annotation information of interest. Taken together, TWAS Atlas is of great value for promoting the utility and availability of TWAS results in explaining the complex genetic basis as well as providing new insights for human health and disease research.
Introduction: Exploring the potential efficacy of a drug is a valid approach for drug development with shorter development times and lower costs. Recently, several computational drug repositioning methods have been introduced to learn multi-features for potential association prediction. However, fully leveraging the vast amount of information in the scientific literature to enhance drug-disease association prediction is a great challenge.Methods: We constructed a drug-disease association prediction method called Literature Based Multi-Feature Fusion (LBMFF), which effectively integrated known drugs, diseases, side effects and target associations from public databases as well as literature semantic features. Specifically, a pre-training and fine-tuning BERT model was introduced to extract literature semantic information for similarity assessment. Then, we revealed drug and disease embeddings from the constructed fusion similarity matrix by a graph convolutional network with an attention mechanism.Results: LBMFF achieved superior performance in drug-disease association prediction with an AUC value of 0.8818 and an AUPR value of 0.5916.Discussion: LBMFF achieved relative improvements of 31.67% and 16.09%, respectively, over the second-best results, compared to single feature methods and seven existing state-of-the-art prediction methods on the same test datasets. Meanwhile, case studies have verified that LBMFF can discover new associations to accelerate drug development. The proposed benchmark dataset and source code are available at: https://github.com/kang-hongyu/LBMFF.
目的:调查食管癌患者临床营养状况,分析其相关影响因素,为提早预防及选择合理营养支持模式提供依据.方法:采用定点连续抽样法,选取食管癌患者137例为研究对象,入院24?h内应用一般资料,PG-SGA评估量表及体质指数给予患者进行营养状况调查.结果:(1)根据PG-SGA评分显示接受调查患者中13.8%不需营养干预,86.2%需不同程度营养干预;(2)根据体质指数BMI分组,与组内PG-SGA分级构成进行比较,具有统计学差异(P<0.05);(3)根据就诊目的进行分组,与各组内PG-SGA分级构成进行对比,具有统计学差异(P<0.05);?(4)根据合并症数量进行分组,对各组内PG-SGA分级构成进行比较,不具有统计学差异(P>0.05).结论:临床食管癌患者普遍存在营养问题;单纯的体质指数不能直接反映患者是否需要给予营养干预;随着体质指数增加,需要给予营养干预的概率会下降;患者拥有合并症数量的多少与是否存在营养不良无直接关系;医务人员给予患者及家属进行健康教育对患者营养改善有重要意义.
Purpose:Researchers have identified gut microbiota that interact with brain regions associated with emotion and mood. Literature reviews of those associations rely on rigorous systematic approaches and labor-intensive investments. Here we explore how knowledge graph, a large scale semantic network consisting of entities and concepts as well as the semantic relationships among them, is incorporated into the emotion-probiotic relationship exploration work.Method:We propose an end-to-end emotion-probiotics relationship exploration method with an integrated medical knowledge graph, which incorporates the text mining output of knowledge graph, concept reasoning and evidence classification. Specifically, a knowledge graph for probiotics is built based on a text-mining analysis of PubMed, and further used to retrieve triples of relationships with reasoning logistics. Then specific relationships are annotated and evidence levels are retrieved to form a new evidence-based emotion-probiotic knowledge graph.Results:Based on the probiotics knowledge graph with 40,442,404 triples, totally 1453 PubMed articles were annotated in both the title level and abstract level, and the evidence levels were incorporated to the visualization of the explored emotion-probiotic relationships. Finally, we got 4131 evidenced emotion-probiotic associations.Conclusions:The evidence-based emotion-probiotic knowledge graph construction work demonstrates an effective reasoning based pipeline of relationship exploration. The annotated relationship associations are supposed be used to help researchers generate scientific hypotheses or create their own semantic graphs for their research interests.
The objective of this study was to develop a hybrid method and perform an initial evaluation of mappings from the International Statistical Classification of Diseases, 10th revision, Chinese version (ICD-10-CN) to the Systematized Nomenclature of Medicine - Clinical Terms (SNOMED-CT). The methods used to perform mapping include reusing existing mappings, term similarity modeling for automatic mapping and manual review. We evaluated the results of automatic mapping and the coverage of the maps between two terminologies. Experimental results demonstrated that fine-tuning the pre-trained biomedical language model of PubmedBERT obtained the optimal performance, with a precision of 0.859, a recall of 0.773, and a F1 of 0.814. 100% 4-digit code ICD-10-CN terms were mapped to SNOMED-CT terms through exsit code mappings. Around 42.41% randomly selected 6-digit code ICD-10-CN terms had exact matches to corresponding SNOMED-CT terms, and we did not find appropriate SNOMED-CT terms for ICD grouping terms.
[目的/意义]探究影响突发公共卫生事件网络虚假信息传播行为的因素,为有针对性地监测预警、阻断传播突发公共卫生事件网络虚假信息提供借鉴参考.[方法/过程]基于S-O-R理论模型,综合考虑外部刺激因素和个体认知因素对传播者信任感知及传播行为的影响,提出相关研究假设,并利用新冠肺疫情期间微博平台的虚假信息进行验证.[结果/结论]突发公共卫生事件的虚假信息质量、信息发布者影响力和事件进展对信息接收者的信息传播行为具有显著促进作用,医护人员主题、呈现积极情感的网络虚假信息更容易获得信任和传播,网络影响力高的信息接收者在虚假信息传播过程中起到辟谣作用.
目的 对近十年甲状腺癌研究现状进行文献计量分析.方法 以中国知网为数据源,检索2011年1月至2020年12月发表的甲状腺癌主题的期刊文献作为实证数据,分析研究主题分布、关键词、基金、机构、作者共现情况,构建多维科研网络,开发原型系统并在学术社会网络分析方面描绘和模拟其应用.结果 共纳入分析文献6790篇,研究主题以甲状腺癌、手术类居多,最高频共现关键词群以甲状腺肿瘤为核心,资助基金以国自然和省自然居多,发文机构主要是知名医院且合作较多,形成多个作者群.结论 研究主题范围较集中,前沿性研究较少,形成多个科研团队且具有地域特点,科研基金资助机构以政府为主.
With the rapid development of bibliographical data of biomedical articles, it is hard for scientists to keep up with the most recent biomedical literatures. Biomedical relation extraction aims to uncover high-quality relations from biomedical literature with high accuracy and efficiency. Of the existing text mining tools and semantic web products for relation extraction, knowledge graph, a large scale semantic network consisting of entities and concepts as well as the semantic relations among them, has enriched information for human annotation and thus has a great potential for assisting the extraction of the new relations. In this paper, we propose a knowledge graph based biomedical relation extraction framework KGBReF and apply the framework to explore emotion-probiotic relations. A probiotics knowledge graph with 40, 442, 404 triples was built and candidate relations in totally 1,453 PubMed articles were further retrieved by reasoning and annotated. Further, the evidence levels of relations were retrieved and visualized. Finally, we got an evidenced emotion-probiotic relation graph. KGBReF demonstrates an effective reasoning based framework of relation extraction by defining top concepts only. The annotated relation associations are supposed be used to help researchers generate scientific hypotheses or create their own semantic graphs for their research interests.
BACKGROUND:The coronavirus disease (COVID-19), a pneumonia caused by severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) has shown its destructiveness with more than one million confirmed cases and dozens of thousands of death, which is highly contagious and still spreading globally. World-wide studies have been conducted aiming to understand the COVID-19 mechanism, transmission, clinical features, etc. A cross-language terminology of COVID-19 is essential for improving knowledge sharing and scientific discovery dissemination.METHODS:We developed a bilingual terminology of COVID-19 named COVID Term with mapping Chinese and English terms. The terminology was constructed as follows: (1) Classification schema design; (2) Concept representation model building; (3) Term source selection and term extraction; (4) Hierarchical structure construction; (5) Quality control (6) Web service. We built open access for the terminology, providing search, browse, and download services.RESULTS:The proposed COVID Term include 10 categories: disease, anatomic site, clinical manifestation, demographic and socioeconomic characteristics, living organism, qualifiers, psychological assistance, medical equipment, instruments and materials, epidemic prevention and control, diagnosis and treatment technique respectively. In total, COVID Terms covered 464 concepts with 724 Chinese terms and 887 English terms. All terms are openly available online (COVID Term URL: http://covidterm.imicams.ac.cn ).CONCLUSIONS:COVID Term is a bilingual terminology focused on COVID-19, the epidemic pneumonia with a high risk of infection around the world. It will provide updated bilingual terms of the disease to help health providers and medical professionals retrieve and exchange information and knowledge in multiple languages. COVID Term was released in machine-readable formats (e.g., XML and JSON), which would contribute to the information retrieval, machine translation and advanced intelligent techniques application.
With the rapid development of bibliographical data of biomedical articles, it is hard for scientists to keep up with the most recent biomedical literatures. Biomedical relation extraction aims to uncover high-quality relations from biomedical literature with high accuracy and efficiency. Of the existing text mining tools and semantic web products for relation extraction, knowledge graph, a large scale semantic network consisting of entities and concepts as well as the semantic relations among them, has enriched information for human annotation and thus has a great potential for assisting the extraction of the new relations. In this paper, we propose a knowledge graph based biomedical relation extraction framework KGBReF and apply the framework to explore emotion-probiotic relations. A probiotics knowledge graph with 40, 442, 404 triples was built and candidate relations in totally 1,453 PubMed articles were further retrieved by reasoning and annotated. Further, the evidence levels of relations were retrieved and visualized. Finally, we got an evidenced emotion-probiotic relation graph. KGBReF demonstrates an effective reasoning based framework of relation extraction by defining top concepts only. The annotated relation associations are supposed be used to help researchers generate scientific hypotheses or create their own semantic graphs for their research interests.