This study focuses on the task of term-domain short-text similarity computation. It addresses two main challenges: dataset scarcity and insufficient deep semantic extraction. To solve these issues, we first construct the multi-domain Terminology Definition Similarity (TDS) dataset using an automated pipeline. This pipeline combines data generation (based on the GPT-4 model) with a manual verification process involving expert quality control. The design ensures the production of high-quality data. We then present an innovative model named SDQKC (Sbert + dynamic QK + contrast). The model optimizes the Siamese BERT network through a dynamic QK co-attention mechanism and enhances its deep-level semantic understanding by incorporating contrastive learning. Experimental results show that the SDQKC model achieves Pearson correlation coefficients of 0.69384 on the BQ Corpus (Bank Question Corpus) and 0.69511 on the TDS dataset. These results significantly outperform other baseline models, demonstrating the effectiveness and superiority of the proposed methodology.
With the rapid development of the Internet and social media, massive amounts of unstructured data have emerged, making event extraction increasingly important for information retrieval. In the financial domain, challenges such as long texts, redundant content, and complex structures hinder extraction tasks. To address this, we propose PosEKE-GPT2, an improved GPT2-based model that reformulates event extraction as a text generation task. The model jointly identifies event types, triggers, and arguments using structured canonical text and a sub-task extraction strategy to reduce error propagation. An expanded positional encoding mechanism enhances event representation in long texts. Furthermore, we introduce a knowledge augmentation module that dynamically selects and integrates external knowledge via prompt mechanisms and attention-based embedding optimization. Experiments on the DuEE-Fin dataset show that PosEKE-GPT2 achieves an average F1-score of 90.61, while on the FewFC dataset it reaches an average F1-score of 88.85, both outperforming baseline models. Ablation studies verify the effectiveness of the positional encoding and knowledge augmentation modules, demonstrating the model’s robustness and suitability for financial event extraction across different datasets.
As internet information expands rapidly, extracting valuable event information from unstructured text has become an important research topic. This paper proposes an improved GPT2 model, termed HACLV-GPT2, which is the initial utilization of a GPT-like architecture for the purpose of event extraction. The model utilizes a generative input template and incorporates a hybrid attention mechanism to enhance the understanding of complex contexts. Additionally, the HACLV-GPT2 model employs a layer-vector fusion strategy to optimize the output of Transformer Blocks, effectively boosting prediction performance. The experimental results show that the HACLV-GPT2 model performs excellently in both event argument extraction and event type detection tasks, with F1 values of 0.8020 and 0.9614, respectively, surpassing several baseline models. This outcome fully validates the effectiveness and superiority of the proposed method. Furthermore, ablation experiments confirm the critical role of the hybrid attention mechanism and layer-vector fusion strategy in performance improvement.
In the research of aspect-level sentiment analysis, traditional neural networks mainly use attention mechanisms to focus on the weight information of words, the grammatical structure of sentences and the dependency information of aspect words as well as part-of-speech information are often ignored in the feature extraction task, and most of the existing mainstream schemes use models such as Word2Vec and GloVe to complete the word embedding work, and the static word vectors generated by these models can only capture the semantic information of words, but cannot reflect the contextual relationship of words. In this paper, we propose an aspect-level sentiment analysis framework based on XLNet-BiLSTM-ATT. The framework uses XLNet to learn the global word contextual representation for word embedding, and adopts an extra-long version of the Transformer feature extractor, which is able to better capture global dependencies; BiLSTM is used to encode the information after word embedding to get the hidden vectors of words; the attention scores of the part-of-speech and dependency labels of the data are computed at the attention layer separately to obtain the feature representation; finally, these two features are fused and the classification output is performed by Softmax layer. Experimental results on three publicly available datasets show that the framework shows significant improvement in both accuracy and F1 value.
Word Sense Disambiguation (WSD) is a significant and challenging task for text understanding and processing. This paper presents an unsupervised approach based on Weighted Co-occurrence bio-Term Graph (WCOTG) for performing WSD in the biomedical domain. The graph is automatically created from biomedical terms that are extracted from a corpus of downloaded scientific abstracts. Two kinds of weights are introduced on the links of the built bio-term graph and are taken as important factors in the process of disambiguation. The modified Personalised PageRank (PPR) algorithm is used for performing WSD. When evaluated on the NLM-WSD and MSH-WSD test datasets, and an acronym test set, the method outperforms the widely used unsupervised ones addressing the same problem, and the average result is almost equal to that of the BlueBERT_LE-based method. In contrast, our method has no additional enhancement or training for BERT-based models. Comparative experiments validate the positive effect of links' weight on disambiguation efficiency. Last, the statistical experiments on the relation among system accuracy, the numbers of medical abstracts in the corpus, and the corresponding extracted terms suggest an excellent minimum corpus scale, when resources are limited.
目的:运用数据挖掘技术,分析阴陵泉穴的优势病症与配伍处方,总结其主治特点与配伍规律.方法:在中国知网、万方学术期刊全文、维普中文期刊资源、SinoMed、Web of Science和PubMed等数据库中检索阴陵泉在1949-10-1到2022-3-31期间的相关资料,建立SQL Server数据库,运用"五输穴主治与配伍数据挖掘软件V1.0"及Gephi可视化软件对数据进行复杂网络、复合关联和聚类分析.结果:共纳入文献548篇.阴陵泉共6种单穴主治病症,优势病症包括肩关节周围炎和尿潴留;133种配伍主治病症,优势病症包括尿潴留、膝关节炎、糖尿病和脑卒中后遗症等19种.高频腧穴包括三阴交、足三里、血海、阳陵泉和关元等30穴.配伍经脉以足阳明胃经、足太阴脾经和任脉等为主.复合关联经筛选后得出45条阴陵泉配伍的腧穴-病症关联规则.聚类分析产生了3类7组高频配伍穴类组合.结论:阴陵泉主治方面主要以腧穴所在部位之膝关节炎、经脉所过部位之肢体关节病症、脾胃系及肾系相关脏腑病症和湿邪所致病症等为主,腧穴配伍主要以本经、交接经和表里经等方式为主,其结果可为临床和科研提供参考.
目的:运用现代统计学和数据挖掘技术,探析阳陵泉穴现代主治优势病症和配伍规律.方法:以中华人民共和国成立为时间节点,通过检索阳陵泉穴主治病症及腧穴配伍等相关文献,建立SQL Server数据库,运用"五输穴主治与配伍数据挖掘软件V 1.0"对数据进行频次、聚类和关联规则分析.结果:中华人民共和国成立后,阳陵泉穴的单穴主治优势病症4种,分别为胆绞痛、胁痛、肩周炎和筋伤;配伍主治优势病症10种,分别为痹证、痿证、中风、坐骨神经痛、中风后遗症、蛇串疮、胁痛、肩周炎、腰痛和腰椎间盘突出症;配伍腧穴归经前3 位分别为足阳明胃经、足太阳膀胱经、足少阳胆经;关联分析得出配伍腧穴以足三里、三阴交和合谷关联程度最高,聚类分析得出3 个有效聚类群.结论:通过对现代文献的挖掘分析,揭示阳陵泉穴主治优势病症和配伍规律,可为今后该穴的临床应用提供参考.
Event extraction stands as a significant endeavor within the realm of information extraction, aspiring to automatically extract structured event information from vast volumes of unstructured text. Extracting event elements from multi-modal data remains a challenging task due to the presence of a large number of images and overlapping event elements in the data. Although researchers have proposed various methods to accomplish this task, most existing event extraction models cannot address these challenges because they are only applicable to text scenarios. To solve the above issues, this paper proposes a multi-modal event extraction method based on knowledge fusion. Specifically, for event-type recognition, we use a meticulous pipeline approach that integrates multiple pre-trained models. This approach enables a more comprehensive capture of the multidimensional event semantic features present in military texts, thereby enhancing the interconnectedness of information between trigger words and events. For event element extraction, we propose a method for constructing a priori templates that combine event types with corresponding trigger words. This approach facilitates the acquisition of fine-grained input samples containing event trigger words, thus enabling the model to understand the semantic relationships between elements in greater depth. Furthermore, a fusion method for spatial mapping of textual event elements and image elements is proposed to reduce the category number overload and effectively achieve multi-modal knowledge fusion. The experimental results based on the CCKS 2022 dataset show that our method has achieved competitive results, with a comprehensive evaluation value F1-score of 53.4% for the model. These results validate the effectiveness of our method in extracting event elements from multi-modal data.
目的:探析解溪穴现代主治优势病症和配伍规律.方法:以中国知网、万方数据知识服务平台、PubMed和Web of Science等数据库为检索源,检索1949年10月-2022年3月期间相关文献,根据纳入、排除标准筛选文献,建立Microsoft Access数据库,将数据导入"五输穴主治与配伍数据挖掘软件V1.0"进行关联分析和聚类分析.结果:共纳入文献102篇.其中主治病症27种,优势病症为中风、踝关节扭伤和痹病等7种;高频配伍腧穴为足三里、阳陵泉、昆仑和三阴交等26穴,其中解溪与足三里、阳陵泉和太冲关联程度最高;得到了以解溪穴为核心的5类配伍处方;配伍经脉以足少阳胆经频次最多.结论:通过对现代文献挖掘分析,揭示了解溪穴主治优势病症和配伍规律,对教学、临床和科研有所裨益.
Biomedical question answering refers to extracting an answer based on given questions and related documents. Existing biomedical question answering research either focuses on a specific stage, such as machine reading comprehension, or uses traditional rule-based methods and ontology with complex construction processes. In this paper, we demonstrate the application of simple but powerful neural-based approaches in improving the end-to-end biomedical question answering system. We employ the BM25-based documents retriever, BERT-based neural ranker, and an answer extraction stage using the BioBERT pre-trained language model. In view of the lack of sufficient training data in the biomedical domain, domain adaptation and data augmentation are adopted to address the question answering task, so as to further reinforce the system performance. Based on our self-built standard large-volume retrieve corpus and neural ranker corpus, we get competitive results on BioASQ8b.
Objective: To explore the clinical application rules of Laog o over line ng (PC8) by analyzing the comparison and contrast of the practice with this acupoint in ancient times and in present.& nbsp;Methods: Relevant literature before October 1, 1949 was selected from Chinese Medical Code (Fifth Edition) and Compilation of Modern Chinese Medicine Journals and literature published after October 1, 1949 was screened from China National knowledge Infrastructure (CNIK), Wanfang Data Knowledge Service Platform (Wanfang), VIP Chinese Journal Service Platform (VIP), Chinese BioMedine Database (CBM), Pubmed Database, and Web of Science. After screening, SQL Server database was established. The complex network analysis was conducted by using Gephi and the cluster analysis was performed by SPSS Statistics.& nbsp;Results: Before October 1, 1949, Laog o over line ng (PC8), on a single use or with a match of Daling (:kPC7), Gu anyuan (jG CV4), and Zus anl i ( AST36), was mainly used for treating ozostomia and edema due to pregnancy. However, Laog o over line ng (APC8), on a single use or with a match of Y ongquan ( it*KI1), Heg u ( p - LI4) and Neigu an (PC6) was commonly for the treatment of toothache and the sequelae after wind stroke after October 1st, 1949.& nbsp;Conclusion: Laog o over line ng (PC8) is mainly used to treat local disorders or diseases pertaining to the pericardium meridian. The single-point therapy was commonly applied before October 1st, 1949, while a combination of other points is more common afterwards and the scope of indication expands, with its application for psychiatry disorders, pain diseases and emergent diseases. Acupoints on the same meridian of Laog o over line ng ( PC8), ying-spring points, the back-shu points and combination between and the upper and the lower are dominated in the selection of supplementary points. The supplementary acupoints with the highest frequency of use were specific acupoints, including five-shu points, yuan-source points, front-mu points, etc. (C)& nbsp;2021 Published by Elsevier B.V. on behalf of World Journal of Acupuncture Moxibustion House.
OBJECTIVE To explore the dominant indications and laws of acupoint compatibility of Yinbai (SP1) by using modern statistics and data mining techniques. METHODS Literature about indications and acupoint prescriptions of SP1 published before October of 1949 were retrieved from books Chinese Medical Dictionary (5th edition) and Collection of Modern Medical Journals of Traditional Chinese Medicine, and those published from October 1st of 1949 to January of 2021 retrieved from databa-ses of CNKI, Wanfang, VIP, CBM, Web of Science and Pubmed by using key words of Yinbai (SP1),"Guilei"(),"Guiyan"() and Jing (Well)-point of Spleen Meridian, followed by screening the data and establishing a SQL Server database after standardized processing. Then, the descriptive analysis, clustering analysis and association rule analysis were conducted by using Gephi visualization software, SPSS Statistics 25.0 and SPSS Modeler, separately. RESULTS Before October of 1949, the single SP1 acupoint was usually used to treat 12 types of diseases (mainly the internal diseases as asthma, abdominal distension, vomiting, etc.), and the compound prescriptions of SP1 were usually used to treat 20 types of diseases (mainly the internal diseases as insomnia and dreamful sleep, blood syndrome, etc.), and its adjunct acupoints belong to the first three meridians: the Foot Yangming Stomach Meridian, Foot Taiyang Bladder Meridian and Foot Taiyin Spleen Meridian. After October of 1949, the single SP1 was used to dominantly treat 2 types of diseases (mainly the gynecological diseases as metrorrhagia and metrostaxis, and hypermenorrhea, etc.), and the compound prescriptions of SP1 were frequently used to treat 10 diseases (metrorrhagia and metrostaxis, sequela of apoplexy, mental disorders, insomnia and dreamful sleep, etc.), and the adjunct acupoints of compound prescriptions belong to the first three meridians, namely the Foot Taiyin Spleen Meridian, Concept Vessel and Foot Yangming Stomach Meridian. Before and after October of 1949, the adjunct acupoints with the highest degree of correlation were Lidui (ST45), Shaoshang(LU11), Zusanli(ST36), Sanyinjiao (SP9), and Guanyuan (CV4). Cluster analysis showed that 9 effective clusters obtained may be used as potential prescriptions of SP1, and association rule analysis displayed that the first three strongly connected acupoint matching groups were: SP1-ST45, SP1-LU11, and SP1-ST36 frequently used before October of 1949, and SP1-SP9, SP1-ST36 and SP1-CV4 employed after October of 1949. CONCLUSION Data mining technology reveals that acupoint SP1 alone is mainly used to treat internal diseases before 1949, and gynecological diseases after 1949; and compound acupoint recipes of SP1 are mainly to treat the internal diseases before 1949, and the gynecological diseases and mental disorders after 1949 in China. The frequently employed adjunct acupoints of SP1 are ST45, LU11, ST36, SP9 and CV4 both before and after 1949.
Similarity calculation is an important part of natural language processing. The higher the accuracy of the similarity, the better the results for downstream tasks. In previous studies, similarity was calculated using a single knowledge source and the similarity calculation was not satisfactory. Later, computing similarity combining multiple knowledge sources can lead to better similarity computation results. In this paper, we calculate similarity based on multiple knowledge sources and use the simulated annealing algorithm to calculate the damping factor between different knowledge sources to obtain the composite similarity. The error between the composite similarity values of the BioBERT model and the Word2Vec model and the standard similarity values is only 0.03. In the data of the official similarity calculation, the composite similarity can improve the accuracy of the similarity very well.
Objective:To explore the clinical application rules of Láogōng(劳宫PC8)by analyzing the comparison and contrast of the practice with this acupoint in ancient times and in present.Methods:Relevant literature before October 1,1949 was selected from Chinese Medical Code(Fifth Edition)and Compilation of Modern Chinese Medicine Journals and literature published after October 1,1949 was screened from China National knowledge Infrastructure(CNIK),Wanfang Data Knowledge Service Plat-form(Wanfang),VIP Chinese Journal Service Platform(VIP),Chinese BioMedine Database(CBM),Pubmed Database,and Web of Science.After screening,SQL Server database was established.The complex net-work analysis was conducted by using Gephi and the cluster analysis was performed by SPSS Statistics.Results:Before October 1,1949,Láogōng(劳宫PC8),on a single use or with a match of Dàlíng(大陵PC7),Guānyuán(关元CV4),and Zúsānlǐ(足三里ST36),was mainly used for treating ozostomia and edema due to pregnancy.However,Láogong(劳宫PC8),on a single use or with a match of Yǒngquán(涌泉KI1),Hegu(合谷LI4)and Nèiguān(内关PC6)was commonly for the treatment of toothache and the sequelae after wind stroke after October 1st,1949.Conclusion:Láogōng(劳宫PC8)is mainly used to treat local disorders or diseases pertaining to the peri-cardium meridian.The single-point therapy was commonly applied before October 1st,1949,while a combination of other points is more common afterwards and the scope of indication expands,with its application for psychiatry disorders,pain diseases and emergent diseases.Acupoints on the same merid-ian of Láogōng(劳宫PC8),ying-spring points,the back-shu points and combination between and the up-per and the lower are dominated in the selection of supplementary points.The supplementary acupoints with the highest frequency of use were specific acupoints,including five-shu points,yuan-source points,front-mu points,etc.
ObjectiveTo explore the clinical application rules of Láogōng (劳宫PC8) by analyzing the comparison and contrast of the practice with this acupoint in ancient times and in present.MethodsRelevant literature before October 1, 1949 was selected from Chinese Medical Code (Fifth Edition) and Compilation of Modern Chinese Medicine Journals and literature published after October 1, 1949 was screened from China National knowledge Infrastructure (CNIK), Wanfang Data Knowledge Service Platform (Wanfang), VIP Chinese Journal Service Platform (VIP), Chinese BioMedine Database (CBM), Pubmed Database, and Web of Science. After screening, SQL Server database was established. The complex network analysis was conducted by using Gephi and the cluster analysis was performed by SPSS Statistics.ResultsBefore October 1, 1949, Láogōng (劳宫PC8), on a single use or with a match of Dàlíng (大陵PC7), Guānyuán (关元CV4), and Zúsānlĭ (足三里ST36), was mainly used for treating ozostomia and edema due to pregnancy. However, Láogōng (劳宫PC8), on a single use or with a match of Yŏngquán (涌泉KI1), Hégŭ (合谷LI4) and Nèiguān (内关PC6) was commonly for the treatment of toothache and the sequelae after wind stroke after October 1st, 1949.ConclusionLáogōng (劳宫PC8) is mainly used to treat local disorders or diseases pertaining to the pericardium meridian. The single-point therapy was commonly applied before October 1st, 1949, while a combination of other points is more common afterwards and the scope of indication expands, with its application for psychiatry disorders, pain diseases and emergent diseases. Acupoints on the same meridian of Láogōng (劳宫PC8), ying-spring points, the back-shu points and combination between and the upper and the lower are dominated in the selection of supplementary points. The supplementary acupoints with the highest frequency of use were specific acupoints, including five-shu points, yuan-source points, front-mu points, etc.
目的:运用现代统计学和数据挖掘技术,探析少商穴主治优势病症和配伍规律.方法:以1949年10月为时间节点,检索少商穴相关文献资料,建立SQL Server数据库,运用"五输穴主治与配伍数据挖掘软件QLZJ1.0版"对数据进行频次分析、关联规则分析、聚类分析.结果:1949年10月前单穴、配伍主治以内科、耳鼻喉科病症为主,1949年10月后单穴、配伍主治以耳鼻喉科病症为主;1949年10月前,配伍腧穴共183个,主要归经为大肠经、心包经,1949年10月后,配伍腧穴共131个,归经主要为大肠经、督脉;关联规则得出1949年10月前频繁项集4阶17条、1949年10月后频繁项集1阶11条;聚类分析得出1949年10月前、后皆有2组有效聚类群.结论:少商穴的主治优势病症以耳鼻喉科、内科疾病为主;配伍腧穴中与商阳、中冲、关冲、合谷等联系较为密切,配伍方式多见本经、表里经、同名经配穴等经典配伍方式,以及鬼针十三穴、五运升降法和手经十二井穴等特殊配伍方式.
目的:探析厉兑穴主治病症和配伍规律.方法:对新中国成立前后有关厉兑穴的文献进行检索,运用数据挖掘软件Gephi、VOSviewer进行统计学分析.结果:1949年前单穴、配伍主治病症分别为29种、36种,配伍腧穴141个,多归属胃经、大肠经,多使用特定穴,尤其是五输穴;1949年后单穴、配伍主治病症分别为10种、30种,配伍腧穴116个,多归属胃经,多使用特定穴,尤其是五输穴.1949年后较1949年前主治优势病症增加了眼部、脑部病症,优势配穴增加了头面部腧穴、交会穴.结论:厉兑穴主治以本经、脏腑和络脉病症为主,优势病症多发挥其远治作用;配伍规律体现了上下、本经、同名经和井井配穴等.
目的:探析关冲主治优势病症和配伍规律.方法:以1949年10月为时间分割点进行文献检索.1949年10月前的文献以《中华医典》(第五版)、《中国近代中医药期刊汇编》为主要检索源.1949年10月后的文献以中国知网、万方数据库、维普期刊资源数据库、中国生物医学文献数据库、PubMed、Web Of Science为主要检索源.根据纳入、排除标准筛选文献,建立关冲穴数据库,采用频次分析和关联分析方法探析关冲穴主治优势病症和配伍规律.结果:1949年10月前,单穴主治病症30种,优势病症为口干、臂痛等10种;配伍主治病症42种,优势病症为喉痹、霍乱等17种,并总结出新处方17首,配伍腧穴127个,与少商、少冲关联性最强.1949年10月后,单穴主治病症3种,配伍后主治病症增多,为35种,优势病症为中风、乳蛾等8种;总结出新处方8首,配伍腧穴85个,与少商、商阳关联性最强.结论:关冲主治病症以耳鼻喉科、内科病症为主,配伍规律以表里经、本经、五输、井原和井络配穴为主,趋向于按经配伍和特定穴间配伍.
文章通过系统整理1949年以前少冲穴相关文献,利用数据挖掘技术,总结和分析在主治病症方面,少冲穴主要以腧穴所在之手挛指麻、经脉所过之胸腋肘臂痛、急症、热病、脏腑病等方面为主;腧穴配伍常见本经配伍、同名经配伍、交接经配伍、表里经配伍,还可见脏腑别通法、大接经法及手十二井配穴等特殊配穴方法等.
Brand name translation is of great importance for international corporations when their goods enter exotic markets. In this chapter, we investigate the strategies and methods of brand name translation from Western languages to Chinese, and propose a hexagonal pyramid brand name translation model, which provides a comprehensive summary of brand name translation methods and makes the classification of some translated brand names from vague to clear. Based on the model and similarity calculating, an efficient automatic translation method has been proposed to provide help of finding adequate translated words in Chinese. And an experiment has been done by the way of a dedicated program with results of a cluster of recommended Chinese brand words with a good potential to be used.