
[Objective]This paper aims to analyse the differences in combating hallucinations in large language models between unstructured knowledge,exemplified by knowledge base resources,and structured knowledge,exemplified by knowledge graph resources,using the Traditional Chinese Medicine(TCM)Q&A domain as a case study.Based on these findings,strategies for improving the ability of large language models to combat hallucinations in vertical domains are discussed.[Methods]The study designs experiments using external knowledge combined with prompt engineering techniques to analyse the differences in prompt effects between knowledge base resources and knowledge graph resources in the TCM Q&A domain.It also investigates the superiority of dynamic triplet strategies and integrated fine-tuning strategies in optimising large language models against hallucinations.[Results]Experimental results show that compared to prompts from unstructured knowledge in the knowledge base,prompts from structured knowledge in the knowledge graph perform better in terms of precision,recall and F1 score,improving by 1.9%,2.42%and 2.2%respectively to reach 71.44%,60.76%and 65.31%.Further analysis of the optimisation strategies shows that the combination of the dynamic triplet strategy and fine-tuning had the best effect against hallucinations,achieving precision,recall and F1 scores of 72.47%,65.87%and 68.62%respectively.[Limitations]This study is limited to a single field,as it was only tested in the field of Traditional Chinese Medicine Q&A,and its generalisability needs to be validated in a wider range of scientific fields.[Conclusions]This study has demonstrated that in the field of Traditional Chinese Medicine,structured knowledge from knowledge graphs outperforms traditional unstructured knowledge in reducing hallucinations and improving the accuracy of model responses.It demonstrates the critical role of structured knowledge in enhancing model comprehension skills.The integration of fine-tuning strategies with knowledge resources provides an effective way to improve performance in large language models.This paper provides a theoretical rationale and methodological support for integrating external knowledge into large language models to improve knowledge performance.
[Objective]This paper aims to quantify sentence alignment scores for parallel corpora of low-resource languages,obtain high-quality parallel corpora,and improve machine translation performance.[Methods]We proposed NeuroAlign,a neural network-based unsupervised sentence embedding alignment scoring method.Parallel sentence pairs were embedded into the same vector space,alignment scores for candidate sentence pairs in the parallel corpus were calculated,and low-scoring sentence pairs were filtered out based on score ranking.Finally,we obtained high-quality bilingual parallel corpora for low-resource languages.[Results]In the BUCC2018 parallel text mining task,the F1 score improved by 0.5%~0.8%.In the CCMT2021 low-resource language neural machine translation task,the BLEU score improved by 0.1-10.9.The sentence alignment scores closely approximated human evaluation.[Limitations]Due to the scarcity of low-resource bilingual parallel corpora,our research was limited to Tibetan-Chinese,Uyghur-Chinese,and Mongolian-Chinese language pairs.[Conclusions]This new method can effectively increase sentence alignment scoring for low-resource language machine translation parallel corpora,improving the corpus quality at the data source level and enhancing machine translation performance.
[Objective]This paper reviews the research progress on semantic novelty in China and abroad.It explores relevant techniques and provides references for future studies.[Coverage]We used keywords such as"novelty of the literature","semantic novelty","literature novelty",and search expressions like"semantic novelty and literature evaluation"to retrieve literature from Web of Science,Elsevier,Springer,Google Scholar,as well as Chinese databases like CNKI,Wanfang,and VIP.A total of 70 representative literature were selected for review.[Methods]We summarized research on semantic novelty,focusing on the definition of novelty,evaluation indicators,and different evaluation methods.We also discussed the current development status and future trends of evaluating semantic novelty in scientific literature.[Results]Semantic novelty evaluation has gradually received widespread attention from the academic community.Related studies have evaluated semantic content without establishing a unified measurement index.[Limitations]Existing evaluation of literature novelty mainly focuses on external features.Fewer research papers directly addressed semantic novelty,limiting support for reviews.[Conclusions]The evaluation of semantic novelty in scientific literature fundamentally lies in the novelty of semantic content.Quantitative research has become the mainstream method,but the calculation method of evaluation indicators needs to be clarified.Future studies on novelty evaluation should combine qualitative and quantitative methods for more comprehensive evaluations.
[Objective]By constructing a three-dimensional research framework of"scenario-problem-method"for altmetrics,this article aims to enrich the research design of altmetrics analysis and promote the healthy and sustainable development of altmetrics.[Methods]Drawing on mature frameworks in science of science and informetrics,and combining the characteristics of altmetrics,a research framework is constructed from four dimensions:application scenarios,research questions,key methods,and exploratory methods.[Results]The application scenarios of altmetrics are summarized in scientific evaluation,scientific communication and knowledge diffusion.Specifically,indicator application,influencing factors and indicator construction are proposed for scientific evaluation scenarios.Communication strategies,communication structures,communication trends,and science-social interaction are proposed for scientific communication scenarios.And research questions on knowledge diffusion include diffusion strategy,diffusion structure and diffusion effect.And we combine three key analytical methods of causal inference,network analysis,and machine learning to explain each research design according to the problem.[Limitations]The frameworks proposed in this study faces challenges in terms of operability and implementation.Further empirical testing is needed in the future.[Conclusions]The framework proposed in this article is conducive to promoting altmetrics to enter the connotative development phase.
[Objective]This paper aims to improve the accuracy of automatic extraction of technical words and function effects of patents.[Methods]First,ChatGPT is used as the Teacher-model,and ChatGLM3 is used as the Student-model.Through knowledge distillation,the training data extracted by ChatGPT are used to fine-tune ChatGLM3,resulting in multiple technical word extraction models and a function word extraction model.These models are performed to extract technical words and function words from the abstract,the first claim,and the technical effect segments of patents,respectively.[Results]Compared to ChatGPT,the fine-tuned technical word extraction models and the function word extraction model show higher accuracy and lower recall rates.The ChatGLM3 fine-tuning model of the first claim has the highest accuracy of 0.734 and Fl values of 0.724,respectively.The accuracy of the function word extraction model reached 0.649,which was higher than the accuracy of the commercial tool's 0.530.[Limitations]This study needs to be further optimized in the following aspects.The technical field and patent language are single,the amount of verification data is small,and the data cleaning rules are not comprehensive enough.[Conclusions]This research scheme improves the accuracy of large language models in automatically extracting technical effects through knowledge distillation operation.Additionally,this study supports mining cutting-edge innovative and hotspot technologies from patents,facilitating higher quality intelligent patent analysis.
近日,一项新的研究表明,来自公众智能手机中的匿名GPS数据可以用于监测公众对公园及城区的其他绿地的使用情况,这有助于为管理决策提供信息。城区的公园和其他绿地具有非常重要的作用,包括促进人类身心健康、保护生态系统生物多样性以及雨水管理和降低热量等。人类与绿地的互动会影响这些功能,但是,以足够精细的分辨率去捕捉人类活动以进行绿地的管理是非常困难的。
疫情带来的不仅仅是疾病的传播,随着感染人数的增加,有关该疾病的信息,包括如何发现它以及如何预防,会在受疫情影响地区的人群中迅速传播。然而,人们对流行病传播过程和流行病相关信息的传播这两者之间的相互作用知之甚少。近期,发表在Chaos上的一篇文章中,研究人员开发了一个模型,通过疾病传播和信息传播两个视角来研究流行病,旨在了解在流行病发生时如何更好地传播可靠的信息。研究人员发现他们所提出的两层模型可以预测大众媒体和感染预防信息对疫情阈值的影响。
互联网时代,大多数人在就医之前都会提前在网上搜索并研究自己的症状,搜索引擎在就医决策过程中发挥着重要作用.未来大语言模型聊天机器人集成到搜索引擎之后,可能会增加用户对聊天机器人所给出的答案的信心.但是,大语言模型已经证明在被提问医学问题时可能提供极其危险的信息.
[目的]高效准确地识别新兴技术,帮助政府、企业等市场各参与主体及时洞察技术前沿并合理配置资源.[方法]本研究以细粒度的技术术语为研究对象,在考虑共词网络结构特征和语义表示的基础上,构建模型进行新兴术语的遴选和新兴分数的量化,并运用Node2Vec图表示学习算法对新兴术语的向量进行编码及语义表示,实现了新兴术语和新兴技术主题的识别.[结果]在数控机床领域进行实证研究,共识别出449个新兴术语以及4个新兴技术主题(机器人自动上下料系统、清洁高效切削加工技术、高速高精度数控加工中心、增减材复合制造技术),验证了所提方法的科学性和合理性.[局限]仅使用专利文献的数据,对其他多源异构文献数据及其中存在的引用、语义相似等其他网络关系利用不足.[结论]运用共词和Node2Vec图表示学习的方法可深入挖掘技术术语间共词网络结构特征和语义表示,实现了新兴技术的细粒度精准量化识别.
[目的]在没有充足标注数据支持模型学习的情况下将对话语言理解任务应用到领域更新频繁的对话系统中.[方法]提出基于信息增强的小样本对话语言理解联合模型(IAM-FSLU),利用小样本学习很好地解决了在新领域和跨领域下的意图种类及数量不同时,数据匮乏和模型适用性差的问题,同时构建了一种更有效的小样本意图识别和小样本槽位提取两个任务间的显式关系.[结果]联合建模与未联合建模模型相比,在1-shot设置下,槽位提取F1分数获得近30个百分点的提升,句准确率有近10个百分点的提升;在3-shot设置下,槽位提取F1分数获得近35个百分点的提升,句准确率有12~16个百分点的提升.[局限]从结果来看,IAM-FSLU在意图识别子任务上仍需要进一步提高性能,同时与隐式关系建模的模型相比,虽然在槽位提取任务上有很大提升,但句准确率提升效果有限.[结论]通过不同的小样本设置的对比实验验证IAM-FSLU的效果,结果表明IAM-FSLU整体效果均优于其他主流模型.
近日,发表在《美国国家科学院院刊》上的一篇文章中,研究人员发现了一种简单可靠的测试方法,可以检测其他行星上过去或现在是否有生命迹象,基于人工智能的方法能将现代和古代生物样本与非生物样本区分开来,准确率达到90%.
[目的]现有少样本知识图谱补全方法在处理复杂关系时不能很好地区分邻居重要性,导致实体预测性能不佳.考虑充分利用实体邻居信息,提高少样本知识图谱补全方法的性能.[方法]通过类型感知邻居编码器学习实体邻居中包含的隐含类型信息,得到类型感知注意力,增强实体表示;利用Transformer编码器捕获任务关系的不同含义;通过联合匹配原型网络聚合参考集得到参考集表示并进行实体预测.[结果]在NELL和Wiki两个公共数据集上进行实体预测任务,实验结果表明,MRR指标分别较Baseline方法提高了1.6和1.2个百分点.[局限]未对与实体相关性较低的邻居进行筛选,使类型感知注意力权重的分配受到噪声影响.[结论]本文方法能够通过学习更丰富的实体邻居信息来有效提高少样本知识图谱补全的性能.
[目的]利用迁移学习和多任务学习解决中文医学文献实体识别冷启动和边界定位难的问题,进一步提高识别准确性.[方法]提出一种基于迁移学习和多任务学习的中文医学文献实体识别方法,构建混合深度学习BERT-BiLSTM-IDCNN-CRF的医学文献实体识别模型,通过实例迁移、模型迁移和特征迁移丰富医学语义特征,利用多任务学习构建粗粒度三分类任务以辅助实体识别任务有效利用实体边界信息,最后引入自注意力机制和Highway网络捕获全局重要信息并优化深层网络训练,提出TLMT-BBIC-HS模型.[结果]TLMT-BBIC-HS模型在中文糖尿病医学文献数据集上F1值达92.98%,较基准模型BERT-BiLSTM-CRF和BERT-IDCNN-CRF分别提高15.99个百分点和16.44个百分点.[局限]未验证模型的领域适应性.[结论]TLMT-BBIC-HS模型可实现医学知识的迁移共享,更适用于中文医学文献实体识别任务,可为医疗健康信息抽取、知识图谱和问答系统构建提供有效支持.
虽然大语言模型聊天机器人ChatGPT的公开发布让人们对这项技术的前景和人工智能的广泛使用感到非常兴奋,但同时也引发了人们的担忧:ChatGPT能在几秒钟内写出一篇合格的大学水平论文,这对未来教育意味着什么?这种恐慌导致了检测程序的激增——其效果各不相同——作弊指控也越来越多.那学生们对这一切的感受如何?德雷塞尔大学的Tim Gorichanaz博士最近发表在Learning:Research andPractice的一项研究首次揭示了被指控使用ChatGPT作弊的大学生的一些反应.
[目的]直观、全面地刻画元宇宙概念所引发的舆情态势及其变迁,为元宇宙相关政策与产业规划提供借鉴.[方法]基于2021年9月-2023年2月元宇宙相关微博文本数据,采用BERT模型和DTM模型抽取其语义和主题特征,借助K-means算法实现主题聚类,解读元宇宙话题的演化规律.[结果]大众对元宇宙的关注焦点发轫于非同质化代币(Non-Fungible Token,NFT)和游戏,随着数字产业的资本炒作,进一步引发文娱产业的跟进以及实体产业的尝试.而ChatGPT的出现则引发了大众对元宇宙产业现状、技术创新和应用展望的进一步探讨.[局限]未结合外文数据(如Twitter)对比分析国内外对元宇宙话题关注点的侧重、趋势等方面的差异.[结论]本研究从定量与宏观的角度解读了元宇宙相关话题的社会关注度特征及演化规律,对正确引导元宇宙网络舆情走向、避免舆论泡沫等工作具有一定参考借鉴意义.
[目的]为促进科研人员之间的交流合作,提出一种融合异质网络与表示学习的科研合作预测方法.[方法]运用学者、机构、论文、期刊等信息构建异质科研合作网络,根据网络中包含的学者之间不同的共现关系,将该异质网络划分为三种同质共现网络,再进一步利用Node2Vec和Doc2Vec算法分别学习学者的网络结构特征向量和内容属性特征向量,并进行融合.最后通过计算学者向量之间的余弦相似度进行合作预测.[结果]采用Web of Science数据库中人工智能领域的论文数据进行对比实验,本文所提预测方法的AUC值和F1值分别达到0.9879和0.9424,优于基线方法.[局限]对学者内容特征的表示没有考虑到学者的研究主题.[结论]本文方法考虑了学者的结构和内容属性,并结合异质网络,融合了机构、论文、期刊等多方面信息,能够得到更好的合作预测效果.
[目的]基于多源异构数据构建中医药知识图谱,辅助研究人员进行中医药领域的创新研究.[方法]从IncoPat专利数据库获取中医药专利数据,从TCMSP、OMIM等数据库获取中药靶点、疾病等数据,利用深度学习信息联合抽取模型抽取中医药专利文本中的实体及关系,采用字符串匹配和词典等方式进行数据规范及实体对齐,进而基于所设计的中医药知识图谱本体结构完成知识图谱构建,在此基础上采用频次分析、关联规则Apriori算法对中药处方优化进行分析.[结果]本文所设计的本体结构共包含31种实体类型、48种语义关系,涵盖中医药领域专利中的解决方案、技术功效等特定实体;选取糖尿病肾病领域具体详解基于多源数据的中医药知识图谱构建及应用过程,验证了本文所构建知识图谱的有效性以及对处方优化提供中医药筛选范围的高效性.[局限]在专利文本信息抽取时,部分标注样本采用人工标注,耗费时间较长.[结论]以中医药专利数据为主、结合多源数据所构建的中医药知识图谱,能够为中医药领域创新研究提供数据支撑,该知识图谱不仅可以实现处方优化研究,也可用于中医药领域的多元研究.
[目的]分析基于大规模语言模型的提示学习方法在学术论文实体识别任务上的可用性.[方法]以ChatGPT这一大规模语言模型为例,将ChatGPT视为实体识别工具、伪标签生成工具以及训练数据生成工具,从性能、价格和时间等维度出发分析以上三个视角下ChatGPT的可用性.[结果]三个视角下基于ChatGPT的方法的F1值高于少量样本训练得到的神经网络基线模型,比如实体识别工具视角的F1宏平均值超过10个学术论文人工标注摘要训练得到的模型21.4个百分点.基于ChatGPT的方法在不同学科领域的学术论文数据集上性能较稳定.[局限]仅在英文学术论文摘要数据集上展开实验,但中文与英文学术论文、学术论文摘要与全文存在逻辑结构和表述上的差异.[结论]当缺少人工标注数据时,将ChatGPT视为实体识别工具可从学术论文摘要中识别出部分实体,但识别结果需进一步过滤以应用到下游任务中.
[目的]评估ChatGPT在中文命名实体识别、关系抽取以及事件抽取等典型中文信息抽取任务中的性能,分析不同任务和领域ChatGPT的表现差异,给出ChatGPT中文场景下的使用建议.[方法]采用Prompt提示的方式,分别依据精确匹配和宽松匹配两种方式,测评ChatGPT在三个典型信息抽取任务、共7个数据集上的性能:在MSRA、Weibo、Resume和CCKS2019数据集评估ChatGPT的命名实体识别效果,并与GlyceBERT和ERNIE3.0模型对比;在FinRE和SanWen数据集测试ChatGPT与ERNIE3.0 Titan的关系抽取效果;在CCKS2020数据集测试ChatGPT与ERNIE3.0的事件抽取效果.[结果]ChatGPT在命名实体识别任务中的表现不及GlyceBERT和ERNIE3.0模型.在关系抽取任务中,ERNIE3.0 Titan优于ChatGPT.在事件抽取任务中,ChatGPT在宽松匹配下的表现优于ERNIE3.0.[局限]以Prompt提示的方式评估ChatGPT的性能表现存在主观性,不同的Prompt会产生效果差异.[结论]ChatGPT在典型的中文信息抽取任务上的表现还有很大改进空间,用户在使用过程中需选择合适的Prompt和问题.
近日,纽约大学发表在Scientific Reports期刊上的一项研究发现,听音乐和喝咖啡这样的日常乐趣会影响人类的大脑活动,提高人类在需要集中注意力和记忆力的任务中的认知能力.该研究利用开创性的大脑监测技术MINDWATCH算法来进行.MINDWATCH算法能通过任何可以监测皮电活动(Electrodermal Activity)的可穿戴设备收集的数据来分析人的大脑活动.这种活动反映了情绪压力引发的电导变化.