[目的/意义]事件自动识别抽取是当前典籍主题挖掘研究中一个新的重要课题,其中事件触发词的识别是一项基础的工作,本研究旨在探索古代典籍中事件触发词自动识别和分类的通用方法.[方法/过程]首先运用LDA模型对动词进行主题聚类,归纳典籍事件触发动词的分类体系;并依据聚类结果与分类体系,初步构建触发动词的种子词集.在此基础上,通过语义相似度计算,对种子词集进行扩展,构建典籍事件触发词语义数据集.在实验阶段,以先秦时期的重要典籍《左传》为例,对分类体系构建和种子词集扩展的方法进行验证.[结果/结论]结果表明,本文所提出的典籍事件触发词识别方法可行有效,据此构建的事件触发词集具有较高可信度,未来可进一步扩大实验的样本数量及范围.
Purpose Digital humanities database is one of the essential tools in digital humanities research area. Therefore, examining the usage of digital humanities database in academic papers is conducive to assessing the value of digital humanities database for scientific research activities and improving the construction of digital humanities infrastructure. Design/methodology/approach This paper constructs an evaluation system of digital humanities database from the perspective of academic influence and social influence, with mention frequency, usage motivation, platform access data, usage region and usage discipline as indicators and takes China Biographical Database Project as the empirical object to explore the usage of digital humanities database in China. Findings The data analysis result demonstrates that digital humanities databases are widely used and recognized in China. However, the problem of low actual usage remains. Originality/value This paper constructs the digital humanistic database's evaluation system and discusses applying the digital humanistic database in China, which provides a new perspective and method for the influence evaluation study of the digital humanistic database.
Against the background that the top-level semantic framework of Chinese traditional culture is not comprehensive and unified, this study aims to preserve and disseminate cultural heritage information about Chinese traditional culture through the development of a domain ontology which is constructed from ancient books. A combination of top-down and bottom-up approaches was used to construct the ontology for Chinese traditional culture (CTCO). An investigation of historians’ needs, and LDA topic clustering model were conducted, understanding the specific needs of historians, collecting the topic, concepts and relationships. CIDOC CRM was reused to construct the basic framework of CTCO. Ontology structure and function were adopted to evaluate the effectiveness of CTCO. Evaluation results show that the ontology meets all the quality criteria of OntoMetrics, and the experts agreed on content representation (average score = 4.30). CTCO contributes to the organization of traditional Chinese culture and the construction of related databases. The study also forms a common path and puts forward proposals for the construction of domain ontology, which has great social relevance.
[目的]在信息交流视角下,探讨我国科技期刊知识服务转型发展的突破路径.[方法]通过文献调研和对比分析法,探讨面向科技期刊知识服务的科学信息交流模式框架,对国内外科技期刊知识服务模式进行分析.[结果]分别从正式交流和非正式交流视角,归纳总结国内外科技期刊所开展的知识服务.通过对比国内外科技期刊在服务模式、手段和渠道的差异,提出我国科技期刊知识服务的发展策略.[结论]我国科技期刊应从集团化经营模式、自动化出版流程、基于微信矩阵拓展服务项目等方面入手,完善自身的知识服务体系,实现知识生产向知识服务的创新转型.
[Purpose/Significance] It is of great significance to carry out research on the recognition and classification of trigger verbs in ancient books oriented to digital humanities for the deep mining and content revealing of ancient texts. This paper uses the deep learning classification algorithm to explore an automated method for multivariate classification of event sentence text based on trigger words in ancient books. [Method/Process] Based on the construction of the classic event trigger word classification system and trigger dictionary, four different types of event sentence texts are selected as experimental data, and the category labels and sentence texts are coded separately using Onehot and Tokenizer, and then the classifier is trained in the Bi-LSTM model, and a comparative experiment is set by adjusting the parameters, and the performance of the classifier is analyzed by using a general evaluation index. [Results/Conclusions] The classifier after many training and adjustments has an accuracy of 0.95 in the evaluation of the test set, which proves that the experimental method based on deep learning and the constructed trigger word data set can effectively help us realize automatic multivariate classification of event sentence text of ancient books.
[目的/意义]针对《左传》中的战争事件展开研究,对先秦历史乃至中华民族文化的研究具有重要参考价值.[方法/过程]基于框架理论构建《左传》战争事件基本框架体系,利用模式匹配法进行战争句识别,选择条件随机场模型、结合特征模板对战争时间、交战双方等7个命名实体进行识别和抽取,最后基于得到的结构化数据对战争事件进行分析和可视化展示.[结果/结论]研究结果表明,条件随机场模型能够较好地应用于《左传》战争事件的抽取;特征选取会影响实体识别的结果;具体内容方面,春秋时期晋国、楚国、齐国、郑国等国参战频率较高,晋国为主要进攻方,郑国为主要防守方.
[目的/意义]针对文化遗产语义组织发展现状展开研究,对我国文化遗产研究具有重要参考价值.[方法/过程]采用系统调研法、案例分析法和统计分析法,以调研数据概括为基础,从语义组织方式和知识服务与工具两个方面对文化遗产项目语义组织研究现状进行梳理,从知识建模、知识抽取和知识挖掘与利用三个维度对文化遗产语义组织关键技术进行剖析.[结果/结论]研究发现,数据互操作、领域本体标准化、个性化语义、自动化工具和数据版权是未来文化遗产语义组织发展的关键.