[Purposes] This paper investigates and analyzes the policies of data journals on humanities and social sciences abroad, aiming at providing a reference for the publication of data papers and the sharing of scientific data on humanities and social sciences in China. [Methods] Based on data life cycle, this paper compared and analyzed data life cycle between humanities and social sciences and natural sciences, and constructed a data life cycle model about humanities and social sciences. [Findings] Policies on data planning, data creation, data processing, data storage, data publishing, and data sharing of data journals on humanities and social sciences are discussed. [Conclusions] Policies for data journals on humanities and social sciences should be developed based on the data life cycle and data quality and evaluation standards in the available policies on data journals of humanities and social sciences should be improved. In addition, we should strengthen the construction and cooperation of data journals and the data storage library and encourage the data journals of this field to adopt international citation and sharing norms.
This paper aims to deepen the researches on databases’ implementation of the FAIR principle,so as to lay a more solid foundation for the further development of scientific data management,and provide more supports for data opening in different disciplines. Through coding and recoding, it makes a textual analysis of the introductory texts of 11 Bioinformatics databases at home and abroad,focusing on their level of being findable,accessible,interoperable and reusable. The result shows that Bioinformatics databases’ implementation of the FAIR principle should be further improved,not only in their management,but also in their application. It is proposed that guided by the FAIR principle, more should be done to improve their standards of design and management and enrich their description of domain data,so as to realize the interconnection of various database platforms in the same discipline,and improve the level of data sharing as well.
准确地梳理古文典籍脉络,抽取典籍中蕴含的事件和事件论元,对古籍从文本数据向智能化数据转化具有重要意义.针对古文事件的抽取研究主要有基于模式匹配、机器学习和神经网络三种方式,本文在现有的基于神经网络的方法中融入机器阅读理解模式,将事件抽取中出现的"事件类型"和"论元角色"糅合为问题形式,由此输出的答案即为事件论元.分别选取编年体史书《左传》和纪传体史书《史记》作为训练和泛化的数据,在具体的泛化过程中引入混淆句以验证模型效果,为古文事件抽取提供了可参照的思路.
探索国内外数据知识库中的科学数据开放共享政策内容,以期为我国数据知识库制定与完善科学数据开放共享政策提出建议.文章首先调研22个国内外数据知识库发布的科学数据开放共享政策文本;其次,通过文本分析构建政策文本分析类目体系;第三,对其中的政策要点进行评估分析;最后,对比国内外数据知识库的政策要点内容,发现相较于国外,国内数据知识库在政策的数量、完整性、差异性方面存在一定的差距.建议国内数据知识库补充数据创建(C)、数据发布(P)、数据访问(A)、数据归档(D)4个阶段的政策内容,并从完整的数据生命周期的每个阶段视角,参考借鉴国外为大学或科研机构服务的机构类型数据知识库中的每个政策要点内容,以完善我国数据知识库的科学数据开放共享政策.
[目的/意义]事件自动识别抽取是当前典籍主题挖掘研究中一个新的重要课题,其中事件触发词的识别是一项基础的工作,本研究旨在探索古代典籍中事件触发词自动识别和分类的通用方法.[方法/过程]首先运用LDA模型对动词进行主题聚类,归纳典籍事件触发动词的分类体系;并依据聚类结果与分类体系,初步构建触发动词的种子词集.在此基础上,通过语义相似度计算,对种子词集进行扩展,构建典籍事件触发词语义数据集.在实验阶段,以先秦时期的重要典籍《左传》为例,对分类体系构建和种子词集扩展的方法进行验证.[结果/结论]结果表明,本文所提出的典籍事件触发词识别方法可行有效,据此构建的事件触发词集具有较高可信度,未来可进一步扩大实验的样本数量及范围.
[目的/意义]对数字人文研究方法的应用情况进行量化分析,有利于加深对数字人文方法体系的理解.[方法/过程]本研究对数字人文国际期刊和会议上发表的3245篇论文进行内容分析,统计分析了研究方法的使用情况、使用方式、主题分布和共现情况.[结果/讨论]研究发现,数字人文领域的学者多使用实证法,对理论法的应用较少,且绝大多数论文对于研究方法的使用还停留在较低层面.数字人文领域应用多种方法的比例高于其他领域,整体来说,数字人文研究偏好使用计算机信息技术相关方法和案例分析法处理问题.以此为基础,对数字人文研究方法的选取、使用与拓展,以及数字人文方法体系的优化与完善提出建议.[创新/局限]本项研究揭示了数字人文领域方法体系的应用与发展现状,对于进一步深化数字人文方法研究具有一定贡献,但数据样本难以全面揭示数字人文领域研究方法的应用情况.
Purpose Digital humanities database is one of the essential tools in digital humanities research area. Therefore, examining the usage of digital humanities database in academic papers is conducive to assessing the value of digital humanities database for scientific research activities and improving the construction of digital humanities infrastructure. Design/methodology/approach This paper constructs an evaluation system of digital humanities database from the perspective of academic influence and social influence, with mention frequency, usage motivation, platform access data, usage region and usage discipline as indicators and takes China Biographical Database Project as the empirical object to explore the usage of digital humanities database in China. Findings The data analysis result demonstrates that digital humanities databases are widely used and recognized in China. However, the problem of low actual usage remains. Originality/value This paper constructs the digital humanistic database's evaluation system and discusses applying the digital humanistic database in China, which provides a new perspective and method for the influence evaluation study of the digital humanistic database.
Against the background that the top-level semantic framework of Chinese traditional culture is not comprehensive and unified, this study aims to preserve and disseminate cultural heritage information about Chinese traditional culture through the development of a domain ontology which is constructed from ancient books. A combination of top-down and bottom-up approaches was used to construct the ontology for Chinese traditional culture (CTCO). An investigation of historians’ needs, and LDA topic clustering model were conducted, understanding the specific needs of historians, collecting the topic, concepts and relationships. CIDOC CRM was reused to construct the basic framework of CTCO. Ontology structure and function were adopted to evaluate the effectiveness of CTCO. Evaluation results show that the ontology meets all the quality criteria of OntoMetrics, and the experts agreed on content representation (average score = 4.30). CTCO contributes to the organization of traditional Chinese culture and the construction of related databases. The study also forms a common path and puts forward proposals for the construction of domain ontology, which has great social relevance.
[目的/意义]数字远读视角下分析历史典籍,将特定时期社会通过可视化等综合技术展现给研究者,以帮助研究者量化史学研究.[方法/过程]以社会发展过程中产生的文本数据为基础,借鉴用户画像概念,提出社会画像的构建方法.根据各发展分面内在逻辑数据构建社会画像描述框架,利用多种文本挖掘技术抽取不同维度的特征标签,形成社会画像,并以先秦时期为例进行实证研究.[结果/结论]借助基于史实的社会画像,能够全景化呈现社会发展状况,可以为研究者快速获得古代社会概貌提供支持,具有一定的实践意义和价值.
[Purpose/Significance] With the deepening of digital humanistic research on ancient books, the requirement for fine-grained classification based on text content is increasing continuously, and reasonable classification has become the key to the research and effective utilization of digital ancient books.[Method/Process] The research uses the concept of faceted classification, takes the text data of ancient books and related dictionary of ancient books as the research object, and combines conceptual semantic information to organize and describe the features of the content of ancient books. [Results/Conclusions] The classification system constructed in this paper breaks through the limitation of the number, genre and type of ancient books. The research selects five dimensions of politics, economy, culture, society and military to organize and reveal the contents of ancient books in an orderly manner, which is of great value to the in-depth development and utilization of digital resources of ancient books.
春秋时代作为中国古代历史发展的重要转型时期,经济、政治、文化等各领域都发生了急剧的变革.研究这一时期社会变迁,对于阐释中华文化的历史渊源、发展脉络、基本走向,建构中国特色社会主义传统文化观皆有非常重要的意义和价值.文章以《左传》为语料来源,以文本挖掘为手段分析春秋社会演变规律,借助社会变迁相关理论对春秋时代进行结构、表现及动力等不同维度的描述,进而从文本分析的角度构建对应的量化指标.通过融合词频分析、聚类分析、时间序列分析、社区结构挖掘等多种文本挖掘技术,实现各项量化指标的计算.实验结果表明,研究设计的文本计算方法较好地描述了春秋社会结构演变、演变动力及演变表现,与人文学者研究结果基本一致.对于人文计算的开展具有一定理论价值与实践意义,但在模型构建、特征挖掘的方法以及结果评价方面仍有待进一步提升.
[目的 /意义]领域学术观点库构建的目的 在于实现领域内各学者、学派学术观点的集成,形成对领域知识更全面的认知.[方法/过程]从柏拉图知识观视角明确学术观点与相关概念区别,分析其与一般观点的特性.通过分析学术观点在学术活动中的作用明确领域学术观点库的功能.设计领域学术观点库系统的模块结构,探讨各个环节构建过程.[结果/结论]领域学术观点库构建可以分为文本获取、结构化表示、关联关系组织和应用服务四个模块.数据搜集需遵循全面性、相关性和及时性原则,结构化表示包括原子型拆解和元素抽取,学术观点间关系组织通过比较实现,领域学术观点库的应用主要包括查询和评估.
[目的]为有效抽取典籍中蕴含的事件信息,构建面向典籍的事件抽取框架,并采用RoBERTa-CRF模型实现事件类型、论元角色和论元的抽取.[方法]选择《左传》的战争句作为实验数据,建立事件类型和论元角色的分类模板.基于RoBERTa-CRF模型,先用多层Transformer提取语料特征,再结合前后文序列标签学习相关性约束,由输出的标记序列识别论元并对其进行抽取.[结果]对比GuwenBERT-LSTM、BERT-LSTM、RoBERTa-LSTM、BERT-CRF、RoBERTa-CRF等5种模型在数据集上的事件抽取实验结果,RoBERTa-CRF的精确度为87.6%、召回率为77.2%、F1值达到82.1%,验证了该模型的有效性和可操作性.[局限]使用的数据集规模较小,无法使主题类别更均衡化.[结论]本文构建的RoBERTa-CRF模型提升了面向《左传》战争句的事件抽取效果.
[目的]在信息交流视角下,探讨我国科技期刊知识服务转型发展的突破路径.[方法]通过文献调研和对比分析法,探讨面向科技期刊知识服务的科学信息交流模式框架,对国内外科技期刊知识服务模式进行分析.[结果]分别从正式交流和非正式交流视角,归纳总结国内外科技期刊所开展的知识服务.通过对比国内外科技期刊在服务模式、手段和渠道的差异,提出我国科技期刊知识服务的发展策略.[结论]我国科技期刊应从集团化经营模式、自动化出版流程、基于微信矩阵拓展服务项目等方面入手,完善自身的知识服务体系,实现知识生产向知识服务的创新转型.
[Purpose/Significance] It is of great significance to carry out research on the recognition and classification of trigger verbs in ancient books oriented to digital humanities for the deep mining and content revealing of ancient texts. This paper uses the deep learning classification algorithm to explore an automated method for multivariate classification of event sentence text based on trigger words in ancient books. [Method/Process] Based on the construction of the classic event trigger word classification system and trigger dictionary, four different types of event sentence texts are selected as experimental data, and the category labels and sentence texts are coded separately using Onehot and Tokenizer, and then the classifier is trained in the Bi-LSTM model, and a comparative experiment is set by adjusting the parameters, and the performance of the classifier is analyzed by using a general evaluation index. [Results/Conclusions] The classifier after many training and adjustments has an accuracy of 0.95 in the evaluation of the test set, which proves that the experimental method based on deep learning and the constructed trigger word data set can effectively help us realize automatic multivariate classification of event sentence text of ancient books.
[目的/意义]针对《左传》中的战争事件展开研究,对先秦历史乃至中华民族文化的研究具有重要参考价值.[方法/过程]基于框架理论构建《左传》战争事件基本框架体系,利用模式匹配法进行战争句识别,选择条件随机场模型、结合特征模板对战争时间、交战双方等7个命名实体进行识别和抽取,最后基于得到的结构化数据对战争事件进行分析和可视化展示.[结果/结论]研究结果表明,条件随机场模型能够较好地应用于《左传》战争事件的抽取;特征选取会影响实体识别的结果;具体内容方面,春秋时期晋国、楚国、齐国、郑国等国参战频率较高,晋国为主要进攻方,郑国为主要防守方.
[目的/意义]构建面向典籍文本的语义本体,能够促进典籍文本的挖掘与分析.然而由于典籍文本与现代文本在语法上存在较大差异,给面向典籍的语义本体构建带来了困难.[方法/过程]本文运用自然语言处理技术探讨针对先秦典籍的本体构建方法.以国际上文化遗产领域通用的CIDOC CRM为框架,设计先秦典籍本体模型.针对典籍文本内容的特点及句法特征,将规则抽取与条件随机场方法相结合,提出一套本体实例自动获取技术,并以《左传》为实验语料进行测试.[结果/结论]实验表明,本文所提出的本体实例抽取技术能够较好地提高面向典籍文本的本体构建效率.基于规则的本体实例抽取实验F值在93%左右,基于条件随机场的本体实例抽取最佳特征模板的F值为82.51%.在本体实例获取中,词性信息和位置信息具有重要作用.
[目的/意义]在人文计算迅速发展的背景下,利用文本挖掘技术对《左传》进行聚类计算,为春秋时期社会发展状况的主题挖掘等定量分析提供参考,同时对典籍文本多维度重组和分析也具有一定的借鉴意义.[方法/过程]采用文本聚类方法对《左传》进行多维度的定量分析,打破《左传》线性的编年体记载顺序,先运用词匹配算法从《左传》特征词语料中得到各个诸侯国语料,再将LDA主题模型先后用于处理《左传》特征词语料和选取的诸侯国语料,最后结合时间信息进行主题强度计算.[结果/结论]实验结果表明,根据主题-词分布可以挖掘出春秋时期社会和诸侯国各方面的发展内容,通过主题强度变化曲线可以总结出春秋时期社会和各诸侯国的各方面发展态势.通过LDA主题聚类方法最终展现出了春秋时期整个社会以及不同诸侯国在战争、政治及外交等的发展变迁.
图书馆出版服务受到学界的广泛关注,近年来我国高校图书馆也开始尝试开展出版服务.本文从出版物类型、服务平台、数字保存及附加服务等方面,对国内42所"双一流"高校图书馆的出版服务现状进行调研、分析,并提出基于出版内容的高校图书馆出版服务升级策略框架,从低到高包括四个层次:(1)与出版相关的辅助性服务和资助支持;(2)基于机构知识库的学术成果发布和基于馆藏资源的数字出版;(3)对经过同行评议的高质量资源进行开放获取出版;(4)基于同行评议的数据出版.同时,以"强基计划"为背景,提出了开展参与教学的高校图书馆出版服务的具体策略.
在本篇笔谈中,特别邀请了图书情报与档案管理、信息与数据管理领域的优秀青年学者,畅谈他们在本领域从业、从教、从研等方面的体会和感想.青年学者们以推动学科发展为己任,沿承百年文华传统,追索无尽知识诸宝.