[目的/意义]吐蕃时期的金石铭刻是了解吐蕃社会政治制度、宗教信仰、对外交往、社会关系、语言状况等的重要依据.本研究致力于构建吐蕃藏文金石铭刻知识图谱,探索民族古文献数字化新途径.1方法/过程]借助数字人文和知识图谱构建技术,通过本体建模分别构建吐蕃金石铭刻概况、研究现状、刻文内容和语法范畴4种本体,抽取概念、属性、关系,并以三元组方式表示;把刻文中的每一个词作为实例,构建实例之间异体、简缩、变形等链接关系以及命名实体之间的各种关联关系,以高度结构化的形式存储知识,形成吐蕃金石铭刻知识图谱.[结果/结论]吐蕃藏文金石铭刻知识图谱是藏文古文献数字人文研究的有益探索.以藏汉双语词级对齐形式呈现实例,使更多的研究者利用该材料开展研究,更好地挖掘吐蕃藏文金石铭刻的学术价值.
Ancient ethnic documents are essential to China’s ancient literature and an indispensable civilizational achievement of Chinese culture. However, few research teams are involved due to language and script literacy limitations. To address these issues, this paper proposes an interlinearized annotation strategy for ancient ethnic literature. This strategy aims to alleviate text literacy difficulties, encourage interdisciplinary researchers to participate in studying ancient ethnic literature, and improve the efficiency of ancient ethnic literature development. Concretely, the interlinearized annotation consists of original, word segmentation, Latin, annotated, and translation lines. In this paper, we take ancient Tibetan literature as an example to explore the interlinearized annotation strategy. However, manually building large-scale corpus is challenging. To build a large-scale interlinearized dataset, we propose a multi-task learning-based interlinearized annotation method, which can generate interlinearized annotation lines based on the original line. Experimental results show that after training on about 10,000 sentences (lines) of data, our model achieves 70.9% and 63.2% F1 values on the segmentation lines and annotated lines, respectively, and 18.7% BLEU on the translation lines. It dramatically enhances the efficiency of data annotation, effectively speeds up interlinearized annotation, and reduces the workload of manual annotation.
汉—藏音译在汉语人名/地名等专有名词翻译为藏文时具有重要价值,对于藏族与其他民族间的交往交流交融具有重要影响。文章对汉—藏音译的现状进行分析,梳理当前汉—藏音译中的主要问题以及这些问题产生的原因;在此基础上提出了五条汉—藏音译的原则,并实现了汉—藏人名/地名自动音译系统,通过浏览器和微信小程序对外提供汉—藏音译的查询与应用。
本文基于藏文全文隔行对照标注数据,结合其他相关文献,详细描写藏文古文献《拔协》中九个存在动词的基本用法、语义功能、句法结构和语法化路径,重点讨论vdug、yod两个存在动词的语义分化、人称搭配以及语法化路径.
构建藏语依存树库是实现藏语句法分析的重要基础,对藏语本体研究和信息处理具有重要价值.基于此,该文提出了一种基于树库转换的藏语依存树库构建方法.该方法首先扩充了前期构建的藏语短语结构树库,然后根据藏语短语结构树和依存树的特征设计树库转换规则,实现藏语短语结构树到依存结构树的初步转换,最后对自动转换结果进行人工校验,得到了2.2万句藏语依存树.为了对转换结果做出量化评价,该文抽取了依存树库中5% 的依存树,对其依存关系进行校验和统计,最终依存关系的准确率达到89.36%,中心词的准确率达到92.09%.此外,该文使用基于神经网络的句法分析模型验证了依存树库的有效性.在该模型上,UAS值和LAS值分别达到83.62% 和81.90%.研究证明,使用半自动的树库转换方法能够有效地完成藏语依存树库构建工作.
Medical machine translation is of great value for cross-border medical translation. Chinese to English neural machine translation has made great progress based on deep learning, powerful modeling ability and large-scale bilingual parallel data. Neural machine translation relies usually on large-scale parallel sentence pairs to train translation models. At present, Chinese-English translation data are mainly in the fields of news, policy and so on. Due to the lack of parallel data in the medical field, the performance of Chinese to English machine translation in the medical field is not compromising. To reduce the size of parallel data for training medical machine translation models, this paper proposes a paraphrase based data augmentation mechanism. The experimental results on a variety of neural machine translation models show that data augmentation through paraphrase augmentation can effectively improve the performance of medical machine translation, and has achieved consistency improvements on mainstream models such as RNNSearch and Transformers, which verifies the effectiveness of paraphrase augmentation method for domain machine translation. Meanwhile, the medical machine translation performances could be further improved based on large-scale pre-training language model, such as MT5.
Frequently corresponding to syntactic components, the Maximal-length Noun Phrase (MNP) possesses abundant syntactic and semantic information and acts a certain semantic role in sentences. Recognition of MNP plays an important role in Natural Language Processing and lays the foundation for analyzing and understanding sentence structure and semantics. By comparing the essence of different MNPs, this article defines the MNP in the Tibetan language from the perspective of syntax tree. A total of 6,038 sentences are extracted from the syntax tree corpus, the structure type, boundary feature, and frequency of MNPs are analyzed, and the MNPs are recognized by applying the sequence tagging model and the syntactic analysis model. The accuracy, recall, and F1 score of the recognition results of applying sequence tagging model are 87.14%, 84.72%, and 85.92%, respectively. The accuracy, recall, and F1 score of the recognition results of applying syntactic analysis model are 87.66%, 87.63%, and 87.65%, respectively.
The research of Tibetan dependency analysis is mainly limited to two challenges: lack of a dataset and reliance on expert knowledge. To resolve the preceding challenges, we first introduce a new Tibetan dependency analysis dataset, and then propose a neural-based framework that resolves the reliance on the expert knowledge issue by automatically extracting feature vectors of words and predicts their head words and type of dependency arcs. Specifically, we convert the words in the sentence into distributional vectors and employ a sequence to vector network to extract feature words. Furthermore, we introduce a head classifier and type classifier to predict the head word and type of dependency arc, respectively. Experiments demonstrate that our model achieves promising performance on the Tibetan dependency analysis task.
Tibetan ancient literature is an important literature material for the study of ancient Tibetan culture, history, and the development of Sino Tibetan language family. However, the lack of work on ancient Tibetan word segmentation tools seriously restricts the research of ancient Tibetan literature. In view of this situation, this paper first utilizes ancient Tibetan interlaced contrast tagging data to extract the ancient Tibetan word segmentation dataset. Based on this dataset, we conduct numerous experiments for the task of ancient Tibetan word segmentation. Experimental results show that BiLSTM + CRF word segmentation algorithm can achieve the best performance, and the performance of ancient Tibetan word segmentation can be further improved through model ensemble. And the results show that the unknown words, insufficient training data and word ambiguity restrict the performance of ancient Tibetan word segmentation.
Tibetan ancient literature is an important literature material for the study of ancient Tibetan culture, history and language, which has important academic value for the study of the development of Sino Tibetan language family. However, the lack of research on ancient Tibetan word segmentation seriously restricts the research of ancient Tibetan literature. In view of this situation, this paper utilizes ancient Tibetan interlaced contrast tagging data to extract the ancient Tibetan word segmentation data set. Based on this dataset, this paper conduct the research on dictionary based word segmentation, statistics based word segmentation and deep learning based ancient Tibetan word segmentation. Experimental results show that BiLSTM + CRF word segmentation algorithm can achieve the best performance by a single model, and the effect of ancient Tibetan word segmentation can be further improved through model ensemble. And the results show that the unknown words, insufficient training data and word ambiguity also restrict the performance of ancient Tibetan word segmentation.
Dependency parsing is an important task for Natural Language Processing (NLP). However, a mature parser requires a large treebank for training, which is still extremely costly to create. Tibetan is a kind of extremely low-resource language for NLP, there is no available Tibetan dependency treebank, which is currently obtained by manual annotation. Furthermore, there are few related kinds of research on the construction of treebank. We propose a novel method of multi-level chunk-based syntactic parsing to complete constituent-to-dependency treebank conversion for Tibetan under scarce conditions. Our method mines more dependencies of Tibetan sentences, builds a high-quality Tibetan dependency tree corpus, and makes fuller use of the inherent laws of the language itself. We train the dependency parsing models on the dependency treebank obtained by the preliminary transformation. The model achieves 86.5% accuracy, 96% LAS, and 97.85% UAS, which exceeds the optimal results of existing conversion methods. The experimental results show that our method has the potential to use a low-resource setting, which means we not only solve the problem of scarce Tibetan dependency treebank but also avoid needless manual annotation. The method embodies the regularity of strong knowledge-guided linguistic analysis methods, which is of great significance to promote the research of Tibetan information processing.
research-article Share on Recent Developments in Tibetan NLP Authors: Long Congjun Chinese Academy of Social Sciences Chinese Academy of Social SciencesView Profile , Nathan W. Hill Trinity College Dublin Trinity College DublinView Profile Authors Info & Claims ACM Transactions on Asian and Low-Resource Language Information ProcessingVolume 20Issue 2March 2021 Article No.: 19pp 1–3https://doi.org/10.1145/3453692Online:23 April 2021Publication History 0citation84DownloadsMetricsTotal Citations0Total Downloads84Last 12 Months69Last 6 weeks4 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
词向量在自然语言处理研究的各个领域发挥着重要作用.该文从语言学角度出发,讨论了词向量技术与语言学理论的关系;根据词向量的特征,提出利用藏文词向量构建语义相似词知识库.该文以哈尔滨工业大学的《词林》为基础,通过汉藏双语词典对译,在获取对译词的词向量的基础上,计算对译词的词向量与原子词群平均词向量的差值,利用不同的差值,自动筛选出与原子词群语义相似度较小的词.该文分别以藏文的词和音节为单位计算词向量,自动筛出不属于原子词群的词,通过对自动筛选结果与人工筛选结果对比,发现两者具有较高的一致性,这说明词向量计算结果与人的语言直觉具有较高的一致性.总体来说,该文所采用的方法有助于提高藏文语义相似词知识库构建效率.
As a shining pearl in traditional Tibetan culture, historical Tibetan documents have received extensive attention from historians, linguists and Buddhist scholars. These documents are converted into digital form using Tibetan document segmentation and recognition methods. The document digitization is of great significance for the research, protection and inheritance of Tibetan history. This paper proposes an overall segmentation and recognition framework for historical Tibetan document images. Firstly, the historical Tibetan document image is preprocessed to correct imbalanced illumination, tilt and noises, and is further transformed into the binarized image. Secondly, we propose a layout segmentation method based on block projection to segment Tibetan document images into texts, lines and frames. Thirdly, in order to solve the problems of touching strokes between text-lines and curvilinear text-lines, we present a text-line segmentation method based on graph model for historical Tibetan text-line segmentation. Lastly, we present a touching segmentation method to segment touching Tibetan character string, and then recognize Tibetan characters. Experimental results show our proposed methods on layout segmentation, text-line segmentation and touching character string segmentation, achieve the satisfactory performance. The proposed methods can also be applied to other fonts in Tibetan font family.
《拔协》是一部重要的古代藏文文献,主要记述了吐蕃王朝后期佛教在西藏的弘传情况,反映了古藏语晚期及中古藏语早期的语言面貌.《拔协》的判断句分为无系词判断句和系词判断句,后者使用系词yin或lags.这些判断句各具特点,它们的使用情况显现出藏语判断表达系统由旧到新的过渡性特点.
在文字识别领域中,手写体识别比印刷体识别更具挑战性.藏文手写体识别已经成为重要的研究课题之一.本文提出了一种基于卷积神经网络LeNet-5模型的藏文手写数字和字母识别方法.分别采集藏文数字手写体样本和字母手写体样本17768和77636例,并对其进行预处理;然后按8:2划分成训练集和测试集,并在CNN(LeNet-5)模型上进行训练.经过测试,数字和字母识别准确率分别达到98.81%和97.89%.
最长名词短语携带着丰富的句法和语义信息,经常与句法成分对应,在句子中充当一定的语义角色.最长名词短语识别在自然语言处理中占重要地位,是分析和理解句子结构 、意义的基础.该文通过梳理不同概念的最长名词短语的含义,从句法树角度界定了藏语最长名词短语的基本概念;从句法树库中抽取6038个句子,分析了最长名词短语的结构类型 、边界特征和出现频次,最后采用序列标注模型和句法分析模型对最长名词短语进行识别.序列标注模型识别结果的正确率 、召回率和F1值分别为87.14% 、84.72% 、85.92%.句法分析模型识别结果的正确率 、召回率 、F1值分别为85.02% 、84.51% 、84.76%.
With the development of information technology,the Tibetan language was widely used on the Internet. To deal with the transliteration issue form Chinese texts to Tibetan,this paper collects five Tibetan website texts and examines the unified forms of transliterations.After analyzing the causes of confusion transliteration between Chinese and Tibetan,this paper proposes some transliteration principles.It also suggests that relevant government agencies should actively promote the standardization of transliteration norms.
本文首先分析了藏文人名的特点以及藏文人名识别的难点,在此基础上,利用条件随机场模型,分别提出了采用基于亚音节标注的藏文人名识别方法和分词与词性标注一体化的藏文人名识别方法.
The Chinese Academy of Sciences launched the Multi-Layer MultiLingual Resource Database (MLLRD) project which aims to collect language resources for natural language processing tasks for low resource languages used in China, such as Mongolian, Tibetan, Uyghur and so on. Tibetan text corpus building is one of the sub projects, in which we have built a Collection of Tibetan Text Corpora(CTTC), including: (1) Tibetan web article corpus which has 440,900 documents. (2)Tibetan text classification corpus. (3) Chinese-Tibetan parallel text corpus which has 773,068 sentence pairs. (4) Part-Of-Speech tagged corpus which has 52,041 sentences. (5) Tibetan tree bank which has 6,040 trees. The paper reports the methods to build these corpora, the contents and scales of each corpus, and applications of them.