This paper proposed five principles for digital resource construction for language education, followed by practical implementation in constructing Teaching Example Sentence (TES) resources. To develop a Teaching Example Sentence Library (TESL) for International Chinese Education, a Grid-based Parsing Framework (GPF) was utilized. The development of TESL effectively addresses how to transform the qualitative requirements for sentences into quantitative conditions and provides a method for sentence extraction. Specific issues with certain sentences in big data unsuitable for language teaching are also outlined. The TESL constructed in this paper was applied by the Integrated Chinese Intelligent Teaching System developed by Beijing Language and Culture University. Although the proposed method is designed for Chinese texts with the aim of language teaching, it can also be applied to sentence extraction in other languages and for different purposes.
The study of syntactic-semantic interface has always been an important part of linguistics and natural language processing (NLP) research, and the introduction of events into it can provide a consistent object for syntactic-semantic research. In the system of Chinese Parataxis Graph (CPG), event is a chunked description of sentence semantics, and the analysis of event structure provides an effective perspective for the semantic analysis. However, the academic research on event structure is mostly theoretical and lacks relevant data statistics. In this paper, we measure and analyse the syntactic representations of event structures based on CPG Annotated Corpus, to reflect the distribution of the syntactic positions of event words when the event structures are mapped to syntactic structures, as well as the presence or absence and co-occurrence of the core arguments of events that are mapped to different syntactic structures, and analyzing some of the data.
This study develops a collocation theory for instructional use, based on theoretical frameworks of linguistics and second language pedagogy. It also creates a comprehensive knowledge base for practical application. Collocations do not exist in isolation. Using the Precision Marker Interconnect concept, collocation resources are constructed based on linguistic ontology and practices in International Chinese Language Education. Collocations are analyzed from syntactic, semantic, and pragmatic perspectives and categorized into logical, formal, and instance levels. Additionally, it formalizes the specific content of collocations through collocation sequences. This paper utilizes the BCC (BLCU Corpus Center, abbreviated as BCC) to extract over 700,000 high-quality collocations, filtered by formalized rules, and adopts a human-in-the-loop approach to validate and optimize the results. This process leads to the development of a knowledge base of over 200,000 collocation pairs to meet teaching needs. The implementation of this knowledge base in International Chinese Language Education has shown promising outcomes.
The interaction between lexis and grammatical patterns, known as lexicogrammar, plays a crucial role in describing a word's usage characteristics and profiles. Lexico-grammatical pattern extraction and further analysis are essential for understanding language features and usage characteristics, particularly in analyzing how words function within specific textual contexts. Given that verbs constitute the syntactic and semantic core of Chinese sentences, investigating verb patterns is crucial, as various grammatical elements frequently co-occur with verbs to form stable and high-frequency patterns, making such research both feasible and necessary for a comprehensive understanding of Chinese linguistic structures and meaning construction. Previous methods of pattern extraction largely rely on linear concordance analyses, which often fail to capture deeper structural relationships, resulting in inaccurate or misleading collocations. To overcome these limitations, this study introduces a chunk-based annotation approach using Chinese legal judgments as a representative case study. We developed a framework to segment predicate chunks into finer-grained sub-chunks, thereby enabling precise extraction of verb patterns. Employing collostructional analysis, we quantitatively assessed the association strength between verbs and their grammatical patterns and identified preferred verb patterns. The findings demonstrate a strong preference for pre-verbal patterns, aligning closely with genre-specific requirements for factual precision and contextual specificity in legal texts. Additionally, distinct pattern preferences emerge across various semantic verb categories (e.g., existential, saying, action), providing empirical insights into the syntactic and functional dimensions of Chinese verb usage. This research contributes to corpus-based studies on the interaction between lexical items and grammatical patterns, and deepens the understanding of the factors influencing verb usage in Chinese. (c) 2025 Elsevier B.V. All rights are reserved, including those for text and data mining, AI training, and similar technologies.
In the field of natural language generation in Chinese, there has been limited attention to question generation tasks, partly due to constraints related to knowledge acquisition. With the increasing popularity of online education, the automatic generation of questions has become a major point in the context of language intelligent education. In this regard, this paper is oriented towards international Chinese language education. This paper constructs a domain knowledge database and utilizes the grid-based language structure parsing framework in conjunction with the domain knowledge database to perform syntactic and semantic analysis on text. In turn, this enables the automatic generation of short-answer questions based on the results of syntactic and semantic analysis, along with predefined question keywords, achieving an accuracy rate as high as 94%. This not only provides a high-accuracy domain-specific research model for automatic question generation but also offers high-quality question-answer pairs for second language learning.
Video resources are among the most crucial digital tools for International Chinese Education. This paper initially presents a practice of video extraction based on subtitles, utilizing the official vocabulary list issued by the Chinese Proficiency Grading Standards for International Chinese Education. Subsequently, the paper explores the potential application of LLM in constructing video resources. Despite the current limitations in the video generation capabilities of these models, their future role in resource construction is promising. Both generation and extraction are anticipated to become fundamental paradigms in creating diverse educational resources. With the advancement of LLM, a new wave of generated resources is expected to emerge.
Patent retrieval is a critical step in patent analysis. Retrievable elements play a key role in constructing search queries and performing accurate searches, and most retrievable elements are created manually. However, the increment of patent applications each year has brought a huge burden on manual extraction of retrievable elements and patent examination, raising the urgent need of automated solutions. As keywords serve as an effective way of expressing retrievable elements in patent retrieval, we explore the automatic extraction of keyword-based retrievable elements from Chinese patent application texts in this study. We employ various keyword extraction methods, including large language model based methods, to identify retrievable elements within these texts. Our experimental results have shown that these methods can effectively extract keywords as retrievable elements from Chinese patent applications, which benefits to manual patent searching and patent examinations.
Non-core argument, an additional component of a sentence, plays an important role in semantic analysis, which can expand the basic predicate structure and form complex propositions. Non-core arguments usually need to be cited by case markers, including not only prepositions but also verbs. This study focuses on verb-guided non-core arguments, as well as an investigation of the expression of non-core argument role relationships based on serial verb construction, which is different from the previous method that mainly focused on preposition-guided non-core arguments. In this study, a new set of non-core argument role classification systems is proposed. At the same time, newspapers and periodicals corpus containing structural tree information in BCC corpus are chosen for obtaining serial verb collocation data. Through disambiguation and non-core argument role identification, the presented method eventually achieved 25517 examples of triple collocations and constructed a verb-guided non-core argument knowledge base, which provides data support for natural language processing and language theory research.
Although Dense Passage Retrieval (DPR) models have achieved significantly enhanced performance, their widespread application is still hindered by the demanding inference efficiency and high deployment costs. Knowledge distillation is an efficient method to compress models, which transfers knowledge from strong teacher models to weak student models. Previous studies have proved the effectiveness of knowledge distillation in DPR. However, there often remains a significant performance gap between the teacher and the distilled student. To narrow this performance gap, we propose MTA4DPR, a Multi-Teaching-Assistants based iterative knowledge distillation method for Dense Passage Retrieval, which transfers knowledge from the teacher to the student with the help of multiple assistants in an iterative manner; with each iteration, the student learns from more performant assistants and more difficult data. The experimental results show that our 66M student model achieves the state-of-the-art performance among models with same parameters on multiple datasets, and is very competitive when compared with larger, even LLM-based, DPR models.
以往的介词知识库构建重视介词语义和介宾的搭配研究,鲜有对介动搭配进行系统研究及知识获取的工作.而汉语介词发达及动词是句子中心的特征决定了介动搭配研究的重要性.该文基于结构检索技术,充分借助短语结构属性和结构信息,从大规模语料中抽取介动搭配16 033对,并提出了介动搭配紧密度的度量方法,初步分析证明该方法远优于依靠绝对频次进行搭配度量的方法.
At present, there are about two shortcomings in the resources of Chinese synonyms. One is that there are no synonym resources specifically for verbs. The other is that scholars usually pay too much attention to the conceptual meaning of synonyms and ignore the grammatical meaning of synonyms. These two shortcomings may lead to inapplicability when applying synonym resources. Therefore, we are oriented toward verbs, consider conceptual and grammatical meanings, and construct grammatical synonym resources for verbs. Firstly, we formulated the Grammatical Synonym Annotation specification for Modern Chinese Disyllabic Verbs . Secondly, we constructed a collection of modern Chinese disyllabic verbs depending on the current research results. Finally, we took the annotation specification as the standard, referred to the related resources and the performance of verbs in the corpus, and built the grammatical synonym resources for modern Chinese disyllabic verbs. After forming the annotation team, it took us fourteen months to complete the annotation task. This paper provides high-quality data for natural language processing and dramatically improves the accuracy of sentence-level retelling.
智慧教育的核心是智慧教学.智慧教学有两个含义:一个是"智能地教学",强调智能技术赋能教学全过程;一个是"智慧的教学",强调教学的结果. "智能地教学"包括两个方面的内容.第一,通过智能技术更好地建设数字化教学资源,提供给教师和学生,推进教育资源的供给侧改革;第二,采用智能技术研发具有教学功能、可以充当教师角色的智能工具,直接赋能学生,实现无师值守的个性化学习.
Verb-object constructions have always been the focus of Chinese grammar research. In Chinese, it has been agreed that transitive verbs take objects, while intransitive verbs taking objects are often confusing. There are very few studies addressing intransitive verbs taking objects. In particular, no investigations have been carried out on the quantitative analysis of the ability of disyllabic intransitive verbs taking objects of different structures and object semantics. Therefore, this study focused on the phenomenon of disyllabic intransitive verbs taking objects and made use of large-scale corpus to investigate the disyllabic intransitive verbs and their objects in order to enrich the understanding of unconventional verb-object constructions in the language. It first proposed a semantic role classification system on the basis of previous studies, and defined the word structure types, as the basis for judging the verb structure types and object semantic types in intransitive verb with object structure. Based on the BCC corpus, this study analyzed the top 4237 high-frequency disyllabic verbs in the fifth edition of Modern Chinese Dictionary. It used structure retrieval to complete the recognition of intransitive verb and its object acquisition, and artificially identified the verb structure types and object semantic roles. It was found that verb-object intransitive verbs have the strongest ability to take objects and the role of “quantity” is the semantic role that most often serves as the object of disyllabic intransitive verbs.
Traditional semantic role labeling is mostly based on the results of syntactic analysis. On the basis of syntactic analysis, argument identification and argument classification are carried out in two steps. Due to the problem of error cascade, the effect of argument identification directly determines the quality of semantic role labeling. However, converting the semantic role labeling into a sequence labeling task tries to ignore syntactic information, which increases the difficulty of the model and relies on limited labeling data. Therefore, this paper focuses on improving the accuracy of Chinese argument identification, that is, identify all candidate arguments given a predicate. Specifically, based on the results of Chinese chunk dependency parsing, an argument identification model is built based on the pretrained language model BERT (bidirectional encoder representations from transformers). The AUC of the model reaches 97.18% and the accuracy reaches 98.10%, which provides a reliable data for subsequent argument classification task. At the same time, large-scale pretrained language model can overcome the problem of sequence dependency, and perform well in argument identification of complex sentences.
In discourse semantic analysis, the primary focus lies on cohesive elements that maintain discourse coherence. Additionally, attention is given to referring components and topic description structures that facilitate articulation. Various terms are used to refer to these cohesive elements, such as connecting components, logical connectives, relational words and cohesive components. However, there is a lack of unified terminology. The concept of cohesive components has been proposed before. In prior research, it primarily encompassed connectives, discourse markers and tone words, all of which contribute to discourse coherence. However, these definitions were relatively imprecise, and previous studies did not systematically categorize discourse cohesive components based on specific semantic functions. Consequently, they were inadequate for chapter-level semantic analysis. This paper extends the definition and scope of cohesive components based on existing research. It categorizes cohesive components according to both their formal characteristics and semantic functions. We extract cohesive components from the BLCU-CST tree bank, filter them based on specific scope and definition, and summarize a total of 247 logical connectives, 118 discourse connective markers and six types of numerical sequence connectives. In terms of semantic function, we analyze the distribution of 14 different logical-semantic relations that cohesive components indicate within discourse. This analysis not only enriches the understanding of discourse semantics for natural language comprehension but also offers valuable insights for linguistic ontology research.
Being an adverbial is a grammatical function of a small number of nouns. The nouns act as the adverbial and the subject are located before the predicate, and there is no formal mark, therefore, it is easy to cause syntax parsing mistakes. Meanwhile, there is a lack of semantic resources of adverbial nouns at this stage which makes it more difficult to do related semantic parsing. We extracted the “noun + verb/adjective” collocation corpora from a large-scale structure tree database. By writing code and manually proofreading, nouns that can act as the adverbial to directly modify verbs or adjectives which are also called the adverbial nouns are listed exhaustively. We conducted a comprehensive classification and quantitative analysis according to the semantics expressed by the adverbial nouns, and constructed an “adverbial noun-verb/adjective” collocation database based on this, so as to improve the accuracy of syntactic parsing and provide corresponding semantic information, and also to provide reference for the related research of linguistics.
一、引言 国际中文智慧教学平台(简称智慧教学平台)建设是国际中文教育集成创新实践的成果.建设智慧教学平台的理念是:用精标互联的教学资源和语言智能技术赋能国际中文教学,以教师为主导,学习者为中心,融课件为载体,开展以联通、互动为主要特征的数据化教学,实现国际中文教学提质增效的目标.
自然语言处理的核心是语义理解,即理解符号后面所表达的意义.如何表示语义知识、如何获取支持语义分析的知识、如何设计和实现应用语义知识的算法或策略,这些都是自然语言处理中重要的问题.
Grammar teaching is the focus of International Chinese Language Education, and grammar collocation resources play an essential role in promoting grammar teaching. Based on the three-dimensional grammar description of the formal structure, function, and typical context of the predecessors, this paper uses the BCC corpus to search the sentences of grammar points. This study explores a novel method to accurately acquire practical collocations and sample sentences in large quantities and constructs a collocational library containing 264 grammar points.
Under current researches on Chinese language, sentences are usually chunked into the components of the same level. However, in actual Chinese environment, the predicate block in a sentence and the subject and object component blocks before and after it constitute the skeleton of the event representation, and the subblocks inside the predicate block modify the core predicates from different perspectives and serve as the related components of the event. Therefore, it is necessary to treat the predicate block and the component inside the predicate block as different levels of block components. Dividing into primary and secondary components plays an important role in understanding the main semantics of a sentence by abstracting the outline and facilitating the event reasoning and calculation. Therefore, this paper first defines predicate as the core of the predicate block and further defines the sub-block inside the predicate block. The predicate-centered predicate blocks are annotated with encyclopedia corpus with relatively high sentence complexity. At the same time, the components inside the predicate block are divided based on the subblock types defined in this paper. As of now, the knowledge base includes 36,360 predicate blocks. Based on this knowledge base, this paper creates an internal boundary recognition task for predicate blocks, and tests it with a sequential labeling model, which provides a baseline for the subsequent researches in the future.