
Textual emotion detection has gained unprecedented interest over the years due to its wide practical applications, but existing emotion corpora are skewed toward English. Malay, a resource-poor language, currently does not have a reliable emotion corpus that can be used as a benchmark, particularly in the variation of the Malay language used in Malaysia. We present EmoTweet-Malay-6, the first carefully curated Malay emotion corpus containing 5526 tweets annotated with six emotion categories: anger, fear, happiness, love, sadness, surprise and none (no emotion). Since Malaysians are typically at least bilingual, a fraction of this corpus is also code-switched (English and Malay). The tweets were initially annotated using a rule-based classifier following the original source, but we enhanced the corpus by providing additional human validation. Three native Malay speakers were trained to annotate these tweets and achieved Fleiss’ kappa of 0.604. We benchmarked EmoTweet-Malay-6 using commonly adopted machine learning, deep learning, and pre-trained language models. The Malay emotion classifiers serving as baselines show good performance with a macro [Formula: see text]1-score ranging between 0.81 to 0.95 across all six emotions. EmoTweet-Malay-6 serves as an essential resource to advance the development and study of emotion detection in the Malay language.
In this paper, we introduce ViFinClass, the first large-scale benchmark dataset for Vietnamese financial news topic classification. Collected from CafeF between 2006 and 2023, the dataset contains over 30,000 curated articles across five major financial topics: Banking and Finance, Stock Market, Real Estate, International Finance, and Macroeconomics. We evaluate multiple baselines, including traditional machine learning models and transformer-based approaches. Experimental results show that PhoBERT achieves the best performance, outperforming multilingual models such as XLM-R. These findings highlight the importance of domain-specific resources in improving financial NLP for low-resource languages. ViFinClass provides a solid foundation for future research on text classification, information extraction, and financial sentiment analysis in Vietnamese, supporting both academic studies and real-world financial applications.
This study compares human and ChatGPT 4o processing of ox (“niu”, 牛) and dragon (“long”, 龍) metaphors in English and Chinese, exploring how cultural values and AI architecture influence interpretation. We used 30 corpus-derived sentences from Academia Sinica Balanced Corpus of Modern Chinese (ASBC) (Chinese Knowledge and Information Processing, 2010) and Corpus of Contemporary American English (COCA) (Davies, 2008), rated by 15 Chinese speakers, 15 English speakers, and ChatGPT 4o on a 5-point Likert scale. A key methodological step involved post-hoc categorization of sentences into daily, literary, and formal context types, which revealed significant quantitative divergence. Findings show statistically significant differences. Human participants demonstrated more context-sensitive and stable understanding; ChatGPT 4o showed high volatility (CV) with complex idioms like “having a cow.” For dragon metaphors, ChatGPT 4o systematically amplified cultural contrasts, leading to a strong cultural polarization effect, particularly in formal contexts. This suggests that while AI is a powerful linguistic tool, its statistical modeling architecture results in a statistically biased representation of cultural knowledge, lacking the nuance and stability of human cognition.
This study provides substantial evidence on the reading behavior of native Mizo speakers’ L1 and L2, which are Mizo and English, respectively. It primarily focuses on reading ambiguous sentences containing homographs. Two experiments observed the reaction time (RT) of Mizo native speakers reading text in L1 and L2, respectively. A comparative analysis of the RT of both L1 and L2 was done to obtain deeper learning about the cognitive processes involved while reading the two languages. Ambiguous sentences were read in significantly longer times in both L1 and L2. Besides, findings reveal that the readers read ambiguous sentences in L1 significantly longer than in L2. This suggests that Mizo speakers can read and process their L2 quicker than their L1 in specific sentences containing homographs. The reason behind this is the absence of tone information of the Mizo homographs, which lengthens reading time and requires the extraction of their meanings and pronunciations.
Language has a profound impact on communication and thought for human beings. It has been instrumental in the dissemination of disaster warning and rescue information during emergencies. Due to the decline in language abilities and the physiological and social disadvantages of not keeping pace with the speed of information technology development, elderly individuals have become a vulnerable group on the linguistic level. Based on field surveys, this paper focuses on the elderly population in rural China and discusses the language barriers faced by this group in emergencies and the emergency language services they received. The research found: (1) The rural elderly’s individual language abilities are so poor that they are less able to cope with emergencies. (2) There is a pronounced barrier for the rural elderly in terms of communication and information acquisition. (3) Current rural emergency language services face issues such as limited service content, lack of professional service teams, insufficient service awareness, and an incomplete service mechanism. Based on these issues, rural emergency language services should provide differentiated services according to the linguistic characteristics of the elderly population and the actual conditions of rural communities. This involves forming specialized, localized emergency language service teams, extending the breadth and depth of emergency language services, and establishing a long-term mechanism to better protect the linguistic rights and emergency needs of the elderly.
The goal of this research is to use deep learning techniques to create a Seq2Seq text summarization model, with Python as the implementation language. The degree to which the algorithm produces summaries that faithfully capture important elements of the source texts is a measure of its effectiveness. The size and caliber of the training dataset, the design of the Seq2Seq model, and hyperparameter modifications all have a big impact on performance. To increase summarization accuracy, we investigate improvements such as attention methods and different model designs through iterative testing and improvement. By collecting contextual information from both input directions and improving the output summary’s richness, stacked LSTM networks are suggested as a way to greatly boost performance. Furthermore, we propose to use the beam search decoding strategy instead of the greedy approach to achieve more coherent results. The performance of the model is objectively assessed using the BLEU score, which also serves as a benchmark for summary quality. Furthermore, we discuss how to address common summarization task problems by combining coverage and pointer-generator networks. All things considered, our findings show how deep learning may be used to automatically generate brief and instructive text summaries, highlighting the need to continue refining the model and expanding the dataset to get the best results. The precision of this machine learning model is 70-85%.
Machine reading comprehension has been an interesting and challenging task in recent years, with the purpose of extracting useful information from texts. To attain the computer ability to understand the reading text and answer relevant information, we introduce ViMMRC 2.0 — an extension of the previous ViMMRC for the task of multiple-choice reading comprehension in Vietnamese Textbooks which contain the reading articles for students from Grade 1 to Grade 12. This dataset has 699 reading passages which are prose and poems, and 5273 questions. The questions in the new dataset are not fixed with four options as in the previous version. Moreover, the difficulty of questions is increased, which challenges the models to find the correct choice. The computer must understand the whole context of the reading passage, the question, and the content of each choice to extract the right answers. Hence, we propose a multi-stage approach that combines the multi-step attention network (MAN) with the natural language inference (NLI) task to enhance the performance of the reading comprehension model. Then, we compare the proposed methodology with the baseline BERTology models on the new dataset and the ViMMRC 1.0. From the results of the error analysis, we found that the challenge of the reading comprehension models is understanding the implicit context in texts and linking them together in order to find the correct answers. Finally, we hope our new dataset will motivate further research to enhance the ability of computers to understand the Vietnamese language.
The synergetic relationship between morphology and meaning in Chinese remains underexplored. Based on the Lancaster Corpus of Mandarin Chinese, we quantitatively analyze the structure of Chinese words by adopting the method of information theory, exploring the interaction between the structural complexity of words (SCW) and polysemy. We also assessed the model parameters’ effectiveness in distinguishing different genres and the textual characteristics reflected by these parameters. The results indicate that, first, an inverse correlation exists between SCW and polysemy (but the increase of one will not cause the infinite reduction of the other), and the model [Formula: see text] effectively captures this relationship. Second, parameters [Formula: see text] and [Formula: see text] reveal differences in lexical forms between ”narrative” and ”expository” texts, with higher values of [Formula: see text] and [Formula: see text] suggesting a tendency toward expository functions; parameter [Formula: see text] reflects distinctions in meaning clarity between the two types of texts, with higher values of [Formula: see text] suggesting a tendency toward narrative functions. Beyond elucidating the interaction mechanism between lexical forms and meanings, our study provides insights into the automatic recognition and classification of texts, as well as into the quantification of lexical morphology and the development of morphological models in other analytic languages.
Automatic speech recognition (ASR) not only enables hands-free use of various electronic gadgets but also facilitates the creation of print-ready dictation. Despite significant advancements in the field of ASR, low-resource languages such as Pashto still face challenges in achieving comparable maturity. While some efforts have been reported in the literature, the lack of publically available datasets and robust models for Pashto ASR remains a significant barrier. This study aims to address these challenges by developing a novel Pashto isolated spoken digit dataset (PISDD) and proposing a hybrid deep learning-based ASR system. The PISDD is developed by first collecting audio data from 500 native speakers of Pashto, followed by preprocessing techniques including denoising, and partitioning each speaker data into ten digits. This individual digit is then converted into 2D histogram images using the Mel-frequency cepstral coefficient (MFCC) feature extraction technique. A hybrid convolutional neural network (CNN) and bidirectional long short-term memory (BLSTM) model is proposed for the classification, which takes this 2D histogram image as input and predict the output digit. The model uses a three-layer CNN for spatial feature extraction, followed by two BLSTM layers for temporal feature modeling. The proposed model achieves an accuracy of 96.42% on the 30% test set, outperforming the two baseline models: CNN and convolutional recurrent neural network (C-RNN), trained on the same dataset, which achieved 94.12% and 95.23% accuracy, respectively.
This paper presents a novel approach to address the scarcity of labeled data in speech de-identification, a critical task for protecting personal privacy. By leveraging a large language model, we propose a fully automated data augmentation strategy that generates synthetic speech text data enriched with diverse personally identifiable information (PII) entities. This augmented dataset is then used to train the speech-de-identification models, significantly improving its performance on spoken language. To further enhance de-identification accuracy, we explore both pipeline and end-to-end models. While the pipeline approach sequentially applies speech recognition and named entity recognition, the end-to-end model jointly learns these tasks. Our experimental results demonstrate the effectiveness of our data augmentation strategy and the superiority of the end-to-end model in improving PII detection accuracy and robustness.
This study evaluates the performance of mainstream large language models (LLMs) in Chinese security generation tasks, examines the potential security risks associated with these models, and proposes strategies for mitigating these risks. To this end, we developed the multidimensional security question answering (MSQA) dataset and the multidimensional security scoring criteria (MSSC). This study compares the performance of three models across six distinct security tasks. Pearson correlation analysis was conducted using GPT-4 and questionnaires, while automatic scoring was implemented using GPT-3.5-Turbo and Llama-3. Experimental results reveal that ERNIE Bot excels in ideology and ethics evaluation, ChatGPT demonstrates strong performance in assessing rumors, false information and privacy security, and Claude performs well in evaluating factual fallacies and social biases. Additionally, the fine-tuned model showed effectiveness in security scoring tasks, and the proposed Security Tips Expert (ST-GPT) successfully mitigates security risks. Despite the promising results, all models exhibit inherent security risks. Based on these findings, we recommend that both domestic and international models adhere to the legal frameworks of their respective jurisdictions, minimize AI hallucinations, continuously expand training corpora, and undergo regular updates and iterations to enhance their reliability and safety.
This paper details the methods proposed by the UIT-DarkCow team during their participation in the 4th Automated Legal Question Answering Competition (ALQAC 2024). Specifically, the team focused on two main tasks: legal document retrieval and legal question answering (LQA). For the legal document retrieval task, the goal was to return articles related to a given question. An article is considered “relevant” if it contains information that helps answer the question. To achieve this, the team applied a Vietnamese sentence parsing method, combined with two prominent retrieval methods: BM25, a traditional and proven method, and sentence transformer, a recently emerged and highly regarded method in natural language processing. This combination significantly improved the accuracy of finding relevant documents for the questions. In the LQA task, we faced three types of questions: true/false questions, multiple-choice questions, and free-text questions. To address these questions, the team used the most advanced technique currently available: prompt engineering with large language models. This technique allows the model to understand and accurately answer legal questions. The results of the ALQAC 2024 competition demonstrated the effectiveness of the methods applied by the team. UIT-DarkCow achieved third in the legal document retrieval task and second in the LQA task.
Chinese, Japanese, and Korean (CJK) Hanzi image recognition still faces significant challenges, particularly with pixel resolution, stroke count, character frequency and structure. This paper introduces the concepts of character and component thresholds to establish foundational parameters for Hanzi recognition, considering Unevenly Distributed Composite Character (UDCC), which account for up to 88% of Chinese characters. We compiled corpora of Simplified, Traditional characters, and Japanese Kanji, analyzing 3395 images with stroke counts ranging from 1 to 64, and developed the ZH-TC-IM965858 database, containing 33,950 images across 10 pixel resolutions. Using Euclidean distance and ResNet50 similarity analysis, we identified a character threshold of 26 strokes for 24×24 pixel images, a component threshold of 14 strokes, and a comprehensive threshold of 16 strokes. We found that 7.61% of 8105 common Chinese characters, 27.41% of 96,585 Traditional characters, and 5.99% of 2163 Japanese Kanji exceed this comprehensive threshold. Character Length and Character Frequency (CLCF) models were employed to explore relationships between stroke count, frequency, and thresholds. Additionally, Scale-invariant Feature Transform (SIFT) was applied to match radicals and components, providing insights for improving recognition accuracy. This research advances Hanzi image recognition and enhances multimodal Large Language Models (LLMs) for ideographic languages.
The current work explores long-term speech rhythm variations to classify Mising and Assamese, two low-resourced languages from Assam, Northeast India. We study the temporal information of speech rhythm embedded in low-frequency (LF) spectrograms derived from amplitude (AM) and frequency modulation (FM) envelopes. This quantitative frequency domain analysis of rhythm is supported by the idea of rhythm formant analysis (RFA), originally proposed by Gibbon [1]. We attempt to make the investigation by extracting features derived from trajectories of first six rhythm formants along with two-dimensional discrete cosine transform-based characterizations of the AM and FM LF spectrograms. The derived features are fed as input to a machine learning tool to contrast rhythms of Assamese and Mising. In this way, an improved methodology for empirically investigating rhythm variation structure without prior annotation of the larger unit of the speech signal is illustrated for two low-resourced languages of Northeast India.
Example sentences serve as a crucial bridge for learners to master language application rules, enhance language skills, and develop a sense of language. These sentences encompass various aspects, including semantics, grammar, and pragmatics, and hold significant importance in the fields of language teaching and publishing. Large Language Model (LLM) have facilitated the construction and development of generative corpora. Empowered by LLM, example sentences are linked with linguistic elements such as parts of speech and meanings. During the generation process, both coarse-grained and fine-grained resources are fully utilized; in the screening process, relevant research findings on example sentences, errors, and corrections are extensively referenced to form screening norms. This approach results in the construction of a generative example sentence corpus that meets educational needs and maintains a high degree of standardization.
Most Named entity recognition (NER) methods can only handle flat entities and ignore nested entities. In Natural language processing (NLP), it is common to contain other entities within entities. Therefore, we propose a Flat-Span contrastive learning (Fla-SpaCL) method for nested NER. This method includes two sub-modules: a flat NER module for outer entities and a candidate span classification module based on contrastive learning. In the flat NER module, we use Star-Transformer and Conditional random field (CRF) to identify the outer entities. In the candidate span classification module, we first generate inner candidate spans based on the identified outer entities. Secondly, to better distinguish entity spans and non-entity spans, we introduce contrastive learning to maximize the similarity between entity spans and use the InfoNEC loss function to handle hard negative samples. Finally, multi-task learning is used to jointly optimize the flat NER module and the candidate span classification module to reduce error propagation and improve model performance. In the experimental analysis, we compared the proposed model with baseline models to verify its effectiveness.
Today the Large Language Model profoundly affects the way we work in all walks of life, as well as the way we teach in the field of education. In this paper, we focus on the Large Language Model we designed for composition education in elementary school language. We focus on the accurate understanding of Chinese vocabulary and the adaptation of language vocabulary and language structures for the domain of elementary school students, which is currently missing in mainstream LLMs. At the same time, we also pay attention to the current educational concerns about the misuse of the LLM, and target the sensitive questioning designed about the direct generation of composition answers. In the process, we collected datasets related to composition tutoring in elementary school language and generated multiple rounds of student-teacher dialogues using ChatGPT-3.5. We obtained a more ideal large-scale language model for essay tutoring in elementary school language by using different datasets and different data input methods.
A significant portion of the global population speaks multiple languages, including many low-resource languages on which current multilingual ASR models perform poorly. To improve the low-resource language performance, the models are adapted to low-resource data and only a small number of (extra) parameters are fine-tuned to prevent overfitting. Multiple models are fine-tuned to support different languages. Then during inference, one of the models is selected for transcription depending on the language to transcribe. However, for applications like smart home devices, the language used each time by multilingual speakers may be different, so the model cannot be selected if the language to transcribe is not known beforehand. To address this, this paper proposes a two-stage regularization-based continual learning method to adapt only one model to transcribe both the source and target languages and to prevent overfitting. Specifically in stage one training, a strong regularization is used to update all the parameters and to prevent overfitting, while in stage two, the regularization is relaxed to improve the model’s learning capacity. By adapting Whisper to 10 hours of data for each of 2 languages from Common Voice, results show that our method can reduce average word error rate from 18.59% to 15.52%.
Video resources are among the most crucial digital tools for International Chinese Education. This paper initially presents a practice of video extraction based on subtitles, utilizing the official vocabulary list issued by the Chinese Proficiency Grading Standards for International Chinese Education. Subsequently, the paper explores the potential application of LLM in constructing video resources. Despite the current limitations in the video generation capabilities of these models, their future role in resource construction is promising. Both generation and extraction are anticipated to become fundamental paradigms in creating diverse educational resources. With the advancement of LLM, a new wave of generated resources is expected to emerge.
Machine reading comprehension has been an interesting and challenging task in recent years, with the purpose of extracting useful information from texts. To attain the computer ability to understand the reading text and answer relevant information, we introduce ViMMRC 2.0 - an extension of the previous ViMMRC for the task of multiple-choice reading comprehension in Vietnamese Textbooks which contain the reading articles for students from Grade 1 to Grade 12. This dataset has 699 reading passages which are prose and poems, and 5,273 questions. The questions in the new dataset are not fixed with four options as in the previous version. Moreover, the difficulty of questions is increased, which challenges the models to find the correct choice. The computer must understand the whole context of the reading passage, the question, and the content of each choice to extract the right answers. Hence, we propose a multi-stage approach that combines the multi-step attention network (MAN) with the natural language inference (NLI) task to enhance the performance of the reading comprehension model. Then, we compare the proposed methodology with the baseline BERTology models on the new dataset and the ViMMRC 1.0. From the results of the error analysis, we found that the challenge of the reading comprehension models is understanding the implicit context in texts and linking them together in order to find the correct answers. Finally, we hope our new dataset will motivate further research to enhance the ability of computers to understand the Vietnamese language.