Background and objective: Mental health problems, including stress, depression, and anxiety are becoming increasingly prevalent, and there is growing interest in expanding the evaluation of agro-healing programs beyond traditional self-report questionnaires and physiological indicators. Linguistic features embedded in open-ended text responses may serve as supplementary indicators of individuals’ emotional and cognitive states. This study examined whether stress, depression, and anxiety levels among adults and older adults are reflected in the structural, lexical, syntactic, and affective characteristics of responses, using a standardized Korean text analysis system.Methods: A total of 268 adults and older adults participated in either online or face-to-face surveys that included both assessments of mental health status and open-ended questions about recent emotional experiences. Stress, depression, and anxiety were measured using the Brief Encounter Psychosocial Instrument-Korean version (BEPSI-K), the Patient Health Questionnaire-9 (PHQ-9), and the Generalized Anxiety Disorder-7 (GAD-7), respectively. Participants provided free-text descriptions of their recent emotional experiences. Responses were analyzed using the Data Evaluation Unified System (DEUS), which quantifies text surface structure, word frequency, lexical diversity, referential cohesion, syntactic complexity, pronoun use, conjunctive adverbs, psycholinguistic lexical indices, and emotional expression intensity. Independent t-tests and Pearson correlation analyses were used to examine linguistic differences and associations with psychological indicators.Results: Higher stress levels were associated with shorter average word length and higher log-transformed word frequency (p < .05). Participants with higher levels of depression produced fewer words and sentences (p < .01), exhibited higher type-token ratios (p < .001), lower noun overlap (p < .05), and greater boredom intensity (p < .05). Participants with anxiety wrote fewer sentences (p < .05), used more modifiers (p < .05), used fewer conjunctive adverbs (p < .01), and showed greater boredom intensity (p < .05). Correlation analyses showed moderate to strong positive associations among stress, depression, and anxiety, and several linguistic features were significantly associated with the psychological indicators.Conclusion: These findings indicate that psychological characteristics are reflected in written language. Text-based linguistic analysis may serve as a complementary tool for mental health assessment and may be applicable to future evaluations of agro-healing programs.
The Test of Proficiency in Korean (TOPIK) stands as a reputable examination capable of assessing Korean language proficiency for various purposes. In order for TOPIK, with its pivotal significance, to be effectively utilized for multifaceted objectives, the validity and reliability of its evaluations must be substantiated. Thus, this study aims to scrutinize the disparities between TOPIK I and TOPIK II, focusing on reading texts, across various factors. Initially, a corpus was constructed comprising reading texts from the 35th session (post-2014 revision) to the most recent 83rd session available as of March 2014, obtained from the TOPIK website's archive. Subsequently, utilizing the Data Evaluation Unified System (DEUS), linguistic indicators including surface features (vocabulary and sentence length), lexical factors (vocabulary frequency, diversity, pronoun usage, emotional vocabulary usage), syntactic factors (proportion of function words, sentence constituent ratio), and discursive factors (referential coherence, patterns of adverbial conjunction usage) were employed to compare the text difficulty. Additionally, recommendations for future test construction were proposed based on the outcomes derived from each factor analysis.
This corpus-driven research aims to examine the intra-grade continuity of reading passages within high school English mock exams, specifically the national academic achievement tests for seniors in 2023. These tests were conducted quarterly by four regional educational offices: Seoul in March, Incheon in April, Gyeonggi Province in July, and Seoul in October. A corpus of 100 reading passages, 25 from each test, was analyzed via a range of Coh-Metrix measures, including word features, lexical diversity, connectives, readability, referential cohesion, syntactic complexity, semantic cohesion, and so on. These Coh-metrix results demonstrated that there was no compelling evidence to support strong intra-grade continuity, highlighting the necessity for meticulous calibration of text difficulty, which should incrementally escalate with successive test administrations. The analysis of basic metrics, lexical frequency and diversity, syntactic complexity, and readability revealed no consistent trends across test administrations. Additionally, lexical attributes such as concreteness, familiarity, imagery, and age of acquisition did not show significant continuity. Moreover, the analysis of syntactic complexity also showed no consistent pattern. The current inconsistency in text difficulty suggests that students may not be receiving the progressive challenge needed to develop their reading comprehension skills effectively. Drawing upon these findings, theoretical and pedagogical implications are elucidated.
This study endeavors to develop a Korean High School English Test (KHET) wordlist that is featured in actual high-stakes exams but not covered by the 2022 revised national curriculum of English (RNCE) wordlist. The KHET corpus consists of reading passages across 6,086 test items from official and mock College Scholastic Ability Tests (CSAT) administered over an expansive 16-year period, spanning 2006 to 2021. This KHET corpus, aggregating to 809,354 tokens, 27,514 types, and 2,991 word families, is categorized based on different question types and their lexical distribution in comparison with the 2022 RNCE wordlist. The findings reveal that the significant dominance of fill-in-the-blank (21.03%), identifying main idea (9.82%), and long reading comprehension (8.65%) question types necessitates a robust contextual vocabulary knowledge in order to excel on the exam. Moreover, the study gauges the alignment of vocabulary within the KHET corpus against the 2022 RNCE wordlist and, with a frequency range threshold set at 3, a total of 5,308 vocabulary items, recurrent across different question types, were identified. Among these, there was a distinct subset of 516 words aligned with CEFR levels C1 and C2, and the Academic Word List. By unveiling previously overshadowed vocabulary items, this study guides future pedagogical resources equipping students not merely for the CSAT but also for broader academic contexts.
Previous studies in corpus-based literary translation have tended to focus on only one or two specific aspects of style. In this study we expand the existing analytical paradigm to show how the style inherent in source texts (STs) is reflected in their translations. We do this using thirty-six multilevel linguistic features. The selected texts are James Joyce's Dubliners and A Portrait of the Artist as a Young Man and their Korean translations. We find that the general stylistic patterns in the STs are mirrored in the target texts (TTs) in terms of several linguistic measures, but that some aspects of style are not reflected in the TTs. The stylistic discrepancies between the STs and TTs may signify the translator's strategic decisions to adhere to the target language (TL) norms and translation conventions as well as to preserve the style in the ST.
This study aims to explore the influence of genre types on the compositions of Korean EFL college learners by examining multiple levels of linguistic features. A computational assessment tool called Coh-Metrix was utilized to analyze a set of 72 compositions, including cause/effect, comparison/contrast expository, and argumentative essays. The findings reveal that low frequency words and syntactically complex sentences were more frequent in expository texts (specifically, cause/effect texts) than in argumentative texts. In addition, compare/contrast essays had the highest density of noun phrases, one measure of syntactic complexity, and argumentative essays had the lowest scores for the number of words before main verbs, indicating lower syntactic complexity. This study offers valuable pedagogical implications for the instruction of diverse genres in EAP writing.
Abstract Previous studies in corpus-based literary translation have tended to focus on only one or two specific aspects of style. In this study we expand the existing analytical paradigm to show how the style inherent in source texts (STs) is reflected in their translations. We do this using thirty-six multilevel linguistic features. The selected texts are James Joyce’s Dubliners and A Portrait of the Artist as a Young Man and their Korean translations. We find that the general stylistic patterns in the STs are mirrored in the target texts (TTs) in terms of several linguistic measures, but that some aspects of style are not reflected in the TTs. The stylistic discrepancies between the STs and TTs may signify the translator’s strategic decisions to adhere to the target language (TL) norms and translation conventions as well as to preserve the style in the ST.
The main objective of this study was to analyze the continuity of reading passages in high school English mock College Scholastic Ability Test (CAST) exams using Coh-Metrix, a multi-level text analysis tool. To achieve this, the study constructed a corpus of reading passages from the 2022 high school English mock exams and subjected them to a broad range of Coh-Metrix measures. These measures included basic measures (the number of words, the number of sentences, average word length, average sentence length), word frequencies (word frequencies for content words), word features (imageability, concreteness, age of acquisition, familiarity), lexical diversity measures (type-token ratios for content words), personal pronouns (first person pronouns, second person pronouns, third person pronouns), connectives, readability indices (Flesch Reading Ease, Flesch-Kincaid Grade Level), syntactic complexity (noun density scores, the number of words before main verbs), coreference cohesion measures, and semantic cohesion measures. The main results revealed that the continuity of high school English mock tests was well established for average word and sentence length measures, the average number of words, word frequencies, word familiarity, second person pronouns, standard readability indices, and syntactic complexity measures. These findings have implications for the development of reading passages in high school English mock exams.
This study aims to investigate to what extent pre-task planning and the source text type(genre) affect Korean English as a Foreign Language (EFL) college learners' summary writings in terms of lexical, sentential, and discourse-level features. A total of 120 summary writings of cause/effect expository texts and argumentative texts in the different modes of planning were collected and analyzed using a computational assessment tool, Coh-Metrix. The results show that the participants' summary writings in the different planning conditions and text types were statistically different according to their lexical-level (the mean word length, word frequency, imageability, concreteness, the third person pronouns), sentential-level (the mean sentence length, causal connectives, temporal connectives, noun density, Flesch-Reading Ease (FRE), Flesch-Kincaid Grade Level (FKGL)), and discourse-level (type-token ratio, LSA cosines for all and adjacent sentences) features. This study provides some pedagogical implications for teaching English summary writings of various planning contexts and source text types.
교과서를 통한 학습의 효율성을 극대화하기 위해서는 교과서 텍스트의 특성이 학습자의 발달수준에 맞추어 조절되어야 한다. 동일 학년 내의 학습자의 발달 수준은 일정할 것이라 가정할 때, 동일 학년 내 여러 교과서 간 텍스트 특성에는 유의미한 차이가 없어야 한다. 본 연구는 과학교과서 텍스트에 이러한 원리가 잘 반영되어 있는지를 알아보기 위하여 수행되었다. 먼저, 중학교 1, 2, 3학년 과학교과서 각 5, 5, 4종(총 14종)의 본문 텍스트를 추출하였다. 그 다음, 각 학년별 여러 교과서 간 텍스트 특성에 차이가 있는지를 알아보기 위하여, 한국어 텍스트 분석 프로그램인 Auto-Kohesion 시스템이 제공하는 20개의 언어적 측정치(e.g., 기본 측정치, 어휘 관련 측정치, 통사적 복잡성 측정치 및 정합성 관련 측정치)에 대해 분석하였다. 분석 결과, 모든 학년에 대해 2-3개의 측정치를 제외한 대부분의 측정치에 대해 출판사 간 변이가 유의미하게 나타나지 않았다. 이러한 결과는 동일 학년 내 학습자들이 어떤 교과서를 통해 학습하더라도 학습 효과에는 유의미한 차이가 나타나지 않을 것이라는 점을 시사한다. 본 연구의 결과는 과학교과서 개발 및 효과적인 과학교육 설계에 대한 함의점을 제시한다.
The comprehension and production of second language (L2) is a key factor in understanding the course of L2 development. Several accounts have been posited to explain the differences in the L1 and L2 processing, including the minimum involvement of morphological parsing and the L1 interference. However, the evidence is not sufficient to be conclusive. This paper investigates and compares L1 and L2 processing of a Korean nominal suffix -tul by Korean L1 speakers and Chinese learners of Korean. Masked and cross-modal priming experiments were performed to examine L1 speakers
This study aims to investigate the text difficulty of the reading materials of Korean middle school English textbooks with Coh-Metrix, a software developed by the Institute for Intelligent Systems at the University of Memphis to analyze the linguistic and psycholinguistic features of English text and textbooks with a wide range of indices on cohesion and language. In this study, the textbook corpus consisted of the text files extracted from 13 English textbooks. These files were used for analyzing the text difficulty among grades with Coh-Metrix. The Coh-Metrix indices selected for this study contained basic counts, word frequency, word features, lexical diversity, pronouns, connectives, readability, syntax complexity, syntax similarity, reference cohesion, semantic cohesion, and situation model measures. The results showed that there were significant differences among grades for basic counts, word features, first pronouns, causal and temporal connectives, readability, reference and semantic cohesion, the number of words before main verbs, syntactic similarity, and situation model measures. The differences among grades, however, were not significant for word frequency, lexical diversity, second and third person pronouns, additive connectives, and NP density measures. The findings have educational implications for textbook design and language learning for English learners.