
Kang, Yein and Jonny Jungyun Kim. 2026. Does AI reflect sound change? A comparison of human and TTS production of Korean stops. Linguistic Research 43(1): 105-140. This study presents an acoustic analysis of speech generated by 60 commercial text-to-speech (TTS) voices to examine whether socially indexed phonetic variation is reflected in artificial speech systems. Focusing on cross-generational variation in phrase-initial Korean stops, we investigated the extent to which voices labeled as "younger" across three TTS platforms approximate the ongoing merger in voice onset time (VOT) between aspirated and lenis stops, as well as a compensatory cue-shift toward fundamental frequency (F0). Younger TTS voices generally showed reduced reliance on VOT cues: for example, IP-initial aspirated and lenis stops differed by 12ms in younger TTS voices, compared to 21ms in older TTS voices. However, the magnitude of this merger in younger TTS voices differed substantially from that observed in younger human speakers, who exhibited an almost complete VOT merger (e.g., a 1ms difference, IP-initially). In contrast, F0 patterns in TTS speech broadly aligned with human speech. The results suggest a potential link between socially indexed phonetic realization and naturalness fidelity in AI-generated speech. We further outline future directions for examining AI speech from a sociophonetic perspective, focusing on fundamental differences between human and machine learning mechanisms. (Pusan National University)
Dita, Shirley N. and Jong-Bok Kim. 2026. Answering polar questions in Tagalog: A discourse-based analysis. Linguistic Research 43(1): 31-59. This paper explores strategies for forming and answering polar questions in Tagalog. While K & ouml;nig and Siemund (2007) identify six strategies for generating polar questions cross-linguistically, only three are evident in Tagalog: special intonation patterns, interrogative particles, and special tags. More significantly, this study reveals that Tagalog exhibits flexibility in its answering system, employing all three response strategies described by Sadock and Zwicky (1985)- yes/no (polarity), agree/disagree (truth), and echo systems. Using authentic examples from the web-based corpus tlTenTen19, we demonstrate that Tagalog response particles oo and hindi function as anaphoric elements referring to salient propositions in discourse, challenging simple binary typologies and supporting recent discourse-based accounts of response particles (Kim 2024). This analysis contributes to our understanding of Philippine languages and cross-linguistic variation in answering systems. (De La Salle University & centerdot;Kyung Hee University)
In sociophonetics and laboratory phonology, researchers often invoke "listener sensitivity," meaning that some individuals or groups notice and differentiate fine-grained phonetic variation more than others, yet this construct is rarely operationalized as a single gradient measure. We introduce a simple, generalizable metric, the sensitivity score, defined as the sum of absolute differences between a participant's mean ratings across all unordered pairs of variants. We demonstrate the metric using a perception experiment on Korean aegyo, a cuteness-indexed register with segmental alternations including nasal insertion, /j/-insertion, and vowel rounding. Forty-three native Korean listeners rated 90 spoken stimuli each (15 variants & times; 6 predicates; 3870 ratings total), produced by two talkers, on a 1-7 aegyofulness scale. We analyze the token-level ratings with a linear mixed-effects model comparable to Jang (2021) and analyze sensitivity scores computed within each predicate (six scores per participant; 258 scores total) with a second mixed-effects model. The token-level model replicates broad social patterning in aegyo evaluations, while the sensitivity score provides a transparent participant-and item-level summary of overall differentiation. Women show higher sensitivity than men, and age shows a trend, such that younger listeners show higher sensitivity than older listeners. Aggregating ratings into sensitivity scores reduces the number of observations relative to token-level modeling, so we treat the metric as complementary to traditional statistical methods. Beyond aegyo, the approach is portable to other sociophonetic, prosodic, and sound-symbolic domains.
This study demonstrates that a muscle-based approach provides an explanatory account of the affrication and palatalization of Japanese /t/ that are difficult to explain using conventional phonological approaches. The central asymmetry is that /t/ becomes [tc] with a posterior place before front high /i/, whereas it becomes [ts] without change in place of articulation before non-front high /w/. The results of the simulations using the 3D tongue model of ArtiSynth show that the shape of the tongue for the letter /t/ varies depending on the coarticulated tautosyllabic vowels. This is because the activated muscle bundle of the tongue differs based on the vowels. In the context of /i/, a muscle activation lowering the tip of the tongue is essential to keep the tip of the tongue from going out of the mouth. The palatalization of /t/ is characterized by a lowered tongue tip and an elevated tongue body. In the coarticulation of tautosyllabic /t/ and /w/, the activated tongue muscle raised the entire tongue without significantly lowering the tongue tip, causing affrication of /t/. The affrication and palatalization of Japanese /t/ are formally captured in an Optimality-Theoretic (OT) analysis as assimilations of constriction orientation (CO) between the tongue tip and body in tautosyllabic onset-vowel sequences, grounded in a muscle-based account and gestural representations. The proposed OT grammar that uses AGREE(CO) and IDENT(CO) can derive the Japanese /t/ alternation patterns. (Sungkyunkwan University)
Recent studies have documented the raising of the mid-back vowel [o] toward the high-back vowel [u] in Seoul Korean (SK), raising questions about whether this shift reflects a merger or a chain shift. This study investigates whether [o]-raising induces a reorganization of the vowel space involving adjacent vowels [u], [Lambda], and [w]. Analyzing vowel productions from 22 SK speakers (16 female, 6 male) across two speech styles (careful vs. conversational), we examined acoustic properties of three vowel pairs: [o]-[u], [w]-[u], and [Lambda]-[o]. Results show that [u] remains distinct from [o] and is produced in a significantly more fronted position, especially in conversational speech. Additionally, [u] was shifted closer to [w], and [Lambda] raised, suggesting a broader vowel chain shift. These findings provide acoustic evidence of an ongoing, systemic restructuring of the SK vowel system. (Chungbuk National University & centerdot; Korea Aerospace University & centerdot; University of Oregon)
In a picture-based comprehension task (similar to an act-out task) used in the original study, L1-English speakers interpreted English definite plurals as referring to (i) a set of previously mentioned objects and (ii) all objects within the context. However, L1-Korean L2-English speakers allowed only the former. This non-targetlike behavior was attributed to L1 transfer from the semantics of Korean demonstratives. Methodologically, the task might reflect participants' preference rather than their representations for definite plurals since participants could express only one possible interpretation in the picture-based comprehension task. We thus modified the task changing it into a picture-based acceptability judgement task wherein participants were shown both interpretations. Participants were thirty-three L1-English and thirty-one L1-Korean L2-English speakers. Ordinal logistic regression models revealed that L1-Korean participants accepted both interpretations. We argue that non-targetlike behavior in the original study is attributable to task effects. We conclude that L1-Korean speakers compute both interpretations and have targetlike representations of English definite plurals. (University of Wisconsin-Madison & centerdot; University of Southern California)
This study investigates how abstract-level discourse features relate to citation impact when considered alongside topical structure. A harmonized set of 436 English-language articles and reviews (1995-2024) from the Web of Science (Linguistics, Language & Linguistics, Education & Educational Research) was analyzed. Topic structure was mapped using bibliometrix through the Biblioshiny interface with all keywords (DE + ID; minimum frequency = 5), producing a co-word network, a thematic map (centrality & times; density), and period-wise evolution. Discourse was profiled with five Coh-Metrix composites-narrativity, syntactic simplicity, concreteness, referential cohesion, and deep cohesion-standardized as z-scores. Citation impact was normalized as average citations per year (TCperYear). Annual production peaked in 2003-2006, whereas TCperYear peaked later, indicating lagged diffusion. Thematic mapping showed a Motor axis centered on language-meaning-cognition, a Basic infrastructure of corpus/syntax/English, and bridging roles for computational linguistics and machine learning. In OLS with HC3 errors, syntactic simplicity and referential cohesion had small but significant negative associations with TCperYear; other indices were not significant. Results were robust in auxiliary checks and showed low multicollinearity (all VIFs < 2). Findings suggest that authors should maintain information density, limit mechanical repetition, and use only as much syntactic complexity as needed for precision-especially in high-centrality topics. (Silla University)
This study examines word-initial prevocalic tensification in English loanword adaptation in Korean from a perceptual perspective. Loanword tensification refers to the phenomenon whereby an initial consonant that would otherwise correspond to a Korean lax stop is realized as its tense counterpart depending on the linguistic context. The goal of the present study is to explain how three conditioning factors reported in prior work-(i) the type of the word-initial consonant, (ii) the height of the following vowel, and (iii) the tenseness of the following syllable onset-are reflected in the perception of tensification. To this end, an ABX perception experiment was conducted with native Korean listeners, using disyllabic English nonce words in which the initial consonant ([b, d, s]), vowel height (high, non-high), and the second-syllable onset ([s], [s']) were systematically manipulated. The results showed that perceived tensification was highest for [s]-initial stimuli, increased in non-high vowel contexts relative to high vowel contexts, and was also higher when the following onset was tense ([s']). These findings can be interpreted as indicating that loanword categorization is not determined solely by a direct mapping of acoustic cues; rather, listeners integrate acoustic evidence with internalized expectations derived from Korean phonological distributions.
This study investigates adjunct constraint (AC) and relative clause constraint (RCC) effects in Korean scrambling constructions using a factorial design that carefully controls relevant functional factors. Across two experiments, both AC-and RCC-violating sentences exhibited uniformly high acceptability, indicating that Korean permits extraction from these domains when functional well-formedness is ensured. Despite the overall high acceptability, RCC effects were still observed, whereas no AC effects emerged. Given previous reports of null RCC effects, the RCC effects found in this study may reflect residual functional influences- one possible source being a violation of the Backgrounded Constructions are Islands (BCI) constraint-though further research is required to determine their precise origin. Overall, the findings underscore the importance of controlling functional factors as rigorously as possible when applying the factorial definition of island effects.
This article aims to offer some unified analysis of a couple of irregular stems in Korean. The l-gemination attested in /li/-irregular stems turns out to be compensatory in nature, and so the phenomenon comes under the purview of moraic theory. The /t/-irregular stems are basically shown to be conditioned in the lexicon. But it is argued that their allomorphic choice becomes phonologically natural once the stem is assumed to carry two underlying allomorphic variants in the lexicon. Those two allomorphs are closely related in that they have the feature value of [voice] unspecified in the input. The value will be filled in with the help of a universal constraint in charge of agreement. Then, what appears to be irregular in the allomorphy of /t/-irregular stems in question is not a phonological irregularity at all, but rather results from a different match between morphemes that under more normal circumstances would not be expected to happen. On the other hand, emergence of nondistinctive [d] and [i] in the conjugation of regular as well as irregular stems should not be something that happens within phonology. It must be an allophonic process taking place in the environment of intervocalic position, so it stands reason to treat it in terms of phonetic implementation.
Ahn, Miyeon and Gwanhi Yun. 2026. An ultrasound study on articulatory variations of Korean [n]. Linguistic Research 43(1): 185-208. The present study aims to explore the articulatory consequences of three types of Korean [n]- canonical, inserted and nasalized-that are derived from phonologically different input representations. We hypothesized that the three [n] sounds are articulatory variants arising from differences in speakers'distinct production planning, shaped by their consideration of phonological rule applications. An ultrasound study was conducted to test this hypothesis, and the data revealed that, for half of the participants, ascending tongue tip gestures were observed in the order of canonical, nasal and inserted, corresponding to increasing phonological processing complexity, while the other half did not show this pattern. Speakers' production of articulatory variants suggests that they make extra articulatory efforts to reach the target more precisely for phonologically more complex variants and that they plan their speech to deliver their rule applications. (Hankyong National University & centerdot;Daegu University)
Prosodic features of second language (L2) speech, encompassing temporal and intonational dimensions, are closely linked to listeners' perceived fluency. In stark contrast to extensive literature on lexical tone production, the prosodic fluency of Korean learners of Mandarin across different proficiency levels remains underexplored. This study aims to fill this gap by analyzing temporal and intonational correlates of prosodic fluency in Korean learners of Mandarin, using sentence-reading data drawn from a large-scale L2 speech corpus developed for AI training that provides balanced proficiency representation across beginner-and advanced-level learners. Results show that articulation rate, the number of non-comma pauses (i.e., pauses not corresponding to orthographic commas), and the degree of pitch declination across an utterance are most closely related to perceived fluency, with articulation rate emerging as the strongest predictor. These findings not only advance our understanding of the development of L2 Mandarin prosody but also shed light on the fluency features that should be prioritized in the design of automated speaking proficiency assessments for Korean learners. (Seoul National University)
This study compares the Korean adjective chakhata ("kind, good-natured") and the Russian adjective dobryj ("kind, good") using collocational and semantic network analyses. Drawing on large-scale web-crawled corpora from Sketch Engine, the study examines how the shared concept of "goodness" is structured in Korean and Russian discourse. The results show clear cross-linguistic differences in collocational distribution and semantic organization. Chakhata predominantly collocates with nouns in the [person] category and forms a tightly connected semantic network centered on normative evaluation. In contrast, dobryj appears across a broader range of conceptual domains, including [emotion], [communication], [cognition], and [quantity], and exhibits a more radial semantic structure extending into abstract evaluative meanings. These patterns point to different evaluative orientations. Chakhata tends to encode norm-based moral judgment focused on socially evaluated persons, whereas dobryj more often conveys affective warmth and communal orientation. Both adjectives also allow paradoxical or ironic uses, in which positive evaluation is contextually inverted by culturally specific expectations. The findings show that evaluative adjectives are organized into culturally specific semantic networks, through which shared notions of "goodness" are structured by distinct moral and affective frameworks in Korean and Russian discourse.
This study examines how case connectivity in Korean fragments is shaped by the interaction between structural licensing and cue-based interpretation. Two acceptability-judgment experiments tested whether case (mis)matching between a remnant and its correlate is modulated by the case-licensing range of the elided predicate. Across both sluicing and why-stripping, we found a robust main effect of case matching and a strong MATCH x VERB interaction: mismatch penalties were substantially smaller when the predicate licensed both dative and accusative case, but sharply degraded when it licensed only dative case. To evaluate competing theories, we compared models representing the silent-structure approach, the interpretive approach, and a hybrid approach. Structural models captured the categorical unacceptability of mismatches with non-alternating verbs, whereas interpretive models captured the general MATCH advantage and the presence of non-categorical mismatch penalties, but did not predict the categorical verb-conditioned interaction. Neither approach alone accounted for the full pattern. Hybrid models, which incorporate both categorical licensing constraints and gradient cue-based effects, provided the best overall fit. The findings show that fragment interpretation in Korean is jointly determined by structural and processing mechanisms: argument-structure identity restricts the set of grammatically licit cases, while cue-based retrieval yields gradience within the structurally permitted domain. These results situate Korean within broader cross-linguistic theories of case connectivity and the syntax-processing interface
Linguistic Research 42(3): 603-625. Contrary to the traditional view that a pronoun is uniformly a DP, recent approaches have shown that a pronoun is not syntactically uniform but is realized as different categories (e.g., Dechaine and Wiltschko 2002; Patel-Grosz and Grosz 2017). I show that the behavior of the 3(rd) person singular pronoun ku in Korean provides support for the non-uniform view. I propose that the pronoun ku instantiates two different categories, pro-DP and pro-phi P, by building on the interpretational differences of ku. The pronoun ku is ambiguous having referential or bound variable reading similar to pro-DP or pro-phi P in other languages such as German or Halkomemlem (Salish). This paper also examines two important characteristics of the proposed structure of the pronoun ku-its determiner use and a null NP-which has not been addressed in the previous studies on ku. The pronoun ku as pro-DP is used as a determiner when the NP is overt, as identified by the non-uniform approach. This paper shows that this is possible as the pronoun and determiner share the same structural core, namely indexP. As for a null NP in the structure of the pronoun, I propose that the phi head is a licensor of an NP ellipsis by providing evidence from the distribution of plural-tul. This paper contributes to the current debate on the syntactic and semantic similarities between pronouns and definite descriptions (e.g., Elborne 2005, 2008). (Chungbuk National University)
(2007) argued that Americans are "persuasive" while Brits are "brutal," based on differences in the transitive into-ing construction, adopting the distinctive collexeme analysis. In the use of this construction, American English typically prefers communication and persuasion verbs (e.g., talk and coax), as in They talked Cassie into breaking up the nuptials, whereas British English more favorably selects force and negative emotion verbs (e.g., pressurize and bully). This study examines whether Wulff et al.'s (2007) generalization "persuasive Americans vs. brutal Brits" also extends to the transitive out of-ing construction (e.g., They talked Cassie out of breaking up the nuptials). Drawing on large-scale corpus data from GloWbE and NOW (2010 to May 2025), we analyze verbs in the V1 slot using collostructional methods (collexeme, distinctive collexeme, and covarying collexeme analyses). Our findings reveal that the generalization largely holds: American English exhibits a broader range of verbs, with communication and persuasion strongly represented (e.g., talk), while British English favors force verbs (e.g., rule and price). At the same time, notable deviations emerge: American English also shows strong distinctive associations with negative emotion verbs (e.g., scare, intimidate, and shame), a preference not observed in British English. These results broaden our understanding of the transitive out of-ing construction and refine the claim that Americans are "persuasive" and Brits are "brutal," demonstrating that although the dialectal divide is robust, its expression shifts in systematic ways across related constructions.
This study investigated the phonetic aspects of the the vowels /o/ and /u/ in Seoul Korean, examining how positional (word-initial vs. word-final), morphological (content vs. function morphemes), and sociolinguistic (gender and age) factors affect the productions of the two vowels. Using an experimental design that compared vowel realizations within matched word positions, the results revealed that overlapping patterns occur more frequently in function morphemes than in content morphemes, particularly along the vowel height dimension (F1). When morphological category was held constant, word-initial positions provided more favorable conditions for vowel convergence than word-final ones. Age-related patterns also emerged, such that the change began with speakers born in the 1950s. Younger speakers exhibited greater acoustic overlap between /o/ and /u/, with female speakers showing consistent /o/-raising, while male speakers displayed more variable and less systematic patterns. Regarding the morphological factor, the merger was most evident in word-final function morphemes, likely due to their high frequency, low semantic weight, and prosodic environments favoring phonetic reduction. Furthermore, results from the Linear Discriminant Analysis (LDA) reliably captured these production realizations, reflecting the observed patterns of vowel convergence across social and linguistic factors. The findings of this study highlight the complex interplay between linguistic structure and sociolinguistic factors in shaping the overlapping patterns observed in the two vowels. (Hannam University)
This research explores the syntactic processing of Large Language Models (LLMs), specifically GPT-3.5 and GPT-4, by comparing them to human processors, focusing on garden-path sentences. These structures are challenging for even proficient human processors, often causing misinterpretations that persist despite reanalysis, revealing the 'good-enough' nature of human syntactic processing. This study aims to determine if LLMs exhibit a similar 'good-enough' syntactic processing as humans and whether more advanced models exhibit a more human-like processing. In a series of experiments, we examined how models handle garden-path sentences such as "While the man hunted the deer ran into the woods," through a comprehension questions task. A key focus was whether misinterpretations in the target phrases ("hunted the deer") erroneously affected the global interpretation of the sentence. Results showed that LLMs display patterns similar to humans, including lingering misinterpretations and the ability to utilize linguistic cues such as plausibility, phrase length, and verb type. This suggests that LLMs mimic human 'good-enough' syntactic processing through probabilistic next-word prediction, including making human-like errors. However, LLMs also showed vulnerability to garden-path structures, showing a higher rate of errors compared to humans, likely due to inherent features of their processing mechanisms. (Korea University Sejong Campus Dongguk University)
As Korean phonotactic constraints do not allow consonant clusters in the syllable coda position, tautosyllabic clusters are realized as a single (first or second) consonant of a cluster when followed by another consonant. However, recent phonetic studies have reported that clusters with a lateral and obstruent have been undergoing a change in preserving the first consonant, or even both consonants. Across two experiments, this study investigated Korean speakers' perceptual and lexical encoding of an innovative variant (i.e., [lp]) of the cluster /lp/ where the on-going change is most noticeable. Experiment 1 involved a speeded AX discrimination task to examine the perception of the innovative variant. Results revealed that Korean speakers had difficulty discriminating word pairs when one member of the pair contained [lp], suggesting that this variant is not categorically perceived. In Experiment 2, a long-term repetition priming lexical decision task was performed to observe the storage of the innovative variant in the lexicon. Significant priming was found for this variant, which was equivalent to that of the identity pairs, indicating that the innovative variant [lp] is stored as a separate category in long-term memory.
Ambisyllabicity has long been central to phonological theory, yet its articulatory basis remains unclear. This study investigates whether American English ambisyllabic retroflex /(sic)/ exhibits intermediate tongue configurations between onset and coda realizations in spatial and/or temporal dimensions. Ultrasound imaging data were collected from four native speakers producing five /(sic)/ types: three intervocalic retroflexes (an ambisyllabic retroflex preceded by a stressed lax vowel; a non-ambisyllabic retroflex preceded by a stressed tense vowel, hereafter non_A; and a non-ambisyllabic retroflex followed by a stressed vowel, hereafter non_B), as well as word-initial onset and word-final coda retroflexes. Tongue contours for each intervocalic retroflex were compared with those of onset and coda retroflexes using generalized additive mixed modeling (GAMM). Across speakers, intervocalic retroflexes followed a consistent trajectory, shifting from intermediate positions between onset and coda toward onset-like configurations. Crucially, this intermediate status was not unique to ambisyllabic retroflexes but was observed across all intervocalic contexts, suggesting that ambisyllabicity lacks a stable articulatory correlate and functions primarily as a theoretical construct. In addition, non_B retroflexes shifted toward onset-like tongue positions as early as the medial time point, whereas ambisyllabic and non_A retroflexes did so only at the final stage of articulation. This pattern indicates that intervocalic retroflexes preceded by a stressed vowel and followed by an unstressed vowel may be regarded as constituting an independent allophone, characterized by intermediate tongue contours that are distinct from onset and coda allophones. (University of Seoul)