지난 70여 년 동안 인공지능 기반 기계번역 기술은 외국어 교육과 애증의 관계를 맺어왔다. 외국어 교육의 일차적인 목표가 해당 학습자의 외국어 사용능력을 기르는 것이므로, 기계번역 시스템의 성능이 고도화되면 해당 외국어 전문가에 대한 수요가 급격하게 준다고 생각했기 때문이다. 꽤 오랜 기간 기계번역 시스템의 성능은 사용자의 기대를 충족하지 못했다. 이는 외국어 교육 현장에서 기계번역 기술의 도입을 반대하거나 주저하는 근거가 되었다. 하지만 2010년대 중후반부터는 이러한 부정적인 시각이 약화되기 시작했고, 외국어 교육뿐만 아니라 번역교육과 전문적인 번역 현장에서 기계번역 시스템의 효과적인 도입과 활용에 대해 활발한 논의가 진행되고 있다. 또한, 2022년 말부터 선보인 거대언어모델 기반 생성형 AI 시스템은 비단 외국어 교육만이 아니라 모든 분야의 교육에 커다란 변화를 초래할 것이다. 필자는 1980년대 말부터 기계번역 기술의 토대가 되는 자연언어처리 분야에서 이론 연구와 더불어 컨텐츠 및 시스템 개발에 참여해 왔고, 그 성과를 외국어 교육 특히 프랑스어 교육 분야에 도입하려고 노력했다. 이 논문에서는 지난 35년간 매 시기에 최신 단계였던 자연언어처리 기술을 프랑스어 교육에 접목했던 필자의 경험을 소개함으로써, 앞으로도 어떤 형태로 나타날지 모르는 미래기술을 교육 분야에 접목하는 작은 방향등 역할이 되고자 한다.
2010년대 중반부터 기계번역 기술이 비약적으로 발전하면서, 새로운 기술과 환경변화를 고려한 다양한 외국어 교수-학습 방법이 모색되고 있다. 본 연구진은 최근 기계번역 기술의 발전 방향을 고려하여, 거시적·미시적 관점에서 프랑스어 교수-학습 방법론을 제안하고, 실제 운영 방식과 효과를 보고해왔다. 이 논문에서는 대학 프랑스어문법 교과과정 중 심화단계 교과목의 교수-학습 방법론 및 새로운 평가방식을 제안하고, 이를 적용한 후, 학습성취도를 세밀하게 분석했다. 특히, 교수-학습 과정뿐만 아니라 모든 시험에서 전자사전과 기계번역기를 사용할 수 있도록 함으로써 완전히 새로운 평가 환경을 적용했다. 본 연구의 결과와 시사점은 크게 3가지로 요약할 수 있다. 첫째, 프랑스어 숙달도 평가에서 전통적으로 중요한 요소로 꼽아온 어휘의 의미, 불규칙 변화형 등 단순한 언어정보의 암기는 최종 학습성취도에 크게 영향을 주지 않는다. 둘째, 학기초에 측정된 학습자의 숙달도가 학기말의 숙달도를 전적으로 좌우할 것이라는 일부 전문가들의 의견과는 달리 두 숙달도는 비례하지 않는다. 즉, 문장구조 분석에 역점을 둔 본 연구의 교수-학습 방식이 학습자의 프랑스어 숙달도를 향상할 수 있다. 셋째, 새로운 세대인 현재 학생들은 본 교과목에 적용한 새로운 교수-학습 및 평가 방식을 매우 긍정적으로 평가한다.
기계번역 기술은 2010년대 중반부터 전세계적으로 급속하게 발전했으며, 국내에도 2018년을 기점으로 기계번역기가 일상생활에 자리잡기 시작했다. 이 기술의 발전은 외국어 학습의 필요성, 외국어 관련 학과의 존재이유와 목표에 대해 끊임없이 재고하도록 요구할 것이다. 현재의 전망으로는, 일반적인 정보 전달은 기계번역기 사용이 보편화될 것이고, 인간번역은 범위가 아주 축소되고, 매우 전문적인 영역에서 지속될 것으로 보인다. 앞으로 통 · 번역 관련 일을 하게 될 현재의 학생들은 고도의 외국어 숙달도와 해당 분야의 전문지식이라는 두 마리 토끼를 모두 잡아야 한다. 또한 당장 통번역을 하지 않더라도 졸업 후 필요할 때 언제라도 프랑스어와 같은 제2 외국어를 사용하려는 전공자들은, 기계번역 결과를 수정할 수 있을 정도의 복원력이 강한 지식 구조를 갖춰야 한다.BR 본 연구는 성인의 외국어 학습에서 효율적으로 알려진 문법 지식에 주목하고, 국내 대학의 프랑스어 교수-학습 상황과 기계번역 기술의 발전방향을 충분히 고려한 문법의 교과과정을 제안했다. 그 특징은 다음과 같이 요약할 수 있다. 첫째, 거시적 관점에서 필수 교과목을 중심으로 선택 교과목을 나선형으로 배열함으로써, 수강하는 교과목의 수와 순서에 따른 모든 경우에 문법 지식을 점층적으로 확장할 수 있도록 했다. 둘째, 미시적 관점에서는 각 교과목의 세부적인 교수-학습 내용 간 치밀한 연계성을 기획하여, 교과과정의 확장성을 강화할 수 있도록 설계했다. 셋째, 암기에 기반한 학습과 평가를 지양하고, 전자사전 등 검색 프로그램과 기계번역 기술을 적극적으로 교수-학습 과정과 성취도 평가에 활용하는 방법론의 적용 원칙을 제안했다.
요 약 1980년대 중반부터 지난 20여 년간 구축해 온 영어 워드넷(PWN)은 인간의 심상어휘집을 재현하려는 목적으 로 개발되기 시작하였으나, 그 활용 가능성에 주목한 것은 자연언어처리와 지식공학 분야다. 컴퓨터 매개 의사 소통(CMC), 인간-컴퓨터 상호작용(HCI)에서 인간 언어를 자연스럽게 사용하여 필요한 정보를 획득하기 위해서 는 의미와 지식의 처리가 필수적인데, 그 해결의 실마리를 어휘라는 실체를 가진 언어단위에서 찾을 수 있기 때문이다. 이후 전 세계적으로 약 50개 언어의 어휘의미망이 PWN을 참조모델로 구축되어 다국어처리의 기반 을 제공할 뿐 아니라, 시맨틱 웹 이후 더욱 주목 받고 다양한 방식으로 활용되고 있다. 본고는 PWN을 참조 모 델로 2004년부터 2007년까지 구축한 한국어 어휘의미망 KorLex 1.5를 소개하는 데 있다. 현재 KorLex은 명 사, 동사, 형용사, 부사 및 분류사로 구성되며, 약 13만 개의 신셋과 약 15만 개의 어의를 포함하고 있다.
Context-sensitive spelling errors (CSSE) are hard to correct, since they are perfect words when analyzed alone. Determined only by considering the semantic and syntactic relations of their context, CSSEs affect largely the performance of spelling and grammar checkers. The existing Korean Spelling and Grammar Checker (KSGC 4.5) adopts a rule-based method, which uses hand-made correction rules for CSSEs. Using rule-based method, the KSGC 4.5 is designed to obtain the very high precision, which results in the extremely low recall. In this paper, we integrate our previous works that control the CSSE correction rules, in order to improve the recall without sacrificing the precision. In addition to the integration, facultative insertion of adverbs and conjugation suffix of predicates are also considered, as for constraint-loosening linguistic features.
The types of errors corrected by a Korean spelling and grammar checker can be classified into isolated-term spelling errors and context-sensitive spelling errors (CSSE). CSSEs are difficult to detect and to correct, since they are correct words when examined alone. Thus, they can be corrected only by considering the semantic and syntactic relations to their context. CSSEs, which are frequently made even by expert wiriters, significantly affect the reliability of spelling and grammar checkers. An existing Korean spelling and grammar checker developed by P University (KSGC 4.5) adopts hand-made correction rules for correcting CSSEs. The KSGC 4.5 is designed to obtain very high precision, which results in an extremely low recall. Our overall goal of previous works was to improve the recall without considerably lowering the precision, by generalizing CSSE correction rules that mainly depend on linguistic knowledge. A variety of rule-based methods has been proposed in previous works, and the best performance showed 95.19% of average precision and 37.56% of recall. This study thus proposes a statistics based method using a conditional probability model with dynamic window sizes. in order to further improve the recall. The proposed method obtained 97.23% of average precision and 50.50% of recall.
Like machine translation between natural languages, good quality dictionary is necessary for automatic translation between Korean and Korean sign language. Korean Sign Language dictionary, which is currently regarded as the most reliable sign language dictionary, contains only 12,000 words including compound words and idiomatic phrases, so it does not provide enough information as a dictionary. In order to solve target word selection problems occurring when Korean is translated into Korean sign language, this study suggests a method of using dictionary definitions in the Korean standard unabridged dictionary along with relational words information of Korean WordNet. When it is impossible to select Korean sign language words for corresponding Korean words due to insufficiency of target words, target words are selected from relational words of Korean WordNet, or main words in dictionary definitions of Korean words are used to translate Korean words into Korean sign language.
The most common method for detection and correction of spelling errors in Korean Spelling and Grammar Checker (KSGC) depends largely on the linguistic rules. Rulebased methods yield very high precision, but extremely low recall because the constituents of the rules should be matched exactly. In this paper, we propose two novel methods that can loosen the case-particle constraints of KSGC’s existing rules, in order to improve the recall of Context-Sensitive Spelling Correction: (1) substitution of each case-particle by adding auxiliary particles or particle-compounds; (2) omission of case-particles.
Error words that appear in Korean texts can be largely categorized into non-word spelling errors and context-sensitive spelling errors. Of the two, context-sensitive spelling errors are shown only when considering the meaning of the word in the given context and its syntactic relation, and they are the most difficult to correct among spelling errors. Context-sensitive spelling errors can be categorized into homophone errors, typographical errors, grammatical errors, and cross-word boundary errors. To correct context-sensitive spelling errors that occur due to typographical errors, this study proposes a statistical context-sensitive spelling check using confusion sets. With confusion sets created in advance, we can find and correct context-sensitive spelling errors using reliability based on the conditional probability between each word of the confusion sets and the context as well as the typing error rate. As a result of applying the proposed method, all 5 confusion sets showed higher precision and recall than the baseline (precision 80%, recall 80%).
이 글은 한국어 어휘 의미망 KorLex 2.0의 구축 방법론과 정보 구조를 비롯한 특성을 소개한다. 영어 워드넷(PWN) 2.0을 모델로 하여, KorLex 1.0에서는 PWN의 신셋/어의에 등가성을 갖는 한국어 어의를 찾아 신셋을 구축하고, 1.5판에서는 영어에 경도된 PWN 신셋의 구조를 수정하고 한국어 어휘 의미 관계를 반영하였으며, 2.0판에서는 한국어라는 개별 언어에 의존적인 정보를 추가한다. KorLex 2.0의 한국어 의존적인 정보로는 첫째, 수분류사와 명사 간 공기 관계를 설정했고, 둘째, 용언의 논항구조와 선택 제약 정보를 구축했다. 또한 이렇게 구축된 정보가 KorLex 내부에서 상호 연결됨으로써, 정합성과 확장성을 갖도록 설계하였다. 특히 KorLex에 『표준』과 『세종』과 같은 기존의 언어 자원 정보를 통합함으로써, 객관성과 일관성을 유지할 수 있을 뿐 아니라, KorLex의 정보를 풍요롭게 만들어 이를 이용하게 될 다양한 사용자의 요구에 부응하고자 하였다. KorLex는 PWN의 장점인 (가) 지식/의미 단위로서의 구체성, (나) 범용 시스템을 개발할 수 있는 크기, (다) 계층 구조의 효율성, (라)다국어 연계성을 공유한다. 한국어에 유효한 정보를 추가한 KorLex 2.0은 의미 처리와 지식 공학의 기반 언어 자원으로 기능하리라고 기대한다.
This paper presents a task model and task ontology based on travelers’ tasks, and an intelligent tourist information service system using them. With the recent advances in Internet and mobile technologies, there has been an increase in the use of intelligent tourist information services via the Web and mobile systems. In addition, many ontologies have been introduced in tourist domain to provide various intelligent tourist information services. However, only a few studies have been undertaken from the perspective of specific tour services for travelers’ generic tasks and activities. Thus, we considered generic tasks and task ontology based on travelers’ perspectives, and intelligent tourist information services using them. Therefore, we propose 1) a task model of travelers’ perspective based on their needs and activities, 2) a task ontology using the generic tasks, their activities, relations, and properties, and 3) an intelligent tourist information system using task ontology based on various tasks and activities of travelers. The system consists of Tourist Contents Service (TCS) and Task-Orient Menu Service (TMS) parts, and can provide various intelligent tourist information services through task-oriented menus.
This paper proposes a statistical-based linguistic methodology for automatic mapping between large-scale heterogeneous languages resources for NLP applications in general. As a particular case, it treats automatic mapping between two large-scale heterogeneous Korean language resources: Sejong Semantic Classes (SJSC) in the Sejong Electronic Dictionary (SJD) and nouns in KorLex. KorLex is a large-scale Korean WordNet, but it lacks syntactic information. SJD contains refined semantic-syntactic information, with semantic labels depending on SJSC, but the list of its entry words is much smaller than that of KorLex. The goal of our study is to build a rich language resource by integrating useful information within SJD into KorLex. In this paper, we use both linguistic and statistical methods for constructing an automatic mapping methodology. The linguistic aspect of the methodology focuses on the following three linguistic clues: monosemy/polysemy of word forms, instances (example words), and semantically related words. The statistical aspect of the methodology uses the three statistical formulae X2, Mutual Information and Information Gain to obtain candidate synsets. Compared with the performance of manual mapping, the automatic mapping based on our proposed statistical linguistic methods shows good performance rates in terms of correctness, specifically giving recall 0.838, precision 0.718 , and F1 0.774. (Pusan National University)