One way to facilitate Multilingual Information Access (MLIA) for digital libraries is to generate multilingual metadata records by applying Machine Translation (MT) techniques. Current online MT services are available and affordable, but are not always effective for creating multilingual metadata records. In this study, we implemented 3 different MT strategies and evaluated their performance when translating English metadata records to Chinese and Spanish. These strategies included combining MT results from 3 online MT systems (Google, Bing, and Yahoo!) with and without additional linguistic resources, such as manually-generated parallel corpora, and metadata records in the two target languages obtained from international partners. The open-source statistical MT platform Moses was applied to design and implement the three translation strategies. Human evaluation of the MT results using adequacy and fluency demonstrated that two of the strategies produced higher quality translations than individual online MT systems for both languages. Especially, adding small, manually-generated parallel corpora of metadata records significantly improved translation performance. Our study suggested an effective and efficient MT approach for providing multilingual services for digital collections.
One way to facilitate Multilingual Information Access (MLIA) for digital libraries is to generate multilingual metadata records by applying Machine Translation (MT) techniques. Current online MT services are available and affordable, but are not always effective for creating multilingual metadata records. In this study, we implemented 3 different MT strategies and evaluated their performance when translating English metadata records to Chinese and Spanish. These strategies included combining MT results from 3 online MT systems (Google, Bing, and Yahoo!) with and without additional linguistic resources, such as manually-generated parallel corpora, and metadata records in the two target languages obtained from international partners. The open-source statistical MT platform Moses was applied to design and implement the three translation strategies. Human evaluation of the MT results using adequacy and fluency demonstrated that two of the strategies produced higher quality translations than individual online MT systems for both languages. Especially, adding small, manually-generated parallel corpora of metadata records significantly improved translation performance. Our study suggested an effective and efficient MT approach for providing multilingual services for digital collections.
ABSTRACTThis paper presents the background, research design, and current progress of a new project on exploring the application of various machine translation strategies working toward multilingual information access for digital collections.
We conducted a research project exploring machine translation performance on digital metadata records. This short paper reports the background, research purposes, research design, experiments, and evaluation results.
ABSTRACTWe present herein a language identifier developed for metadata records. Trained with European Parallel Corpora, this program can distinguish English metadata records from those of other languages. We evaluated the program's accuracy with a test collection of 800 metadata records and compared the program's accuracy to that of an open‐source language identification program.
This paper describes the role of machine translation (MT) for multilingual information access, a service that is desired by digital libraries that wish to provide cross-cultural access to their collections. To understand the performance of MT, we have developed HeMT: an integrated multilingual evaluation platform (http://txcdk-v10.unt.edu/HeMT/) to facilitate human evaluation of machine translation. The results of human evaluation using HeMT on three online MT services are reported. Challenges and benefits of crowdsourcing and collaboration based on our experience are discussed. Additionally, we present the analysis of the translation errors and propose Multi-engine MT strategies to improve translation performance.
PurposeThe purpose of this study is to evaluate freely available machine translation (MT) services' performance in translating metadata records.Design/methodology/approachRandomly selected metadata records were translated from English into Chinese using Google, Bing, and SYSTRAN MT systems. These translations were then evaluated using a five point scale for both fluency and adequacy. Missing count (words not translated) and incorrect count (words incorrectly translated) were also recorded.FindingsConcerning both fluency and adequacy, Google and Bing's translations of more than 70 percent of test data received scores equal to or greater than three, representative of “non‐native Chinese” and “much coverage,” respectively. SYSTRAN scored lowest in both measures. However, these differences were not statistically significant. A Pearson correlation analysis demonstrated a strong relationship (r=0.86) between fluency and adequacy. Missing count and incorrect count strongly correlated with fluency and adequacy.Originality/valueMost existing digital collections can be accessed in English alone. Few digital collections in the USA support multilingual information access (MLIA) that enables users of differing languages to search, browse, recognize and use information in the collections. Human translation is one solution, but it is neither time nor cost effective for most libraries. This study serves as a first step to understand the performance of current MT systems and to design effective and efficient MLIA services for digital collections.
This paper reports a study conducted to understand students collaborative e-learning behavior within a team project in a graduate-level class. The effectiveness of collaborative e-learning is measured, and the factors affecting it are identified. The study finds that cognition, information communication, IT skill, member relationship, instructor guidance and supervision have a significant impact on collaborative e-learning. This research proposes a primitive model for collaborative e-learning. It will help educators to better design collaborative e-learning courses. © 2011 IEEE.
Notice of Retraction After careful and considered review of the content of this paper by a duly constituted expert committee, this paper has been found to be in violation of IEEE's Publication Principles. We hereby retract the content of this paper. Reasonable effort should be made to remove all past references to this paper. The presenting author of this paper has the option to appeal this decision by contacting TPII@ieee.org. This paper reports a study conducted to understand students collaborative e-learning behavior within a team project in a graduate-level class. The effectiveness of collaborative e-learning is measured, and the factors affecting it are identified. The study finds that cognition, information communication, IT skill, member relationship, instructor guidance and supervision have a significant impact on collaborative e-learning. This research proposes a primitive model for collaborative e-learning. It will help educators to better design collaborative e-learning courses.
This paper describes different instructional design strategies for teaching computer technological courses online.
This paper presents the results of the team of the University of North Texas in the Wikipedia image retrieval track of Image-CLEF-2010. Our approach is based on performing translation of the French and German image captions to English and using of Language Models for generating our runs. We also explore the use of complex queries by asking two users to manually build queries based on the original topics distributed. Our results indicate that the approach of translating the image captions is feasible and yields results that are quite competitive with other teams that participated in the same track. This paper presents the results of the UNT team participation in the Wikipedia retrieval task. Traditionally, the most common approach to solve the cross language retrieval problem is to perform automatic translation of the user queries into the language of the document to be retrieved. However, in the presence of short queries the automatic translation might not have enough context to generate an appropriate translation. Our main goal was to explore the efficacy of using the captions associated with the Wikipedia images and providing automatic translations of them in English. We also address the effectiveness of using this approach using automatic queries as well as manual queries constructed by real users. Section 2 of this paper presents a short background of the CLIR retrieval problem in image retrieval. Section 3 presents the methods used to conduct our experiments. Section 4 presents our results and preliminary analysis of results. The last section of this paper presents our conclusion and plans for future work.