The quality of e-Commerce services largely depends on the accessibility of product content as well as its completeness and correctness. Nowadays, many sellers target cross-country and cross-lingual markets via active or passive cross-border trade, fostering the desire for seamless user experiences. While machine translation (MT) is very helpful for crossing language barriers, automatically matching existing items for sale (e.g. the smartphone in front of me) to the same product (all smartphones of the same brand/type/colour/condition) can be challenging, especially because the seller’s description can often be erroneous or incomplete. This task we refer to as item alignment in multilingual e-commerce catalogues. To facilitate this task, we develop a pipeline of tools for item classification based on cross-lingual text similarity, exploiting recurrent neural networks (RNNs) with and without pre-trained word-embeddings. Furthermore, we combine our language agnostic RNN classifiers with an in-domain MT system to further reduce the linguistic and stylistic differences between the investigated data, aiming to boost our performance. The quality of the methods as well as their training speed is compared on an in-domain data set for English–German products.
Organisations seeking competitive advantage in the age of big data often adopt the strategy of becoming data-driven. The paper describes research in progress with an organisation pursuing this stra ...
Iacer Calixto, Daniel Stein, Evgeny Matusov, Pintu Lohar, Sheila Castilho, Andy Way. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers. 2017.
In this paper, we study how humans perceive the use of images as an additional knowledge source to machine-translate user-generated product listings in an e-commerce company. We conduct a human evaluation where we assess how a multi-modal neural machine translation (NMT) model compares to two text-only approaches: a conventional state-of-the-art attention-based NMT and a phrase-based statistical machine translation (PBSMT) model. We evaluate translations obtained with different systems and also discuss the data set of user-generated product listings, which in our case comprises both product listings and associated images. We found that humans preferred translations obtained with a PBSMT system to both text-only and multi-modal NMT over 56% of the time. Nonetheless, human evaluators ranked translations from a multi-modal NMT model as better than those of a text-only NMT over 88% of the time, which suggests that images do help NMT in this use-case.
This article describes the process of the conception and implementation of a digital repository which is based on the software framework Fedora, Islandora, and Drupal. The main data type that is to be organized, maintained, and distributed with the repository is spoken language corpora, with a focus on multilingualism, which means that the content of the repository may be of special interest to researchers of linguistic, translational and socio-psychological studies. We concentrate on the technical requirements that had to be fulfilled with regard to our status as an official center of the Common Language Resource Infrastructure Network (CLARIN), a pan-European initiative in the context of the Digital Humanities. In particular, the requirements include preparation of information for the description of resources providing a source for metadata harvesting, registration of persistent identifiers for the long-term access and citeability of resources, implementation of a federated content search module, dealing with privacy restrictions, and the possibility to provide a single sign-on method via Shibboleth. The HZSK Repository ultimately passed two assessment processes: the assessment of the Data Seal of Approval, and a CLARIN-internal assessment procedure.
The paper presents the LinkedTV approaches for the Search and Hyperlinking (S&H) task at MediaEval 2014. Our submissions aim at evaluating 2 key dimensions: temporal granularity and visual properties of the video segments. The temporal granularity of target video segments is dened by grouping text sentences, or consecutive automatically detected shots, considering the temporal coherence, the visual similarity and the lexical cohesion among them. Visual properties are combined with text search results using multimodal fusion for re-ranking. Two alternative methods are proposed to identify which visual concepts are relevant to each query: using WordNet similarity or Google Image analysis. For Hyperlinking, relevant visual concepts are identied by analysing the video anchor.
In this paper we describe the large-scale German broadcast corpus (GER-TV1000h) containing more than 1,000 hours of transcribed speech data. This corpus is unique in the German language corpora domain and enables significant progress in tuning the acoustic modelling of German large vocabulary continuous speech recognition (LVCSR) systems. The exploitation of this huge broadcast corpus is demonstrated by optimizing and improving the Fraunhofer IAIS speech recognition system. Due to the availability of huge amount of acoustic training data new training strategies are investigated. The performance of the automatic speech recognition (ASR) system is evaluated on several datasets and compared to previously published results. It can be shown that the word error rate (WER) using a larger corpus can be reduced by up to 9.1% relative. By using both larger corpus and recent training paradigms the WER was reduced by up to 35.8% relative and below 40% absolute even for spontaneous dialectal speech in noisy conditions, making the ASR output a useful resource for subsequent tasks like named entity recognition also in difficult acoustic situations.
Finding the optimal decoding parameters in speech recognition is often done manually in a rather tedious manner, although automatic gradient-free optimization techniques have been shown to perform quite well for this task. While there have been recent scientific contributions in this field, no thorough comparison of possible methods, in terms of convergence speed and performance, has been undertaken. In this paper, we conduct a series of experiments with three decoding paradigms and four different optimization techniques found in recent literature, both on unconstrained and time-constrained decoder optimization. We offer our findings on the German Difficult Speech Corpus and on the LinkedTV test sets.
Spoken languages are often rich in regional accents and dialects. These local variations often pose challenges to automatic speech recognition. In this study, we analyse the influence of German regional accents on the performance of a large vocabulary continuous speech recogniser trained on standard German data. The experiments show a large variation in the error rate over different regions. We investigate the influence of phoneme-level variations on recognition errors and use our findings to build a dialect recognition system.
This is the user’s manual for Jane, RWTH’s statistical machine translation toolkit. Jane supports state-of-the-art techniques for phrase-based [Wuebker & Huck+ 12] and hierarchical phrase-based machine translation [Vilar & Stein+ 10, Stein & Vilar+ 11, Vilar & Stein+ 12] as well as for system combination [Freitag & Huck+ 14]. Many advanced features are implemented in the toolkit, as for instance forced alignment phrase training for the phrase-based model and several syntactic extensions for the hierarchical model.
This paper introduces a framework for establishing links between related media fragments within a collection of videos. A set of analysis techniques is applied for extracting information from different types of data. Visual-based shot and scene segmentation is performed for defining media fragments at different granularity levels, while visual cues are detected from keyframes of the video via concept detection and optical character recognition (OCR). Keyword extraction is applied on textual data such as the output of OCR, subtitles and metadata. This set of results is used for the automatic identification and linking of related media fragments. The proposed framework exhibited competitive performance in the Video Hyperlinking sub-task of MediaEval 2013, indicating that video scene segmentation can provide more meaningful segments, compared to other decomposition methods, for hyperlinking purposes.
Enriching linear videos by offering continuative and related information via, e.g., audio streams, web pages, as well as other videos, is typically hampered by its demand for massive editorial work. While a large number of analysis techniques that extract knowledge automatically from video content exists, their produced raw data are typically not of interest to the end user. In this paper, we review our analysis efforts as defined within the LinkedTV project and present the recent advances in core technologies for automatic speech recognition and object-redetection. Furthermore, we introduce our approach for an automatically generated localized person identification database. Finally, the processing of the raw data into a linked resource available in a web compliant format is described.
In scenarios with multiple input single output systems, the stochastic constrained least mean-squares (LMS) algorithm has been proven to be an effective approach. However, when only two input channels are available, it is unclear whether this approach still yields improvements. In this paper, we investigate the stableness and the robustness of the constrained LMS algorithm on “Track 1” of “2 CHiME Challenge” [1] and show that it leads to small yet consistent improvements on all signal-to-noise settings.
Event recognition systems have high potential to support crisis management and emergency response. For large-scale scenarios, however, the sheer amount of possible audio and video channels requires adequate processing of the material by automatic means. In this article, the authors focus on automatic audio and video event recognition, by means of detecting abnormalities both in train noise as well as surveillance videos, and by conducting automatic speech recognition on fire fighter communication. All components are integrated in an overall intelligent resource management system. The authors elaborate on the challenges expected from real life data and the solutions that the authors applied. The overall system, based on Event-Driven Service-Oriented Architecture, has been implemented and partly integrated into the end users' infrastructures. The system has been continuously running for more than two years, collecting data for research purposes.
In this paper we describe the construction of a parallel corpus between the standard and a non-standard language variety, specifically standard Austrian German and Viennese dialect. The resulting parallel corpus is used for statistical machine translation (SMT) from the standard to the non-standard variety. The main challenges to our task are data scarcity and the lack of an authoritative orthography. We started with the generation of a base corpus of manually transcribed and translated data from spoken text encoded in a specifically developed orthography. This data is used to train a first phrasebased SMT. To deal with out-of-vocabulary items we exploit the strong proximity between source and target variety with a backoff strategy that uses character-level models. To arrive at the necessary size for a corpus to be used for SMT, we employ a boot-strapping approach. Integrating additional available sources (comparable corpora, such as Wikipedia) necessitates to identify parallel sentences out of substantially differing parallel documents. As an additional task, the spelling of the texts has to be transformed into the above mentioned orthography of the target variety.
While both the acoustic model and the language model in automatic speech recognition arc typically well-trained on the target domain, the free parameters of the decoder itself are often set manually. In this paper, we investigate in how far a stochastic approximation algorithm can be employed to automatically determine the best parameters, especially if additional time constraints are given on unknown machine architectures. We offer our findings on the German Difficult Speech Corpus, and present significant improvements over both the spontaneous and planned clean speech task.
This paper aims at presenting the results of LinkedTV’s rst participation to the Search and Hyperlinking task at MediaEval challenge 2013. We used textual information, transcripts, subtitles and metadata, and we tested their combination with automatically detected visual concepts. Hence, we submitted various runs to compare diverse approaches and see the improvement when adding visual information.
The attempt to translate meaning from one language to another by formal means traces back to the philosophical schools of secret and universal languages as they were originated by Ramon Llull or Johann Joachim Becher. Until today, machine translation (MT) is known as the crowning discipline of natural language processing. Due to current MT approaches, the time needed to develop new systems with similar power to the older ones, has decreased enormously. However, when comparing current achievements to those of thirty years ago, only a minor dierence in the number and type of errors can be observed. In this article, the history of MT, the dierence to computer aided translation and the current approaches are discussed.
Onno Crasborn合作论文数Centre for Language Studies, Radboud University Nijmegen, The Netherlands4