Word segmentation plays an important role in speech recognition as a text pre-processing step that helps decrease out-of-vocabulary items and lowers language model perplexity. Segmentation is applied mainly for agglutinative languages, but other morphologically rich languages, such as German, can also benefit from this technique. Using a relatively small, manually collected broadcast corpus of 134k tokens, the current study investigates how Finite-State Transducers (FSTs) can be applied to perform word segmentation in German. It is shown that FSTs incorporating word-formation rules can reach high segmentation performance with 0.97 precision and 0.93 recall rate. It is also shown that FSTs incorporating n-gram models of manually segmented data can reach even higher performance with accuracy and recall rates of 0.97. This result is remarkable considering the fact that the bottom-up approach performs on par with the expert system without requiring explicit knowledge about morphological categories or word formation rules.
We present a portable minimally invasive system to determine the state of the velum (raised or lowered) at a sampling rate of 40 Hz that works both during normal and silent speech. The system consists of a small capsule containing a miniature loudspeaker and a miniature microphone. The capsule is inserted into one nostril by about 10 mm. The loudspeaker emits chirps with a power band from 12-24 kHz into the nostril and the microphone records the signal reflected from the nasal cavity. The chirp response differs between raised and lowered velar positions, because the velar position determines the shape of the nasal cavity in the posterior part and hence its acoustic behaviour. Reference chirp responses for raised and lowered velar positions in combination with a spectral distance measure are used to infer the state of the velum. Here we discuss critical design aspects of the system and outline future improvements. Possible applications of the device include the detection of the velar state during silent speech recognition, medical assessment of velar mobility and speech production research.
In this paper we present a soft toy named Lingufino for preschool children that uses speech input and output for communication. It takes the child onto a journey to an adventure world: Lingunia. Based on a story that is shown in a picture book the toy explains different topics like animals, colours, numbers, seasons etc. and involves the child into the fictional situation with the help of question-and-answer games. By asking for words and facts which previously have been mentioned, the child becomes part of the adventure and ‐ on the fly ‐ improves it’s active vocabulary. Index Terms: human machine interface, children 1. Embedded Speech Dialogue System Lingufino (project is based on previous studies [1, 2, 3] and funded by [4, 5]) implements a speech dialogue system (SDS) running on an embedded microcontroller platform mounted hidden insight of the soft toy. The human machine interface (HMI) communicates via speech only, no additional modes such as keys, touch functions or displays are applied. The system consists of three parts (Fig. 1): • TTS ‐ speech production, output is performed by the limited word/phrase based text-to-speech system using prompts that are prerecorded from a professional. • ASR ‐ speech input, an automatic speech recognition (ASR) system named Picard ASR is able to recognise words, phrases and sentences (64 in parallel). Its very low memory consumption (RAM: 15 KB, FLASH: 90 KB, CPU: 40 MHz) enables Picard ASR to run on very low price microcontrollers, which is necessary for the toy market but also for other consumer markets. It is a phoneme based Hidden Markov Model (HMM) recogniser using 64 shared GMMs, configurable tristate mono- or triphones. Feature extraction runs with 13 primary MFCCs and additional and features. • DMS ‐ the dialogue management system, runs dialogues that are given by dialogue description files. These include fixed dialogues but also dynamic structures driven by random processes to make the toy dialogue more alive.
In this paper, a large-scale evaluation of open-source speech recognition toolkits is described. Specifically, HTK in association with the decoders HDecode and Julius, CMU Sphinx with the decoders pocketsphinx and Sphinx-4, and the Kaldi toolkit are compared in terms of usability and expense of recognition accuracy. The evaluation presented in this paper was done on German and English language using respective the Verbmobil 1 and the Wall Street Journal 1 corpus. We found that, Kaldi providing the most advanced training recipes gives outstanding results out of the box, the Sphinx toolkit also containing recipes enables good results in short time. HTK provides the least support and requires much more time, knowledge and effort to obtain results of the order of the other toolkits.
This paper is intended to survey the state of the art of automatic speech recognition (ASR) for children's speech.Investigating ASR for children is a current trend in research.Therefore databases of children's speech are needed for training and testing of ASR systems.In the first part of this paper the most relevant databases of children's speech are described.There are less speech data of children available than of adults and speech of preschool children is even more rarely available.In the second part of this paper the common techniques for recognizing children's speech are summarized.Most investigations about children's ASR focus on the acoustic model.The common methods are described and approaches regarding the lexical and speech model are mentioned subsequently.In an extensive literature research we collected papers investigating ASR for children.Several studies have been carried out investigating children's ASR.Due to the lack of data from preschool children only a few investigations for this age group have been accomplished.This is illustrated by presenting a statistic on the age of the children in past studies.
Der Artikel gibt einen Uberblick uber den Stand der Technik im automatischen Erkennen von Kindersprache. Begonnen wird mit einer kurzen Einfuhrung und einer Darstellung verfugbarer Kindersprachdatenbasen. Anschliesend werden verschiedene Verfahren zum Erkennen von Kindersprache betrachtet und mit Beispielen unterlegt. Der Fokus richtet sich dabei auf das Erstellen des akustischen Modells, indem die drei Falle betrachtet werden: Training der Spracherkenner mit Kindersprache, Training mit Erwachsenensprache und anschliesende Vokaltraktlangennormierung sowie Training mit Erwachsenensprache und anschliesende Adaption an Kindersprache. Danach werden einige Beispiele fur Systeme gegeben, in denen Spracherkennung von Kindersprache angewendet wird und als Abschluss folgt eine kurze Zusammenfassung.
This paper reports comparative evaluations of conventional voice activity detection (VAD) methods in reverberant environments. Both conventional and standard (G.729) methods are discussed. In general, these methods work well under clean conditions, but their performance is drastically affected by reverberation. Preliminary comparative evaluations showed that the false acceptance rate (FAR) is significantly increased due to the false rejection rate (FRR) being moderately increased by reverberation. We therefore developed a method using MTFbased power envelope restoration to improve the robustness of VAD in reverberant environments. This restoration method can blindly restore the power envelope of reverberant speech based on the MTF concept. The proposed method consists of an MTFbased restoration method as the front end and a conventional VAD method as the final decision. Experimental results demonstrated that the proposed method is superior to conventional methods with regard to robustness and providing accurate VAD (reducing both FAR and FRR) in reverberant environments.
Der Artikel beschreibt Herausforderungen, die bei der Entwicklung eines Sprachinterfaces fur Kinder entstehen. Kindliche Sprachinterfaces stellen eine wissenschaftliche Neuerung dar. Demnach ergeben sich auch fur die benotigten traditionellen Forschungsgebiete im Bereich Sprachinterfaces (Spracherkennung, Sprachsynthese und Dialog) neue Herausforderungen. Der Artikel beschreibt dies fur die Gebiete Spracherkennung und Dialog. Da Kinder andere Verhaltens- und Interaktionsmuster aufweisen als Erwachsene, sind herkommliche Dialoggestaltungsprinzipien zu uberdenken. Fur die Spracherkennung ergeben sich neue Herausforderungen sowohl durch die anatomischen Unterschiede als auch durch die sich mit dem Alter entwickelnden sprachlichen Fahigkeiten. Erkennungsexperimente zeigen das Verhalten von Spracherkennungssystemen mit Kindersprache.
In this article the authors continue previous studies regard-ing the investigation of methods that aim to improve the de-creased recognition rate (RR) in reverberant environments of automatic speech recognition (ASR) systems. Previously three robust front-end methods are tested, the harmonicity based feature analysis (HFA), the temporal power envelope feature analysis (TPEFA) and their combination (HFA+TPEFA). This paper additionally introduces two well-known methods into the comparison. These are the dereverberation method using the inverse modulation transfer function (IMTF) and the delay-and-sum beamformer (DSB). Recognition experiments are accom-plished for command word recognition, the reverberant environments are comprehensive chosen as functions of the reverberation time 𝑇 60 and the speaker to microphone distance (SMD) as the most important parameters to describe reverberant dis-tortions. The results of this first comparison of such methods prove experimentally some drawn assumptions, e. g. the IMTF method achieves robustness only in the far field, the DSB improves the RR slightly but is outperformed by the HFA due to its indirectivity at low frequencies.
Mehrmikrofonanordnungen erweitern die Moglichkeiten der Vorverarbeitung von Audiosignalen fur die Spracherkennung. Sie ermoglichen die Nutzung von Beamforming-Algorithmen, mit deren Hilfe Storgerausche und Raumeinflusse reduziert werden konnen. In diesem Beitrag werden verschiedene Beamformer vorgestellt und deren Vor- und Nachteile fur die Spracherkennung diskutiert. Es wurden Messungen in praktisch relevanten Umgebungen vorgenommen, deren Auswirkung auf die Erkennungsrate hier gezeigt werden. Des Weiteren erfolgt die Vorstellung eines aktuellen Projektes des Institut fur Akustik und Sprachkommunikation der TU Dresden mit mehreren Partnern zur Entwicklung eines public Terminals mit robuster Spracherkennung.
This contribution describes the robustness evaluation and optimization steps for a speech interface which is suitable for embedded language tutoring with special focus on children's speech. The baseline algorithms are derived from the pronunciation tutoring system AzAR directed to adult learners of German. The first prototype LiSA (2008) - directed to young children starting at 3 years - is currently evaluated and optimized, mainly addressing following issues: (a) the challenge of ASR-based pronunciation assessment for children's speech, (b) the handling of noise and reverberation in an embedded application scenario, and (c) the extraction of additional information such as age or gender. The article summarizes evaluation results of the speech recognizer in laboratory and real-world room environment.
Automatic speech recognition (ASR) systems used in real-world indoor scenarios suffer from performance degradation if noise and reverberation conditions differ from the training conditions of the recognizer. This thesis deals with the problem of room reverberation as a cause of distortion in ASR systems. The background of this research is the design of practical command and control applications, such as a voice controlled light switch in rooms or similar applications. Therefore, the design aims to incorporate several restricting working conditions for the recognizer and still achieve a high level of robustness. One of those design restrictions is the minimisation of computational complexity to allow the practical implementation on an embedded processor. One chapter comprehensively describes the room acoustic environment, including the behavior of the sound field in rooms. It addresses the speaker room microphone (SRM) system which is expressed in the time domain as the room impulse response (RIR). The convolution of the RIR with the clean speech signal yields the reverberant signal at the microphone. A thorough analysis proposes that the degree of the distortion caused by reverberation is dependent on two parameters, the reverberation time T60 and the speaker-to-microphone distance (SMD). To evaluate the dependency of the recognition rate on the degree of distortion, a number of experiments has been successfully conducted, confirming the above mentioned dependency of the two parameters, T60 and SMD. Further experiments have shown that ASR is barely affected by high-frequency reverberation, whereas low frequency reverberation has a detrimental effect on the recognition rate. A literature survey concludes that, although several approaches exist which claim significant improvements, none of them fulfils the above mentioned practical implementation criteria. Within this thesis, a new approach entitled ’harmonicity-based feature analysis’ (HFA) is proposed. It is based on three ideas that are derived in former chapters. Experimental results prove that HFA is able to enhance the recognition rate in reverberant environments. Even practical applicable results are achieved when HFA is combined with reverberant training. The method is further evaluated against three other approaches from the literature. Also combinations of methods are tested. In a last chapter the two base technologies fundamental frequency (F0) estimation and voiced unvoiced decision (VUD) are evaluated in reverberant environments, since they are necessary to run HFA. This evaluation aims to find one optimal method for each of these technologies. The results show that all F0 estimation methods and also the VUDmethods have a strong decreasing performance in reverberant environments. Nevertheless it is shown that HFA is able to deal with uncertainties of these base technologies as such that the recognition performance still improves.
This paper proposes two methods for robust automatic speech recognition (ASR) in reverberant environments. Unlike other methods which mostly apply inverse filtering by blindly estimated room impulse responses to achieve dereverberation, the proposed methods are based on the utilization of the characteristics of speech. The first method - Harmonicity based Feature Analysis takes advantage of the harmonic components of speech, which are assumed to be undistorted. The second method - Temporal Power Envelope Feature Analysis utilizes the temporal modulation structure of speech, representing the phoneme level temporal events which contain most intelligibility information. Both methods increase the recognition performance remarkably in a different way. Combining both of them connects their individual advantages. In order to examine the performance of utilizing harmonicity and modulation temporal structure for reverberant ASR, the methods are tested in clean and reverberant training. As results show, even in strong reverberant conditions both methods obtain practical applicable performance for reverberant training. In addition, besides testing their performance in dependency on the reverberation time, their performance considering the speaker-to-microphone distance is tested, which is another new contributions in this paper.
The paper is concerning some problems which are posed by the growing interest in social interaction research as far as they can be solved by engineers in acoustics and speech technology. Firstly the importance of nonverbal and paraverbal modalities in two prototypical scenarios are discussed: face-to-face interactions in psychotherapeutic consulting and side-by-side interactions of children cooperating in a computer game. Some challenges in processing signals are stated with respect to both scenarios. The following technologies of acoustic signal processing are discussed: (a) analysis of the influence of the room impulse response to the recognition rate, (b) adaptive two-channel microphone, (c) localization and separation of sound sources in rooms, and (d) single-channel noise suppression.
This paper takes note of the major influences of room acoustic effects on the fundamental frequency F0 of speech and its determination. A detailed description of room acoustic measures and effects is given. As a conclusion of those the dependency of the speaker to microphone distance (SMD) is to be studied combined with the reverberation time T60. Evaluation experiments aiming to find dependencies of the room acoustic effects are accomplished. In contrast to most of the studies which deal with reverberation, this paper proves that T60 cannot be the only measure to describe the behavior of systems in reverberant environments. Experiments studying the dependency on the SMD are new contributions within this paper. The experiments are an extension of the previous studies on the dependency on artificial reverberation, where twelve F0 estimation methods are compared. Here, two further methods are added and the systematic use of real measured room impulse responses (RIR) is applied. Apart from the SMD dependency other main contributions are the experiments of different disturbing effects on high and low F0. The results show male F0 estimation is significantly more sensitive to reverberation than female.
This article proposes a new signal analysis method for automatic speech recognition designed to aim high robustness against distortions caused by room reverberation. The method is initially named Harmonicity based Feature Analysis (HFA) and implements the following three ideas: (i) reconstruction of a spectrum from the harmonic components (assumed to be undistorted) of a voiced speech spectrum. (ii) suppression of disturbing reverberation in unvoiced spectra coming from previous voiced sections. (iii) high frequency regions are not affected by HFA since they have negligible effect on the recognition rate. HFA works on the basis of fundamental frequency estimation and voiced/unvoiced decision. Evaluation results show significant improvement of the recognition performance over a wide range of reverberant conditions while using HFA in connection with reverberant training. Apart from good performance, advantages of HFA compared to state of the art dereverberation approaches are real time processing (no adaptation time) and robustness against changes of the room impulse response.
Automatic speech recognition (ASR) systems used in real indoor scenarios suffer from different noise and reverberation conditions compared to the training conditions. This article describes a study which aims to find out what are the most harming parts of reverberation to speech recognition. Noise influences are left out. Therefore different real room impulse responses in different rooms and different speaker to microphone distances are measured and modified. The results of the recognition experiments with the related convoluted impulse responses clearly show the dependency of early and late as well as high and low frequency reflections. Conclusions concerning the design of a dereverberation method are made.
The article deals with a novel speech recognizer technology which has the potential to overcome some problems of in-car speech control. The verbKEY recognizer bases on the Associative-Dynamic (ASD) algorithm which differs from established techniques as HMM or DTW. The speech recognition technology is designed to run on a 16 bit, fixed point DSP platform. It enables high recognition performance and robustness. At the same time, it is highly cost efficient due to its low memory consumption and its less calculation complexity. Typical applications such as dialling, word spotting or menu structures for the device control are processed by the continuous, real-time recognition engine with an accuracy higher 98% for a 20 words vocabulary. The article describes a hardware prototype for command & control applications and the measures taken to improve the robustness against environmental noises. Finally, the authors discuss some ergonomic aspects to obtain a higher level of traffic safety.