
In diesem Beitrag stellen wir eine Klasse von neuen Algorithmen zur blinden Enthallung von Sprachsignalen und gleichzeitiger Trennung von Signalgemischen (blind source separation, BSS) vor, basierend auf TRINICON, einem allgemeinen Konzept fur breitbandige adaptive MIMO-Signalverarbeitung. Um alle grundlegenden stochastischen Signaleigenschaften von Sprache fur den Enthallungs- und Separationsprozess ausnutzen zu konnen und um die bekannten Whitening-Artefakte in bisherigen Verfahren zu vermeiden, schlagen wir die Einfuhrung eines speziell entworfenen Signalmodells vor, welches auf einer Entwicklung mittels multivariater Tschebyscheff-Hermite-Polynome basiert. Das multivariate Modell beinhaltet auch inharent eine lineare Pradiktion, welche bekanntermasen direkt mit der Modellierung des menschlichen Vokaltrakts in Verbindung gebracht werden kann. Das vorgestellte Konzept ist anwendbar sowohl fur Einzelsprecher-Szenarien, als auch fur mehrere, simultan aktive Sprecher. Im letzteren Fall beinhaltet es zusatzlich zur Enthallung auch eine blinde Quellentrennung (BSS).
Communicative actions are specific movements or gestures which are accomplished by vocal tract articulators (lips, tongue, velum etc.) for speech, by facial articulators (eye brows, eye lids etc.) for co-verbal facial expression and by other bodily articulators (hands, arms etc.) for co-verbal gesturing. While action-based approaches already exist for spoken language processing it is the aim of this paper to adopt action theory for signed language processing (i.e. production and perception of sign language). Method: An action-based method for the phonetic annotation of sign language has been developed and a 100 sentence American Sign Language corpus has been analyzed using this method. Results: Five basic types of sign actions were identified, all indicating the importance of movement phases even if the goal of a gesture is to reach a specific target (e.g. specific hand shape, orientation, location, and/or direction). Conclusion: This study is a starting point for investigating sign language production quantitatively in terms of a unified action theory.
This paper presents a novel method for rescoring the n-best recognition hypotheses using intonation knowledge. The model synthesizes the f0 contours for each of the n-best hypotheses and estimates an intonative matching index between the synthetic shapes and the real f0 contour. This index is applied in the rescoring process, and can be viewed as a degree of intonation compatibility between the hypotheses and the input sentence. The f0 prediction is based on classification and regression trees and the Fujisaki model. We evaluate our approach using a single speaker of the Buenos Aires Spanish LIS-SECYT database under clean and babblenoisy conditions. Considering the systems under no grammar condition, the proposed model reduces the mean absolute word error rate in 3.1% with respect to the baseline system, in a consistent manner and under different noise conditions.
In the present work we observe two subjects interacting in a collaborative task on a shared environment. One goal of the experiment is to measure the change in behavior with respect to gaze when one interactant is wearing dark glasses and hence his/her gaze is not visible by the other one. The results show that if one subject wears dark glasses while telling the other subject the position of a certain object, the other subject needs significantly more time to locate and move this object. Hence, eye gaze - when visible - of one subject looking at a certain object speeds up the location of the cube by the other subject. The second goal of the currently ongoing work is to collect data on the multimodal behavior of one of the subjects by means of audio recording, eye gaze and head motion tracking in order to build a model that can be used to control a robot in a comparable scenario in future experiments.
In verschiedenen Anwendungen ist die Bestimmung der Laufzeit akustischer Signale von einem Sende- zu einem Empfangsort von Bedeutung. Dazu werden ublicherweise Korrelationsverfahren verwendet. Die zur Detektion des Maximums in der Korrelationsfunktion verwendeten Schwellwertverfahren liefern jedoch bei Signalen mit geringem Signal-Storabstand Ergebnisse, deren Zuverlassigkeit bzw. Genauigkeit fur viele Anwendungen nicht ausreichend ist. Bei dem hier vorgestellten Ansatz soll durch die Verwendung von Verfahren der Mustererkennung, wie sie z.B. in der automatischen Spracherkennung Verwendung finden, eine zuverlassigere bzw. genauere Laufzeitmessung ermoglicht werden. Das Verfahren wurde in verschiedenen Einsatzszenarien getestet und erreichte eine mittlere Detektionsrate der Laufzeit (Abweichung innerhalb vorgegebener Grenzen vom Sollwert) von ca. 70 %, wobei insbesondere bei schwachen Signalen am Empfangsort mit entsprechend geringem Signal-Rauschabstand eine Verbesserung gegenuber einer ausschlieslich Schwellwert-basierten Bestimmung der Laufzeit erreicht werden konnte.
In the study a simple conversion of the voice character using a modification of the glottal pulse is shortly described. The glottal signal is estimated by homomorphic speech deconvolution of the speech signal into the maximum- and minimum-phase parts. The maximum-phasepart is an approximation of the glottal signal. For speech reconstruction the parametric mixed phase speech generation model based on the complex cepstrum is used, which takes into account not only the magnitude spectrum of the modeled speech frame, but also the phase spectrum.
Researchin the field of neural processing proposesa bidirectional computation scheme amongthe hierarchical organized levels of the brain. This scheme is called cortical algorithm. In speech processing a similar approach is applied to the distinct linguistic levels, where a powerful technologyis given by the weighted Finite State Transducers. FSTs translate a symbol sequence of one level into a symbol sequence of another level. In this approach either a top-down process or a bottom-up process takes place, whereas in the cortical algorithm both processes simultaneously occur. To utilize this property in the area of signal- and speech processing, we formulate the cortical algorithm for HMMs and FSTs.
The statistical properties of segments [8] using a specific acoustic model called the hidden chunk model (HCM) is investigated. We call the sequence of feature vectors assigned to a segment a chunk of length l. The HCM still assumes that the feature vectors are statistically independent. In contrast to hidden Markov model (HMM) we introduce emission probabilities which depend on l. Segment error rates (SERs) are calculated on a database with over 33 million chunks aligned to 607 segments. The HCM achieves more than 10%absolute improvement in SER compared to the HMM. Based on the estimated Shannon’s entropy, the proposed HCM model paves the way to create acoustic models which are heading towards the lowest possible SER.
A system for mobile office and entertainment applications was developed. The input modalities include touch screen and automatic speech recognition, the output modalities screen output and speech synthesis. Services comprise an e-mail reading and responding application, a configurable news feed (RSS) reader and, as entertainment, a license plate information application. Being conceptually eyes- and hands-free, the main target application is the car domain, but as a native android application it runs also on every Android based mobile phone. We report on the software structure and a framework for the integration of speech resources into Android based software, describe the functionality of the services e-mail, RSS and license number information, report on the design concept of multimodal usage, explain integration, evaluation and tuning of speech recognition as well as synthesis and conclude with a description of the evaluation user tests.
Angestrebt ist eine Vorverarbeitung von Audiosignalen aus Mediadaten durch Klassifikation und Segmentation. Die in den Signalen vorkommende Sprache soll von den anderen Elementen wie Musik und Hintergrundgerauschen getrennt werden, um eine ressourcen-optimierte Spracherkennung mit einer hohen Erkennungsrate zu gewahrleisten. Mediadaten bestehen haufig aus Elementen, die nicht eindeutig einer Klasse zuzuordnen sind, sondern Mischformen aus den eben genannten Audioklassen darstellen. Die gangig fur die Unterscheidung von Sprache und Musik eingesetzten akustischen Merkmale wurden daher in einer Mixed-Type- Multi-Class-Klassifikation angewendet, um darin ihre Aussagekraft zu uberprufen. Bei den verwendeten Merkmalen handelt es sich um statistische Kennwerte der zeitlichen und spektralen Struktur der Signale: Zero-Crossing-Rate, RMS-Energie, Spectral Centroid, Spectral Rolloff-Point, Spectral Flux und MFCCs. Wahrend Sprache auch in den Mischformen mit Musik oder Hintergrundgerauschen gut als diese klassifiziert werden konnte,stellt die Diskrimination der Hintergrundgerausche ein Problem und eine Herausforderung an zukunftige Losungsansatze dar.
In diesem Beitrag wird uber erste Erfahrungen mit der Realisierung von Sprachdialogen mit dem Web-Entwicklungs-Framework Grails berichtet. Als Beispiel werden Anfragen zu einem Bestand von Buchern betrachtet. In eine entsprechende Grails-Anwendung werden gemas des X+V Konzepts zusatzliche VoiceXML-Teile in die generierten Webseiten eingefugt. Mit einem entsprechenden Browser kann man dann die Webanwendung auch per Spracheingabe steuern. Dank der guten Unterstutzung durch entsprechende Konzepte in Groovy ist die Erweiterung um VoiceXML-Teile sehr einfach. Insgesamt gesehen erweist sich der Weg von der Spezifikation des Datenmodells bis zu ersten Sprachdialogen als kurz und gradlinig.
Digitale Filterbanke sind wesentliche Bestandteile von aktuellen Algorithmen der Sprach- und Audiosignalverarbeitung. Sie sind ein wichtiger Bestandteil von Systemen zur Echounterdruckung und Bandbreitenerweiterung. Ein weiteres Anwendungsgebiet ist die Gerauschreduktion zur Sprachsignalverbesserung. Die Wahl der verwendeten Filterbank hat einen signifikanten Einfluss auf die Leistungsfahigkeit eines solchen Systems in Bezug auf Signalqualitat, Rechenkomplexitat und Signalverzogerung. Bei der Signalverarbeitung eines Innenraumkommunikationssystems mussen alle diese Kriterien berucksichtigt werden. Eine overlap-save-basierte FFT-Filterbank kann fur eine Gerauschreduktion mit geringer Verzogerung verwendet werden. Dabei mussen die Filterkoeffizienten der Gerauschreduktion durch ein Projektionsfilter modifiziert werden, um zyklische Faltungseffekte zu unterdrucken. Der Nachteil einer solchen Filterbankstruktur ist die erhohte Rechenkomplexitat. Durch die Wahl einer geeigneten Projektionsfilternaherung kann diese bei nahezu gleichbleibender Signalqualitat erheblich reduziert werden.
In this talk I will start by defining what I mean by a ‘model’ and by emphasizing the importance of creating a proper model for the phenomenon under study. What is meant by a proper model is not just an ad hoc mathematical approximation of what one superficially observes, but one that is based on understanding of the underlying principle. I will then show four examples of such models I developed during my research career of over the past 50 years. 1. A model for the humanuse of language in sending and receiving information By looking into the processes involved in the use of language in sending and receiving information, this work shows that the two aspects are quite different and their characteristics require separate mathematical formulations. A model for the human processesof controlling the fundamental frequency of speech By examining the underlying mechanisms, this work explains why formulating the fundamental frequency contour in terms of logarithmic frequency leads to a precise model and separation of two types of components related to phrasing and accentuation, respectively, and why the two components have shapes of response curves of second-order linear systems. A model for the humanprocesses involved in identification and discrimination By critically analyzing the cognitive processes involved in identification and discrimination experiments, this work shows the mathematical relationship between the two processes and hence between the two performance curves, and proves that the so-called phenomenon of categorical perception is an artifact coming from the experimental paradigm. A modelfor the process of dialogue By looking into the process of dialogue, this work clearly shows that the finite-state automaton often used in modeling a dialogue is inadequate, and presents a novelalternative way of modeling a dialogue in terms of two interacting finite-state automata that are different from conventional finite-state automata. In summary, it is the author’s belief that one needs to go beyond conventional views, but proper models are often based on the commonsense.
Umfang und Qualitat gesammelter Sprachdaten haben einen masgeblichen Einfluss auf die Entwicklung von Sprachverarbeitungssystemen. Aufgrund der notwendigen Datenvolumen wird eine automatische Segmentierung und Annotation der Sprachdaten angestrebt. In der Praxis ist in der Regel eine manuelle Kontrolle und Korrektur der Daten erforderlich. Das Prosodisch-Phonetische Annotationssystem PROPHANO stellt weitgehend automatisierte Werkzeuge zur Verfugung, deren Zusammenschaltung zu einer zuverlassigen Segmentierung und Annotation aufgezeichneter Sprachdaten fuhrt und gleichzeitig das fur die manuelle Korrektur notwendige Expertenwissen in den Algorithmen abbildet. Somit kann auch der weniger erfahrene Bearbeiter hochwertige Sprachdatenbasen erstellen. PROPHANO basiert auf Vorarbeiten des Instituts im Bereich hochqualitativer Sprecherdaten fur die Sprachsynthese (TC-STAR [1]) sowie bei der webgestutzten Datensammlung von SMS-Texten und dementsprechenden Studioaufnahmen [2]. Vorhandene Werkzeuge zur Aufnahme und Phonemsegmentierung, z. B. WiGE [3], werden verwendet und mit einer prosodischen Analysemoglichkeit erweitert. PROPHANO integriert die Sprachdatenannotation in eine ubersichtliche XMLStruktur, welche eine synchrone oder asynchrone Weiterverarbeitung der Sprachdaten erleichtert und fur neue Sprachmerkmale erweiterbar ist.
This paper reports on the continued activities towards the development of a computer-aided language learning (CALL) system for German learners of Mandarin. In this experiment the method for detecting the pronunciation errors which was presented in a previous experiment was tested on two different databases in order to study the effect of complexity of corpus on the results of pronunciation error detection. The first corpus is simple and consists of monosyllabic and disyllabic words and read from German students of Mandarin in the first year of language education. The second corpus is more complex and consists of whole sentences and read from German students from three different years of language education. The data are perceptually evaluated by human judges as well as processed by two Automatic Speech Recognition (ASR) systems. Acoustic model of the first ASR system trained on data of native speakers of Mandarin. The second ASR system used an adapted acoustic model that considers the errors expected from the German learners of Mandarin. The experimental results show that the performance of the modified ASR system is better. The ratings of strength of foreign accent and intelligibility are strongly correlated with the correctness of tones than with the correctness of initials and finals. The ratio of correct initials and finals in the complex corpus is greater than in the simple corpus, but the number of correct tones is lower in the complex corpus.
A language learning system for guiding a student on how to pronounce words of a second language must provide meaningful feedback while locating the learner’s errors with a high accuracy. We propose a method to detect pitch accent errors in the speech of learners of Japanese. In previous methods, Tokyo accent type recognition of the learner utterance was focused on. However, the learner may produce pitch patterns outside of these accent types so it is necessary to be able to identify all possible pitch level combinations. This paper presents a technique to identify all patterns. Employing either a 2 mora model or a 3 mora model, this method identifies pitch level for contiguous two or three mora unit sets. Then it determines the most likely combination of these units to determine the pitch level pattern of the word. Through this method we achieved a 93.4% correct pitch level identification rate at the mora level.
Synthetic speech needs prosody to get the right structure and to sound natural. Therefore, the emerging speech technology pushed the development of prosody models. Today, prosody research is well established with an own conference series, and powerful tools are available for investigating prosodic effects. The 80th birthday ofthe pioneer of quantitative prosody modeling, Professor Hiroya Fujisaki, is an excellent occasion to look at the situation in earlier times of speech technology. The authors give an outline using mainly the material which is available from the history in Dresden and Berlin. The oral presentation will be accompanied by numerous historic audio examples.
Die Integration von Spracherkennung und -synthese auf Systeme mit begrenzten Hardwareressourcen wird immer haufiger benotigt. Wir haben ein Dialogsystem entwickelt, welches auf einer Kombination aus einem digitalen Signalprozessor (DSP) und einem Field Programmable Gate Array (FPGA) lauffahig ist. Es wurde versucht die Verluste in der Erkennungsleistung sowie der Synthesequalitat moglichst gering zu halten. Um Speicherplatz zu sparen, verwenden Erkenner und Synthese die selben, sprecherunabhangigen Hidden-Markov-Modelle. Der Spracherkenner ist phonembasiert und kann beliebige kontextfreie Grammatiken verarbeiten sowie zwischen verschiedenen Grammatiken wechseln. Die Synthese basiert auf der Verkettung von HMM-kodierten Einheiten. Die Einheiten bestehen aus einer Folge von Indizes der Verteilungsdichten der HM-Modelle des Spracherkenners, einer FO- sowie einer Intensitats-Kontur. Zur Generierung einer Zielstimme werden die sprecherunabhangigen Syntheseparameter auf Merkmalebene konvertiert. Im Rahmen dieser Veroffentlichung werden wir Ergebnisse zur Leistungsfahigkeit, Performancemessung sowie zum Speicherbedarf des Dialogsystems vorstellen. Auserdem soll das entstandene System vorgefuhrt werden.
Whereas methods for synthesizing speech signals from written text have made considerable advances in the past decade, methods for assessing the performance of speech synthesizers and evaluating their fitness for particular applicationsare still cumbersome. A reasonis that quality assessment and evaluation require perception and judgmentprocesses to take place, which ultimately happen only inside a human assessor. Thus, auditory test methods are currently the only wayto validly and reliably assess and evaluate synthesized speech quality. Still, advances in speech transmission quality prediction show up ways to model perception and judgmentprocessesto a limited extent, and thusto predict speech transmission quality on the basis of instrumental measurements only. Weare thus interested in the question whether such an approach is also feasible with synthesized speech signals and their corresponding degradationsoriginating from the synthesis process. In this talk, we will discuss several such ways and analyze their (expected) performance on different synthesized speech databases. First, we will briefly review approaches which rely on the availability of a natural reference speech signal, and compare synthesized speech signals to such natural references [1]. This approach is limited by the (non-) availability of natural speech data usually required from the same speaker the synthesis inventory has been built from. Second, we will address approaches which rely on a model of natural speech, and derive quality predictions on thebasis of the similarity of the synthesized speech signalto this model[2, 3, 4]. Third, we will review approaches extracting parameters from the synthesized speech signals which coincide with particular types of degradations [5]. Such approaches have been successful with transmitted speech signals, but the parameters heavily depend of the speech databases and speaker gender. Finally, we will show that improvements can be reached by combining different types of approaches [6]. We will justify our claims on the basis of empirical data from typical German synthesizers as well as from the Blizzard Challenges organized as a controlled comparative assessment of synthesized speech. Wewill address the deficiencies of the current approaches by showingthat their performance heavily depends on the used databases, and will identify research which has to be carried outjointly by the synthesis and evaluation communities to overcomethe currentlimitations.