The historical acoustic-phonetic collection (HAPS) of the TU Dresden documents the development of speech technology from the emergence of experimental phonetics to the introduction of computers in laboratories for spoken language processing. While the focus so far has been on recording the inventory of ca. 1.000 exhibits of the HAPS, the accompanying archive holdings are now being evaluated. This paper describes the development of technical speech communication at the TU Dresden, where an own chair for this subject was set up at a very early stage. The chair, headed by W. Tscheschner, was one of the most important in this subject in the Eastern Bloc. The paper offers insights into the successful development of a subject under the special conditions of the Cold War in the Eastern German state until its end in 1990.
At the TU (formerly TH) Dresden, acoustics is part of the faculty of electrical engineering. Its development started in 1911 when Heinrich Barkhausen was appointed Professor for “low-current technology", which was an umbrella for both, acoustics and communications engineering. Barkhausen contributed to the field of acoustics, e.g., with the first device for loudness measurement. After the war and the retirement of Barkhausen, several new institutes were established from which we mention: (1) the Institute of Electro- and Building Acoustics led by Walter Reichardt, contributing to many fields of technical acoustics, and (2) the Institute of Telecommunications Engineering supervised by Kurt Freitag, contributing to speech acoustics with the design of a vocoder and the measurement of speech quality. When the GDR performed a “higher education reform" in 1969, the acoustical activities were concentrated in a laboratory for “communications and data acquisition" which included five chairs in acoustics, sensors, speech, and measurement. This step took into account the growing role of computer technology. After the political changes in 1990, the number of chairs was reduced to two which is expressed by the today’s name “Institute of Acoustics and Speech Communications". The paper is finished by an overview on the recent activities of the institute.
Machines that can talk and listen have been part of historical tradition since antiquity. Dresden also boasts one of the legendary “speaking heads”, which is said to have been made around 1700 by the versatile vice-rector of the “Kreuzschule”, Johann Valentin Merbitz. However, it was not until the Age of Enlightenment, under the influence of Leonhard Euler, that serious scientific attempts were made to build “speaking machines”. One of these, constructed by Wolfgang von Kempelen and described in 1791, has rightly become famous. Its creator presented it on one of his lecture tours in Dresden in 1784, and today a replica from 2009 enhances the Acoustic-Phonetic Collection of the TUD Dresden University of Technology. It was a long road, then, to the establishment of the field that is referred to today as natural language human-machine interaction or more simply as speech. Its development would not have been possible at all without the accompanying preliminary work in electronics and computer technology. At universities, the field was institutionalized in different ways, with its interdisciplinary nature rooted in very different disciplines, such as phonetics, physiology, perceptual psychology, acoustics, communications engineering, and computer science. In Dresden, speech technology emerged as part of weak current engineering. At present, this field is called information technology and has its own collection: the “elektron” Collection.
The historic acoustic-phonetic collection (HAPS) of the Technical University Dresden includes more than 1.000 objects from the history of experimental phonetics and speech technology. The rather young collection is a composition of historic devices from mainly three places: the TU Dresden and the former Phonetic Institutes of the universities in Berlin and Hamburg. It was a challenge to explore these inhomogeneous convolutes scientifically. After two decades, this process is now finished for the material exhibits. The paper summarizes the results which are hitherto achieved in cataloguing and describes the remaining tasks with respect to the audiovisual and the archival materials.
: Speech Communication is an interdisciplinary topic, but in the academic practice it has to be assigned to disciplines like linguistics and phonetics, electrical engineering, etc. In the case of engineering, the interest of the faculties originated in the first years mainly from scientific problems in telecommunication. But around the end of the 1960s, the focus widened to more general aspects of human-human and human-machine communication. The paper investigates how this process ad-vanced at Dresden Technical University (TU) with special respect to external influences, originating from phonetics and communication sciences.
Among the numerous activities of the younger Ernst Mach in the development of psychophysics, there is a period of cooperation with the German otologist Johannes Kessel (1839-1907) in Prague from 1871 to 1874, which was not investigated in detail before. This cooperation was important because a number of essential findings in the psychophysics of hearing were published by both authors. Preparing a biography of Kessel, we collected new material about his cooperation with Mach from hitherto unpublished letters, archive material, and from Mach's unpublished diaries from the corresponding years. This paper describes the previous activities of Mach in psychophysics of hearing including the development of the required methods, the grant from the Vienna Academy for the investigation of the sound conduction in the human hearing organ through the middle ear and especially through the ossicles, the curriculum vitae of Kessel before he came to Prague, the contents and the results of the common research in Prague, and the influence of this period on the further development of Mach and Kessel.
This study is concerned with the prosodic properties of English speech produced by Chinese students. The database used in this study is a reasonably large-scale, multilingual speech corpus designed for prosody research. The speech data contains three recordings: 1) 40 English passages read by 10 native English speakers; 2) the same English passages read by 10 Chinese native speakers; 3) 40 Chinese passages read by the same Chinese speakers. Momel algorithm was employed to calculate the melody metrics on this database, and 14 prosodic parameters were compared among the three types of speech. The discrimination of Chinese English from native English and native Chinese regardless of gender reached approximately 76.5%. The results demonstrated that the pitch movements in Mandarin Chinese were greater and faster than in English, and the English produced by Chinese speakers showed greater and faster pitch movements than native English speech, but smaller and slower pitch movements than Mandarin Chinese speech. However, there was a slight difference between Chinese male and female speakers, but larger differences between English male and female speakers. Chinese female speakers even displayed smaller magnitudes of pitch movements on the basis of Momel targets than English female speakers in their English speech. Further investigations should be carried out to find out the reasons of such pitch movement patterns.
: Speech recognition and speech synthesis have been fascinating research disciplines since a long time, but potential applications suffered from low acceptance mainly due to poor quality. This situation has rapidly changed with the introduction of “virtual assistants” and similar speech-controlled products, and speech technology has arrived at the daily life in the recent decade. This is an opportune moment to look back into the long history of speech synthesis as an engineering discipline, which starts with early heuristic attempts, continues with the elaboration of a scientific foundation, and culminates with the availability of electronics and computing. It is the aim of this paper, not only to trace this path by selected examples, but also to identify the driving forces which were acting in the different periods of this development. The oral presentation will be illustrated by material from the Historic Acoustic-phonetic Collection (HAPS) of the TU Dresden.
This paper demonstrates by means of an example, how historic collections of universities can be utilized in modern research and teaching. The project refers to the Historic Acoustic-phonetic Collection (HAPS) of the TU Dresden. Two "guiding fossils" from the history of speech technology are selected to present a selection of results.
A comparison was made among the fundamental frequency (F0) patterns of continuous speech in English, Mandarin Chinese and L2 English produced by Chinese speakers.Ten adult native Chinese speakers were asked to read narrative text written in both English and Chinese.The comparative analysis of 300 sentences was performed in the following aspects: F0 mean, pitch range, pitch change rate and pitch change amount.It is found that in terms of both pitch range on the phoneme level and pitch change amount on the utterance level, L2 English speech by Chinese subjects displayed a significantly larger value than the English speech by native speakers.Moreover, the same Chinese subjects demonstrated still a larger value in these two pitch-related variables in their Chinese speech.The dynamic characteristic of L2 English can be attributed to the negative transfer of L1 Chinese.The findings can shed some light on the understanding of the difference in F0 patterns between a tone language and a non-tone language, and can also provide some implications for L2 speech learning.
This study is concerned with the organization of inter-lexical pauses of L2 English speech produced by Chinese students. Foreign language learners (L2 learners) have often been observed to lack a natural “rhythm” when reading aloud. This investigation analyzed read speech of 5 English sentences at three different speech tempos by 18 Chinese native speakers. We compared the patterns of inter-lexical pauses of the Chinese subjects in the collected speech with those of the native English speakers reported in previous literature. It was found that the Chinese speakers could hardly read English at a very fast speech rate. Their patterns of inter-lexical pauses were more similar to those of the native speakers at a slow or normal speech rate than at a fast speech rate. However, most interlexical pause patterns of L2 English by the Chinese speakers deviate from those of the native speakers, which may result in foreign-accented speech with degraded intelligibility. We suggest that the importance of the organization of pauses between words should be taken into account in L2 English speech learning and teaching.
We examine a 10:1 scale model of a recently proposed type of integrated sensor-actuator system to be used as part of a middle-ear hearing implant, and show that this inherently unstable system can be controlled by conventional methods of digital feedback suppression. Sensor and actuator of the model are mechanically coupled by a free-floating hard metal frame about 30 mm in length, corresponding to 3 mm original scale. Due to the direct mechanical coupling, the resulting electroacoustic system exhibits strong feedback and high-frequency oscillation will occur even at small gain. We show that the feedback of this system can be successfully controlled by delaying the signal in the forward path and then applying adaptive and active filtering. Through the use of a least-mean-square adaptive filter combined with active digital filtering, an increase in maximum stable gain (MSG) of up to 43 dB and a functional gain of up to 35 dB (maximum) and 24.7 dB (mean) are achieved.
When the collection of phonetic instruments of the Phonetic Institute of the Hamburg University came to the HAPS Dresden in 2005, it included also some small mechanic voices, which can produce single sounds as well as few simple words. These voices are well-known in the phonetic literature as an early attempt to provide hard-hearing people with automatic training tools, following a proposal of the otologist Johannes Kessel in 1899. On the other hand, the phonetic literature never took any notice from the real origin of these interesting pieces. Therefore the author started an investigation some years ago [1], which guided him not only to the interesting field of mechanical voices in the manufacturing of toys and dolls, but moreover back to the roots of mechanical speech synthesis at the end of the 18th century. This paper gives a rough overview about the recent state of this investigation. 1 From Kempelen to Mälzel Apart from other reasons, the speaking machine of Wolfgang von Kempelen received its fame by a good marketing, which was mainly effected by demonstrations of Kempelen’s automata throughout Europe at his journey in the years 1783/84 [2]. The Saxon major-domo Joseph Freiherr zu Racknitz reports about the presentation of the automata in Dresden 1784, that “the speaking machine aroused admiration, while the chess player also produced curiosity” [3]. In 1791, von Kempelen published his summarizing book about the speaking machine [4]. The chess player, however, was stored in the Schönbrunn castle for two decades. After von Kempelen’s death in 1804, the chess player (called “the Turk”) came into the ownership of Johann Nepomuk Mälzel (1772–1838). He was a famous German musician, engineer, automata constructor, and entertainer, who is known today mainly as the eponym of the metronome. He restored Kempelen’s chess player and demonstrated it together with his own constructions throughout Europe and, from 1825, in America [5]. It is not completely clear whether Mälzel also acquired a copy of the speaking machine from the estate of Kempelen. Anyhow, he came to Vienna already in 1992, where he certainly got in touch with Kempelen and his work. He was educated very well in constructing mechanical musical instruments like the famous “Panharmonicon” (1805) and instrument-playing automata like a spectacular trumpeter (1808). Therefore it appears logically, that he started to equip his automata with voices. As an instance, he presented an automatic tightrope walker, which spoke words like “Oh là là”. Best known, the chess player obtained in the winter season 1819/20 the ability to pronounce the word “échec” (check) [6]. Edgar Allan Poe mentions in his famous essay on “Maelzel’s Chess-Player” from 1836 [7]: “During the progress of the game, the figure now and then rolls its eyes, as if surveying the board, moves its head, and pronounces the word ‘echec’ (check) when necessary. [. . . ] The making the Turk pronounce the word ‘echec’, is an improvement by M. Maelzel. When in possession of Baron Kempelen, the figure indicated a ‘check’ by rapping on the box with his right hand.” 60 ISCA Archive http://www.isca-speech.org/archive First International Workshop on the History of Speech Communication Research (HSCR 2015) Dresden, Germany September 4-5, 2015
Background: Up to 89% of the individuals with Parkinson's disease (PD) experience speech problem over the course of the disease.Speech prosody and intelligibility are two of the most affected areas in hypokinetic dysarthria.However, assessment of these areas could potentially be problematic as speech prosody and intelligibility could be affected by the type of speech materials employed.Objective: To comparatively explore the effects of different types of speech stimulus on speech prosody and intelligibility in PD speakers.Methods: Speech prosody and intelligibility of two groups of individuals with varying degree of dysarthria resulting from PD was compared to that of a group of control speakers using sentence reading, passage reading and monologue.Acoustic analysis including measures on fundamental frequency (F0), intensity and speech rate was used to form a prosodic profile for each individual.Speech intelligibility was measured for the speakers with dysarthria using direct magnitude estimation.Results: Difference in F0 variability between the speakers with dysarthria and control speakers was only observed in sentence reading task.Difference in the average intensity level was observed for speakers with mild dysarthria to that of the control speakers.Additionally, there were stimulus effect on both intelligibility and prosodic profile.Conclusions: The prosodic profile of PD speakers was different from that of the control speakers in the more structured task, and lower intelligibility was found in less structured task.This highlighted the value of both structured and natural stimulus to evaluate speech production in PD speakers.
The historic acoustic-phonetic collection (HAPS) of the TU Dresden documents the history of experimental phonetics and speech technology with a remarkably high degree of completeness. Due to some construction work, the collection could move to new and improved showrooms in the Barkhausen Building. This paper is a slightly extended version of the welcome speech, presented at the occasion of the re-opening of the collection in the framework of the First International Workshop on the History of Speech Communication Research (HSCR 2015), a satellite event of the Interspeech, Dresden 2015.
The present study investigates the possible prosodic deviance due to foreign accent in the German speech by Chinese speakers. German has lexical stress and has been described as stress-timed, while Mandarin Chinese has lexical tone and has been described as syllable-timed. It is by now well documented that the prosody of the second language can be influenced by the learner's native language. In the present investigation, we compare the speech by 18 Chinese learners of German at the low–intermediate level with six native German speakers. Ten sentences were selected for the analysis. The results of the investigation show that: (a) Chinese speakers of German have both a higher proportion of vocalic intervals (%V) and a higher standard deviation of consonantal intervals (ΔC) than German native speakers, resulting from their vowel epentheses and non-reduction of vowels, and their slow speaking rate respectively; (b) Chinese speakers produce a larger pitch range within the vocalic intervals and can hardly vary the intonation patterns to match different sentence types in German in order to express different intonational meanings. Their prosodic organization of German speech is more syllable-oriented rather than stress-oriented. All these deviant prosodic behaviours can be traced back to the characteristics of their native language. The findings of the present investigation can have implications for cross-language studies and foreign language education.