The long-term objective of this research is the development of a one-hand-controlled speech synthesizer, to give laryngectomees and other speech-impaired persons a means of producing higher-quality speech with less effort than currently available methods such as an electrolarynx or a text-to-speech system. To demonstrate the feasibility of a one-hand-controlled speech synthesizer, a system was constructed using a hand-held device similar to a pen connected to an articulated arm for measuring six degrees of freedom (three Cartesian and three rotational dimensions) as the user interface to an HLsyn-based speech synthesizer. Through this interface, the user controls parameters for the first three formants, pitch, subglottal pressure, and glottal area. Parameter control was introduced progressively in that order to four participants who underwent training to produce synthesized speech composed of a subset of English phonemes: vowels, semivowels, diphthongs, /h/, and the glottal stop. The complexity of the synthesized speech targets also grew from monosyllabic utterances to short phrases over the training. After training, a separate group of four listeners compared the naturalness and intelligibility of the synthesized speech to the same utterances produced by the participants with a text-to-speech system. [Work supported by NIDCD Grant Number R43 DC006134-01.]
This report completes the project entitled “Concept and Technology Exploration for Transparent Hearing Systems”, funded by the US Air Force Research Laboratory at Wright-Patterson Air Force Base in collaboration with Natick Soldier Systems of the US Army. The document outlines the project as planned and details the project as executed. Given the importance and time criticality of determining a solution to the problem addressed, the project team exploited knowledge gained during the project, redirecting the plan as necessary to maximize exploration. This document outlines the goals of the project, provides an overview of previous relevant work, discusses the work planned for the project, details the work and its findings, and describes how a solution system could be integrated into a dismounted soldier’s personal information system.The intended audience for this document includes the project sponsors, the intermediate contract managers, designated reviewers, and future helmet system designers. Additionally, the report authors assume the document may be published to a wider audience. The designated reviewers may encompass professionals in the fields of hearing, signal processing, sensors, warfighting equipment, hearing enhancement/augmentation, and aural displays, who can give feedback and guidance to extensions of the project.
Sets of pellet coordinates from the X-ray Microbeam Speech Production Database, each corresponding to a static articulatory configuration, are submitted to a principal components analysis. Several talkers are analyzed separately, and various criteria are used to select the times in the continuous speech data from which to extract pellet data. In some instances, all of a talker’s speech data are submitted to an automatic procedure intended to select syllable nuclei, based on Mermelstein’s algorithm [P. Mermelstein, J. Acoust. Soc. Am. 58, 880–883 (1975)]. In other instances specific data are chosen by hand according to utterance type (e.g., vowel, glide, and consonant–vowel transition). These analyses are intended for use in a project to recover articulation from speech acoustics. In particular, we examine the relationship between pellet positions from data, the tongue shape reconstructed from the principal components analysis, and the parameters of a simple acoustic tube model, which is derived from the Stevens and House model [K. N. Stevens and A. S. House, J. Acoust. Soc. Am. 27, 484–493 (1955)]. [Work supported by Grant NIDCD-01247 to Sensimetrics Corporation.]
HLsyn is a quasiarticulatory speech synthesizer in which a small set of parameters control a Klatt synthesizer [Stevens and Bickley, J. Phon. 19 (1991)]. Originally, ten parameters were used. In this paper three new physiologically based parameters, together with some additional modifications, are described. The first is a time-varying subglottal pressure parameter, which provides the user with additional control of the voice-source amplitude. It can also be used to turn on voicing, and it has an influence on fundamental frequency. The second parameter is a time-varying percentage change in the compliances of the vocal-tract walls and vocal folds. This parameter can, for example, be employed when synthesizing voiced obstruents: increasing it results in facilitation of glottal vibration and lowering of the fundamental frequency. The third parameter is the time-varying cross-sectional area of a posterior glottal chink. Because it is independent of the area at the membranous folds, significant noise-source amplitudes can be achieved during voiced speech. Thus one can synthesize aspirated voiced stops, as well as breathy speech, in a natural manner. Finally, a default female voice and an intrinsic pitch feature have been added to the system. Examples of copy synthesis where the new parameters play a role are presented. [Work supported by NIH Grant MH52358.]
We show that Martin’s axiom for countable partial orders implies the existence of a countable dense homogeneous Bernstein subset of the reals. Using Martin’s axiom we derive a characterization of the countable dense homogeneous spaces among the separable metric spaces of cardinality less thanc. Also, we show that Martin’s axiom implies the existence of a subset of the Cantor set which isλ-dense homogeneous for everyλ <c.