In the modern era, mouse control has become an important part of human computer interaction which becomes difficult for physically disabled people. This dissertation presents a system called as Vocal Mouse (VM). This device will allow users to continuously control the mouse pointer using words as well as sounds by varying vocal parameters such as vowel quality, loudness and pitch. Traditional method of using only standard spoken words was inefficient for performing continuous tasks and they are often recognized poorly by automatic speech recognizers. Now, VM will allow users to work on both continuous and discrete motion control. This includes commands given as words or regular sounds consisting of vowels and consonants. Low-level acoustic features are extracted in real time using LPC (Linear Predictive Coding). Pattern recognition is performed using a new proposed technique called “minimum feature distance technique”. This proposed technique is based on calculating distances between the spoken word and each stored word in the library during training process. Features from pattern recognition module are processed to produce output in the form of cursor’s 2-D movement. VM can be used by novice users without extensive training and it presents a viable alternative to existing speech-based cursor control methods. KeywordsLPC, VM, Speech recogniser
The goal of Human Computer interaction (HCI) is to design a system that improves the interaction between computers and users, mainly for people with disabilities. For this reason, this research describes how to make computers more usable by using speech based interaction. Proposed system will use speech and non speech characteristics of human voice for hands free computing. System will use human vocal characteristics to command mouse pointer. Words and sounds will be used to perform mouse 2-D movement and performing operating system’s operations. By doing this people with disabilities can use the computer systems normally as the normal users use it. Current speech recognition systems provide discrete interaction because they process user’s vocal utterances at word level, proposed system will be able to perform mouse-like operations fluidly.
Despite the variety of statistical classifiers available, many inductive classifiers are ultimately based on exponential functional forms. These models are quite versatile and perform well in a variety of situations, but they also have a strong tendency towards overconfidence. They will typically be quite sure of their predictions, whether correct or not, except when quite close to a decision boundary. This dissertation, motivated by the Vocal Joystick project, an assistive device allowing individuals with motor impairments to operate mouse pointers or other electromechanical devices using non-verbal vocalizations, examines a family of statistical classification models designed to allow smoother transitions between classes. We first present a classification model formulated as a ratio of semi-definite polynomials, which we call the ratio semi-definite classifier (RSC). The RSC allows smoother transitions across class boundaries, and does not demonstrate the overconfidence bias typical of models based on ratios of exponentials Testing on several corpora of various sizes, we find that the RSC performs well, but often has slightly lower accuracy than an exponential model such as a multi-layer perceptron (MLP). To improve the accuracy of the RSC while retaining its other properties, we propose several extensions to the model, creating a family of RSC-based classifiers. We test two other members of that family, the multi-layer RSC (ML-RSC) which adds a hidden layer to the model allowing us to learn a kernel-like feature transformation in a data-driven manner. Experimental results show that the ML-RSC provides superior accuracy compared with the original RSC and an MLP.We also propose a semi-supervised learning (SSL) framework suitable for any differentiable, parametric and optionally multiclass model. With this framework, we demonstrate improved results with semi-supervised versions of the RSC, ML-RSC and MLP on several data sets.Finally, we present a comprehensive look at the Vocal Joystick engine, describing its architecture and the various signal processing and machine learning techniques used in the system. This includes a discussion of where the RSC and ML-RSC, and adaptation of the parameters for those models, are used in the Vocal Joystick.
We present the Vocal Joystick engine, a real-time software library which can be used to map non-linguistic vocalizations into realizable continuous control signals. The system is designed to strike a balance between low latency and accurate recognition while simultaneously taking advantage of the rich complexity of sounds producible by the human vocal tract. By developing a modular, cross-platform library, we aim to provide a robust but simple means of incorporating such controls into any application, thereby producing a new form of accessible technology. This is demonstrated by the various applications that so far have used the Vocal Joystick engine. Unlike previous discussions of parts of the Vocal Joystick, this paper presents a detailed view of the inner workings of the current version of the engine.
Seeking classifier models that are not overconfident and that better represent the inherent uncertainty over a set of choices, we extend an objective for semi-supervised learning for neural networks to two models from the ratio semi-definite classifier (RSC) family. We show that the RSC family of classifiers produces smoother transitions between classes on a vowel classification task, and that the semi-supervised framework provides further benefits for smooth transitions. Finally, our testing methodology presents a novel way to evaluate the smoothness of classifier transitions (interpolating between vowels) by using samples from classes unseen during training time.
We address the issue of learning multi-layered perceptrons (MLPs) in a discriminative, inductive, multiclass, parametric, and semi-supervised fashion. We introduce a novel objective function that, when optimized, simultaneously encourages 1) accuracy on the labeled points, 2) respect for an underlying graph-represented manifold on all points, 3) smoothness via an entropic regularizer of the classifier outputs, and 4) simplicity via an ‘2 regularizer. Our approach provides a simple, elegant, and computationally efficient way to bring the benefits of semi-supervised learning (and what is typically an enormous amount of unlabeled training data) to MLPs, which are one of the most widely used pattern classifiers in practice. Our objective has the property that efficient learning is possible using stochastic gradient descent even on large datasets. Results demonstrate significant improvements compared both to a baseline supervised MLP, and also to a previous non-parametric manifold-regularized reproducing kernel Hilbert space classifier.
We conducted a 2.5 week longitudinal study with five motor impaired (MI) and four non-impaired (NMI) participants, in which they learned to use the Vocal Joystick, a voice-based user interface control system. We found that the participants were able to learn the mapping between the vowel sounds and directions used by the Vocal Joystick, and showed marked improvement in their target acquisition performance. At the end of the ten session period, the NMI group reached the same level of performance as the previously measured "expert" Vocal Joystick performance, and the MI group was able to reach 70% of that level. Two of the MI participants were also able to approach the performance of their preferred device, a touchpad. We report on a number of issues that can inform the development of further enhancements in the realm of voice-driven computer control.
We develop a novel extension to the Ratio Semi-definite Classifier, a discriminative model formulated as a ratio of semi-definite polynomials. By adding a hidden layer to the model, we can efficiently train the model, while achieving higher accuracy than the original version. Results on artificial 2-D data as well as two separate phone classification corpora show that our multi-layer model still avoids the overconfidence bias found in models based on ratios of exponentials, while remaining competitive with state-of-the-art techniques such as multi-layer perceptrons.
We address the issue of learning multi-layered perceptrons (MLPs) in a discriminative, inductive, multiclass, parametric, and semi-supervised fashion. We introduce a novel objective function that, when optimized, simultane- ously encourages 1) accuracy on the labeled points, 2) respect for an underlying graph-represented manifold on all points, 3) smoothness via an entropic regularizer of the classifier outputs, and 4) simplicity via an '2 regularizer. Our approach provides a simple, elegant, and computationally efficient way to bring the benefits of semi-supervised learning (and what is typically an enormous amount of unlabeled training data) to MLPs, which are one of the most widely used pattern classifiers in practice. Our objective has the property that efficient learning is possible using stochastic gradient descent even on large datasets. Results demonstrate significant improvements compared both to a baseline supervised MLP, and also to a previous non-parametric manifold-regularized reproducing kernel Hilbert space classifier.
We present a system whereby the human voice may specify continuous control signals to manipulate a simulated 2D robotic arm and a real 3D robotic arm. Our goal is to move towards making accessible the manipulation of everyday objects to individuals with motor impairments. Using our system, we performed several studies using control style variants for both the 2D and 3D arms. Results show that it is indeed possible for a user to learn to effectively manipulate real-world objects with a robotic arm using only non-verbal voice as a control mechanism. Our results provide strong evidence that the further development of non-verbal voice controlled robotics and prosthetic limbs will be successful.
We present a novel classification model that is formulated as a ratio of semi-definite polynomials. We derive an efficient learning algorithm for this classifier, and apply it to two separate phoneme classification corpora. Results show that our disciminatively trained model can achieve accuracies comparable with state-of-the-art techniques such as multi-layer perceptrons, but does not posses the overconfident bias often found in models based on ratios of exponentials.
The Vocal Joystick is a novel human-computer interface mechanism designed to enable individuals with motor impairments to make use of vocal parameters to control objects on a computer screen (buttons, sliders, etc.) and ultimately electro-mechanical instruments (e.g., robotic arms, wireless home automation devices). We have developed a working prototype of our "VJ-engine" with which individuals can now control computer mouse movement with their voice. The core engine is currently optimized according to a number of criterion. In this paper, we describe the engine system design, engine optimization, and user-interface improvements, and outline some of the signal processing and pattern recognition modules that were successful. Lastly, we present new results comparing the vocal joystick with a state-of-the-art eye tracking pointing device, and show that not only is the Vocal Joystick already competitive, for some tasks it appears to be an improvement
We explore the use of the Vocal Joystick (VJ) for robotic limb control using a small-scale robotic arm. The purpose of this research is to allow individuals with mobile disabilities to obtain greater independence by continuously controlling a robotic arm with their voice. The VJ-Voicebot relies only on continuous and discrete non-verbal vocal sounds to interact with objects in its environment. The demonstration will allow users to experience real-time control of a 5 degrees-of-freedom (DOF) robotic arm using sounds produced from their own vocal system.
Vocal Joystick is a mechanism that enables individuals with motor impairments to make use of vocal parameters to control objects on a computer screen (buttons, sliders, etc.) and ultimately will be used to control electro-mechanical instruments (e.g., robotic arms, wireless home automation devices). In an effort to train the VJ-system, speech data from the TIMIT speech corpus was initially used. However, due to problematic issues with co-articulation, we began a large data collection effort in a controlled environment that would not only address the problematic issues, but also yield a new vowel corpus that was representative of the utterances a user of the VJ-system would use. The data collection process evolved over the course of the effort as new parameters were added and as factors relating to the quality of the collected data in terms of the specified parameters were considered. The result of the data collection effort is a vowel corpus of approximately 11 hours of recorded data comprised of approximately 23500 sound files of the monophthongs and vowel combinations (e.g. diphthongs) chosen for the Vocal Joystick project varying along the parameters of duration, intensity and amplitude. This paper discusses how the data collection has evolved since its initiation and provides a brief summary of the resulting corpus.
The Vocal Joystick (VJ) is an assistive device that uses the rich complexity of the human voice to drive a human-computer interface. The system has previously been shown to work well for control of a computer mouse, yet it does not currently make full use of the continuous nature of the vowel space. This work examines the potential use of the Gaussian process latent variable model (GP-LVM), which provides a non-linear, smooth probabilistic mapping from latent to data space. The results show promise for some speakers but do not currently generalize well across speakers. The GP-LVM provides a well-motivated approach, but additional work is still needed to translate the existing potential into an effective new control scheme for VJ.
This paper introduces a novel adaptive direction filtering algorithm in the Vocal Joystick (VJ) setting that utilizes context information and applies real-time inference in a continuous space. The VJ system using this algorithm is endowed with the ability to produce movements in arbitrary directions and the ability to draw smooth curves. This is in contrast to previous VJ settings whereby vowel quality was used to determine mouse movement in only a finite discrete set of directions [1].
This paper introduces a novel adaptive direction filtering algorithm in the Vocal Joystick (VJ) setting that utilizes context information and applies real-time inference in a continuous space. The VJ system using this algorithm is endowed with the ability to produce movements in arbitrary directions and the ability to draw smooth curves. This is in contrast to previous VJ settings whereby vowel quality was used to determine mouse movement in only a finite discrete set of directions [1].
With the skyrocketing popularity of mobile devices, new processing methods tailored to a specific application have become necessary for low-resource systems. This work presents a high-speed, low-resource speech recognition system using custom arithmetic units, where all system variables are represented by integer indices and all arithmetic operations are replaced by hardware-based table lookups. To this end, several reordering and rescaling techniques, including two accumulation structures for Gaussian evaluation and a novel method for the normalization of Viterbi search scores, are proposed to ensure low entropy for all variables. Furthermore, a discriminatively inspired distortion measure is investigated for scalar quantization of forward probabilities to maximize the recognition rate. Finally, heuristic algorithms are explored to optimize system-wide resource allocation. Our best bit-width allocation scheme only requires 59 kB of ROMs to hold the lookup tables, and its recognition performance with various vocabulary sizes in both clean and noisy conditions is nearly as good as that of a system using a 32-bit floating-point unit. Simulations on various architectures show that, on most modern processor designs, we can expect a cycle-count speedup of at least three times over systems with floating-point units. Additionally, the memory bandwidth is reduced by over 70% and the offline storage for model parameters is reduced by 80%.
Vocal Joystick is a mechanism that enables individuals with motor impairments to make use of vocal parameters to control objects on a computer screen (buttons, sliders, etc.) and ultimately will be used to control electro-mechanical instruments (e.g., robotic arms, wireless home automation devices). In an effort to train the VJ-system, speech data from the TIMIT speech corpus was initially used. However, due to problematic issues with co-articulation, we began a large data collection effort in a controlled environment that would not only address the problematic issues, but also yield a new vowel corpus that was representative of the utterances a user of the VJ-system would use. The data collection process evolved over the course of the effort as new parameters were added and as factors relating to the quality of the collected data in terms of the specified parameters were considered. The result of the data collection effort is a vowel corpus of approximately 11 hours of recorded data comprised of approximately 23500 sound files of the monophthongs and vowel combinations (e.g. diphthongs) chosen for the Vocal Joystick project varying along the parameters of duration, intensity and amplitude. This paper discusses how the data collection has evolved since its initiation and provides a brief summary of the resulting corpus.