We present a framework for the induction of semantic frames from utterances in the context of an adaptive command-and-control interface. The system is trained on an individual user's utterances and the corresponding semantic frames representing controls. During training, no prior information on the alignment between utterance segments and frame slots and values is available. In addition, semantic frames in the training data can contain information that is not expressed in the utterances. To tackle this weakly supervised classification task, we propose a framework based on Hidden Markov Models (HMMs). Structural modifications, resulting in a hierarchical HMM, and an extension called expression sharing are introduced to minimize the amount of training time and effort required for the user. The dataset used for the present study is PATCOR, which contains commands uttered in the context of a vocally guided card game, Patience. Experiments were carried out on orthographic and phonetic transcriptions of commands, segmented on different levels of n-gram granularity. The experimental results show positive effects of all the studied system extensions, with some effect differences between the different input representations. Moreover, evaluation experiments on held-out data with the optimal system configuration show that the extended system is able to achieve high accuracies with relatively small amounts of training data.
In command-and-control applications, a vocal user interface (VUI) is useful for handsfree control of various devices, especially for people with a physical disability. The spoken utterances are usually restricted to a predefined list of phrases or to a restricted grammar, and the acoustic models work well for normal speech. While some state-of-the-art methods allow for user adaptation of the predefined acoustic models and lexicons, we pursue a fully adaptive VUI by learning both vocabulary and acoustics directly from interaction examples. A learning curve usually has a steep rise in the beginning and an asymptotic ceiling at the end. To limit tutoring time and to guarantee good performance in the long run, the word learning rate of the VUI should be fast and the learning curve should level off at a high accuracy. In order to deal with these performance indicators, we propose a multi-level VUI architecture and we investigate the effectiveness of alternative processing schemes. In the low-level layer, we explore the use of MIDA features (Mutual Information Discrimination Analysis) against conventional MFCC features. In the mid-level layer, we enhance the acoustic representation by means of phone posteriorgrams and clustering procedures. In the high-level layer, we use the NMF (Non-negative Matrix Factorization) procedure which has been demonstrated to be an effective approach for word learning. We evaluate and discuss the performance and the feasibility of our approach in a realistic experimental setting of the VUI-user learning context.
This paper issues in the design of a vocal interface for a robot that can learn to understand spoken utterances through demonstration. Weakly supervised non-negative matrix factorization (NMF) is used as a machine learning algorithm where acoustic data are augmented with semantic labels representing the meaning of the command. Many parameters that the robot needs in order to execute the commands have an ordinal structure. Constrained subspace NMF (CSNMF) is proposed as an extension to NMF that aims to better deal with ordinal data and thus increase the learning rate of the grounding information with an ordinal structure. Furthermore automatic relevance determination is used to deal with model order selection. The use of CSNMF yields a significant improvement in the learning rate and accuracy when recognising ordinal parameters.
In this paper, we investigate unsupervised acoustic model training approaches for dysarthric-speech recognition. These models are first, frame-based Gaussian posteriorgrams, obtained from Vector Quantization (VQ), second, so-called Acoustic Unit Descriptors (AUDs), which are hidden Markov models of phone-like units, that are trained in an unsupervised fashion, and, third, posteriorgrams computed on the AUDs. Experiments were carried out on a database collected from a home automation task and containing nine speakers, of which seven are considered to utter dysarthric speech. All unsupervised modeling approaches delivered significantly better recognition rates than a speaker-independent phoneme recognition baseline, showing the suitability of unsupervised acoustic model training for dysarthric speech. While the AUD models led to the most compact representation of an utterance for the subsequent semantic inference stage, posteriorgram-based representations resulted in higher recognition rates, with the Gaussian posteriorgram achieving the highest slot filling F-score of 97.02%.
Speech technology is firmly rooted in daily life, most notably in command-and-control (C&C) applications. C&C usability downgrades quickly, however, when used by people with non-standard speech. We pursue a fully adaptive vocal user interface (VUI) which can learn both vocabulary and grammar directly from interaction examples, achieving robustness against non-standard speech by building up models from scratch. This approach raises feasibility concerns on the amount of training material required to yield an acceptable recognition accuracy. In a previous work, we proposed a VUI based on non-negative matrix factorisation (NMF) to find recurrent acoustic and semantic patterns comprising spoken commands and device-specific actions, and showed its effectiveness on unimpaired speech. In this work, we evaluate the feasibility of a self-taught VUI on a new database called domotica-3, which contains dysarthric speech with typical commands in a home automation setting. Additionally, we compare our NMF-based system with a system based on Gaussian mixtures. The evaluation favours our NMF-based approach, yielding feasible recognition accuracies for people with dysarthric speech after a few learning examples. Finally, we propose the use of a multi-layered semantic frame structure and demonstrate its effectiveness in boosting overall performance.
This research is situated in a project aimed at the development of a vocal user interface (VUI) that learns to understand its users specifically persons with a speech impairment. The vocal interface adapts to the speech of the user by learning the vocabulary from interaction examples. Word learning is implemented through weakly supervised non-negative matrix factorization (NMF). The goal of this study is to investigate how we can improve word learning when the number of interaction examples is low. We demonstrate two approaches to train NMF models on scarce data: 1) training word models using smoothed training data, and 2) training word models that strictly correspond to the grounding information derived from a few interaction examples. We found that both approaches can substantially improve word learning from scarce training data.
In this work we describe research aimed at developing an assistive vocal interface for users with a speech impairment. In contrast to existing approaches, the vocal interface is self-learning, which means it is maximally adapted to the end-user and can be used with any language, dialect, vocabulary and grammar. The paper describes the overall learning framework and the vocabulary acquisition technique, and proposes a novel grammar induction technique based on weakly supervised hidden Markov model learning. We evaluate early implementations of these vocabulary and grammar learning components on two datasets: recorded sessions of a vocally guided card game by non-impaired speakers and speech-impaired users engaging in a home automation task.
A self-learning vocal user interface learns to map user-defined spoken commands to intended actions. The voice user interface is trained by mining the speech input and the provoked action on a device. Although this generic procedure allows a great deal of flexibility, it comes at a cost. Two requirements are important to create a user-friendly learning environment. First, the self-learning interface should be robust against typical errors that occur in the interaction between a non-expert user and the system. For instance, the user gives a wrong learning example to the system by commanding “Turn on the television!” and pushing a power button on the wrong remote control. The spoken command is then supervised by a wrong action and we refer to these errors as label noise. Secondly, the mapping between voice commands and intended actions should happen fast, i.e. require few examples. To meet these requirements, we implemented learning through supervised NMF. We tested keyword recognition accuracy for different levels of label noise and different sizes of training sets. Our learning approach is robust against label noise, but some improvement regarding fast mapping is desirable.
This paper gives an overview of research within the ALADIN project, which aims to develop an assistive vocal interface for people with a physical impairment. In contrast to existing approaches, the vocal interface is trained by the end-user himself, which means it can be used with any vocabulary and grammar, and that it is maximally adapted to the — possibly dysarthric — speech of the user. This paper describes the overall learning framework, the user-centred design and evaluation aspects, database collection and approaches taken to combat problems such as noise and erroneous input.
In distinguishing individual shapes (defined by their contours), older children (6.5 years of age on average) performed better than younger children (4 years of age on average), and, although the task did not involve any categorization or generalization, the error pattern was qualitatively affected by shape differences that are generally common distinctions between objects belonging to different categories. The influence of these shape differences was also observed for unfamiliar shapes, demonstrating that the influence of categorization experience was not modulated by the retrieval of shape features from known categories but rather related to a different perception of shape by age. The results suggest a direct influence of categorization experience on more abstract shape processing. When children were distinguishing shapes, new words were paired with the target shapes, and in 2 additional tasks, the acquired name–shape associations were tested. The younger age group was able to remember more words correctly.
Rules and similarity are at the heart of our understanding of human categorization. However, it is difficult to distinguish their role as both determinants of categorization are confounded in many real situations. Rules are based on a number of identical properties between objects but these correspondences also make objects appearing more similar. Here, we introduced a stimulus set where rules and similarity were unconfounded and we let participants generalize category examples towards new instances. We also introduced a method based on the frequency distribution of the formed partitions in the stimulus sets, which allowed us to verify the role of rules and similarity in categorization. Our evaluation favoured the rule-based account. The most preferred rules were the simplest ones and they consisted of recurrent visual properties (regularities) in the stimulus set. Additionally, we created different variants of the same stimulus set and tested the moderating influence of small changes in appearance of the stimulus material. A conceptual manipulation (Experiment 1) had no influence but all visual manipulations (Experiment 2 and 3) had strong influences in participants' reliance on particular rules, indicating that prior beliefs of category defining rules are rather flexible.
A shape bias for extending names to objects that look visually similar has been commonly accepted but it is hard to define which kind of shape dissimilarities are diagnostic for the identity of an object. Here, we present a transformational approach to describe shape differences that can incorporate many significant shape features. We introduce two kinds of transformations: one kind concerns linear transformations of the image plane (affine transformations), generally limiting shape variations within the borders of basic-level categories; the other kind concerns nonlinear continuous transformations of the image plane (topological transformations), allowing all kinds of shape variation crossing and not crossing the borders of basic-level categories. We administered stimulus pairs differing in these shape transformations to children of 3 years to 7 years old in a delayed match-to-sample task. With increasing age, especially between 5 years and 6 years, children became more sensitive to the topological deformations that are relevant for between-category distinctions, indicating that acquired categorical knowledge in early years induces perceptual learning of the relevant generic shape differences between categories.
Visual anisotropy has been demonstrated in multiple tasks where performance differs between vertical, horizontal, and oblique orientations of the stimuli. We explain some principles of visual anisotropy by anisotropic smoothing, which is based on a variation on Koenderink's approach in [1]. We tested the theory by presenting gaussian elongated luminance profiles and measuring the perceived orientations by means of an adjustment task. Our framework is based on the smoothing of the image with elliptical gaussian kernels and it correctly predicted an illusory orientation bias towards the vertical axis. We discuss the scope of the theory in the context of other anisotropies in perception.
Plateau’s irradiation phenomenon in particular describes what one sees when observing a brighter object on a darker background and a physically congruent darker object on a brighter background: the brighter object is seen as being larger. This phenomenon occurs in many optical visual illusions and it involves some fundamental aspects of human vision. We present a general geometrical model of human visual sensation and perception, hereby taking into account the law of Fechner in addition to the anisotropic smoothing that was introduced in [1], and explicitly illustrate its meaning for irradiation illusions of Helmholtz and Kitaoka.
There is a long-standing debate concerning the holistic or analytic nature of perceptual processing and whether the representation builds up from global to local or from local to global. Despite many fruitful contributions demonstrating the global and local levels in object recognition, little systematic attention has been directed at the local-global dominance in generalization and categorization. We investigated the global and local dominance of visual cues in a name generalization task. Two obvious features, one on a global scale and one on a more local scale, were implemented in the stimulus set. Stimuli were always 2-D shapes, generated by Boundary Descriptors. Participants were asked to infer new members of a category and could freely choose new instances in the stimulus set. By evaluating the inferred exemplars, the global-local dominance in generalization was studied. The results demonstrate that general conclusions are hard to draw based on the theoretically presumed allocation of the visual cues to the local or the global level of processing. However, by a thorough evaluation of the visual cues in terms of their regularities in the stimulus set, we were able to draw a general conclusion: Participants derive category membership by relying on regularities on the most global scale. Regularities on a more global scale are conceived as less coincidental features in the stimulus set, and therefore, they are preferred as a basis to infer new instances. Additionally, we tested the moderating influence of particular context variables like superordinate category ownership, planar rotations, and complexity on participants' reliance on global or local scale shape properties in generalization. Superordinate category ownership had no influence but more complex shapes and arbitrary planar rotations of the stimuli led to huge differences in participants' reliance on visual cues. We will discuss the results in light of some contemporary models of object recognition.
The shape of an object is fundamental in object recognition but it is still an open issue to what extent shape differences are perceived analytically (i.e., by the dimensional structure of the shapes) or holistically (i.e., by the overall similarity of the shapes). The dimensional structure of a stimulus is available in a primary stage of processing for separable dimensions, although it can also be derived cognitively from a perceived stimulus consisting of integral dimensions. Contrary to most experimental paradigms, the present study asked participants explicitly to analyze shapes according to two dimensions. The dimensions of interest were aspect ratio and medial axis curvature, and a new procedure was used to measure the participants' interpretation of both dimensions (Part I, Experiment 1). The subjectively interpreted shape dimensions showed specific characteristics supporting the conclusion that they also constitute perceptual dimensions with objective behavioral characteristics (Part II): (1) the dimensions did not correlate in overall similarity measures (Experiment 2), (2) they were more separable in a speeded categorization task (Experiment 3), and (3) they were invariant across different complex 2-D shapes (Experiment 4). The implications of these findings for shape-based object processing are discussed.