A biological model of cortical learning based on neural fields is elaborated and analyzed. In an effort to establish the method as a viable option in today's machine learning world, the method is evaluated based on the MNIST dataset. The approach relates to established work in dynamic systems and neural field models, with an innovation that prevents runaway feedback between neural fields. The method is advanced through four studies. The first three studies illustrate the mechanics of the method by way of supervised learning. Results are compared to analogous results achieved through other methods. The final study builds on the first three studies to illustrate and evaluate how unsupervised learning is accomplished. Concluding discussion considers advantages of the approach over other approaches, and how large neural field networks may be constructed and applied to temporal pattern learning in domains such as robotics.
In building robotic systems that interact with people through speech, many robotics engineers are obliged to treat artificial speech recognition and synthesis as a black-box problem best left to speech engineers to solve. Yet speech engineers today typically do not have access to the kinds of expensive robots needed for this development. Progress on the human-robot speech interface thus suffers from something of a diffusion of responsibility. In an attempt to remedy the situation, we have developed a low-cost interactive embodied speech device. The device is constructed from off-the-shelf components and from 3D-printed and laser-cut parts. We make the files for the 3D and laser-cut parts freely available for download. In addition to offering basic assembled devices and kits for self-assembly, we provide an . assembly guide and a shopping list of components a user will need in order to build, maintain, and customize their own device. We supply a basic software framework (in both Matlab 'and in C/C++), and template code for a ROS node for interfacing with the device. The idea is to establish a standard and accessible hardware platform with an open-source foundation for the sharing of ideas and research.
Previous attempts at modeling the neuro-cognitive mechanisms underlying word processing have used connectionist approaches, but none has modeled spoken word architectures as the input is presented in real-time. Hence, such models rely on the ingenuity of the modeler to establish a mapping of realtime stimulus to the model’s input which may not preserve processing that happens during each time step. We present a neural field model which successfully replicates the effect of immediate auditory repetition of monosyllabic words and fits it to a component of a well-studied mechanism for analyzing language processing, the event-related potential (ERP). This represents a new modeling approach to studying the neurocognitive processes, one that is based on the bottom-up interaction of real-time sensory information with higher-level categories of cognitive processing.
The development of large-scale automatic classroom dialog analysis systems requires accurate speech-to-text translation. A variety of automatic speech recognition (ASR) engines were evaluated for this purpose. Recordings of teachers in noisy classrooms were used for testing. In comparing ASR results, Google Speech and Bing Speech were more accurate with word accuracy scores of 0.56 for Google and 0.52 for Bing compared to 0.41 for AT&T Watson, 0.08 for Microsoft, 0.14 for Sphinx with the HUB4 model, and 0.00 for Sphinx with the WSJ model. Further analysis revealed both Google and Bing engines were largely unaffected by speakers, speech class sessions, and speech characteristics. Bing results were validated across speakers in a laboratory study, and a method of improving Bing results is presented. Results provide a useful understanding of the capabilities of contemporary ASR engines in noisy classroom environments. Results also highlight a list of issues to be aware of when selecting an ASR engine for difficult speech recognition tasks.
A theory of motor-feedback learning based on a bidirectional-edge graph is summarized. The graph is inspired as a model of the cortico-thalamic circuit. Output from the network drives the motors of an articulatory speech synthesizer. The network is first exposed to an ongoing stream of speech sounds where neural fields tune themselves to respond to repetitive co-occurring patterns in the input (phonetic features). The network is then made to actuate motors with babble feedback learning in an effort to train the network to produce the speech sounds and strings of sounds that the network was originally exposed to. The model furthermore makes use of a bank of resonators, intended to depict the functional role of the cerebellum. This cerebellar model helps in coordinating the longer-term timing of network dynamics and helps to smoothly string together patterns of speech sounds through time.
We evaluate a variety of audio recording techniques for a project on the automatic analysis of speech dialog in middle school and high school classrooms. In our scenario, the teacher wears a headset microphone or a lapel microphone. A second microphone is then used to collect speech and related sounds from students in the classroom. Various boundary microphones, omni-directional microphones, and cardioid microphones are tested as this second classroom microphone. A commercial microphone array [Microsoft Xbox Kinect] is also tested. We report on how well digital source-separation techniques work for segregating the teacher and student speech signals from one another based on these various microphones and placements. We also test the recordings using various automatic speech recognition engines for word recognition error rates under different levels of background noise. Preliminary results indicate one boundary microphone, the Crown PZM-30, to be superior for the classroom recordings. This is based on its performance at capturing near and distant student signals for ASR in noisy conditions, as measured by ASR error rates across different ASR engines.
A method for the analysis of prosodic-level temporal structure is introduced. The method is based on measured phase angles of an oscillator as that oscillator is made to synchronize with reference points in a signal. Reference points are the predicted peaks of acoustic change as determined by the output of a bank of tuned resonators. A framework for articulatory resynthesis is then described. Jaw movements of a robotic vocal tract are made to replicate the mean phase portrait of an utterance with reference to a production oscillator. These jaw movements are modeled to inform the dynamics of withinsyllable phonemic articulations. Index Terms: suprasegmental timing analysis, dynamic time warping, articulatory synthesis, speech robotics.
Two related questions posed to biology-inspired cognitive robotics are: (1) what form might a motor plan take in the brain, and (2) how might such a plan be converted into coordinated motor behavior? One approach to modeling the motor plan is to assume a stream of discrete instructions. This conceptualization has been attractive in cognitive science because it allows theorists to discuss motor planning (and perception) in terms of notation systems. However, many brain and robotics researchers have come to view the idea of the discrete motor plan to be wanting. SCA-Net provides a concrete alternative based on established brain circuitry. The approach allows for the coding of perception and production representations in a way that is useful to notation-oriented cognitive architectures.
A mechanical vocal tract with a pneumatic sound source, silicone tongue, and lip rounding mechanism is introduced. The tract is designed to make controlled transitions between static articulatory vowel configurations. The focus on transitions is important because many argue that it is the change between steady-state sounds that the nervous system is tuned for in extracting information from the speech signal. I draw on examples and experimental results as I review this steady state versus transition distinction. The notion of articulatory transitions as the motor control targets of robotic speech production is then discussed. Video demonstrations of the mechanical tract helps to illustrate how transitions as control targets may be implemented. In conclusion, I argue for why an analog vocal tract with its true aerodynamics (as opposed to synthesis using a cascade of digital filters) is called for in generating articulator-transition-based speech sounds.
Gaining a command of the timing flow of a language is essential to being able to speak and understand what is spoken. It has been claimed that a ‘proactive memory’ for when the transitions into and out of vowels will occur in a speech signal can help us to attend to and parse the information in a speech signal [1, 2]. This paper describes and evaluates a method for analysing temporal structure in Japanese speech based on vowel onset patterns. In this method, vowel onset locations are first detected in a speech signal, and then a spectral timing model is used to evaluate periodicity in the event sequences. The spectral model is built of a bank of resonators in where each resonates with a distinct natural period in response to the input signal [3]. Output from the model is the sum of responses from the resonators. When local periodic structure exists in the input signal, the resonators with a related natural frequency become more active. If the signal continues to have periodic structure, newly arriving events are better anticipated and the corresponding resonators are reinforced and remain active. However, if an event occurs at a time when active resonators are out of phase, those resonators will absorb the input rather than respond to it, and be correspondingly weakened. The spectral model boosts periodic events and filters non-periodic events.
A control framework for a robotic vocal tract is introduced. I incorporate a timing orchestration mechanism into a basic ART network and the network is situated within a speech perception-production loop. The network is first trained to auto- associate its motor output with its resulting sound input during a 'babble learning' phase. Further training involves alternating between babble learning and learning for a set of pre-synthesized rhythmic vowel patterns. Preliminary analysis indicates that the network is successful at distinguishing and reproducing the training patterns with their correct temporal structures.
The regularity of mora timing in Japanese has remained controversial over the years [Warner and Arai, Phonetica 58, 1–25 (2001)]. It is possible that the degree of regularity varies with speaking style. Four Japanese subjects spoke six-mora proper names with varying syllable structures where the second name had either three simple syllables (e.g., Tomiko) or two syllables with a long vowel (e.g., Tooko) or a long consonant (e.g., Tokko). These were read either in formal sentences, read in conversational style sentences, or used in spontaneous description of pictures of characters having these names. Moras were measured as the intervals between vowel onsets using an automatic vowel-onset detection algorithm. Although all styles suggested regular mora timing, the results show that the styles differ in the degree of temporal compensation for the moras constituted by long vowels and long consonants which are shorter than consonant-vowel syllables. The compensation was clearest in the most formal style of speech.
An analog speech synthesizer is constructed based on the sourcefilter model of the human vocal tract. The artificial tract was developed in part for studies on emotional and paralinguistic content in speech synthesis and for eventual studies on speech interaction. In this study, the glottal open quotient (OQ) of the waveforms generated as the source (weak to strong high frequency spectrum), the fundamental frequency (F0), and a jitter parameter were varied as the filter shape or tongue posture was randomly altered. Subjects were first asked to rate caricature facial expressions on affective attributes and then to rate goodness of fit between the faces and the sound stimuli. Association scores between attributes and sounds based on these ratings were calculated. Results indicate that in the context of isolated sound segments, OQ is the strongest determinant of perceived android arousal and F0 is the strongest determinant of perceived android valence. Further results point to dependencies between the three varied parameters. Findings are consistent with related research and help to establish a foundation for future work with the artificial tract. A concluding discussion remarks on how this study and the project itself speaks to other research on speech and to android science in general. The rich visual cues and the more ecologically realistic acoustic radiation characteristics of an analog synthesizer are strategic to a science where naturalistic mechanisms for social interaction are needed.
The conventional approach to speech production assumes that a linguistic control signal feeds down into an execution module where vocal articulators are coordinated. The linguistic signal takes the form of a stream of phonological units or discrete symbolic commands. This characterization reflects how a variety of control architectures in cognitive robotics are also based on symbolic commands. There are problems with symbolic motor control and in robotics there are alternatives to the assumption of symbols. This paper focuses on one such alternative. A minimal neural field model for speech motor planning and production is introduced. The model illustrates how some simple words may be represented for perception and production without coding the words in terms of phonological units. Concluding discussion considers how a scaled version of the model supports a construction grammar account of speech and language.