Cognitive load refers to the mental demand experienced while performing a cognitive task. A cognitive load measurement system can potentially be a useful tool for monitoring and enhancing human task performance. In the area of speech-based cognitive load classification, while there are various spectral and vocal tract-based features proposed for classification purposes, there is still a lack of studies that investigate how cognitive load affects the voice source, and whether glottal features are effective in cognitive load classification systems. This work introduces a set of databases that contains both speech and electroglottograph (EGG) data. Using these databases, we present results that provide arguably the first direct insight into how cognitive load affects the voice source. Additionally, we show that glottal-based features carry complementary information with respect to formant-based features, and that fusion between glottal and formant-based systems produces classification results that are comparable with (if not better than) existing baseline systems across three out of five evaluation databases. (C) 2015 Elsevier B.V. All rights reserved.
High cognitive load arises from complex time and safety-critical tasks, for example, mapping out flight paths, monitoring traffic, or even managing nuclear reactors, causing stress, errors, and lowered performance. Over the last five years, our research has focused on using the multimodal interaction paradigm to detect fluctuations in cognitive load in user behavior during system interaction. Cognitive load variations have been found to impact interactive behavior: by monitoring variations in specific modal input features executed in tasks of varying complexity, we gain an understanding of the communicative changes that occur when cognitive load is high. So far, we have identified specific changes in: speech, namely acoustic, prosodic, and linguistic changes; interactive gesture; and digital pen input, both interactive and freeform. As ground-truth measurements, galvanic skin response, subjective, and performance ratings have been used to verify task complexity. The data suggest that it is feasible to use features extracted from behavioral changes in multiple modal inputs as indices of cognitive load. The speech-based indicators of load, based on data collected from user studies in a variety of domains, have shown considerable promise. Scenarios include single-user and team-based tasks; think-aloud and interactive speech; and single-word, reading, and conversational speech, among others. Pen-based cognitive load indices have also been tested with some success, specifically with pen-gesture, handwriting, and freeform pen input, including diagraming. After examining some of the properties of these measurements, we present a multimodal fusion model, which is illustrated with quantitative examples from a case study. The feasibility of employing user input and behavior patterns as indices of cognitive load is supported by experimental evidence. Moreover, symptomatic cues of cognitive load derived from user behavior such as acoustic speech signals, transcribed text, digital pen trajectories of handwriting, and shapes pen, can be supported by well-established theoretical frameworks, including O'Donnell and Eggemeier's workload measurement [1986] Sweller's Cognitive Load Theory [Chandler and Sweller 1991], and Baddeley's model of modal working memory [1992] as well as McKinstry et al.'s [2008] and Rosenbaum's [2005] action dynamics work. The benefit of using this approach to determine the user's cognitive load in real time is that the data can be collected implicitly that is, during day-to-day use of intelligent interactive systems, thus overcomes problems of intrusiveness and increases applicability in real-world environments, while adapting information selection and presentation in a dynamic computer interface with reference to load.
Speech is a promising modality for the convenient measurement of cognitive load, and recent years have seen the development of several cognitive load classification systems. Many of these systems have utilised mel frequency cepstral coefficients (MFCC) and prosodic features like pitch and intensity to discriminate between different cognitive load levels. However, the accuracies obtained by these systems are still not high enough to allow for their use outside of laboratory environments. One reason for this might be the imperfect acoustic description of speech provided by MFCCs. Since these features do not characterise the distribution of the spectral energy within subbands, in this paper, we investigate the use of spectral centroid frequency (SCF) and spectral centroid amplitude (SCA) features, applying them to the problem of automatic cognitive load classification. The effect of varying the number of filters and the frequency scale used is also evaluated, in terms of the effectiveness of the resultant spectral centroid features in discriminating between cognitive loads. The results of classification experiments show that the spectral centroid features consistently and significantly outperform a baseline system employing MFCC, pitch, and intensity features. Experimental results reported in this paper indicate that the fusion of an SCF based system with an SCA based system results in a relative reduction in error rate of 39% and 29% for two different cognitive load databases.
Previous work in speech-based cognitive load classification has shown that the glottal source contains important information for cognitive load discrimination. However, the reliability of glottal flow features depends on the accuracy of the glottal flow estimation, which is a non-trivial process. In this paper, we propose the use of acoustic voice source features extracted directly from the speech spectrum (or cepstrum) for cognitive load classification. We also propose pre-and post-processing techniques to improve the estimation of the cepstral peak prominence (CPP). 3-class classification results on two databases showed CPP as a promising cognitive load classification feature that outperforms glottal flow features. Score-level fusion of the CPP-based classification system with a formant frequency-based system yielded a final improved accuracy of 62.7%, suggesting that CPP contains useful voice source information that complements the information captured by vocal tract features.
Cognitive load measurement systems measure the mental demand experienced by human while performing a cognitive task, which is useful in monitoring and enhancing task performance. Various speech-based systems have been proposed for cognitive load classification, but the effect of cognitive load on the speech production system is still not well understood. In this work, we study formant frequencies under different load conditions and utilize formant frequency-based features for automatic cognitive load classification. We find that the slope, dispersion, and duration of vowel formant trajectories exhibit changes under different load conditions; slope and duration are found to be useful features in vowel-based classification. Additionally, 2-class and 3-class utterance-based classification results, evaluated on two different databases, show that the performance of frame-based formant features was comparable, if not better than, baseline MFCC features.
Recent results seem to cast some doubt over the assumption that improvements in fused recognition accuracy for speaker recognition systems based on different acoustic features are due mainly to the different origins of the features (e.g. magnitude, phase, modulation information). In this study, we utilize clustering comparison measures to investigate acoustic and speaker modelling aspects of the speaker recognition task separately and demonstrate that front-end diversity can be achieved purely through different 'partitioning' of the acoustic space. Further, features that exhibit good 'stability' with respect to repeated clustering are shown to also give good EER performance in speaker recognition. This has implications for feature choice, fusion of systems employing different features, and for UBM data selection. A method for the latter problem is presented that gives up to an 11% relative reduction in EER using only 20-30% of the usual UBM training data set.
Cognitive load measurement is important when designing adaptive interfaces that optimize the performance of users working on high mental load tasks. Recent research on automatic speech-based measurement system indicates that cognitive load information is more prominent in the frequency region below 1 kHz. This study investigates the effects of cognitive load on glottal parameters (open quotient, normalized amplitude quotient and speed quotient), and proposes a system employing these parameters as features for cognitive load classification. Analysis of the glottal parameter distributions suggests that an increase in cognitive load can be related to a more creaky voice quality. Additionally, three-class classification results show that score-level fusion of systems based on the glottal features and baseline features (MFCCs, pitch, intensity and shifted delta cepstra) improves the baseline accuracy from 79% to 84%.
The ability to automatically classify different cognitive load levels can be very useful, especially in the field of human computer interaction, as human task performance is related to the cognitive load experienced. Although Mel-frequency cepstral coefficients (MFCCs) are commonly used in current speech-based cognitive load classification systems, they offer relatively little insight into how cognitive load affects the speech spectrum and physical speech production system-an area of research which remains poorly understood. Since formants are directly related to the physical characteristics of the vocal tract, we propose the novel use of formant frequencies, bandwidths and formant-based regression coefficients for cognitive load classification. Three-class classification results showed that formant frequencies performed comparably to MFCCs. Additionally, formant frequency-based regression coefficients outperformed MFCC-based regression coefficients by a relative improvement of about 11%. These results imply that formants contain important cognitive load information.
Speech has been recognized as an attractive method for the measurement of cognitive load. Previous approaches have used mel frequency cepstral coefficients (MFCCs) as discriminative features to classify cognitive load. The MFCCs contain information from both the voice source and the vocal tract, so that the individual contributions of each to cognitive load variation are unclear. This paper aims to extract speech features related to either the voice source or the vocal tract and use them to discriminate between cognitive load levels in order to identify the individual contribution of each for cognitive load measurement. Voice source-related features are then used to improve the performance of current cognitive load classification systems, using adapted Gaussian mixture models. Our experimental result shows that the use of voice source feature could yield around 12% reduction in relative error rate compared with the baseline system based on MFCCs, intensity, and pitch contour.
The current automatic cognitive load measurement system based on MFCC and prosodic features does not take into account phase based speech information. This paper aims to improve the performance of the baseline system by introducing phase based features into the system. The additional features proposed are group delay features, all-pole model based FM features and zero crossing count based FM features. Decrease in performance is observed when phase based features are considered individually or when concatenated with baseline features. However, significant performance improvement is observed when group delay features are fused with baseline features using linear combination score level fusion.
This paper focuses on tone classification for the Vietnamese speech. Traditionally, tone was classified or recognized by the fundamental frequency F0. However, our experimental results indicate that along with the fundamental frequency, Mel Frequency Cepstrum Coefficients and frequency modulation also carry a significant amount of tone information in the Vietnamese speech. Therefore, the proposed method takes into account these two types of features to improve the classification accuracy. The experimental results show that the proposed classification system provides an improvement of 7.5% in accuracy, compared to the conventional system based on F0 alone.
With applications of new information and communication technologies, computer-based information systems are becoming more and more sophisticated and complex. This is particularly true in large incident and emergency management systems. The increasing complexity creates significant challenges to the design of user interfaces (UIs). One of the fundamental goals of UI design is to provide users with intuitive and effective interaction channels to/from the computer system so that tasks are completed more efficiently and user's cognitive work load or stress is minimized. To achieve this goal, UI and information system designers should understand human cognitive process and its implications, and incorporate this knowledge into task design and interface design. In this chapter we present the design of CAMI, a cognition-adaptive multimodal interface, for a large metropolitan traffic incident and emergency management system. The novelty of our design resides in combining complementary concepts and tools from cognitive system engineering and from cognitive load theory. Also presented in this chapter is our work on several key components of CAMI such as real-time cognitive load analysis and multimodal interfaces.
Group delay is proposed as an effective means of representing spectral phase information as a feature in speaker recognition. Robustness of group delay features is difficult to achieve, since the spiky nature of the group delay masks the fine structure of the group delay. In this paper, two features based on group delay are proposed by reducing the effect of spikes with two different approaches. The first is log compression, to address the masking effects of the spikes, and the second is to use a sub-band based approach, where masking is restricted within certain bands containing the spikes. The purpose of this paper is to introduce different types of group delay feature extraction methods. The two features are evaluated on the cellular NIST 2001 database.
In this paper we propose a novel language identification system which utilizes fused phonotactic information. The phase spectrum of speech signals is used with the magnitude spectrum in order to obtain a more robust feature representation. Parallel Broad Phoneclass Recognition followed by Language Model (PBPRLM) is used in order to remove the bias of the likelihood scores introduced by the size inequality of phone inventories in traditional PPRLM systems. The likelihood scores from the MFCC-based and group-delay-based PPRLM and PBPRLM systems are fused together by using a Gaussian Mixture Model. Furthermore, a pre-classification based on Kohonen's map is used in order to maintain the system robustness while handling a large number of target languages. Using this proposed novel system we achieve an EER of 6.7% on the 2005 NIST LRE, and a LID recognition rate of 83.9% on a 22-language task.
Speech has recently been recognized as an attractive method for the measurement of cognitive load. Current speech-based cognitive load measurement systems utilize acoustic features derived from auditory-motivated frequency scales. This paper aims to investigate the distribution of speech information specific to cognitive load discrimination as a function of frequency. We found that this distribution is neither uniform nor very similar to the Mel auditory scale and based on our experiments, we propose a novel non-uniform filterbank for acoustic feature extraction to classify cognitive load. Experimental results showed that the use of the proposed filterbank provided a relative improvement of about 10%, compared with the classification accuracy of the traditional cognitive load classification system based on a Mel-scale filterbank.
This paper presents two novel contributions to automatic language identification. The first one is the use of the modified multi-layer Kohonen self-organizing feature map (MLKSFM) as a pre-classification for language identification (LID). Secondly, we discuss the novel application of empirical mode decomposition (EMD) to generate features for the LID pre-classification task. The use of instantaneous frequency (IF) and instantaneous amplitude (IA) of a speech signal as features for the pre-classifier is investigated. The experiment results on a 16-language speech database indicates that, the EMD by itself cannot perform well in the LID task, however it helps to improve the pre-classification rate when concatenated with other cepstral features. The overall LID performance is also increased when pre-classification is applied. We achieve LID rates of 85.2% and 62.3% for 45-sec and 10-sec test utterances, respectively.
A device is provided that begins an electrical starting mode in a parallel circuit configuration providing an improved speed of response and after an interval of time, current is switched to a series circuit configuration for operation during a run mode. The device is initially energized at several times the power normally required to activate a coil and switching means is thereafter utilized to reduce power to a normal requirement soon after the coil is energized. Specific hardware is provided in the form of schematic diagrams for defining the operation of the invention.
LifeLog can be used as a stand-alone consumer device to serve as a powerful automated multimedia diary and scrapbook. By using a search engine interface, the user can easily retrieval a specific thread of past transactions, or recall a few seconds ago or from many years earlier in as much detail as is desired, including imagery , audio, or video event. In this paper, we have studied and implemented three audio classification algorithms. The testing results show that KNN is the best classifier to classify the multimedia data.