In a multimodal interface, a user can use multiple modalities, such as speech, gesture, and eye gaze etc., to communicate with a system. As a critical component in a multimodal interface, multimodal input fusion explores the ways to effectively interpret the combined semantic interpretation of user's multimodal inputs. Although multimodal inputs may contain spare information, few multimodal input fusion approaches have tackled how to deal with spare information in multimodal inputs. This paper proposes a novel multimodal input fusion approach to flexibly skip spare information in multimodal inputs and derive semantic interpretation of them. The evaluation about the proposed approach confirms that the approach makes human-computer interaction more natural and smooth.
A multimodal system is a system equipped with a multimodal interface through which a user can interact with the system by using his/her natural communication modalities, such as speech, gesture, eye gaze, etc. To understand a user's intension, multimodal input fusion, a critical component of a multimodal interface, integrates a user's multimodal inputs and finds the combined semantic interpretation of them. As powerful, yet affordable input and output technologies becoming available, such as speech recognition and eye tracking, it becomes possible to attach recognition technologies to existing applications with a multimodal input fusion module; therefore, a practical multimodal system can be built. This paper documents our experience about building a practical multimodal system with our multimodal input fusion technology. The pilot study has been conducted over the multimodal system. By outlining observations from the pilot study, the implications on multimodal interface design are laid out.
A multimodal interface provides multiple modalities for input and output, such as speech, eye gaze and facial expression. With the recent progresses in multimodal interfaces, various approaches about multimodal input fusion and output generation have been proposed. However, less attention has been paid to how to integrate them together in a multimodal input and output system. This paper proposes an approach, termed as THE HINGE, in providing agent-based multimodal presentations in accordance with multimodal input fusion results. The analysis of experiment result shows the proposed approach enhances the flexibility of the system while maintains its stability.
Recent music retrieval systems can handle only the symbolic query on a symbolic database such as MIDI file. In such systems, it can perform only the polyphonic acoustic query and convertthe polyphonic acoustic files into symbolic ones. This process will cause the loss ofdata and affect the accuracy of retrieval. In this paper a new Content-Based Music Retrieval (CBMR) system based on vector quantization has been proposed and implemented. The proposed new system can allow the users to search the waveform data based on the query musicsamples. Without converting the audio wave file into a MIDI format, histograms are generated from the query data by vector quantization directly. By using the vector quantization, we can achieve the high accuracy in audio retrieval rate. In the proposed CBMR system, the use of 128 clusters in Kmeans-clustering quantization algorithm can achieve 87% retrieval accuracy …
As the demands and usage for audio classification and retrieval increase, the classification methods need to be improved to be more automatic and effective. The traditional text-based audio classification fails to recognize the underlying content of audio files. In this paper we propose a new hybrid audio classification algorithm based on SVM weight factor and Euclidean distance in order to improve the accuracy for audio classification. The proposed algorithm can extract the weight factor from Supporting Vector classification Model (SVM) and apply it to the Euclidean measurement. The experimental results show that it can improve theaudio classification accuracy by 28% at the maximum and 7% in the overall performance. By using the new proposed algorithm, some mis-classified audio data from a conventional Euclidean distance classifier can be classified.
The text-based classification dominates in the conventional audio classification systems, in which tedious manual work is used to notate the name, class, or sample rate. However, on most occasions, this method is not satisfying due to its opaque to real content. In order to retrieve the audio files effectively and efficiently, content-based audio classification becomes more and more necessary. In this paper, a new hybrid approach of audio classification algorithm is proposed to improve the performance of some misclassified audio data. The proposed method has been compared with the traditional Euclidean-based K-Nearest Neighbor classifier. As to improve the accuracy for specific problems, a weight factor based on Supporting Vector will apply to the Euclidean distance and K-NN rule to achieve better accuracy. The proposed new weighted Euclidean algorithm has been proved to be more sensitive to the classification criteria. The experimental results show that it can improve the audio classification accuracy by 28% at the maximum and 7% in the overall performance. By using the new proposed algorithm, some mis-classified audio data from a conventional Euclidean distance classifier can be classified.
In this paper, we have proposed and implemented a new music retrieval system based on content of music wave files. We have investigated different quantization methods by constructing them into the music data histograms as the feature vectors for the music files. There are three important aspects that will affect implementation of the system: audio feature extraction, quantization and distance computation. The proposed new system can allow the users to search the waveform data based on the query music samples. Without converting the audio wave file into a MIDI format, histograms are generated from the query data by vector quantization directly. By using the vector quantization, we can achieve the high accuracy in audio retrieval rate. In the proposed CBMR system, the use of 128 clusters in Kmeans-clustering quantization algorithm can achieve 87% retrieval accuracy and 90% high retrieval accuracy rate with the tree-based quantization.
Content-based retrieval is being widely studied. The colour-based retrieval is one of the most popular retrieval algorithms used in this area. It analyses the colour distribution of a query image and returns a list of videos that contain video sequence frames that match the colour features of the query. These systems use low-level image properties, and the retrieval accuracy rate will decrease when dealing with video sequence images. This paper has proposed and implemented a content-based video retrieval system using Haar wavelet, Daubechies D4 wavelet transform and five different types of clustering techniques. The experimental results show that the Haar wavelet with 3-Level transform can perform best with ahigh retrieval accuracy rate (89%).
Multimodal user interfaces (MMUI) allow users to control computers using speech and gesture, and have the potential to minimise users. experienced cognitive load, especially when performing complex tasks. In this paper, we describe our attempt to use a physiological measure, namely Galvanic Skin Response (GSR), to objectively evaluate users. stress and arousal levels while using unimodal and multimodal versions of the same interface. Preliminary results show that users. GSR readings significantly increase when task cognitive load level increases. Moreover, users. GSR readings are found to be lower when using a multimodal interface, instead of a unimodal interface. Cross-examination of GSR data with multimodal data annotation showed promising results in explaining the peaks in the GSR data, which are found to correlate with sub-task user events. This interesting result verifies that GSR can be used to serve as an objective indicator of user cognitive load level in real time, with a very fine granularity.
Multimedia Retrieval is one of the most important and fastest growing research areas in the field of multimedia technology. Large collections of scientific, artistic and commercial data comprising image, text, audio and video abound in the present information based society. There must be an effective and precise method of assisting users to search, browse and interact with these collections. This paper has proposed and implemented a content based video retrieval system using Haar wavelet, Daubechies D4 wavelet transform and five different types of clustering techniques. The experimental results show that the Haar wavelet with 3- Level transform can perform best with a high retrieval accuracy rate (89%).
The large and growing amount of digital data, and the development of the Internet highlight the need to develop sophisticated access methods that provide more than just simple text-based queries. Many programs have been developed with complex mathematics algorithms to allow the transformation of image or audio data in a way that enchances searching accuracy. However it becomes difficult when dealing with large sets of multimedia data. This paper proposes and demonstrates a Content Based Multimedia Retrieval System (CBMRS). The proposed CBMRS includes both video and audio retrieval systems. The Content Based Video Retrieval System (CBVRS) based on DCT and clustering algorithms. The audio retrieval system based on Mel-Frequency Cepstral Coefficients (MFCCs), the Dynamic Time Warping (DTW) algorithm and the Nearest Neighbor (NN) rule.
Multimodal User Interaction (MMUI) technology aims at building natural and intuitive interfaces allowing a user to interact with computer in a way similar to human-to-human communication, for example, through speech and gestures. As a critical component in MMUI, Multimodal Input Fusion explores ways to effectively interpret the combined semantic interpretation of user inputs through multiple modalities. This paper presents a novel approach to multi-sensory data fusion based on speech and manual deictic gesture inputs. The effectiveness of the technique has been validated through experiments, using a traffic incident management scenario where an operator interacts with a map on a large display at a distance and issues multimodal commands through speech and manual gestures. The description of the proposed approach and preliminary experiment results are presented.
User interface technology is an integral element of modern information and communications systems. This paper describes the outcomes of a field study that sought to design, develop, and evaluate a user interface suitable for an Incident Management System. Initially, the study comprised interviews and questionnaires used to examine how current systems are used, determine key issues facing current users, and identify the functions, features and behaviours a new system should exhibit. Subsequently, a mock-up was tested by potential end-users. All user feedback was incorporated into a set of design guidelines for the multimodal user interface of the new system. Preliminary analysis of the mock-up interface suggests that a 37% improvement in task time-to-completion could be achieved, together with 59% reduction in missed calls from a variety of contacts.
Operators of traffic control rooms are often required to quickly respond to critical incidents using a complex array of multiple keyboards, mice, very large screen monitors and other peripheral equipment. To support the aim of finding more natural interfaces for this challenging application, this paper presents PEMMI (Perceptually Effective Multimodal Interface), a transport management system control prototype taking video-based manual gesture and speech recognition as inputs. A specific theme within this research is determining the optimum strategy for gesture input in terms of both single-point input selection and suitable multimodal feedback for selection. It has been found that users tend to prefer larger selection areas for targets in gesture interfaces, and tend to select within 44% of this selection radius. The minimum effective size for targets when using 'device-free' gesture interfaces was found to be 80 pixels (on a 1280x1024 screen). This paper also shows that feedback on gesture input via large screens is enhanced by the use of both audio and visual cues to guide the user's multimodal input. Audio feedback in particular was found to improve user response time by an average of 20% over existing gesture selection strategies for multimodal tasks.