The latest advances in speech processing technology have allowed the development of automated reading tutors (ART) for improving children's literacy. An ART is a computer-assisted learning system based on oral reading fluency (ORF) instruction and automated speech recognition (ASR) technology. However, the design of an ART system is language-specific, and thus, requires developing a system specifically for the Filipino language. In a previous work, the authors have presented the development of the children's Filipino speech corpus (CFSC) for the purpose of designing an ART in Filipino. In this paper, the authors present the evaluation of the ART in Filipino which integrates a reference verification (RV)- and word duration analysis-based reading miscue detector (RMD), a user interface, and a feedback and instruction set. The authors also present the performance evaluation of the RMD in offline tests, and the effectiveness of the ART as shown by the results of the intervention program, a month-long pilot study that involved the use of the ART by a small group of students. Offline test results show that the RMD's performance (i.e., FA rate approximate to 3% and MDerr rate approximate to 5%) is at par with those from state-of-the-art RMDs reported in the literature. The results of the ART intervention experiment showed that the students, on the average, have improved in their words correct per minute (WCPM) rate by 4.66 times, in their ORF-16 scores by 6.0 times, and in their reading comprehension exam scores by 4.4 times, after using the ART.
A Philippine LID system has not been previously created because of the limited amount of recorded speech data. This research initiates the LID research using the Philippine Language Database (PLD) collected by the Digital Signal Processing Laboratory of the University of the Philippines Diliman (DSP-UPD). Mel Frequency Cepstral Coefficients (MFCC), Perceptual Linear Prediction (PLP), Shifted Delta Cepstra (SDC) and Linear Predictive Cepstral Coefficients (LPCC) features are extracted from the speech segments. Gaussian Mixture Model (GMM) using Expectation Maximization (EM) and Universal Background Model (UBM) approach is used to model the acoustic characteristics of the language. Maximum a Posteriori (MAP) probability is then used to determine the language of a speech utterance based on the language GMMs. PLP using a 16 Mixture GMM-EM has been found to produce the best performance among the four feature vectors in discriminating the languages.
In this paper, a new database suitable for HMM-based automatic Filipino speech recognition is described for the purpose of training a domain-independent, large-vocabulary continuous speech recognition system. Although it is known that high-performance speech recognition systems depend on a superior speech database used in the training stage, due to the lack of such an appropriate database, previous reports on Filipino speech recognition had to contend with serious data sparsity issues. In this paper we alleviate such sparsity through appropriate data analysis that makes the evaluation results more reliable. The best system is identified through its low word-error rate to a cross-validation set containing almost three hours of unknown speech data. Language-dependent problems are discussed, and their impact on accuracy was analyzed. The approach is currently data driven, however it serves as a competent baseline model for succeeding future developments.
It is widely known that database quality has a huge impact on speech recognition system performance, most especially when the expected domain is well represented. In this paper, we use this idea as leverage for a data-driven solution to the problem of code-switching in Filipino. Practical Filipino conversations often contain English and other loan words in varying frequencies, demanding better training of parameters and models for its speech recognition system. We alleviate the underrepresentation of loan words through the development of a new speech database for training, and applied appropriate data analysis to make reliable evaluation results. The best system was searched via lattice rescoring from a cross-validation set containing almost three hours of unknown speech data. The description and results of our experiments serve as a new and competent baseline model for succeeding future developments.
Voltage regulator modules (VRMs) have to maintain perfect current sharing between modules in both static and dynamic loading. In this study, sliding mode control was applied to both the voltage loop and the current share loop of the two-phase VRM with the aim of providing an effective current sharing scheme. Unlike other sliding mode control implementation that uses infinite switching frequency in its control method, fixed switching frequency was used. Derivation for the fixed frequency sliding mode control was done in the analog domain and transformed into the digital domain. A simulation model using PSIM was developed for both the analog and the digital control implementations. Results show that current sharing is achieved in both the analog and the digital control.
Philippine Indigenous Music is slowly disappearing partly because of insufficient accessibility to Philippine indigenous musical instruments. This project aims to give musicians access to the sounds of the Philippine instrument known as the kulintang through a synthesizer plug-in. This paper ventures into the implementation of a Virtual Studio Technology Instrument (VSTi) plug-in containing kulintang sounds analyzed using the Modal Distribution and synthesized using Sum of Sinusoid (SOS) Synthesis. The VSTi plug-in was implemented in C++ using the VST2.4 SDK (Software Development Kit). The plug-in is controlled by Musical Instrument Digital Interface (MIDI) events. The Filipino Kulintang MIDI (FKM1.0) was proposed as an initial MIDI mapping standard for Filipino kulintang instruments. A VST Dynamic Link Library (DLL) containing the kulintang synthesizer was successfully ported and controlled by a Digital Audio Workstation (DAW). Twenty three tuning presets with eight gongs each were included in the library.
We explore the use of manifold learning as a suitable representation for recognizing visual signs found in Filipino Sign Language. During the learning phase, a reference manifold is derived from a training set of visual signs using Isomap, a nonlinear manifold learning algorithm. Individual signs are then projected onto this reference manifold transforming them into trajectories which are compiled into a library. For recognition, the manifold trajectory of an unknown sign is computed and compared with the trajectories found in the library using either Dynamic Time Warping or Longest Common Subsequence Similarity Matching. Experiments using individual uninflected signs achieve recognition rates exceeding 80%.
The increasing market for mobile devices requires us to investigate the viability of writing computationally expensive digital signal processing applications for these devices. Modal Distribution was used to acquire a time-frequency representation to evaluate singing ability using a karaoke application. Pitch and tempo were both used as basis for grading. The trade-off between the time-frequency resolution and the execution time of the algorithm especially for devices with slower processors was studied to observe the optimal parameters for Modal Distribution. A resolution of 10 milliseconds for tempo and 31.25 Hz were thus obtained.
Recognizing the potential benefit that the current speech processing technology offers to improve children's literacy, researchers in the past few years have devoted their efforts in developing reading miscue detectors (RMDs) and automated reading tutors (ARTs). A primary challenge however in developing speech technologies for children may be the unavailability of a dedicated children's speech corpus that can be used for system design and test. In the past few years, children's speech corpora have been developed for languages such as English, Dutch, Chinese Mandarin, Italian, German and Swedish. But since Filipino has features and orthography that are distinct from other languages, the focus of this study is the development of a children's Filipino speech corpus (CFSC). In this paper, we present the CFSC design, reading text, data collection procedure and speech transcription method. We also performed initial analysis of the reading miscues and disfluencies found in the CFSC. The results of the miscue analysis suggest possible ways for modeling the reading miscues and possible methods for detecting them. Among these methods are acoustic model likelihood calculation and analysis of duration-based prosodic features. The CFSC presented in this study will be used for the development of an RMD and an ART for Filipino.
In this study, the feature set which brought about the highest classification accuracy for sorting Philippine Gong Music clips by indigenous group was sought. The features reflected Timbre, Loudness, Rhythm and Melody-and-Pitch. Two classifiers were used: Support Vector Machines and Neural Networks. Sequential Feature Selection was used to optimize the feature set. The highest accuracy achieved was 90.83% when the combination of SVM, 30s clips and the full Timbre feature set (64 features) was used. K-means clustering was also done to find similarities among the gong styles of the different groups.
This paper presents an attempt to experimentally find the general characteristics of interrogatives in Filipino speech in terms of pitch, duration and intensity patterns. The speech corpus developed in this study consisted of recordings of 275 interrogative sentences of different types uttered by 88 individuals. Analysis showed that out of the 5 different types of interrogatives considered, only the yes-no type exhibits a rising intonation towards the end of the sentence. Another important observation from the experiments was the existence of duration lengthening pattern on either the penultimate or the final syllable of the sentence. Although intensity patterns do not show a strongly generalizable result, a minor finding was that the interrogative particle ‘ba’ often receives the highest intensity throughout the construction.
This paper proposes the first Automated Essay Grader (AEG) for the Filipino language with results that are competitive with human checkers. In this study, the authors focused on content evaluation using Concept Indexing (CI) and Latent Semantic Indexing (LSI). The two NLP algorithms were compared based on their accuracy and speed. The effects of spell checking, stop word removal, stemming, sub-clustering and normalization weighting schemes were also examined.
Measurements show that the transmission delay in Next Generation Networking (NGN) typically exceeds the limit for acceptable voice quality. This paper presents the design and implementation of a parameter-based Voice Enhancement Device (VED) that operates on coded speech rather than on speech waveforms. This VED utilizes the information contained in speech parameters to achieve noise reduction and acoustic echo control. The approach needs no transcoding resulting in lower transmission delay, thus achieving acceptable voice quality in NGN. For objective test using ITU-T G.160, results show that the developed VED passed both the Acoustic Echo Control (AEC) and Noise Reduction (NR) tests. For subjective listening test, however, the NR section resulted in lower listening-effort Mean Opinion Score (MOS LE ). Nevertheless, increase in MOS LE was achieved when the NR and AEC sections were combined to work together in the listening test.
The Bantay-Wika (Language Watch) project was started in 1994 by the University of the Philippines (UP) - Sentro ng Wikang Filipino 1 (SWF) in order to track for long periods of time how the Philippine national language is being used and how it develops, particularly in the Philippine media. The first phase of this project, from 1994 to 2004, involved the manual collection and tallying of frequency counts for all the words in eleven major Philippine tablods. With increasing online presence of Philippine news organizations, the project was revived in March 2010, with UP-SWF partnering with UP - Digital Signal Processing (DSP) laboratory. The project objectives were also re-drafted to include the development of software that would automate the process of downloading of Filipino news articles. In this paper, we further detail the goals and the history and of the Bantay-Wika project, its accomplishments and plans for future work. The project ultimately endeavors to build a computational model for language development that can guide language policy makers in a multi-lingual country such as the Philippines, in drafting policies that can effectively promote the use and development of their national language.
This study incorporates computational and perceptual methods to classify Filipino speech rhythm. Speech rhythm may be described as a language‟s distinguishing durational sound pattern, resulting from the complexity of the language‟s syllable inventory. 1 Computational methods involve the correlation of rhythm-types to acoustic features such as the vocalic and consonantal intervals, one of which is the implementation of Multivariate Discriminant Analysis (MDA). Perceptual methods involve contrasting the rhythm of an unclassified language from prototype syllable-timed and stress-timed sentences. In order to isolate rhythm from speech, a data-stripping technique called flat sasasa resynthesis was implemented wherein the consonants are replaced with /s/ and vowels with /a/, producing a resynthesized alternating “sasasa” sounds at a constant pitch (F0). The rhythm discrimination and classification were closely examined for consistency between the data modeling and listening test results. The computational experiment was able to show that an MDA classifier trained to distinguish English and Japanese sentences tend to label Filipino sentences as Japanese 67% of the time, vis-a-vis the perceptual experiment showing that the listeners perceive Filipino to be more similar with Japanese, this study shows computational and perceptual validation that Filipino is syllable-timed, just like Japanese.
Emotion recognition can help improve humanmachine interfaces (HMI). To develop a stable emotion recognition system, there should be a basis of characteristics to compare with the speech sample. Since the system should be independent of the actual word or sentence uttered, only prosodic features were considered as parameters. The performance accuracy of each feature was evaluated using multivariate discriminant analysis. The emotions considered were anger, boredom, happiness or satisfaction, and neutral. In a previous study, it was determined that among the features tested, the median derivative of pitch had the highest percentage accuracy of 69.62% [1]. Neutral was the most successfully recognized emotion while the non-neutral emotions had very low accuracy rates. This pattern can be related to the distribution of emotion in the data used since most of the samples were classified as neutral, with the non-neutral emotions having a combined distribution of less than 10%. For this study, acted call center speech data was collected to remove the biases introduced by having unequal emotion distribution. To form the data, call center scripts were written. Each actor portrayed the scripts into the four emotions. The same set of prosodic features was extracted from the acted call center speech data and multivariate discriminant analysis was performed. Features with high performance accuracy constitute the feature set for the emotion recognition system. In this study, all four emotions were recognized with better accuracy.
In this paper, a prosody-based emotion monitoring system for call centers is proposed. It aims to simplify the tracking and management of emotions extracted from call center agent-client conversations. The system is composed of four modules: Emotion Detection, Emotion Analysis and Report Generation, Database Manager, and User Interface. The Emotion Detection module uses Digital Signal Processing algorithms to extract the low-level features necessary for reliable emotion detection. A Support Vector Machines (SVM) classifier then detects the emotions contained in an utterance. The four emotions detected by the system are happiness, anger, boredom and neutral. The Emotion Analysis and Report Generation module performs computational analysis of the detected emotions during calls. This module also polishes the output of Emotion Detection module to provide a more presentable output of sequence of emotions of the agent and the caller. The Database Manager is responsible for the management of the database wherein it handles the creation, deletion and update of data. The Interface module serves as the view and user interface for the whole system. The system is comprised of a stand alone application and a web application. The stand alone application was developed using Eclipse Rich Client Platform to maintain the modularity and flexibility of the system. It displays the detected previous and current emotions of both the caller and the agent. On the other hand, the web application was constructed using Google Web toolkit to maintain its modularity and abstraction. It provides reports and analysis of the emotions expressed by the agents during conversations. Using the Model View Controller (MVC) approach, the emotion monitoring system is scalable, reusable and modular.
We present a new approach to essay content analysis using the dimensionality reduction algorithm called Concept Indexing (CI). Experiments were conducted to compare the performance of CI K-means and CI Fuzzy C-means with Latent Semantic Indexing (LSI). Both versions of CI outperform LSI in Exact Agreement Accuracy and Pearson's Product-Moment Correlation Coefficient measures on sample essays taken from high school English classes.