SLVideo is a video moment retrieval system for Sign Language videos that incorporates facial expressions, addressing this gap in existing technology. The system extracts embedding representations for the hand and face signs from video frames to capture the signs in their entirety, enabling users to search for a specific sign language video segment with text queries. A collection of eight hours of annotated Portuguese Sign Language videos is used as the dataset, and a CLIP model is used to generate the embeddings. The initial results are promising in a zero-shot setting. In addition, SLVideo incorporates a thesaurus that enables users to search for similar signs to those retrieved, using the video segment embeddings, and also supports the edition and creation of video sign language annotations. Project web page: https://novasearch.github.io/SLVideo/
Acoustic monitoring of cetaceans is crucial for studying and conserving these animals and their environment. With the rising interest in deciphering dolphin and whale communication, and the promise shown by machine learning solutions in this field, the demand for gathering and processing large vocalization datasets is only increasing. In this paper, we propose an entropy-based spectrogram filtering method that removes noise naturally present in recordings of cetacean vocalizations. This method enhances the visual clarity of cetacean spectrograms, aiding biologists, and improves vocalization detection when using image-based convolutional neural networks. Its implementation focuses on efficiency, processing 256 times faster than similar filters. The resulting spectrograms achieve on average a 98.73% reduction in required storage space and allow for a segmentation technique that reduces labelling and classification times in machine learning solutions.
Identifying small dolphin species based on their vocalizations remains a challenging task due to their similar vocal signatures and frequency modulation patterns, particularly when the available data sets are relatively limited. To address this issue, a new feature set has been introduced that focuses on capturing both the predominant frequency range of the vocalizations and other higher level details in the spectral contour, which are valuable for distinguishing between small dolphin species. These features are computed from two distinct representations of the vocalizations: the short time Fourier transform and Mel frequency cepstral coefficients. By utilizing these features with two popular classifiers (K-Nearest Neighbors and Support Vector Machines), a model accuracy of 95.47% has been achieved, representing an improvement over previous studies.
Dementia is an uncurable neurodegenerative disease that leads to a gradual loss of cognitive capacities and negatively affects emotional state, quality of life, and ability to autonomously perform activities of daily living. Although pharmaceutical approaches can mitigate symptoms in people with dementia, their effect is still limited. Complementary approaches, such as music and reminiscence-related activities, have been proposed for stimulation purposes. Here we present the results of a pilot 14-session longitudinal study with an interactive platform called Musiquence, which allows the incorporation of music and reminiscence elements in cognitive stimulation activities, with 8 participants with dementia. In general, the results of the intervention show improvements in all the assessed domains: cognition, anxious and depressive symptomatology, functionality, and quality of life. Preliminary results appear to support the platform’s feasibility while providing positive outcomes of clinical efficacy.
In this paper, we introduce Immerscape, an interactive tool for composing immersive soundscapes. Immerscape can be used to support the study and creation of historical soundscapes, enhancing the experience for visitors at culture and heritage sites. The tool allows non-technical users to compose immersive auditory scenes, by simply combining different sound recordings that simulate moving sound sources, as well as sound sources positioned in different locations in the scene. A preliminary evaluation of the tool’s usability and acceptance has been conducted, involving domain users and other volunteers.
Children with speech sound disorders should attend speech and language therapy and should practice the speech exercises regularly to surpass their speech difficulties. Since doing the speech exercises often may be tedious, there is the need to motivate children to practice them. During the COVID-19 pandemic, speech and language pathologists had the need to adapt their procedures to others with less physical contact. Here, we propose two serious games to motivate children with sigmatism on doing the speech exercises, which can be used at home and during face-to-face and online speech therapy sessions. The games use automatic speech recognition to classify speech productions. Visual and auditory feedback are used to help children understand their performance, and a hint system is used to help them perform the exercises correctly. A dynamic difficulty adjustment system is used to change the level of difficulty according to the child's speech performance in previous trials.
People with Parkinson's disease (PD) can have dysarthria, a voice disorder that affects speech intelligibility. To fight this disorder people may resort to speech and language therapy. Unfortunately, weekly speech therapy sessions may not be enough, because to achieve and maintain good voice quality, intensive training is required. Additionally, the COVID-19 pandemic brought attention to the need for alternative speech therapy treatments that complement face-to-face appointments. Here, we propose a serious therapy game to improve voice loudness that can be used for intensive therapy or when face-to-face appointments are not possible. The game integrates three voice exercises used in speech therapy sessions for people with PD and aims to provide motivation for patients to perform the exercises on a daily basis. This application evaluates the vocal intensity, vocal frequency and maximum phonation time, offering real-time visual feedback. It also allows pathologists to customize the exercises difficulty to the needs of each patient.
Here we propose a platform for speech and language therapy that includes speech therapy games for children, and that allows the speech and language pathologist to customize the games’ settings according to the child’s speech difficulties. This allows a better adaptation of the games to each child’s needs. The platform can be used during face-to-face and online speech therapy session. In addition, it can also be used at home for intensive training. The platform provides post-training information about the child’s speech performance during the games’ trials. This feature is quite relevant when the platform is used for home intensive training, as it allows the speech and language pathologist understand the child’s performance and be able to make better choices when planning future therapy sessions and future home training plans. The platform was accessed by four speech and language pathologists. The validation confirmed that the platform is suitable for home intensive training and provides relevant settings to adjust the games to the needs of different children. ACM Reference Format: Sofia Martins, Sofia Cavaco. 2021. The BioVisualSpeech speech therapy games platform. In Proceedings of ACM Conference (Conference’21). ACM, New York, NY, USA, 2 pages. https://doi.org/10.1145/nnnnnnn.nnnnnnn
Speech therapy games present a relevant application of business intelligence to real-world problems. However many such models are only studied in a research environment and lack the discussion on the practical issues related to their deployment. In this article, we depict the main aspects that are critical to the deployment of a real-time sound recognition neural model. We have previously presented a classifier of a serious game for mobile platforms that allows children to practice their isolated sibilants exercises at home to correct sibilant distortions, which was further motivated by the Covid-19 pandemic present at the time this article is posted. Since the current classifier reached an accuracy of over 95%, we conducted a study on the ongoing issues for deploying the game. Such issues include pruning and optimization of the current classifier to ensure near real-time classifications and silence detection to prevent sending silence segment requests to the classifier. To analyze if the classification is done in a tolerable amount of time, several requests were done to the server with pre-defined time intervals and the interval of time between the request and response was recorded. Deploying a program presents new obstacles, from choosing host providers to ensuring everything runs smoothly and on time. This paper proposes a guide to deploying an application containing a neural network classifier to free- and controlled-cost cloud servers to motivate further deployment research.
We developed a new cueing system that capitalises on the spared musical processing abilities of people with dementia. We studied the cueing efficacy by testing different musical distortions (pitch, rhythm, and pitch-rhythm) while using different augmented reality-based interaction modalities in multiple cognitive stimulation tasks. A total of 18 volunteers participated in this study (8 with non-organic dementia) caused by substance abuse, such as alcoholism (control group) and 10 with organic dementia such as Alzheimer's disease (experimental group). We evaluated their performance using three augmented reality tasks: (1) a Knowledge Quiz activity, which is a general quiz knowledge game, (2) a Search Objects activity, which is a game that involves finding hidden images using a virtual magnifying glass and (3) an Association activity, which consists of a categorisation task. A within-subject experimental design has been used so that all participants could be exposed to distortions and activities. Results show that (1) participants, in general, were faster in completing the activities using rhythm and pitch distortions than non-distortion condition; (2) the experimental group reacted to the sound throughout the activities and, therefore, mitigating erroneous decision making; (3) despite the experimental group reacted to the distortions, the control group had better performance in the search and association activities and (4) music distortions compensated dementia related deficits, such as age, schooling and cognitive state. The results suggest that this novel cueing approach can improve people's performance with organic dementia and, more specifically, that such people can effectively process and exploit musical distortions.
Many children with speech sound disorders cannot pronounce the sibilant consonants correctly. We have developed a serious game, which is controlled by the children's voices in real time, with the purpose of helping children on practicing the production of European Portuguese (EP) sibilant consonants. For this, the game uses a sibilant consonant classifier. Since the game does not require any type of adult supervision, children can practice producing these sounds more often, which may lead to faster improvements of their speech. Recently, the use of deep neural networks has given considerable improvements in the classification of a variety of use cases, from image classification to speech and language processing. Here, we propose to use deep convolutional neural networks to classify sibilant phonemes of EP in our serious game for speech and language therapy. We compared the performance of several different artificial neural networks that used Mel frequency cepstral coefficients or log Mel filterbanks. Our best deep learning model achieves classification scores of 95.48% using a 2D convolutional model with log Mel filterbanks as input features. Such results are then further improved for specific classes with simple binary classifiers.
Background Serious games (SGs) are used as complementary approaches to stimulate patients with dementia. However, many of the SGs use out-of-the-shelf technologies that may not always be suitable for such populations, as they can lead to negative behaviors, such as anxiety, fatigue, and even cybersickness. Objective This study aims to evaluate how patients with dementia interact and accept 5 out-of-the-shelf technologies while completing 10 virtual reality tasks. Methods A total of 12 participants diagnosed with dementia (mean age 75.08 [SD 8.07] years, mean Mini-Mental State Examination score 17.33 [SD 5.79], and mean schooling 5.55 [SD 3.30]) at a health care center in Portugal were invited to participate in this study. A within-subject experimental design was used to allow all participants to interact with all technologies, such as HTC VIVE, head-mounted display (HMD), tablet, mouse, augmented reality (AR), leap motion (LM), and a combination of HMD with LM. Participants’ performance was quantified through behavioral and verbal responses, which were captured through video recordings and written notes. Results The findings of this study revealed that the user experience using technology was dependent on the patient profile; the patients had a better user experience when they use technologies with direct interaction configuration as opposed to indirect interaction configuration in terms of assistance required (P=.01) and comprehension (P=.01); the participants did not trigger any emotional responses when using any of the technologies; the participants’ performance was task-dependent; the most cost-effective technology was the mouse, whereas the least cost-effective was AR; and all the technologies, except for one (HMD with LM), were not exposed to external hazards. Conclusions Most participants were able to perform tasks using out-of-the-shelf technologies. However, there is no perfect technology, as they are not explicitly designed to address the needs and skills of people with dementia. Here, we propose a set of guidelines that aim to help health professionals and engineers maximize user experience when using such technologies for the population with dementia.
In order to develop computer tools for speech therapy that reliably classify speech productions, there is a need for speech production corpora that characterize the target population in terms of age, gender, and native language. Apart from including correct speech productions, in order to characterize the target population, the corpora should also include samples from people with speech sound disorders. In addition, the annotation of the data should include information on the correctness of the speech productions. Following these criteria, we collected a corpus that can be used to develop computer tools for speech and language therapy of Portuguese children with sigmatism. The proposed corpus contains European Portuguese children’s word productions in which the words have sibilant consonants. The corpus has productions from 356 children from 5 to 9 years of age. Some important characteristics of this corpus, that are relevant to speech and language therapy and computer science research, are that (1) the corpus includes data from children with speech sound disorders; and (2) the productions were annotated according to the criteria of speech and language pathologists, and have information about the speech production errors. These are relevant features for the development and assessment of speech processing tools for speech therapy of Portuguese children. In addition, as an illustration on how to use the corpus, we present three speech therapy games that use a convolutional neural network sibilants classifier trained with data from this corpus and a word recognition module trained on additional children data and calibrated and evaluated with the collected corpus.
Children with fricative distortion errors have to learn how to correctly use the vocal folds, and which place of articulation to use in order to correctly produce the different fricatives.Here we propose a virtual tutor for fricatives distortion correction.This is a virtual tutor for speech and language therapy that helps children understand their fricative production errors and how to correctly use their speech organs.The virtual tutor uses log Mel filter banks and deep learning techniques with spectral-temporal convolutions of the data to classify the fricatives in children's speech by place of articulation and voicing.It achieves an accuracy of 90.40% for place of articulation and 90.93% for voicing with children's speech.Furthermore, this paper discusses a multidimensional advanced data analysis of the first layer convolutional kernel filters that validates the usefulness of performing the convolution on the log Mel filter bank.
The development of reliable speech therapy computer tools that automatically classify speech productions depends on the quality of the speech data set used to train the classification algorithms. The data set should characterize the population in terms of age, gender and native language, but it should also have other important properties that characterize the population that is going to use the tool. Thus, apart from including samples from correct speech productions, it should also have samples from people with speech disorders. Also, the annotation of the data should include information on whether the phonemes are correctly or wrongly pronounced. Here, we present a corpus of European Portuguese children's speech data that we are using in the development of speech classifiers for speech therapy tools for Portuguese children. The corpus includes data from children with speech disorders and in which the labelling includes information about the speech production errors. This corpus, which has data from 356 children from 5 to 9 years of age, focuses on the European Portuguese sibilant consonants and can be used to train speech recognition models for tools to assist the detection and therapy of sigmatism.
Dementia is a neurodegenerative disease that leads to impairment of cognitive and emotional faculties and makes patients dependent on performing activities of daily living. Due to limited effects of pharmaceutical approaches, the seek of alternatives has been growing. Music and reminiscence related activities appear to be promising approaches to stimulate people with dementia (PwD). The usage of serious games (SG) has also been proposed as a valid manner to stimulate PwD. However, most SG are not designed and adapted to the profile of individuals who are diagnosed with dementia. Thus, here we will demonstrate a framework that allows customization of activities in terms of content and technology while capitalizing the benefits of music and reminiscence related approaches within gamified activities.
The possibility of using serious games to stimulate people with dementia (PwD) has gained much attention in recent years. However, most of such games are not adapted to individual needs of such population in terms of the design of technology and its content. Thus, the desired therapeutic outcomes may not be achieved. Alternatively, more traditional approaches, such as the usage of music and reminiscence, have been shown to be able to lead to positive outcomes. Here, we propose a framework for serious games that allows healthcare professionals to customize music and reminiscence-based activities to stimulate PwD. It runs on an augmented reality setup, but also on PC, interactive table and tablet. Results from a usability study show that participants (1) were efficient in using the framework, (2) therapists are very interested in using it for stimulation purposes in PwD and (3) the usage of the framework was adequate in terms of effort and workload for PwD. Future deployments will be discussed in this article regarding the usage of the framework.
Many children suffering from speech sound disorders cannot pronounce the sibilant consonants correctly. We have developed a serious game that is controlled by the children's voices in real time and that allows children to practice the European Portuguese sibilant consonants. For this, the game uses a sibilant consonant classifier. Since the game does not require any type of adult supervision, children can practice the production of these sounds more often, which may lead to faster improvements of their speech. Recently, the use of deep neural networks has given considerable improvements in classification for a variety of use cases, from image classification to speech and language processing. Here we propose to use deep convolutional neural networks to classify sibilant phonemes of European Portuguese in our serious game for speech and language therapy. We compared the performance of several different artificial neural networks that used Mel frequency cepstral coefficients or log Mel filterbanks. Our best deep learning model achieves classification scores of 95.48% using a 2D convolutional model with log Mel filterbanks as input features.
Studies on childhood dysphonia have revealed considerable rates for voice disorders in 4 - 12 year-old children. The sustained vowel exercise is widely used as a technique in the vocal (re)education process. However this exercise can become tedious after a short practice. Here, we propose a novel dynamic difficulty adjustment model to be used in a serious game with the sustained vowel exercise to motivate children on practicing this exercise often. The model automatically adapts the difficulty of the challenges in response to the child's performance. The model is not exclusive to this game and can be used in other games for dysphonia treatment. In order to measure the child's performance, the model uses parameters that are relevant to the therapy treatment. The proposed model is based on the flow model in order to balance the difficulty of the challenges with the child's skills.
Problems in vocal quality are common in 4 to 12-year-old children, which may affect their health as well as their social interactions and development process. The sustained vowel exercise is widely used by speech and language pathologists for the child's voice recovery and vocal re-education. Nonetheless, despite being an important voice exercise, it can be a monotonous and tedious activity for children. Here, we propose a computer therapy game that uses the sustained vowel exercise to motivate children on doing this exercise often. In addition, the game gives visual feedback on the child's performance, which helps the child understand how to improve the voice production. The game uses a vowel classification model learned with a support vector machine and Mel frequency cepstral coefficients. A user test with 14 children showed that when using the game, children achieve longer phonation times than without the game. Also, it shows that the visual feedback helps and motivates children on improving their sustained vowel productions.