BACKGROUND:Although speech and language therapy (SLT) is central to post-stroke aphasia rehabilitation, global SLT provision often falls short of recommended dosages. In response, interest has grown in technology-based interventions, including therapy software, virtual reality (VR) and artificial intelligence (AI) tools. However, current evidence for the effectiveness of technology is fragmented, and no recent review offers a comprehensive synthesis across these three modalities. AIM:This review examined the range of technologies used in aphasia assessment and therapy and summarised their effectiveness across different intervention targets. The two research questions were: (1) What types of technology have been investigated for assessing and treating people with aphasia (PWA)? (2) How effective is the use of technology in the assessment and treatment of PWA? METHODS:A systematic search of four databases (PubMed, PsycINFO, Web of Science and Scopus) covering the period 2013 to May 2026 identified 67 included studies, of which 14 were randomised controlled trials. Studies reporting quantitative outcomes, were peer-reviewed, and focused on technology-based intervention for PWA were eligible. Quality was appraised using the NICE checklist. The GRADE framework was applied to evaluate certainty of evidence for each intervention target. Findings were then synthesised narratively due to heterogeneity across study designs, and outcome measures. RESULTS:Three technology types were identified: computerised speech and language therapy (CSLT) (38 studies), VR (17 studies) and AI (13 studies). AI was used predominantly for aphasia assessment and classification. The strongest and most consistent evidence related to word-finding, where high certainty of evidence was supported by multiple RCTs delivering therapy at or above the recommended 20-h threshold. For language production and comprehension, functional communication, and reading, outcomes were more variable, reflecting moderate certainty of evidence, and inconsistent dose adherence. Writing interventions received a low certainty rating, reflecting small samples, limited blinding and task-specific rather than generalised gains. Across domains, higher-dose studies were consistently associated with better outcomes, which may suggest that technology functions primarily as a tool to enable high-intensity practice rather than as an independently effective treatment ingredient. CONCLUSION:CSLT, VR and AI tools show promise as adjuncts to face-to-face SLT for aphasia assessment and rehabilitation. Word-finding interventions delivered at recommended doses have the strongest evidence base. Some studies did not use technology to support the recommended therapy dose. For other intervention targets, larger, higher-dose trials are needed. Future research should also examine whether integrating different technology types could offer additional clinical benefit. WHAT THIS PAPER ADDS:What is already known about the subject Previous systematic reviews have demonstrated the emerging role of technology in aphasia rehabilitation, with earlier work focusing primarily on computer-based therapy or AI technologies. However, these reviews were either narrow in scope (targeted specific technology type), or outdated. What this study adds to the existing knowledge This review provides an updated, cross-technology synthesis encompassing AI, virtual reality, and computerised speech-and-language therapy. It outlines how these tools were applied within the studies in the literature. The review also identifies persistent limitations in therapy dosage across studies, underscoring the need for future higher-dose trials to confirm the certainty of evidence across different intervention targets. What are the clinical implications of this study? The growing evidence for computerised speech and language therapy, virtual reality, and artificial intelligence tools continues to support their role as adjuncts to face to face SLT, particularly for language assessment and targeted word-finding interventions. These technologies may extend therapy provision beyond clinical hours, enable therapeutic doses of practice to be achieved, and improve consistency in assessment procedures.
Foley artistry is an essential part of the audio post-production process for film, television, games, and animation. By extension, it is as crucial in emergent media such as virtual, mixed, and augmented reality. Footsteps are a core activity that a Foley artist must undertake and convey information about the characters and environment presented on-screen. This study sought to identify if characteristics of age, gender, weight, health, and confidence could be conveyed, using sounds created by a professional Foley artist, in three different 3D humanoid models, following a single walk cycle. An experiment was conducted with human participants (n=100) and found that Foley manipulations could convey all the intended characteristics with varying degrees of contextual success. It was shown that the abstract 3D models were capable of communicating characteristics of age, gender, and weight. A discussion of the literature and inspection of related audio features with the Foley clips suggest signal parameters of frequency, envelope, and novelty may be a subset of markers of those perceived characteristics. The findings are relevant to researchers and practitioners in linear and interactive media and demonstrate mechanisms by which Foley can contribute useful information and concepts about on-screen characters.
Background The motor speech disorder, dysarthria, is common in cerebral palsy. The Speech Systems Approach therapy programme, which focuses on controlling breath supply and speech rate, has increased children’s intelligibility. Objective To ascertain if increased intelligibility is due to better differentiation of the articulation of individual consonants in words spoken in isolation and in connected speech. Design Secondary analysis. Setting University. Participants Forty-two children with cerebral palsy and dysarthria aged 5–18 years, Gross Motor Function Classification System I–V. Intervention The Speech Systems Approach is a motor learning therapy delivered to individuals by a speech and language therapist in 40-minute sessions, three times per week for 6 weeks. Intervention focuses on production of a strong, clear voice and speaking at a steady rate. Practice changes from single words to increasingly longer utterances in tasks with increasing cognitive load. Main outcome measures Unfamiliar listeners’ identification of singleton consonants (e.g. nap) and clusters of consonants (e.g. stair, end) at the start and end of words when hearing single words in forced choice tasks and connected speech in free transcription tasks. Acoustic measures of sound intensity and duration. Data sources Data collected at 1-week pre- and 1-week post-therapy from three studies: two interrupted time series design, one feasibility randomised controlled trial. Results Word initial and word final singleton consonants and consonant clusters were better identified post-therapy. The extent of improvement differed across word initial and word final singleton consonant subtypes. Improvement was greater for single words than connected speech. Change in sound identification varied across children, particularly in connected speech. Sound intensity and duration increases also were inconsistent. Limitations The small sample size did not allow for analysis of cerebral palsy type. Acoustic data were not available for all children, limiting the strength of conclusions that can be drawn. The different but phonetically balanced word lists, used in the original research, created variability in single words spoken across recordings analysed. Low frequencies of plosives, fricatives and affricates necessitated their combination for analysis preventing investigation of the effect of specific consonants. Connected speech was spontaneous, again creating variability within the data analysed. The estimated effects of therapy may therefore be partially explained by differences in the spoken language elicited. Conclusions The Speech Systems Approach helped children generate greater breath supply and a steady rate, leading to increased intensity and duration of consonant sounds in single words, thereby aiding their identification by listeners. Transfer of the motor behaviour to connected speech was inconsistent. Future work Refining the Speech Systems Approach to focus on connected speech early in the intervention. Personalisation of cues according to perceptual and acoustic speech measures. Creation of a battery of measures that can be repeated across children and multiple recordings. Study registration This trial is registered as Research Registry 6117. Funding This project was funded by the National Institute for Health and Care Research (NIHR) Efficacy and Mechanism Evaluation programme (NIHR130967) and will be published in full in Efficacy and Mechanism Evaluation; Vol. 10, No. 4. See the NIHR Journals Library website for further project information.
The Dysarthric Expressed Emotional Database (DEED) is a novel, parallel multimodal (audio-visual) database of dysarthric and typical emotional speech in British English which is a first of its kind. It is an induced (elicited) emotional database that includes speech recorded in the six basic emotions: "happiness", "sadness", "anger", "surprise", "fear", and "disgust". A "neutral" state has also been recorded as a baseline condition. The dysarthric speech part includes recordings from 4 speakers: one female speaker with dysarthria due to cerebral palsy and 3 speakers with dysarthria due to Parkinson's disease (2 female and 1 male). The typical speech part includes recordings from 21 typical speakers (9 female and 12 male). This paper describes the collection of the database, covering its design, development, technical information related to the data capture, and description of the data files and presents the validation methodology. The database was validated subjectively (human performance) and objectively (automatic recognition). The achieved results demonstrated that this database will be a valuable resource for understanding emotion communication by people with dysarthria and useful in the research field of dysarthric emotion classification. The database is freely available for research purposes under a Creative Commons licence at: https://sites.google.com/sheffield.ac.uk/deed.
Few studies have investigated how individuals with partially intelligible speech choose to communicate, including how, when, and why they might use a speech-generating device (SGD). This study aimed to add to the literature by exploring how this group of individuals use different communication strategies. Qualitative interviews were carried out with 10 participants with partially intelligible speech with the aim of investigating participants' perceptions of modes of communication and communication strategies. Transcripts were analyzed using Framework Analysis to investigate the role of SGDs alongside other communication strategies. Factors that influence why, when, and how a person chooses to communicate were identified and these were interpreted as an explanatory model of communication with partially intelligible speech. Participants described how they made the decision whether to attempt to communicate at all and then which communication method to use. Decision-making was influenced by the importance of the message, how much time is available, past experience, and the communication partner. Each communication attempt adds to an individuals' experience of communicating and influences subsequent decisions. This study suggests that individuals with partially intelligible speech are at risk of reduced communication environments and networks and that current SGDs may not be designed in a way that recognizes their particular needs.
Foley artistry is an essential part of the audio post-production process for film, television, games, and animation. By extension, it is as crucial in emergent media such as virtual, mixed, and augmented reality. Footsteps are a core activity that a Foley artist must undertake and convey information about the characters and environment presented on-screen. This study sought to identify if characteristics of age, gender, weight, health, and confidence could be conveyed, using sounds created by a professional Foley artist, in three different 3D humanoid models, following a single walk cycle. An experiment was conducted with human participants (n=100) and found that Foley manipulations could convey all the intended characteristics with varying degrees of contextual success. It was shown that the abstract models were capable of communicating characteristics of age, gender, and weight. The findings are relevant to researchers and practitioners in linear and interactive media and demonstrate mechanisms by which Foley can contribute useful information and concepts about on-screen characters.
The esports industry has seen enormous growth in popularity. With increased viewership and revenue, further investment has been made to improve professional players’ competitive strength. The modern esports team is a hierarchical business fuelled by investors and sponsorship. This paper is focused on the professional competitions in League of Legends esports. In existing real-world sports such as football or baseball, there is great attention paid to statistic driven analysis of the competition, and these stats are used to quantify player and team performance. These statistics hold significant value for competitive improvement, the gambling industry, and market influence within the esports industry. This paper presents an analysis of data and metrics gathered from professional games during 2020 in several League of Legends international competitions. The objective was to build a predictive model through the combination of existing data analysis and machine learning that can rate team and player performance. The best performing model was able to correctly predict 67% of 306 games. Results indicate that while it is possible to predict the outcome of a competitive League of Legends game, to do so with a higher degree of accuracy would require substantially more data and contextual information.
The teaching and study of User Experience (UX) is relatively new compared to more established fields. Additionally, the interdisciplinary nature of UX may make a consensus on teaching pedagogies and methods difficult to reach. To add to the complexity of teaching UX, is the emergence of the degree apprenticeship programme where apprentices study at a higher education university for 20% of their time while 80% is spent with their employer. This paper describes the process of launching a programme to respond to this unique and emerging context during a global pandemic. It focuses on reflections from lecturers on the challenges and reception of the programme and discusses the initial contributions this work can bring to conceptions of UX pedagogies, especially those delivered in an online delivery environment.
This chapter discusses practice-led and interdisciplinary methods in sound design research. Sound design can be undertaken as a creative activity in combination with theoretical investigation, so that theory informs practice; while conversely, the generation of new practical approaches provides new theoretical insights. The authors discuss how this may provide a suitable methodology for research in sound design, where sound artifacts themselves may transmit and advance knowledge in the field, but may also be complemented and disseminated through means of papers and publications. Furthermore, the authors also argue that practice-led methods can be used in combination with other empirical approaches, which may help to reinforce the outcomes. For instance, it is possible to develop artifacts of sound design or prototype systems, which are informed by, or provide a basis for, quantitative, qualitative or mixed methods evaluations. How such practice-led and interdisciplinary studies may be formulated, is specifically explored in this chapter with reference to the authors' previous research in affective sound design, which considers aspects of human emotion and perception as they relate to sound. Through the discussion of various projects, the chapter explores how practice-led and interdisciplinary approaches may lead to positive research outcomes, and may also have an enriching effect on the skillsets, knowledge and attributes of individual researchers and their teams.
The use of robot companion pets for people in care homes has been extensively studied. The results are largely positive and suggest that they are valuable in enhancing wellbeing, communication and behavioural aspects. However, there has been little research in people's own homes, possibly due to the cost and complexity of some of the robot pets currently available. As dementia affects people in different ways, this study explores the effects of a robot cat for people in their own homes, without specifically investigating the effects on a particular symptom. We utilised a case study design to investigate the proposition that various factors influence the impact of a robot cat on the person living with dementia and their carer, including acceptability of the robot pet and acceptance of dementia and its symptoms. The qualitative analysis explores the similarities and differences within the data which were gathered during interviews with people with dementia and their families. This analysis revealed four themes: Distraction, Communication, Acceptance and rejection, and Connecting with the cat and connecting with others. These themes were synthesised into two overarching themes: the effect of the cat on mood and behaviour, and The interaction with the cat. We present the acceptability and impact of the robot cat on symptoms of dementia, with data presented across and within the group of participants. Our analysis suggests that benefits of the robot pet were evident, and although this was a small-scale study, where they were accepted, robot pets provided positive outcomes for the participants and their families.
Communicating emotion is essential in building and maintaining relationships. We communicate our emotional state not just with the words we use, but also how we say them. Changes in the rate of speech, short-term energy and intonation all help to convey emotional states like 'angry', 'sad' and 'happy'. People with dysarthria, the most common speech disorder, have reduced articulatory and phonatory control. This can affect the intelligibility of their speech, especially when communicating with unfamiliar conversation partners. However, we know little about how people with dysarthria convey their emotional state, and whether they are having to make changes to their speech to achieve this. In this study, we investigated the ability of people with dysarthria, caused by cerebral palsy and Parkinson's disease, to communicate emotions in their speech, and we compared their speech to that of speakers with typical speech. A parallel database of emotional speech was collected. One female speaker with dysarthria due to cerebral palsy, 3 speakers with dysarthria due to Parkinson's disease (2 female and 1 male), and 21 typical speakers (9 female and 12 male) produced sentences with 'angry', 'happy', 'sad', and 'neutral' emotions. A number of acoustic features were analysed using linear multi-level modeling. The results show that people with dysarthria were able to control some aspects of the suprasegmental and prosodic features when attempting to communicate emotions. For most speakers the changes they made are consistent with the changes made by speakers with typical speech. Even when the changes might be different to that of typical speakers, acoustic analysis shows these were consistent for different emotions. The analysis shows that variation in energy and jitter (local absolute) are major indicators of emotion in the study.
This contribution presents two conferences, GameAbilitation and ArtAbilitation, representing a trans-disciplinary research platform that emerged from a mature body of work investigating ICT across functional diversity of ability. The work questioned requirements of cybertherapy systems based upon gameplay and creative expression experiences. Supplementing traditional intervention whilst planning ahead to address societal demographic predictions of increased aged and disabled persons was explored. The need for ICT solutions parallels forecasts of service industries’ predicted inability to cope with such increases. The conferences offer a podium for sharing such work with an aim to inform whilst inspiring collaborations and advancements. The research platform was coined by the author as Ludic Engagement Designs for All (LEDA).
Effective communication relies on the comprehension of both verbal and nonverbal information. People with dysarthria may lose their ability to produce intelligible and audible speech sounds which in time may affect their way of conveying emotions, that are mostly expressed using nonverbal signals. Recent research shows some promise on automatically recognising the verbal part of dysarthric speech. However, this is the first study that investigates the ability to automatically recognise the nonverbal part. A parallel database of dysarthric and typical emotional speech is collected, and approaches to discriminating between emotions using models trained on either dysarthric (speaker dependent, matched) or typical (speaker independent, unmatched) speech are investigated for four speakers with dysarthria caused by cerebral palsy and Parkinson's disease. Promising results are achieved in both scenarios using SVM classifiers, opening new doors to improved, more expressive voice input communication aids.
This chapter explores a selection of practical approaches for designing video game audio based on the subjective perception of a player avatar. The authors discuss several prototype video game systems developed as part of their practice-led research, which provide interactive audio systems that represent the aural experience of a virtual avatar undergoing an altered state of consciousness. A variety of possible approaches for sound design are described for representing the subjective perceptual experiences of a player avatar. Building upon this work, they argue for "avatar-centered subjectivity," as a generalized concept applicable for first-person perspective video games and simulations.
The number of people living with dementia is growing, leading to increasing pressure upon care providers. The mechanisms to reduce symptoms of dementia can take many forms and have the aim of improving the wellbeing and quality of life of the person living with dementia and those who care for them. Besides the person who has dementia, the condition has a profound impact upon their loved ones and carers. One therapeutic approach is the use of music, an area recognised as having potential benefit, but requiring further research. The present paper reports upon a mixed methods cohort study that examines the use of a musical mobile app as a way to promote song-task association in people living with dementia. The study took place in care home environments in the UK. A total of fourteen participants (N = 14) were recruited. Quantitative measurements were taken on a daily basis prior to, and during, use of the mobile app over several weeks. Metrics came from the complete Self-Assessment Manikin scale (arousal, valence, and dominance), and a subset of three from the Quality of Life in Alzheimer's Disease questionnaire (physical health, memory, and life as a whole). Subsequently, semistructured interviews were conducted with staff at the care home to assess the impact of the app upon their role and the residents they care for. No significant differences were found in the combined quantitative measures for the ten (n = 10) sets of responses sufficient to be analysed. However, the qualitative results suggest that use of the mobile app produced positive changes in terms of behaviour, ability, and routine in the life of residents living with dementia. These findings contribute to the growing body of evidence-based research in the field of musical therapies for reducing symptoms of dementia and highlight elements where further study is warranted.
Changes in human sensory perception can occur for a variety of reasons. In the case of distortions or transformations in the human auditory system, the aetiology may include factors such as medical conditions affecting cognition or physiology, interaction of the ears with mechanical waves, or stem from chemically induced sources, such the consumption of alcohol. These changes may be permanent, intermittent, or temporary. In order to communicate such effects to an audience in an accessible, and easily understood manner, a series of electroacoustic compositions were produced. This concept follows on from previous work on the theme of representing auditory hallucinations. Specifically, these compositions relate to auditory impairments that humans can experience due to tinnitus or through the consumption of alcohol. In the case of tinnitus, whilst much is known about the causes and symptoms, the experience of what it is like to live with tinnitus is less explored and those who have acquired the condition may often feel frustration when trying to convey the experience of ‘what it is like’ for them. In terms of impairment from alcohol consumption, whilst there is much hearsay, little research exists on the immediate and short-term effects of alcohol consumption on the human auditory system, despite over half of the UK population reported as consuming alcohol in 2017. The methodology employed to design these compositions draws upon scientific research findings, including experimental and explorative studies involving human participants, coupled with electroacoustic composition techniques. The pieces are typically constructed by mixing field recordings with synthesised materials and incorporating a range of temporal and frequency domain manipulations to the elements therein. In this way, the listener is able to experience the phenomenon in a recognisable context, where distortions of reality can be emulated to varying degrees. It is intended that these compositions can serve as easily accessible and understood examples of auditory impairments and that they might find utility in the communication of symptoms to those who have never experienced the underlying causes or conditions. This presents opportunities for pieces like these to be used in scenarios such as education and public health awareness campaigns.
Altered states of consciousness (ASC) can be represented in video games through appropriate use of sound and computer graphics. Our research seeks to establish systematic methods for simulating ASC using computer sound and graphics, to improve the realism of ASC representations in video game engines. Quake Delirium is a prototype ‘ASC Simulation’ that we have created by modifying the video game Quake. Through automation of various graphical parameters that represent the conscious state of the game character, hallucinatory ASC are represented. While the initial version of Quake Delirium utilised a pre-determined automation path to produce these changes, we propose that immersion may be improved by providing the user with a ‘passive’ method of control, using a brain-computer interface (BCI). In this initial trial, we explore the use of a consumer-grade electroencephalograph (EEG) headset for this purpose. Keywords—Altered States of Consciousness; Video Games; Brain-Computer Interface (BCI)
A simplified model of a lighting process applied in theatrical productions is one that involves two key players. The first is that of the lighting designer, to produce a set of intentions and plans for the scenes that define the show. The second, the lighting technician, has the job of translating these designs into practice using control equipment, luminaires, and other technical instruments. The lighting design often becomes a ‘working document’ subject to change and adaptation as the physical reality of the design becomes apparent, and the input of other stakeholders is considered. This process can be a valuable creative tool, and also a difficult technical hurdle to overcome, depending on a varied number of factors. A common frustration with this process is that either the complexity of the task, or difficulty in communication can make it difficult for the final creative vision to be effectively realised. Strains may also arise in the case of small, often touring, theatre companies where the lighting designer and technician may be the same person, and frequently one of the performers as well. Considering the design aspect, there can be challenges in ensuring efficacy of lighting plans between venues in touring productions, with 2D lighting sketches or even 3D computer simulations confined to the paper or screen. From a technical perspective, the role of the lighting technician in theatres and performance situations has included the operation of lighting control equipment during shows. The equipment has evolved over time but has, until recently, been grounded upon the basis of faders and the mixing desk. It is argued that this paradigm has failed to keep pace with the change in other interactive technologies. The on-going research described in this paper explores existing and upcoming technologies in the field, whilst also seeking to understand the roles and communication workflows of those involved in theatrical lighting to find the best areas to seek improvement, adopting principles of user-centred design. The intention of this research is to develop a new paradigm, and manifestation of it, using a control method for lighting or projection that allows a more intuitive form of operation in theatre productions, which will be scalable and flexible.