Musical (MSS) source separation of western popular music using non-causal deep learning can be very effective. In contrast, MSS for classical music is an unsolved problem. Classical ensembles are harder to separate than popular music because of issues such as the inherent greater variation in the music; the sparsity of recordings with ground truth for supervised training; and greater ambiguity between instruments. The Cadenza project has been exploring MSS for classical music. This is being done so music can be remixed to improve listening experiences for people with hearing loss. To enable the work, a new database of synthesized woodwind ensembles was created to overcome instrumental imbalances in the EnsembleSet. For the MSS, a set of ConvTasNet models was used with each model being trained to extract a string or woodwind instrument. ConvTasNet was chosen because it enabled both causal and non-causal approaches to be tested. Non-causal approaches have dominated MSS work and are useful for recorded music, but for live music or processing on hearing aids, causal signal processing is needed. The MSS performance was evaluated on the two small datasets (Bach10 and URMP) of real instrument recordings where the ground-truth is available. The performances of the causal and non-causal systems were similar. Comparing the average Signal-to-Distortion (SDR) of the synthesized validation set (6.2 dB causal; 6.9 non-causal), to the real recorded evaluation set (0.3 dB causal, 0.4 dB non-causal), shows that mismatch between synthesized and recorded data is a problem. Future work needs to either gather more real recordings that can be used for training, or to improve the realism and diversity of the synthesized recordings to reduce the mismatch...
The Cadenza machine learning challenges are improving the processing of music in hearing aids and consumer devices for those with hearing loss. There are two tasks in the current round (CAD2), which is organized within the IEEE SPS program. The tasks are motivated by the problem people with hearing loss can have when trying to hear out lyrics and instruments. Task 1 is to improve lyric intelligibility for pop/rock music without compromising audio quality. The objective audio quality is evaluated using HAAQI, the Hearing Aid Audio Quality Index, and intelligibility by a system based on the Whisper ASR (automatic-speech recognition). Subjective audio quality and intelligibility are also evaluated via a panel of listeners with hearing loss. Task 2 is to rebalance instruments within small classical ensembles (duets to quintets of strings and woodwinds) to enable personalized remixing. The objective metric is HAAQI. Both tasks start with stereo recordings. We will present the challenge design, a summary of the approaches taken by challenge entrants, and an overview of the objective and perceptual evaluation of the systems entered. We will also discuss how CAD2 will inform the third Cadenza challenge, launching later in 2025.
Listening to music can be an issue for those with a hearing impairment, and hearing aids are not a universal solution. This paper details the first use of an open challenge methodology to improve the audio quality of music for those with hearing loss through machine learning. The first challenge (CAD1) had 9 participants. The second was a 2024 ICASSP grand challenge (ICASSP24), which attracted 17 entrants. The challenge tasks concerned demixing and remixing pop/rock music to allow a personalized rebalancing of the instruments in the mix, along with amplification to correct for raised hearing thresholds. The software baselines provided for entrants to build upon used two state-of-the-art demix algorithms: Hybrid Demucs and Open-Unmix. Objective evaluation used HAAQI, the Hearing-Aid Audio Quality Index. No entries improved on the best baseline in CAD1. It is suggested that this arose because demixing algorithms are relatively mature, and recent work has shown that access to large (private) datasets is needed to further improve performance. Learning from this, for ICASSP24 the scenario was made more difficult by using loudspeaker reproduction and specifying gains to be applied before remixing. This also made the scenario more useful for listening through hearing aids. Nine entrants scored better than the best ICASSP24 baseline. Most of the entrants used a refined version of Hybrid Demucs and NAL-R amplification. The highest scoring system combined the outputs of several demixing algorithms in an ensemble approach. These challenges are now open benchmarks for future research with freely available software and data.
The Clarity (Speech in noise) and Cadenza (music) projects are two large, complementary research projects that are exploiting the latest in machine learning to create improved listening experiences for those with a hearing loss.In both, we are running a series of open competitions, for which entrants are challenged to improve and personalise the audio for listeners with a hearing loss.This challenge methodology fosters a new research community devoted to making music and speech more accessible, as well as creating open-source tools and databases to facilitate future investigations.The challenges pose a variety of dilemmas to the competitors: for instance, while a hearing aid must manipulate live speech with low latency and limited computing power, recorded music from consumer devices can be pre-processed with non-causal techniques using cloud computing.In this presentation we will update the latest news on the third Clarity challenge and the first Cadenza challenge and report on the open-access computational tools and rating scales we have developed.
This study introduces a novel real-time, gaze-directed audio-visual speech enhancement (AVSE) framework for hearing aids designed to improve speech intelligibility for individuals with hearing loss (pHL) in noisy environments. Existing gaze estimation methods often rely solely on eye angle, leading to reduced accuracy. Our approach addresses this limitation by combining eye angle and nose position with head pose estimation to enhance target speaker identification and facilitate noise reduction within the AVSE framework. We utilize a novel eye gaze estimation algorithm that leverages the listener's nose position for improved accuracy. Head pose estimation is also used to capture the overall direction of attention. This combined information is utilized in real-time to steer a beamformer towards the target speaker, effectively enhancing their voice and suppressing background noise. Pilot trials with pHL users demonstrated high accuracy (99.55% - 99.88%) in estimating target speaker direction using the proposed algorithm. This research presents a promising approach for improving communication accessibility and social interaction for pIH users by potentially enhancing speech recognition in challenging listening situations. Future studies will quantify the improvement in speech intelligibility achieved by the gaze directed AVSE framework.
This paper reports on the design and results of the 2024 ICASSP SP Cadenza Challenge: Music Demixing/Remixing for Hearing Aids. The Cadenza project is working to enhance the audio quality of music for those with a hearing loss. The scenario for the challenge was listening to stereo reproduction over loudspeakers via hearing aids. The task was to: decompose pop/rock music into vocal, drums, bass and other (VDBO); rebalance the different tracks with specified gains and then remixing back to stereo. End-to-end approaches were also accepted. 17 systems were submitted by 11 teams. Causal systems performed poorer than non-causal approaches. 9 systems beat the baseline. A common approach was to fine-tuning pretrained demixing models. The best approach used an ensemble of models.
Rodent models of tinnitus are commonly used to study its mechanisms and potential treatments. Tinnitus can be identified by changes in the gap-induced prepulse inhibition of the acoustic startle (GPIAS), most commonly by using pressure detectors to measure the whole-body startle (WBS). Unfortunately, the WBS habituates quickly, the measuring system can introduce mechanical oscillations and the response shows considerable variability. We have instead used a motion tracking system to measure the localized motion of small reflective markers in response to an acoustic startle reflex in guinea pigs and mice. For guinea pigs, the pinna had the largest responses both in terms of displacement between pairs of markers and in terms of the speed of the reflex movement. Smaller, but still reliable responses were observed with markers on the thorax, abdomen and back. The peak speed of the pinna reflex was the most sensitive measure for calculating GPIAS in the guinea pig. Recording the pinna reflex in mice proved impractical due to removal of the markers during grooming. However, recordings from their back and tail allowed us to measure the peak speed and the twitch amplitude (area under curve) of reflex responses and both analysis methods showed robust GPIAS. When mice were administered high doses of sodium salicylate, which induces tinnitus in humans, there was a significant reduction in GPIAS, consistent with the presence of tinnitus. Thus, measurement of the peak speed or twitch amplitude of pinna, back and tail markers provides a reliable assessment of tinnitus in rodents.
Understanding lyrics is a major barrier to enjoying music for people with hearing loss. To improve lyric understanding through machine-learning, metrics need to be informed by the experiences of the target population. Currently, there are no data on lyric-recall ability of older individuals with hearing loss. Twelve older participants with mostly mild-sloping hearing loss listened to and recalled 100 segments of popular music that varied in genre, duration and word count. In each trial, participants heard a randomly chosen segment twice with a 5-s interstimulus interval over headphones at an A-weighted level of 65 dB plus individualised frequency-dependent nonlinear gain. The proportion of words heard correctly varied greatly across samples, from 0 to 100%, and varied as a function of genre, similar to past studies with different populations. Individual intelligibility across samples was correlated with age; sample intelligibility across individuals, however, was not correlated with word count or rate. Improving sung-lyric intelligibility is a different challenge from spoken-speech enhancement for those with hearing loss. Not only does the enhancement need to be considered within the overall enjoyment of the music, but the variation in results—as seen in the current study—reflects variations in genre, orchestration and vocal quality in sung music.
We update an earlier publication (Akeroyd MA et al., Int J Audiol. 2014) to derive new estimates of the number of adults in the UK with a hearing loss, using population data from the 2021/2022 UK censuses and prevalence data from the UK National Study of Hearing (Akeroyd MA et al. Trends In Hearing 2019). Setting a criterion of a four-frequency hearing level >= 35 dB in the better ear gives an estimate of 4.6 million adults aged 18 to 80, which is a 17% increase on the earlier value from the 2011 census. With a criterion of >= 20 dB hearing level, the estimate is 12.3 million in total, or 1 in 4 of the population aged 18-80. The revised numbers highlight just how common hearing problems are in the population. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This research was supported by the National Institute for Health and Care Research (NIHR) Nottingham Biomedical Research Centre and Nottingham Biomedical Research Centre. The views expressed in this article are those of the authors and not necessarily those of the NHS, the NIHR, or the Department of Health and Social Care. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: Data was obtained from recent UK censuses and the UKL National Study of Hearing https://www.nisra.gov.uk/publications/census-2021-main-statistics-demography-tables-age-and-sex https://www.ons.gov.uk/peoplepopulationandcommunity/populationandmigration/populationestimates/datasets/2011censuspopulationestimatesbysingleyearofageandsexforlocalauthoritiesintheunitedkingdom https://www.ons.gov.uk/peoplepopulationandcommunity/birthsdeathsandmarriages/livebirths/articles/trendsinbirthsanddeathsoverthelastcentury/2015-07-15 https://www.ons.gov.uk/datasets/RM121/editions/2021/versions/1 https://www.scotlandscensus.gov.uk/2022-results/scotland-s-census-2022-rounded-population-estimates/ https://doi.org/10.3109/14992027.2013.850539 I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data produced in the present study are available upon reasonable request to the authors
The Cadenza project is an ongoing project that aims to improve music quality for those with a hearing loss. The project is running signal-processing and machine-learning challenges to address different listening issues and scenarios. During the first round, the challenge focused on non-causal music source separation to allow remixing for those with hearing loss. This fed into an ICASSP 2024 challenge, which had crosstalk from loudspeaker reproduction included. There are three potential arms to our upcoming 2024 challenge based on reported issues from hearing-impaired music listeners: (1) low-latency causal audio source separation, (2) lyric intelligibility enhancement without loss of timbre or instrumental balance, and (3) loudness/dynamic range control. Each of these potential challenges raise questions as to the appropriate reference signals, as well as the practicalities of deriving the appropriate signals for an open-source machine-learning challenge. [Work supported by UK EPSRC Grant No. EP/W019434/1]
IntroductionPrevious work on audio quality evaluation has demonstrated a developing convergence of the key perceptual attributes underlying judgments of quality, such as timbral, spatial and technical attributes. However, across existing research there remains a limited understanding of the crucial perceptual attributes that inform audio quality evaluation for people with hearing loss, and those who use hearing aids. This is especially the case with music, given the unique problems it presents in contrast to human speech.MethodThis paper presents a sensory evaluation study utilising descriptive analysis methods, in which a panel of hearing aid users collaborated, through consensus, to identify the most important perceptual attributes of music audio quality and developed a series of rating scales for future listening tests. Participants (N = 12), with a hearing loss ranging from mild to severe, first completed an online elicitation task, providing single-word terms to describe the audio quality of original and processed music samples; this was completed twice by each participant, once with hearing aids, and once without. Participants were then guided in discussing these raw terms across three focus groups, in which they reduced the term space, identified important perceptual groupings of terms, and developed perceptual attributes from these groups (including rating scales and definitions for each).ResultsFindings show that there were seven key perceptual dimensions underlying music audio quality (clarity, harshness, distortion, spaciousness, treble strength, middle strength, and bass strength), alongside a music audio quality attribute and possible alternative frequency balance attributes.DiscussionWe outline how these perceptual attributes align with extant literature, how attribute rating instruments might be used in future work, and the importance of better understanding the music listening difficulties of people with varied profiles of hearing loss.
This paper reports on the design and outcomes of the 2nd Clarity Prediction Challenge (CPC2) for predicting the intelligibility of hearing aid processed signals heard by individuals with a hearing impairment. The challenge was designed to promote new approaches for estimating the intelligibility of hearing aid signals that can be used in future hearing aid algorithm development. It extends an earlier round (CPC1, 2022) in a number of critical directions, including a larger dataset coming from new speech intelligibility listening experiments, a greater degree of variability in the test materials, and a design that requires prediction systems to generalise to unseen algorithms and listeners. This paper provides a full description of the new publicly available CPC2 dataset, the CPC2 challenge design, and the baseline systems. The challenge attracted 12 systems from 9 research teams. The systems are reviewed, their performance is analysed and conclusions are presented, with reference to the progress made since the earlier CPC1 challenge. In particular, it is seen how reference-free, non-intrusive systems based on pre-trained large acoustic models can perform well in this context.
This paper presents the Cadenza Woodwind Dataset . This publicly available data is synthesised audio for woodwind quartets including renderings of each instrument in isolation. The data was created to be used as training data within Cadenza's second open machine learning challenge (CAD2) for the task on rebalancing classical music ensembles. The dataset is also intended for developing other music information retrieval (MIR) algorithms using machine learning. It was created because of the lack of large-scale datasets of classical woodwind music with separate audio for each instrument and permissive license for reuse. Music scores were selected from the OpenScore String Quartet corpus. These were rendered for two woodwind ensembles of (i) flute, oboe, clarinet and bassoon; and (ii) flute, oboe, alto saxophone and bassoon. This was done by a professional music producer using industry-standard software. Virtual instruments were used to create the audio for each instrument using software that interpreted expression markings in the score. Convolution reverberation was used to simulate a performance space and the ensembles mixed. The dataset consists of the audio and associated metadata.
BACKGROUND:Tinnitus is a common problem in patients with a cochlear implant (CI). Between 4% and 25% of CI recipients experience a moderate to severe tinnitus handicap. However, apart from handicap scores, little is known about the real-life impact tinnitus has on those with CIs. We aimed to explore the impact of tinnitus on adult CI recipients, situations impacting tinnitus, tinnitus-related difficulties and their management strategies, using an exploratory sequential mixed-method approach.METHODS:A 2-week web-based forum was conducted using Cochlear Ltd.'s online platform, Cochlear Conversation. A thematic analysis was conducted on the data from the forum discussion to develop key themes and sub-themes. To quantify themes and sub-themes identified, a survey was developed in English with face validity using cognitive interviews, then translated into French, German and Dutch and disseminated on the Cochlear Conversation platform, in six countries (Australia, France, Germany, New Zealand, the Netherlands and United Kingdom). Participants were adult CI recipients experiencing tinnitus who received a Cochlear Ltd. CI after 18 years of age.RESULTS:Four key themes were identified using thematic analysis of the discussion forum: tinnitus experience, situations impacting tinnitus, difficulties associated with tinnitus and tinnitus management. Among the 414 participants of the survey, tinnitus burden on average was a moderate problem without their sound processor and not a problem with the sound processor on. Fatigue, stress, concentration, group conversation and hearing difficulties were the most frequently reported difficulties and was reported to intensify when not wearing the sound processor. For most CI recipients, tinnitus seemed to increase when performing a hearing test, during a CI programming session, or when tired, stressed, or sick. To manage their tinnitus, participants reported turning on their sound processor and avoiding noisy environments.CONCLUSION:The qualitative analysis showed that tinnitus can affect everyday life of CI recipients in various ways and highlighted the heterogeneity in their tinnitus experiences. The survey findings extended this to show that tinnitus impact, related difficulties, and management strategies often depend on sound processor use. This exploratory sequential mixed-method study provided a better understanding of the potential benefits of sound processor use, and thus of intracochlear electrical stimulation, on the impact of tinnitus.
Interior car noise refers to the general noise generated by the engine transmission, the interaction between road and types, and weather conditions such as turbulent wind. For drivers or passengers with hearing loss, these can create especially challenging listening situations. The Cadenza Project is organising a series of machine learning challenges to advance signal processing of music for listeners with a hearing loss, and a key scenario in its first challenge is listening in a car to music in the presence of noise. To create enough machine-learnable training materials we need to simulate typical car noises rather than just use one particular recording. We are systematically reviewing the literature on real-world recordings to determine the range of parameters for these simulations. We searched Web of Science with the terms “(car noise, car noise interior, interior noise) AND (speed OR FFT OR spectr*).” A total of 126 studies have been found so far and 12 papers retained on the basis that a frequency spectrum for interior car noise was provided that was suitable for numerical analysis. Results will be presented.
This paper reports on the design and outcomes of the 2nd Clarity Enhancement Challenge (CEC2), a challenge for stimulating novel approaches to hearing-aid speech intelligibility enhancement. The challenge was for a listener attending to a target speaker in a noisy, domestic environment. The challenge extends the previous edition, CEC1, in a number of key respects: scenes have multiple interferers including speech, noise and music; ambisonics are used to model listener head movement; target speaker identity is provided to encourage speaker extraction approaches. Systems are evaluated both via the HASPI intelligibility metric and with listening tests using a panel of hearing-impaired listeners. The paper reviews the 18 systems that were submitted describing them in terms of their enhancement and amplification stages. HASPI is seen to be a good predictor of listener performance. The top system, using carefully engineered neural approaches, produces highly intelligible signals for complex scenes with SNRs down to -12 dB while obeying the challenges 5 ms latency constraint.