Human auditory systems resolve spatial ambiguity through individual spectral cue integration, a process influenced by both acoustic and non-acoustic factors. We propose a Bayesian sagittal-plane localization model with five spectral weighting schemes, including an adaptive approach that dynamically adjusts frequency contributions based on the reliability of the main notch region, alongside four fixed schemes from previous studies. Model parameters were individually calibrated to capture listener-specific localization behavior. Evaluations were performed in the median plane and lateral sagittal planes, as well as under reduced perceptual confidence using broadband click-train stimuli at multiple spectral resolutions. Results show that no single weighting scheme exhibits a consistent group-level preference for broadband noise stimuli, although the adaptive scheme most closely approximates human responses at low confidence levels. As spectral cues decrease, human responses increasingly reflect the bimodal prior distribution, a trend captured by the model. These findings underscore the importance of spatial priors and adaptive weighting in modeling human sagittal-plane localization.
This study provides a systematic characterization of Mandarin consonant perception in adult cochlear implant (CI) users and benchmarks the resulting perceptual organization against a condition-matched hearing-aid (HA) dataset and a previously reported normal-hearing (NH) reference. Thirty-five Mandarin-speaking CI listeners completed a 21-alternative closed-set identification task using 21 /C(i)ā/ syllables presented in quiet. Data were analyzed using accuracy, hierarchical clustering, multidimensional scaling, and feature-based information transmission. CI listeners achieved moderate overall accuracy (68%), with the greatest difficulty for aspirated stops (notably p and t) and several high-frequency sibilants. Furthermore, confusion analyses revealed a structured but compressed perceptual organization under device-mediated hearing. Both CI and HA data were well described by a two-dimensional perceptual space (NH required at least three), anchored by robust contrasts of sonorancy and aspiration, with other distinctions (including place, manner, and sibilance) collapsed into a narrow band of partial separability. Within this shared compression, CI and HA diverged in the perceptual organization supported by residual cues: HA showed stronger sonorant and aspiration-aligned organization, whereas CI showed a relatively clearer broad obstruent-manner organization. Together, these findings provide new insight into how device-mediated hearing restructures the perceptual organization of Mandarin consonants beyond overall accuracy.
OBJECTIVES:The antiphasic digit-in-noise test has demonstrated the advantages of antiphasic presentation in hearing screening. Inspired by this, this study explores whether incorporating antiphasic stimuli in the Chinese zodiac-in-noise (ZIN) test can enhance its sensitivity to detect hearing loss. Furthermore, this study investigates the relation between the binaural intelligibility level difference (BILD), calculated as the difference between antiphasic and diotic results, and hearing thresholds, and evaluates the potential of BILD as an indicator for hearing loss. DESIGN:Normative data for the antiphasic ZIN test were established based on data from 117 normal-hearing listeners. Subsequently, the performance of the antiphasic ZIN test was evaluated in 195 listeners with varying degrees and types of hearing loss. Participants were classified into four groups based on their audiograms: normal hearing (n = 115), symmetric hearing loss (n = 37), asymmetric hearing loss (n = 14), and unilateral hearing loss (n = 29). BILD was further analyzed to assess its relation with high- and low-frequency hearing thresholds. RESULTS:The antiphasic ZIN test has a cutoff value of -13.7 dB and a reference speech reception threshold of -19.0 ± 3.2 dB. The measurement error, as estimated from the test-retest reliability, was 1.2 dB. Partial correlation analysis controlling for age revealed comparable associations of hearing thresholds across ears with BILD ( ρpoorer-antiphasic = 0.65, 95% confidence interval [CI]: 0.54 to 0.74 versus ρpoorer-diotic = 0.43, 95% CI: 0.29 to 0.55). Compared with the diotic ZIN test, the antiphasic ZIN demonstrates higher sensitivity (0.88, 95% CI: 0.78 to 0.93 versus 0.49, 95% CI: 0.38 to 0.60) in detecting hearing loss of >25 dB HL in the poorer ear, with comparable specificity (0.86, 95% CI: 0.79 to 0.91 versus 0.97, 95% CI: 0.93 to 0.99). Compared with high-frequency hearing loss, BILD showed a stronger association with low-frequency hearing loss, with a more pronounced decrease as hearing impairment in the poorer ear worsened. CONCLUSIONS:The antiphasic ZIN test demonstrates higher screening sensitivity compared with the diotic ZIN, particularly in detecting asymmetric and unilateral hearing loss. BILD is more sensitive to low-frequency hearing impairments and shows substantial promise as an indicator for hearing loss.
Speech perception has been extensively studied using degradation algorithms such as channel vocoding, mosaic speech, and pointillistic speech. Here, an "atomic speech model" is introduced to generate unique sparse time-frequency patterns. It processes speech signals using a bank of bandpass filters, undersamples the signals, and reproduces each sample using a Gaussian-enveloped tone (a Gabor atom). To examine atomic speech intelligibility, adaptive speech reception thresholds (SRTs) are measured as a function of atom rate in normal-hearing listeners, investigating the effects of spectral maxima, binaural integration, and single echo. Experiment 1 showed atomic speech with 4 spectral maxima out of 32 bands remained intelligible even at a low rate under 80 atoms per second. Experiment 2 showed that when atoms were nonoverlappingly assigned to both ears, the mean SRT increased (i.e., worsened) compared to the monaural condition, where all atoms were assigned to one ear. Individual data revealed that a few listeners could integrate information from both ears, performing comparably to the monaural condition. Experiment 3 indicated higher mean SRT with a 100 ms echo delay than that with shorter delays (e.g., 50, 25, and 0 ms). These findings demonstrate the utility of the atomic speech model for investigating speech perception and its underlying mechanisms.
Previous research has demonstrated the negligible impact of harmonicity on English speech perception for normal hearing (NH) listeners inquiet environments. This study aims to bridge the gap in understanding the role of harmonicity in Mandarin speech perception for cochlear implant (CI) users. Speech perception inquiet was tested in both CI simulation group and actual CI user group using harmonic and inharmonic Mandarin speech. Furthermore, speech-on-speech perception was tested in NH, CI simulation, and actual CI user groups. For speech perception in quiet, results show that, compared to harmonic speech, inharmonic speech decreased the mean recognition rate for both actual CI user and CI simulation groups by about 10 percentage points. For speech-on-speech perception, all groups (i.e., NH, CI simulation, and actual CI user) performed worse with inharmonic stimuli compared to harmonic stimuli. The findings of this study, along with previous studies in NH listeners, indicate that harmonicity aids target speech recognition for NH listeners in speech-on-speech conditions but not speech perception inquiet. In contrast, harmonicity plays an important role in CI users' Mandarin speech recognition in both quiet and speech-on-speech conditions. However, under speech-on-speech conditions, CI users could only understand target speech at positive SNRs (often > 5 dB), suggesting that their performance depends on the intelligibility of the target speech. The contribution of harmonicity to masking release in CI users remains unclear.
Diffusion-based speech enhancement has demonstrated remarkable performance, but existing models still lack precise time-frequency modeling capabilities and fall short in aligning with perceptual metrics. To address these issues, we propose TF-NCSN++, which integrates a TF-GridNet module into the NCSN++ backbone to enable explicit cross-frequency and temporal modeling; at the same time, we introduce MD-NCSN++, which leverages normalized PESQ scores to train a lightweight metric discriminator, providing auxiliary perceptual guidance and imposing a weak constraint on the diffusion objective. Experiments on the VoiceBank+DEMAND and WSJ0-CHiME3 datasets show that TF-NCSN++ significantly improves intelligibility and distortion metrics, while MD-NCSN++ effectively enhances perceptual metrics. When combined, the two achieve optimal or near-optimal results across six evaluation measures, fully demonstrating the synergistic advantages of structural modeling and objective alignment.
Deep learning has become a popular approach for improving acoustic echo cancellation (AEC) in communication systems. However, existing systems mostly rely on traditional delay estimation methods, which often result in performance degradation due to inaccurate delay estimation. Furthermore, most deep learning-based methods use single-stage networks, the limited learning capability of which hinders the performance under harsh echo conditions. To address these challenges, this paper proposes a two-stage system with dual-path alignment. The two-stage strategy performs echo suppression on the magnitude spectrum in the first stage, followed by phase correction in the second one. In addition, the first stage processes two parallel features: the magnitude spectrum and its exponential compressed version, with which dual-path alignment is conducted to improve delay estimation. Experiment results demonstrate the effectiveness of the proposed system for echo suppression in challenging scenarios involving long delay, double-talk and nonlinear distortion.
Cochlear implants (CIs) have limited spectral resolution due to a limited number of electrodes and inter-electrode current interaction, causing difficulties in speech-in-noise perception even for bilateral CI (BiCI) users. Previous studies have suggested alternately stimulating odd and even electrodes between the two sides to reduce current interaction in several ways, showing promising results if dichotic stimuli could be effectively integrated. To utilize the total electrode number of a pair of BiCIs, we propose Doubling the signal analysis Band Density (DBD) and encoding odd and even bands alternately for each side. Two preliminary vocoder-simulation experiments in spectral-temporally modulated ripple discrimination and speech-in-noise perception were carried out to compare DBD with the default setting of the advanced combination encoder (ACE) strategy. The proposed method showed promising benefits as well as limitations to be further resolved in theory and evaluated in BiCI users.
Convolution-augmented transformer (Conformer) has shown impressive performance in speech enhancement. However, it employs a single fixed large kernel convolutional block for local feature modeling, which limits its attention to local features. Also, modeling the entire frequency band in the same manner hinders the model from focusing on more critical sub band information, resulting in low performance. To address these two issues, we propose a GAN-based approach called Multi-Scale Convolution-augmented Transformer GAN (MSCTGAN) for speech enhancement. MSCTGAN incorporates several stacks of MSCT modules to better capture local features. Taking into account the varying importance of speech sub-bands, these features are introduced in frequency domain relative position encoding and loss function, which facilitated the fusion of both full-band and sub-band information. The objective evaluation demonstrates the competitive performance of MSCTGAN compared to models utilizing conventional Transformer and Conformer architectures.
Objectives: Motivated by the growing need for hearing screening in China, the present study has two objectives. First, to develop and validate a new test, called the Chinese Zodiac-in-noise (ZIN) test, for large-scale hearing screening in China. Second, to conduct a large-scale remote hearing screening in China, using the ZIN test developed. Design: The ZIN test was developed following a similar procedure as the digits-in-noise test but emphasizes the importance of consonant recognition by employing the 12 zodiac animals in traditional Chinese culture as speech materials. It measures the speech reception threshold (SRT) using triplets of Chinese zodiac animals in speech-shaped noise with an adaptive procedure. Results: Normative data of the test were obtained in a group of 140 normal-hearing listeners, and the performance of the test was validated by comparisons with pure-tone audiometry in 116 listeners with various hearing abilities. The ZIN test has a reference SRT of −11.0 ± 1.6 dB in normal-hearing listeners with a test-retest variability of 1.7 dB and can be completed in 3 minutes. The ZIN SRT is highly correlated with the better-ear pure-tone threshold ( r = 0.82). With a cutoff value of −7.7 dB, the ZIN test has a sensitivity of 0.85 and a specificity of 0.94 for detecting a hearing loss of 25 dB HL or more at the better ear. A large-scale remote hearing screening involving 30,552 participants was performed using the ZIN test. The large-scale study found a hearing loss proportion of 21.0% across the study sample, with a high proportion of 57.1% in the elderly study sample aged over 60 years. Age and gender were also observed to have associations with hearing loss, with older individuals and males being more likely to have hearing loss. Conclusions: The Chinese ZIN test is a valid and efficient solution for large-scale hearing screening in China. Its remote applications may improve access to hearing screening and enhance public awareness of hearing health.
This study investigates the effects of hearing loss and amplification on Mandarin consonant perception. 44 listeners with varying degrees of hearing loss were tested, both with and without the use of hearing aids. Consonant recognition was strongly correlated with the hearing threshold (r = -0.87), and was significantly improved by hearing-aid amplification (by more than 20% in group means) but was still not perfect. The underlying reasons are discussed. Furthermore, confusion patterns were analyzed and compared with those of normal-hearing listeners in the literature. The most challenging Mandarin consonants for hearing-impaired listeners include consonants with a spectral center of gravity in the high-frequency range (such as s and z), consonants with short duration (such as b, d, and g), and aspirated stops (such as p, t, and k). The findings of this study contribute to a better understanding of the difficulties experienced by Mandarin-speaking listeners with hearing loss.
Cochlear implant (CI) recipients face great challenges in speech-in-noise recognition, partially due to the fact that only temporal envelopes from a limited number of bands are preserved in most CI signal processing strategies. In "n-of -m " strategies (e.g., the Advanced Combinational Encoder, ACE), the number of maxima (nmax) and elec-trical dynamic range (EDR) are two essential parameters that may affect the envelope representation and further influence speech perception. Speech recognition can be improved by optimizing parameter settings in CI pro-gramming. To investigate the effects of nmax and EDR on speech-in-noise perception, Mandarin speech reception thresholds (SRTs) in babble noise were measured in CI recipients using ACE. The nmax was set to 2, 4, 6, 8, and 16. The EDR was set to the base EDR (i.e., participants' clinical EDR) and 50 % EDR (i.e., 50 %-compressed base EDR). Results showed that: 1) there was no significant interaction effect between nmax and EDR, 2) SRTs with nmax = 2, 4, and 16 were significantly higher (or worse) than those with moderate nmax (6-8), 3) narrower EDRs significantly lead to higher SRTs. Simulation experiments using a Gaussian-Enveloped Tones Vocoder in normal -hearing listeners were also conducted and provided both supportive and additional observations to the CI results. This study suggests that, in CI programming, nmax and EDR are two independent influencing factors. Large nmax (e.g., 16) is not recommended as it may harm speech intelligibility in noisy environments, and inaccurate EDR measurements should be avoided.
Periodicity is one of the main cues for pitch-related speech perception but is poorly encoded in cochlear implants (CIs). Most efforts towards CI periodicity enhancement involve explicit fundamental frequency (F0) detection, which requires additional computational loads and may not be reliable in real settings. Here we propose a new strategy, namely F0inTFS, which encodes F0 information without explicit F0 detection. Our idea is inspired by the fact that temporal fine structures (TFS) at the low-frequency channels inherently contain strong F0-related periodicity cues. In F0inTFS, the TFS of the lowest-frequency channel is integrated into the temporal envelopes at all higher-frequency channels using a specifically designed algorithm. F0inTFS can be lightly implemented in the FFT-based framework of one clinically used strategy, i.e., the Advanced Combination Encoder (ACE) strategy. The benefits of F0inTFS are supported by a lexical tone perception experiment in simulated CI users.
Channel vocoders with noise and sine-wave carriers are widely used to simulate modern multi-channel cochlear implants (CIs) in psychoacoustic experiments with normal hearing (NH) subjects. NH subjects perceive vocoded speech as impoverished and unnatural, but how CI listeners perceive vocoded sounds has not been systematically investigated. This letter reports that CI listeners could equally recognize both noise and sine-wave vocoded speech, albeit less well than NH listeners, and the recognition performance would not significantly increase beyond 8 channels. Nevertheless, they can easily discriminate upto-80-channel vocoded speech from the original natural speech.
Perception with electric neuroprostheses is sometimes expected to be simulated using properly designed physical stimuli. Here, we examined a new acoustic vocoder model for electric hearing with cochlear implants (CIs) and hypothesized that comparable speech encoding can lead to comparable perceptual patterns for CI and normal hearing (NH) listeners. Speech signals were encoded using FFT-based signal processing stages including band-pass filtering, temporal envelope extraction, maxima selection, and amplitude compression and quantization. These stages were specifically implemented in the same manner by an Advanced Combination Encoder (ACE) strategy in CI processors and Gaussian-enveloped Tones (GET) or Noise (GEN) vocoders for NH. Adaptive speech reception thresholds (SRTs) in noise were measured using four Mandarin sentence corpora. Initial consonant (11 monosyllables) and final vowel (20 monosyllables) recognition were also measured. NaÏve NH listeners were tested using vocoded speech with the proposed GET/GEN vocoders as well as conventional vocoders (controls). Experienced CI listeners were tested using their daily-used processors. Results showed that: 1) there was a significant training effect on GET vocoded speech perception; 2) the GEN vocoded scores (SRTs with four corpora and consonant and vowel recognition scores) as well as the phoneme-level confusion pattern matched with the CI scores better than controls. The findings suggest that the same signal encoding implementations may lead to similar perceptual patterns simultaneously in multiple perception tasks. This study highlights the importance of faithfully replicating all signal processing stages in the modeling of perceptual patterns in sensory neuroprostheses. This approach has the potential to enhance our understanding of CI perception and accelerate the engineering of prosthetic interventions. The GET/GEN MATLAB program is freely available athttps://github.com/BetterCI/GETVocoder.
Modern cochlear implants (CIs) generate electric current pulsatile stimuli from real-time incoming to stimulate residual auditory nerves of deaf ears. In this unique way, deaf people can (re)gain a sense of hearing and consequent speech communication abilities. The electric hearing mimics the normal acoustic hearing (NH), but with a different physical interface to the neural system, which limits the performance of CI devices. Simulating the electric hearing process of CI users through NH listeners is an important step in CI research and development. Many acoustic modelling methods have been developed for simulation purposes, e.g., to predict the performance of a novel sound coding strategy. Channel vocoders with noise or sine-wave carriers are mostly popular among the methods. The simulation works have accelerated the procedures of re-engineering and understanding of the electric hearing. This paper presents an overview of the literature on channel-vocoder simulation methods. Strengths, limitations, applications, and future works about acoustic vocoder simulation methods are introduced and discussed.
Matlab和Python程序设计课程包含两门编程语言.这门课的授课对象主要是大二大三的学生,授课对象已在大一大二学习过C语言程序设计以及面向对象程序设计(C++程序设计),因此学生需要面对Matlab、Python、C以及C++等多门编程语言所构成的知识编织网,存在类似知识相互纠缠的困惑难题.针对该特点,文章对该课程的教学做大胆探索,精心设计课程教学内容,提出"比较式教学、案例式培养、对象式传授、深入式巩固"为主线的Matlab与Python课程教学模式.通过比较式教学解决相似知识之间模糊纠缠的困惑;案例式培养根据问题培养学生的整体编程思维;对象式传授和互动式提高措施因材施教,激活学生学习热情,培养编程兴趣,提高教学效率;深入式巩固进一步加固学生课堂学习到的内容.
Despite pitch being considered the primary cue for discriminating lexical tones, there are secondary cues such as loudness contour and duration, which may allow some cochlear implant (CI) tone discrimination even with severely degraded pitch cues. To isolate pitch cues from other cues, we developed a new disyllabic word stimulus set (Di) whose primary (pitch) and secondary (loudness) cue varied independently. This Di set consists of 270 disyllabic words, each having a distinct meaning depending on the perceived tone. Thus, listeners who hear the primary pitch cue clearly may hear a different meaning from listeners who struggle with the pitch cue and must rely on the secondary loudness contour. A lexical tone recognition experiment was conducted, which compared Di with a monosyllabic set of natural recordings. Seventeen CI users and eight normal-hearing (NH) listeners took part in the experiment. Results showed that CI users had poorer pitch cues encoding and their tone recognition performance was significantly influenced by the “missing” or “confusing” secondary cues with the Di corpus. The pitch-contour-based tone recognition is still far from satisfactory for CI users compared to NH listeners, even if some appear to integrate multiple cues to achieve high scores. This disyllabic corpus could be used to examine the performance of pitch recognition of CI users and the effectiveness of pitch cue enhancement based Mandarin tone enhancement strategies. The Di corpus is freely available online: https://github.com/BetterCI/DiTone.
The deep learning (DL)-based speech enhancement (SE) has demonstrated its advantage over the classical methods. In most DL-based SEs, however, the systems are optimized based on the minimum mean squared error (MSE), which could result in poor performance in severe noise conditions, e.g., very low signal-to-noise ratios. This paper presents a speech-noise-equilibrium loss function, i.e., a weighted combination of the speech distortion and the noise residue, for network training. Furthermore, based on the observation of the non-Gaussian distribution of the prediction error, a mean absolute error (MAE) criterion is adopted for speech distortion, and hybrid training, i.e., MSE followed by MAE, is proposed for network optimization. Experiment results demonstrate that long-short term memory networks (LSTM)-based SE systems with the proposed loss function achieve better performance than the baselines, particularly, improving both speech quality and intelligibility at low signal-to-noise ratios.
This paper presents a new feature extraction method for Electroencephalogram (EEG)-based motor imagery (MI) classification. Current researches mostly classify different MIs by detecting the event-related desynchronization (ERD) phenomenon from the EEG signals. Due to the poor spatial resolution of the MI-EEG signals, the cortical area (source) activating the MI cannot be located accurately with the EEG (sensor) signal, which might degrade the classification accuracy. This study adopts the EEG source imaging (ESI) technique to estimate the cortical area where source ERD happens from the EEG signal. An improved ESI method based on the linearly constrained minimum variance (LCMV) algorithm, in which an average LCMV filter and an average baseline covariance are constructed for the ESI, is proposed to locate the activated cortical area from the noisy EEG signals. The source ERD features are then extracted. Analytical results show that, for subjects with obvious average source ERD phenomenon, their activated cortical area in a single-trial MI can be well located. MI classification results also support the feasibility of the proposed method for MI-EEG signal processing.