A great deal of information is extracted from spoken exchanges, independent of content. Indeed, the same information can be conveyed in different manners across time and space in light of the goals of the interlocutors, potentially leading individuals to emphasize, modulate, or mask certain vocal characteristics for social advantages. However, no research has examined our perceptions of the same person speaking two closely related dialects separated by a relatively small geographic distance. The current experiment examined perceptions of the same men asked to speak Scottish Standard English (majority national dialect) versus Dundonian (minority national dialect), focusing on an array of well-studied traits in research on social perceptions. We observed robust positive effects of Scottish Standard English on perceived attractiveness and competence, with almost identical large effect sizes. By shaping perception on some of our measured trait dimensions, dialect may impact perceptions of these men as potential dating partners, allies, and/or leaders. In sum, the same man can be treated differently based on the dialect they switch to, both inadvertently and of the speaker's own volition. These findings can motivate further research into geographic differences in social perception and communication/interaction at both narrow and broad geographic bandwidths.
Recent empirical studies have highlighted the large degree of analytic flexibility in data analysis that can lead to substantially different conclusions based on the same data set. Thus, researchers have expressed their concerns that these researcher degrees of freedom might facilitate bias and can lead to claims that do not stand the test of time. Even greater flexibility is to be expected in fields in which the primary data lend themselves to a variety of possible operationalizations. The multidimensional, temporally extended nature of speech constitutes an ideal testing ground for assessing the variability in analytic approaches, which derives not only from aspects of statistical modeling but also from decisions regarding the quantification of the measured behavior. In this study, we gave the same speech-production data set to 46 teams of researchers and asked them to answer the same research question, resulting in substantial variability in reported effect sizes and their interpretation. Using Bayesian meta-analytic tools, we further found little to no evidence that the observed variability can be explained by analysts' prior beliefs, expertise, or the perceived quality of their analyses. In light of this idiosyncratic variability, we recommend that researchers more transparently share details of their analysis, strengthen the link between theoretical construct and quantitative system, and calibrate their (un)certainty in their conclusions.
OBJECTIVES:A range of professions experience high demands on their voices and are potentially at risk of developing voice disorders. Teachers have been studied extensively in this respect, while voiceover artists are a growing professional group with unknown levels of voice training, voice problems and voice care attitudes. To better understand profession-specific voice care requirements, we compared voice training, voice care habits and self-reported voice problems of these two professional groups and measured attitudes to voice care, informed by the Health Belief Model (HBM). STUDY DESIGN:The study was a cross-sectional survey study with two cohorts. METHODS:We surveyed 264 Scottish primary school teachers and 96 UK voiceover artists . Responses were obtained with multiple-choice and free-text questions. Attitudes to voice care were assessed with Likert-type questions that addressed five dimensions of the HBM. RESULTS:Most voiceover artists had some level of voice training, compared to a minority of teachers. Low numbers of teachers reported regular voice care, compared to over half of voiceover artists. Higher numbers of teachers reported work-related voice problems. Voiceover artists reported greater awareness for vocal health and perceived potential effects of voice problems on their work as more severe. Voiceover artists also saw voice care as more beneficial. Teachers perceived barriers to voice care as substantially higher and felt less confident about voice care. Teachers with existing voice problems showed increased perceptions of voice problem susceptibility and severity and saw more benefit in voice care. Cronbach's alpha was below 0.7 for about half of the HBM-informed survey subsets, suggesting that reliability could be improved. CONCLUSIONS:Both groups reported substantial levels of voice problems, and different attitudes to voice care suggest that the two groups require different approaches to preventative intervention. Future studies will benefit from the inclusion of further attitude dimensions beyond the HBM.
Smartphone technology is continuously being updated through software and hardware changes. At present, a limited number of studies have been undertaken to assess the impact of these changes on data collection for linguistic research. This paper discusses the potential of smartphones to gather reliable recordings, along with ethical considerations for storing additional personal information when working in other contexts (i.e. healthcare settings). A pilot study was undertaken using the Fitvoice (TM) account-based application to analyse articulatory proficiency in depressed and healthy participants. Results suggest that phonetic differences exist between these groups in terms of plosive production, and that smartphones are capable of adequately recording these minute aspects of the speech signal for analysis.
BACKGROUND:Occupational voice problems constitute a serious public health issue with substantial financial and human consequences for society. Modern mobile technologies such as smartphones have the potential to enhance approaches to prevention and management of voice problems. This paper addresses an important aspect of smartphone-assisted voice care: the reliability of smartphone-based acoustic analysis for voice health state monitoring.AIM:To assess the reliability of acoustic parameter extraction for a range of commonly used smartphones by comparison with studio recording equipment.METHODS & PROCEDURES:Twenty-two vocally healthy speakers (12 female, 10 male) were recorded producing sustained vowels and connected speech under studio conditions using a high-quality studio microphone and an array of smartphones. For both types of utterance, Bland-Altman analysis was used to assess overall reliability for mean F0, cepstral peak prominence (CPPS), Jitter (RAP) and Shimmer %.OUTCOMES & RESULTS:Analysis of the systematic and random error indicated significant bias for CPPS across both sustained vowels and passage reading. Analysis of the random error of the devices indicated that that mean F0 and CPPS showed acceptable random error size, while jitter and shimmer random error was judged as problematic.CONCLUSIONS & IMPLICATIONS:Confidence in the feasibility of smartphone-based voice assessment is increased by the experimental finding of high levels of reliability for some clinically relevant acoustic parameters, while the use of other parameters is discouraged. We also challenge the practice of using statistical tests (e.g., t-tests) for measurement reliability assessment.
Smartphones have become powerful tools for data capture due to their computational power, internet connectivity, high quality sensors and user-friendly interfaces. This also makes them attractive for the recording of voice data that can be analysed for clinical or other voice health purposes. This however requires detailed assessment of the reliability of voice parameters extracted from smartphone recordings. In a previous study we analysed reliability of measures of periodicity and periodicity deviation, with very mixed results across parameters. In the present study we extended this analysis to measures of added noise and spectral tilt. We analysed systematic and random error for six frequently used acoustic parameters in clinical acoustic voice quality analysis. 22 speakers recorded sustained [a] and a short passage with a studio microphone and four popular smartphones simultaneously. Acoustic parameters were extracted with Praat and smartphone recordings were compared to the studio microphone. Results indicate a small systematic error for almost all parameters and smartphones. Random errors differed substantially between parameters. Our results suggest that extraction of acoustic voice parameters with mobile phones is not without problems and different parameters show substantial differences in reliability. Careful individual assessment of parameters is therefore recommended before use in practice.
Smartphone mediated voice monitoring has the potential to support voice care by facilitating data collection, analysis and biofeedback. To field-test this approach we have developed a smartphone app that allows recording of voice samples alongside voice self-report data. Our longterm aim is convenient and accessible voice monitoring to prevent voice problems and disorders. Our current study focussed on the automatic detection of voice changes in healthy voices that result from common transient illnesses like colds. We have recorded a database of approximately 700 voice samples from 62 speakers and selected a subset of 225 voice samples from 8 speakers who had submitted at least 10 recordings and reported at least one instance of a moderate cold. We extracted 12 acoustic parameters and applied multivariate statistical process control procedures (Hotelling’s T2) to detect whether instances of cold caused violations of distributional control limits. Results showed significant association between control limit violations and reporting of a cold. While there is scope for further improvement of sensitivity and specificity of the procedure, it could already support early detection of voice problems, especially if mediated by voice experts.
Given the importance of voice quality in signalling personal identity and social group membership, effective control of voice features may become especially important during adolescence, yet this has to be achieved in the context of significant physical changes within the speech production system. Most previous research has focussed on phonation, but this study used Vocal Profile Analysis (VPA) [11] for perceptual analysis of both laryngeal and vocal tract voice settings in Scottish adolescents, in order to identify voice quality markers of gender and geographical background in this age group. VPA analysis was carried out for 76 speakers (31 male; 45 female), drawn from three geographically distinct areas of Scotland. Some of the observed variation in voice quality (especially phonatory settings) may be attributable to physical changes associated with puberty, but other setting adjustments seem more likely to be sociophonetic in origin.
A common feature of voice disorders is the impairment of the ability to initiate and sustain adequately periodic vocal fold vibrations. Traditional acoustic approaches that use sustained vowels in which initial/final portions are excluded have been criticised for poor validity and for exclusion of factors that may be a rich source of clinically relevant data e.g. regarding the onset of vocal fold vibration. The aim of this study was to establish if phonation stabilisation time (PST), as determined by cepstal peak prominence (CPP), is useful as an indicator of voice disorders in connected speech. Disordered voices from all groups showed a significantly longer mean PST than normal voices from the same group. The proportion of voiced segments that reached the stable threshold of periodicity were significantly higher for normal voices in all groups. Our results indicate that PST using CPP has potential to differentiate between the normal and disordered voices. The results for the ‘below threshold’ groups for both male and female are of particular interest. These results suggest that PST using CPP may be a potential indicator of voice disorder in cases where traditional acoustic analysis measures of sustained vowels do not show any pathological findings.
There is increasing emphasis on use of connected speech for acoustic analysis of voice disorder, but the differential impact of disorder on initiation, maintenance and termination of phonation has received little attention. This study introduces a new measure of dynamic changes at onset of phonation during connected speech, phonation stabilisation time (PST), and compares this measure with conventional analysis of sustained vowels. Voice samples obtained from the KayPENTAX Disordered Voice Database were analysed (202 females, 128 males) including ‘below threshold’ voices where there was a clinical diagnosis but acoustic parameters for sustained vowels were within the normal range. Female disordered voices showed significantly longer PST duration than normal voices, including those in the ‘below threshold’ group. Overall differences for male voices were also significant. Results suggest that, at least for females, PST measurement from connected speech could provide a more sensitive indicator of disorder than traditional analysis of sustained vowels.
This paper presents articulatory data on silent preparation in a standard Verbal Reaction Time experiment. We have reported in a previous study [6] that Reaction Time is reliably detectable in Ultrasound Tongue Imaging and lip video data, and between 120 to 180 ms ahead of the standard acoustics-based measurements. The aim of the current study was to investigate in more detail how silent speech preparation is timed in relation to faster and slower Reaction Times, and faster and slower articulation rates of the verbal response. The results suggest that the standard acoustic-based measurements of Reaction Time may not only routinely underestimate fastness of response but also obscure considerable variation in actual response behaviour. Particularly tokens with fast Reaction Times seem to exhibit substantial variation with respect to when the response is actually initiated, i.e. detectable in the articulatory data.
There is a sizeable delay between any formulation of an intention to speak and the audible vocalisation that results. Silent articulatory movements in preparation for audible speech comprise a proportion of this phase of speech production. The extensive literature on Reaction Time (RT) is based on the delay between a stimulus and the acoustic onset to speech that is elicited, ignoring the preceding silent elements of speech production in what is an utterance-initial position. We used a standard Snodgrass and Vanderwart picture-naming task to elicit speech in a standard Reaction Time protocol, but recorded the behaviour of two typical speakers with audio plus Ultrasound Tongue Imaging (201 frames per second) and de-interlaced NTSC video of the mouth and lips (60fps). On average, acoustic RT occurred between 120 to 180 ms later than a clearly observable articulatory movement, with no consistent advantage for lip or tongue-based measures.
This study examines pitch range production in the read speech of female German second language (L2) learners of English of moderate to advanced proficiency. The study set out to identify to what extent the learners deviated from or adopted the language-appropriate pitch range values of the target language. Two potential ways in which the learners could deviate from or approximate the target were recognized: (a) by globally expanding their pitch range or (b) by adjusting their pitch range in a position-sensitive way that is linked to the phonetic realization patterns of underlying high and low tones at different points in intonation contours. Results showed that the L2 speakers produced pitch range values that were often language appropriate or approximated the target, although some deviations from the target were also identified. Deviations and target approximation were found to be position sensitive; that is, L2 learners were found to adjust their pitch range differently at the beginning as compared to later parts of intonational phrases.
This paper presents a systematic comparison of various measures of f0 range in female speakers of English and German. F0 range was analyzed along two dimensions, level (i.e., overall f0 height) and span (extent of f0 modulation within a given speech sample). These were examined using two types of measures, one based on "long-term distributional" (LTD) methods, and the other based on specific landmarks in speech that are linguistic in nature ("linguistic" measures). The various methods were used to identify whether and on what basis or bases speakers of these two languages differ in f0 range. Findings yielded significant cross-language differences in both dimensions of f0 range, but effect sizes were found to be larger for span than for level, and for linguistic than for LTD measures. The linguistic measures also uncovered some differences between the two languages in how f0 range varies through an intonation contour. This helps shed light on the relation between intonational structure and f0 range.
This study examined whether rapid temporal auditory processing, verbal working memory capacity, non-verbal intelligence, executive functioning, musical ability and prior foreign language experience predicted how well native English speakers (N=120) discriminated Norwegian tonal and vowel contrasts as well as a non-speech analogue of the tonal contrast and a native vowel contrast presented over noise. Results confirmed a male advantage for temporal and tonal processing, and also revealed that temporal processing was associated with both non-verbal intelligence and speech processing. In contrast, effects of musical ability on non-native speech-sound processing and of inhibitory control on vowel discrimination were not mediated by temporal processing. These results suggest that individual differences in non-native speech-sound processing are to some extent determined by temporal auditory processing ability, in which males perform better, but are also determined by a host of other abilities that are deployed flexibly depending on the characteristics of the target sounds.
This study examined how speakers with ataxic dysarthria produce sentence stress and how these findings relate to other measures of speech performance. Ten speakers with ataxia and ten control speakers performed maximum performance, sentence stress, and passage reading tasks. Perceptual analyses established intelligibility levels and accuracy of stress production. Acoustic analyses included F-0, intensity, and duration measures for sentence stress targets and MPTs, as well as acoustic rhythm measures for the sentence and passage reading tasks. Results showed that 60% of speakers experienced problems in signalling sentence stress irrespective of the severity of their dysarthria. Intensity and duration were most impaired, with F-0 and pause insertion being used as compensatory strategies. The results highlighted the need for a detailed examination of speaker abilities in a variety of tasks in order to inform selection of the most effective treatment strategies.
While it is well known that languages have different phonemes and phonologies, there is growing interest in the idea that languages may also differ in their ‘phonetic setting’. The term ‘phonetic setting’ refers to a tendency to make the vocal apparatus employ a language-specific habitual configuration. For example, languages may differ in their degree of lip-rounding, tension of the lips and tongue, jaw position, phonation types, pitch range and register. Such phonetic specifications may be particularly difficult for second language (L2) learners to acquire, yet be easily perceivable by first language (L1) listeners as inappropriate. Techniques that are able to capture whether and how an L2 learner’s pronunciation proficiency in their two languages relates to the respective phonetic settings in each language should prove useful for second language research. This article gives an overview of a selection of techniques that can be used to investigate phonetic settings at the articulatory level, such as flesh-point tracking, ultrasound tongue imaging and electropalatography (EPG), as well as a selection of acoustic measures such as measures of pitch range, long-term average spectra and formants.