Active Perception perspectives claim that action is closely related to perception. An empirical approach that supports these theories is the minimalist, in which participants perform a task using an interface that provides minimal information. Their exploratory movements are crucial to generating a meaningful sequence of information. Previous studies analyzed sensorimotor trajectories describing qualitative strategies and linear quantification of participants' movement performance, but that approach struggles to capture the behavior of non-stationary data. In the present study, we applied the recurrence plot (RP) and recurrence quantification analysis (RQA) to study the structure of sensorimotor trajectories developed by participants trying to discriminate between two invisible geometric shapes (Triangle or Rectangle). The exploratory movements were made using a computer mouse and sonification-mediated feedback was provided, which depended exclusively on whether the pointer was inside or outside the shape. We applied RP and RQA to the sensorimotor trajectories, with the aim of studying their fine structure characteristics, focusing on their repetitive patterns. Recurrence analysis proved to be useful for quantifying differences in dynamic behavior that emerge when participants explore invisible virtual geometric shapes. The differences obtained in RQA-based measures associated with the vertical structures allowed to postulate the existence of particular exploration strategies for each figure. It was also possible to determine that the complexity of the dynamics changed according to the shape. We discuss these results in light of antecedents in haptic and visual perceptual exploration.
La restauración de la audición es posible mediante el uso de audífonos, que son dispositivos de tecnología asistiva que amplifican los sonidos para compensar la disminución de la sensibilidad auditiva. Sin embargo, la calidad sonora que proporcionan actualmente es subóptima. Por otro lado, la hipoacusia degrada las habilidades de localización espacial de los sonidos y la conservación de todas las claves de audición espacial es uno de los principales problemas de los audífonos actuales.En este trabajo se presenta un plan de investigación y desarrollo que busca realizar aportes al estado del arte en aspectos aún no resueltos en los audífonos digitales actuales, como son, la mejora de la inteligibilidad del habla en ambientes sonoros complejos y la localización espacial de fuentes sonoras junto con la mejora de la naturalidad auditiva. Para ello se propone desarrollar un modelo de audición binaural para audífonos que incluya algoritmos de realce del habla y clasificación de escenas sonoras, e incorpore una función de corrección individualizada dependiente de la dirección de incidencia de los sonidos.
The COVID-19 pandemic has significantly modified the behavior of societies. The application of isolation measures during the crisis resulted in changes in the acoustic environment. The aim of this work was to characterize the perception of the acoustic environment during the COVID-19 lockdown of people residing in Argentina in 2020. A descriptive cross-sectional correlational study was carried out. A virtual survey was conducted from April 14 to 26, 2020, and was answered mainly by social network users. During this period, Argentina was in a strict lockdown. The sample was finally composed of 1371 people between 18 and 79 years old. It was observed that most of the participants preferred the new acoustic environment. Mainly in the larger cities, before the isolation, mechanical sounds predominated, accompanied by the perception of irritation. Confinement brought a decrease in mechanical sounds and an increase in biological sounds, associated with feelings of tranquility and happiness. The time window opened by the lockdown offered an interesting scenario to assess the effect of anthropogenic noise pollution on the urban environment. This result offers a subjective approach, which contributes to understanding the link between individuals and communities with the environment.
This paper addresses the issue of local disturbances in the fundamental frequency contour of speech, caused by the articulation of voiced/unvoiced consonant phonemes. Depending on the intended use of the F0 contour, these disturbances are usually eliminated by a filtering, smoothing or stylization procedure. These procedures that seek to preserve only the F0 points perceptually relevant, are generally applied roughly at a global level, which may not completely eliminate micro intonation in some cases or distort macro intonation in others. In this work we propose a local filtering algorithm based on a fine level analysis of the microprosodic morphologies. The performance of the algorithm is validated by a perceptual experiment. Assuming the algorithm allows partial/total disturbance elimination, we perform a statistical description of the perturbation morphologies. Statistics were collected from a corpus of 741 sentences designed to study Argentine Spanish prosody. The corpus was recorded by four professional announcers native speakers from Buenos Aires city. The results show that perturbation morphologies are affected by: consonant phoneme identity; global F0 contour shape; and speaker identity. As an application case, we use the proposed filtering algorithm as a pre-processing stage in our automatic prominent syllable detection system, with a statistically significant improvement in its performance.
The goal is to determine using acoustic measurements, which information is the most relevant to listeners at the time of categorizing the overall degree of dysphonia. Eight voice signals were chosen (4 female voices and 4 male voices). Each voice was perceptually evaluated through the item G of GRBAS scale by 10 experienced listeners and acoustically by aperiodicity, noise and chaos measures. The statistical study by discriminant analysis shows the importance of GNE, Jit and Lyapunov Jitter_cc as parameters and predictors of overall degree of dysphonia. The application of the k-means evidence there are features in the acoustic parameters that allow us to objectively group the voices studied with 100% accuracy for class 0,96% for class 2 and 79% for class 3. A greater number and variability of cases are need to verify those preliminary results.
Bernd J. Kroger, Jim Kannampuzha, Dominik Bauer, Peter Birkholz, Philippe Dreuw, Hermann Ney An Action-Based Concept for the Phonetic Annotation of Sign Language Gestures 33 Sascha Fagel, Gérard Bailly Speech, Gaze and Head Motion in a Face-to-Face Collaborative Task 40 Ralf Winkler, Gunter Uhlmann, Gerd Schneider Maschinelle Klassifikation von Artikulationsbewegungen im Rahmen einer visuellen Artikulationsschulung für gehörlose und schwerhöriger Kinder 48 Prosody and Affect Benjamin Weiss, Sebastian Möller, Tim Polzehl Wirkung menschlicher Stimme auf die wahrgenommene SympathieEinfluss der Stimmanregung anhand von Laryngogrammen 56 Jürgen Trouvain Affektäußerungen in Sprachkorpora 64 Sören Wittenberg, Oliver Jokisch Das Prosodisch-Phonetische Annotationssystem PROPHANO 71
This paper introduces Emilia, a speech corpus created to build a female voice in Spanish spoken in Buenos Aires for the Aromo text-to-speech system. Aromo is a unit selection text-to-speech system, which employs diphones as units of synthesis. The key requirements and design criteria for Emilia were: to synthesize any text in Spanish into high-quality speech with a minimum corpus size. The text corpus was designed to guarantee the phonetic and prosodic coverage. A three-stage strategy was used: in the first stage, 741 sentences were designed with all of the syllables of Spanish spoken in Argentina, with and without stress, and in all positions within the word; in the second stage, 852 sentences were added to balance out the distribution of the diphones; and after a perceptual evaluation of the quality of synthesized speech, in the third and final stage, 625 sentences were added to achieve the specified unit coverage, and to introduce sentences with more complex syntactic and prosodic structures. Issues from all three corpus building stages are reported. The paper also presents the results from the quality perceptual evaluations of the synthesized voice. Emilia has a duration of three hours and 15 minutes; its speech quality synthesized with Aromo system is similar to the level obtained with commercial systems, with a real-time ratio less than one.
Prominence is a perceptual attribute employed to communicate focus, contrasts and expressive nuances. This article ex-plores the automatic detection of segments considered prominent by native listeners, using a corpus of Argentinean Spanish. The prominence detection is modeled as a binary classification problem over syllabic units. From perceptual assessments by a group of native listeners, we obtained a set of prominent syllable annotations, which are used as the gold standard to train and evaluate automatic classifiers. We study the performance of the classifiers under different sets of acoustic features, under various combinations of syllabic contexts, and using different classification algorithms. The best overall performance using leave-one speaker out cross validation had a mean precision rate of 94.75%, and was obtained using an SVM classifier, with two context syllables around each side of the central syllable, and applying the complete set of acoustic features considered.
This paper describes the implementation of a real-time, speaker-independent isolated speech recognition system using Hidden Markov Models on a 32-bit ARM Cortex-M4F microcontroller. We introduce the theory, the requirements and details of the embedded implementation. The evaluation was made using a multi-speaker isolated-digit corpus of Argentinian Spanish, and its performance in terms of accuracy, speed and required memory was compared against a baseline Dynamic Time Warping recognizer, implemented previously on the same architecture. Test results show that the proposed HMM system outperforms the baseline system, and exhibits a recognition accuracy of 96.21% under a clean acoustic environment.
This paper describes the development and validation of an Embedded Isolated Word Recognition System (IWR) for the Argentinian Spanish language, implemented on the STM32F4-Discovery platform. Its front-end extracts Mel Frequency Cepstral Coefficients (MFCC), while its classification step is based on the Dynamic Time Warping (DTW) algorithm. Since the system was conceived as a base platform for the research and development of speech based command and control applications, it was designed to be modular and to meet real-time performance. The system includes a Real Time Operating System (RTOS) to manage various processing and control tasks, which can be easily reconfigured with different acquisition, processing and recognition parameters using a single file. The validation was done using a scenario of robotic control, achieving performance rates which demonstrates the practical usefulness of the system.
The segmentation of anatomical and pathological structures plays a key role in the characterization of clinically relevant evidence from digital images. Recently, plenoptic imaging has emerged as a new promise to enrich the diagnostic potential of conventional photography. Since the plenoptic images comprises a set of slightly different versions of the target scene, we propose to make use of those images to improve the segmentation quality in relation to the scenario of a single image segmentation. The problem of finding a segmentation solution from multiple images of a single scene, is called segmentation fusion. This paper reviews the issue of segmentation fusion in order to find solutions that can be applied to plenoptic images, particularly images from the ophthalmological domain.
Fil: Guirao, Miguelina. Consejo Nacional de Investigaciones Cientificas y Tecnicas. Oficina de Coordinacion Administrativa Houssay. Instituto de Inmunologia, Genetica y Metabolismo. Universidad de Buenos Aires. Facultad de Medicina. Instituto de Inmunologia, Genetica y Metabolismo; Argentina. Universidad de Buenos Aires. Facultad de Medicina. Hospital de Clinicas General San Martin; Argentina
The unbalanced of data is a common problem in many domains. Using unbalanced data for standard machine learning classifiers significantly affect the obtained performance. In this paper is presented a description of the problem and a review of the main alternatives to solve it. It is also proposed an alternative model, illustrating its application through a case of the medical field. The proposed model manages to get a balanced distribution of instances per class. It is based on the automatic selection of a subset of cases from the majority classes, using the natural groupings of these classes through self-organizing maps. The model is applied to the recognition of heartbeat types and the results are compared with others methods. The results show the feasibility of using this model to address this problem.
This paper explores the relationship between perceived syllable prominence and the acoustic properties of a speech utterance. It is aimed at establishing a link between the linguistic meaning of an utterance in terms of sentence modality and focus and its underlying prosodic features. Applications of such knowledge can be found in computer-based pronunciation training as well as general automatic speech recognition and understanding. Our acoustic analysis confirms earlier results in that focus and sentence mode modify the fundamental frequency contour, syllabic durations and intensity. However, we could not find consistent differences between utterances produced with noncontrastive and contrastive focus, respectively. Only one third of utterances with broad focus were identified as such. Ratings of syllable prominence are strongly correlated with the amplitude Aa of underlying accent commands, syllable duration, maximum intensity and mean harmonics-to-noise ratio.
This paper describe a monitoring model for intelligent systems in the assessment tasks of patient's critical states. The model was designed according to temporal abstraction dimensions and state interpretation levels, particularly the monitored physiological state. Monitoring arrhythmias was used as the reference case for model development. This model is based on the Smith-Waterman algorithm, with a modified substitution matrix. It was developed for the state space assessment model, for said application. Matrix modification was oriented for giving more influence to the principal state for the temporal pattern represented. The performance of the state interpretation model was similar than other models reported in the literature. The state space assessment model allow to get extra temporal information about a possible evolution, not previously available in classic monitors.
This paper gives an integrated view of developing a model for an assessment systems and monitoring tasks of critical patients. This model was designed according to the Temporal Abstraction dimensions and state interpretation levels, particularly the monitored physiological state. Monitoring arrhythmias was used as reference case for models development. A model based on the algorithm Smith-Waterman with the matrix substitution modified, was developed for the assessment state interpretation model, for this application. The modification was oriented to give more influence to the principal state for the temporal pattern represented. The assessment state space model, allow to obtain additional temporal information about a possible evolution, not previously available in classic monitors.
The effect of ethanol in modulating the intensity and duration of the perceived sourness induced by citric acid was studied. Magnitude Estimation-Converging Limits method was applied to rate the sourness of seven solutions (3–70 mM) of citric acid in aqueous solution presented alone and mixed with 8% V/V or 15% V/V ethanol. Dynamic sourness ratings of 5, 15, and 45 mM citric acid alone and mixed with the same two ethanol levels were assessed by the Time Intensity Method (TI). Results were consistent with both methods. Sourness changed with citric acid concentration and ethanol levels. From TI measurements, a similar interactive pattern was obtained for parameters as duration, area under the curve, peak and average intensity.