Progress in automatic sign language recognition and translation has been hindered by the scarcity of datasets available for the training of machine learning algorithms, a challenge that is even more acute for languages with smaller signing communities, such as Spanish. In this paper, we introduce a dataset of 300 isolated signs in Spanish Sign Language, collected online via a web application with contributions from 124 participants, resulting in a total of 8000 instances. This dataset, which is openly available, includes keypoints extracted using MediaPipe Holistic. The goal of this paper is to describe the construction and characteristics of the dataset and to provide a baseline classification method using a spatial–temporal graph convolutional network (ST-GCN) model, encouraging the scientific community to improve upon it. The experimental section offers a comparative analysis of the method’s performance on the new dataset, as well as on two other well-known datasets. The dataset, code, and web app used for data collection are freely available, and the web app can also be used to test classifier performance on-line in real-time.
Isolated Sign Language Recognition (ISLR) fits nicely in the domain of problems that can be handled by graph-structured spatial-temporal algorithms. A recent multi-scale spatial-temporal graph convolution operator, MS-G3D, takes advantage of the semantic connectivity among non-neighbor nodes of the graph in a flexible temporal scale, which results in improved performance in classical Human Action Recognition datasets. In this work, we present a solution for ISLR using a skeleton graph that includes body and finger joints and makes use of this specific property of MS-G3D, which seems crucial to capture the internal relationship among semantically connected distant nodes in sign language dynamics. To complete the analysis, we compare the results with a 3D-CNN architecture, S3D, already used for SLR, and fuse it with MS-G3D. The performance achieved on the AUTSL dataset shows that MS-G3D alone stands out as a viable technique for ISLR. In fact, the improvement after fusing with a 3D-CNN approach, at least on this medium-scale dataset, appears marginal. The transfer learning capability of the trained models is also explored using pre-training with the larger WLASL dataset and post-training with the smaller LSE_UVIGO dataset. The classification performance based on the MS-G3D model over AUTSL does not benefit from pre-training with WLASL, but the performance on the more similarly acquired LSE_UVIGO dataset improves significantly from fine-tuning the MS-G3D AUTSL model.
Poster presentado no XIII Congreso Internacional de Linguistica Xeral (CILX 2018), celebrado en Vigo do 13 ao 15 de xuno de 2018.
Depression is a common mental disorder that is usually addressed by outpatient treatments that favour patients’ inclusion in the society. This raises the need for tools to remotely monitor the emotional state of the patients, which can be carried out via telephone or the Internet using speech processing approaches. However, these strategies lead to privacy concerns caused by the transmission of the patients’ speech and its subsequent storage in servers. The use of speech de-identification to protect the privacy of these patients seems straightforward, but the influence of this procedure in the manifestation of the disease in the patients’ speech has not been addressed yet. Hence, this study evaluates the performance of an automatic depression level estimation system when dealing with original and de-identified speech, in order to analyse the influence of the de-identification procedure in the detection of depression. Two de-identification approaches based on voice transformation via frequency warping and amplitude scaling are assessed, which can be applied to any speaker without additional training. Experiments carried out in the framework of the audio/visual emotion challenge 2014 show that the proposed de-identification approaches achieve promising de-identification results at the expense of a slight degradation of depression detection.
Speaker de-identification approaches must accomplish three main goals: universality, naturalness and reversibility. The main drawback of the traditional approach to speaker de-identification using voice conversion techniques is its lack of universality, since a parallel corpus between the input and target speakers is necessary to train the conversion parameters. It is possible to make use of a synthetic target to overcome this issue, but this harms the naturalness of the resulting de-identified speech. Hence, a technique is proposed in this paper in which a pool of pre-trained transformations between a set of speakers is used as follows: given a new user to de-identify, its most similar speaker in this set of speakers is chosen as the source speaker, and the speaker that is the most dissimilar to the source speaker is chosen as the target speaker. Speaker similarity is measured using the i-vector paradigm, which is usually employed as an objective measure of speaker de-identification performance, leading to a system with high de-identification accuracy. The transformation method is based on frequency warping and amplitude scaling, in order to obtain natural sounding speech while masking the identity of the speaker. In addition, compared to other voice conversion approaches, the proposed method is easily reversible. Experiments were conducted on Albayzin database, and performance was evaluated in terms of objective and subjective measures. These results showed a high success when de-identifying speech, as well as a great naturalness of the transformed voices. In addition, when making the transformation parameters available to a trusted holder, it is possible to invert the de-identification procedure, hence recovering the original speaker identity. The computational cost of the proposed approach is small, making it possible to produce de-identified speech in real-time with a high level of naturalness.
The main drawback of speaker de-identification approaches using voice conversion techniques is the need for parallel corpora to train transformation functions between the source and target speakers. In this paper, a voice conversion approach that does not require training any parameters is proposed: it consists in manually defining frequency warping (FW) based transformations by using piecewise linear approximations. An analysis of the de-identification capabilities of the proposed approach using FW only or combined with FW modification and spectral amplitude scaling (AS) was performed. Experimental results show that, using the manually defined transformations using only FW, it is not possible to obtain de-identified natural sounding speech. Nevertheless, when modifying the FW, both de-identification accuracy and naturalness increase to a great extent. A slight improvement in de-identification was also obtained when applying spectral amplitude scaling.
In a previous work we presented a method to combine the acoustic characteristics of a speech synthesis model with the linguistic characteristics of another one. This paper presents a more extensive evaluation of the method when applied to cross-lingual adaptation. A large number of voices from a database in Spanish are adapted to Basque, Catalan, English and Galician. Using a state-of-the-art speaker identification system, we show that the proposed method captures the identity of the target speakers almost as well as standard intra-lingual adaptation techniques.
This paper presents a web-based multimedia search engine built within the Buceador (www.buceador.org) research project. A proof-of-concept tool has been implemented which is able to retrieve information from a digital library made of multimedia documents in the 4 official languages in Spain (Spanish, Basque, Catalan and Galician). The retrieved documents are presented in the user language after translation and dubbing (the four previous languages + English). The paper presents the tool functionality, the architecture, the digital library and provide some information about the technology involved in the fields of automatic speech recognition, statistical machine translation, text-to-speech synthesis and information retrieval. Each technology has been adapted to the purposes of the presented tool as well as to interact with the rest of the technologies involved.
In this article we describe HiFST, a lattice-based decoder for hierarchical phrase-based translation and alignment. The decoder is implemented with standard Weighted Finite-State Transducer (WFST) operations as an alternative to the well-known cube pruning procedure. We find that the use of WFSTs rather than k-best lists requires less pruning in translation search, resulting in fewer search errors, better parameter optimization, and improved translation performance. The direct generation of translation lattices in the target language can improve subsequent rescoring procedures, yielding further gains when applying long-span language models and Minimum Bayes Risk decoding. We also provide insights as to how to control the size of the search space defined by hierarchical rules. We show that shallow-n grammars, low-level rule catenation, and other search constraints can help to match the power of the translation system to specific language pairs.
In this paper we present results for the Europarl Spanish-to-English translation task. We use HiFST, a novel hierarchical phrase-based translation system implemented with finite-state technology that creates target lattices rather than k-best lists.
Comunicacio presentada a: Paper presented at: 4th Language and Technology Conference celebrat del 6 al 8 de novembre de 2009 a Poznan, Polonia.
This paper focuses on the relationship between intonation and syntactic and morphosyntactic information. Although intonation and syntax are both related to dependencies between the different parts of a sentence, and are therefore related to meaning, the precise influence of grammar on intonation is not clear. We describe a novel method that uses syntactic and part-of-speech features in the framework of corpus-based intonation modelling, and which integrates part of the phrasing algorithm in the unit selection stage. Subjective tests confirm an improvement in the quality of the resulting synthetic intonation: 75% of the sentences synthesised with the new intonation model were considered to be better or much better than the sentences synthesised using the old model, while only 7.5% of sentences were rated as worse or much worse.
We describe refinements to hierarchical translation search procedures intended to reduce both search errors and memory usage through modifications to hypothesis expansion in cube pruning and reductions in the size of the rule sets used in translation. Rules are put into syntactic classes based on the number of non-terminals and the pattern, and various filtering strategies are then applied to assess the impact on translation speed and quality. Results are reported on the 2008 NIST Arabic-to-English evaluation task.
In this paper we present results for the Europarl Spanish-to-English translation task. We use HiFST, a novel hierarchical phrase-based translation system implemented with finite-state technology that creates target lattices rather than k-best lists.
This paper describes the ECESS evaluation campaign of voice activity and voicing detection. Standard VAD classifies signal into speech and non-speech, we extend it to VAD+ so that it classifies a signal as a sequence of non-speech, voiced and unvoiced segments. The evaluation is performed on a portion of the Spanish SPEECON database with manually labeled segmentation. To avoid errors caused by the limited precision of manual labeling we introduce “dead zones” –tolerance intervals ±5 ms around label changes in the data set. In these tolerance intervals we don't evaluate the signal.
In this article, we present the main linguistic and phonetic features of Galician which need to be considered in the development of speech technology applications for this language. We also describe the solutions adopted in our text-to-speech system, also useful for speech recognition and speech-to-speech translation. On the phonetic plane in particular, the handling of vocal contact and the determination of mid-vowel openness are discussed. On the linguistic plane we place special emphasis on the handling of clitics and verbs. It should be noted that in Galician there is high interrelation between phonetics and grammatical information. Therefore, the task of morphosyntactic disambiguation is also addressed. Moreover, this task is fundamental for a higher level linguistic analysis.
This paper presents a novel approach for avoiding f0 dis- continuities in the framework of a unit selection model of into- nation with the accent group contour as the basic unit for con- catenation. Taking as input the resulting contour of the model, it is modified on an accent group basis according to the typica l decomposition into phrase and accent curves of the superposi- tional models. Results are presented that show a good perfor- mance of the proposed method, what implies that it can be a good approach for the treatment of the discontinuities, the main drawback of the concatenative models.