This paper presents a methodology for rapidly generating FST-based verbalizers for ASR and TTS systems by efficiently sourcing language-specific data. We describe a questionnaire which collects the necessary data to bootstrap the number grammar induction system and parameterize the verbalizer templates described in Ritchie et al. (2019), and a machine-readable data store which allows the data collected through the questionnaire to be supplemented by additional data from other sources. This system allows us to rapidly scale technologies such as ASR and TTS to more languages, including low-resource languages.
We describe a new approach to converting written tokens to their spoken form, which can be shared by automatic speech recognition (ASR) and text-to-speech synthesis (TTS) systems. Both ASR and TTS need to map from the written to the spoken domain, and we present an approach that enables us to share verbalization grammars between the two systems while exploiting linguistic commonalities to provide simple default verbalizations. We also describe improvements to an induction system for number names grammars. Between these shared ASR/TTS verbalizers and the improved induction system for number names grammars, we achieve significant gains in development time and scalability across languages.
Cet article présente un dispositif de calcul de la longueur des vers dans les textes versifiés en langue française. Pour compter les syllabes, il néglige leurs frontières et ne prend en considération que leur noyau vocalique, qu'il projette directement sur la graphie selon un processus en deux temps. Un système expert de moins de 50 règles, fondées linguistiquement sur la notion de graphème , segmente chaque séquence de voyelles graphiques et se prononce sur sa participation dans le décompte de la longueur, à l'exception des cas difficiles dont l'ambiguïté est conservée. Une heuristique tire alors parti des longueurs décrites par l'étape précédente comme des intervalles et des régularités contextuelles des textes versifiés pour établir la bonne mesure de chaque vers. Mots-clés : versification, mètre, syllabisme, décompte syllabique, noyau vocalique
This paper describes a device for computing the length of lines in French versified texts. Being a syllable counter, it does not take the syllabic boundaries into account but considers their vocalic nucleus, which is mapped onto the orthographic string in two steps. By using grapheme features, a rule-based expert system of less than 50 items breaks up sequences of graphic vowels and states how the resulting pieces should be counted, except for ambivalent exceptions which are left unresolved. Then a heuristic chooses the best length for each line based on the intervals produced by the previous process and the regular nature of metrical patterns.