Data-Driven Parametric Text Normalization: Rapidly Scaling Finite-State Transduction Verbalizers to New Languages
Sandy Ritchie, Eoin Mahon, Kim Heiligenstein, Nikos Bampounis, Daan van Esch, Christian Schallhart, Jonas Mortensen, Benoit Brard
Abstract
This paper presents a methodology for rapidly generating FST-based verbalizers for ASR and TTS systems by efficiently sourcing language-specific data. We describe a questionnaire which collects the necessary data to bootstrap the number grammar induction system and parameterize the verbalizer templates described in Ritchie et al. (2019), and a machine-readable data store which allows the data collected through the questionnaire to be supplemented by additional data from other sources. This system allows us to rapidly scale technologies such as ASR and TTS to more languages, including low-resource languages.- Anthology ID:
- 2020.sltu-1.30
- Volume:
- Proceedings of the 1st Joint Workshop on Spoken Language Technologies for Under-resourced languages (SLTU) and Collaboration and Computing for Under-Resourced Languages (CCURL)
- Month:
- May
- Year:
- 2020
- Address:
- Marseille, France
- Venues:
- LREC | SLTU | WS
- SIG:
- Publisher:
- European Language Resources association
- Note:
- Pages:
- 218–225
- URL:
- https://www.aclweb.org/anthology/2020.sltu-1.30
- DOI:
- PDF:
- https://www.aclweb.org/anthology/2020.sltu-1.30.pdf
You can write comments here (and agree to place them under CC-by). They are not guaranteed to stay and there is no e-mail functionality.