Reading recommendations regarding computational tools for language documentation/description
Linguistic description and documentation involves a lot of data annotation and management. I am interested in how we can use computational tools to assist in this process, naturally only as an aid and not a replacement for human work. It is necessary to have humans in the loop and guards agains biases in tools that cause smaller languages to be misrepresented. At a basic level of data annotation, I have made posts in the past about my ELAN workflow for segmenting speakers into tiers , creating tiers from search results in ELAN and another way of doing segmentation in ELAN via PRAAT (kudos to Kashima and Ellis).
I think there's a common misconception that computational tools in language documentation/description is something very new, that before the dawn of LLMs there wasn't much done. That's not the case, researchers have long been interested in how to effective part of their labour - especially the language technology branch of the ARC Centre of Excellence for the Dynamics of Language.
Below are some reading recommendations, with special focus on Australian linguistics which I think is in the very forefront in this area. One of the things that make Australian linguistics so great in this respect is that they tend to work respectfully together with language communities and work with a wide diversity of languages.
Pre-November 2022 (release of ChatGPT)
- Gonzalez, S., Grama, J., & Travis, C. E. (2020). Comparing the performance of forced aligners used in sociophonetic research. Linguistics Vanguard, 6(1), 20190058. https://openresearch-repository.anu.edu.au/server/api/core/bitstreams/448cdabd-83d9-4d0b-80f3-595b178e4622/content
- Foley, B., Arnold, J. T., Coto-Solano, R., Durantin, G., Ellison, T. M., Van Esch, D., ... & Wiles, J. (2018, August). Building Speech Recognition Systems for Language Documentation: The CoEDL Endangered Language Pipeline and Inference System (ELPIS). In SLTU (pp. 205-209). https://www.researchgate.net/profile/Ben-Foley/publication/328067867_Building_Speech_Recognition_Systems_for_Language_Documentation_The_CoEDL_Endangered_Language_Pipeline_and_Inference_System_ELPIS/links/5beb91704585150b2bb4ebcb/Building-Speech-Recognition-Systems-for-Language-Documentation-The-CoEDL-Endangered-Language-Pipeline-and-Inference-System-ELPIS.pdf
- Adams, O., Galliot, B., Wisniewski, G., Lambourne, N., Foley, B., Sanders-Dwyer, R., ... & Hill, N. (2021, March). User-friendly automatic transcription of low-resource languages: Plugging ESPnet into ELPIS. In Proceedings of the 4th Workshop on the Use of Computational Methods in the Study of Endangered Languages Volume 1 (Papers) (pp. 51-62). https://aclanthology.org/2021.computel-1.7.pdf
- San, N., Bartelds, M., Browne, M., Clifford, L., Gibson, F., Mansfield, J., ... & Jurafsky, D. (2021, December). Leveraging pre-trained representations to improve access to untranscribed speech from endangered languages. In 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) (pp. 1094-1101). IEEE. https://arxiv.org/abs/2103.14583
- Michaud, A., Adams, O., Cohn, T. A., Neubig, G., & Guillaume, S. (2018). Integrating automatic transcription into the language documentation workflow: Experiments with Na data and the Persephone toolkit. https://scholarspace.manoa.hawaii.edu/server/api/core/bitstreams/a8ad6c4b-6dd1-4eee-a65c-01c7193a6659/content
Post-November 2022
- Mahmudi, Aso and Dang, Ting and Vylomova, Ekaterina and Thieberger, Nick (2026) Easper: An Accessible ASR Pipeline for Language Documentation. Interspeech 2026 https://arxiv.org/pdf/2608.11629
Also, several exciting things at this years langdoc conference, see book of abstracts here.
Comments
Post a Comment