Reading recommendations regarding computational tools for language documentation/description

Linguistic description and documentation involves a lot of data annotation and management. I am interested in how we can use computational tools to assist in this process, naturally only as an aid and not a replacement for human work. It is necessary to have humans in the loop and guards agains biases in tools that cause smaller languages to be misrepresented. At a basic level of data annotation, I have made posts in the past about my ELAN workflow for segmenting speakers into tiers , creating tiers from search results in ELAN and another way of doing segmentation in ELAN via PRAAT (kudos to Kashima and Ellis).

I think there's a common misconception that computational tools in language documentation/description is something very new, that before the dawn of LLMs there wasn't much done. That's not the case, researchers have long been interested in how to effective part of their labour - especially the language technology branch of the ARC Centre of Excellence for the Dynamics of Language.

Below are some reading recommendations, with special focus on Australian linguistics which I think is in the very forefront in this area. One of the things that make Australian linguistics so great in this respect is that they tend to work respectfully together with language communities and work with a wide diversity of languages.

Pre-November 2022 (release of ChatGPT)
Post-November 2022
  • Mahmudi, Aso and Dang, Ting and Vylomova, Ekaterina and Thieberger, Nick (2026) Easper: An Accessible ASR Pipeline for Language Documentation. Interspeech 2026 https://arxiv.org/pdf/2608.11629

Also, several exciting things at this years langdoc conference, see book of abstracts here.

Comments

Popular posts from this blog

That infographic on languages of the world - some context to help you understand what's going on

A Global Tree of Languages

Language family maps