Tag
This paper presents Easper, an open-source no-code ASR pipeline that lets field linguists fine-tune models like Whisper from ELAN annotations, and evaluates data selection strategies for bootstrapping ASR on low-resource Vanuatu languages.
This paper introduces a pipeline that uses large language models to extract grammatical rules, example sentences, and lexicons from grammar books to generate synthetic parallel corpora for fine-tuning machine translation models, achieving ChrF++ gains of up to +8.8 on three low-resource languages.
GlossAssist is a tool for creating interlinear glossed text (IGT) corpora in low-resource language documentation settings, built around the CWoMP retrieval-based architecture with an active learning feedback loop that improves predictions as annotators make corrections without retraining the model.