Tag
This paper introduces a pipeline that uses large language models to extract grammatical rules, example sentences, and lexicons from grammar books to generate synthetic parallel corpora for fine-tuning machine translation models, achieving ChrF++ gains of up to +8.8 on three low-resource languages.
Large language models can improve translation for low-resource languages through structured linguistic reasoning traces, with the most significant benefits occurring during inference rather than training.