Tag
This paper presents the first systematic study of word segmentation for the extinct Tangut language, combining traditional lexicons, unlabeled text, and a pretrained character encoder to achieve high performance despite extreme resource scarcity.
TRACE-BN introduces a curriculum-guided dataset for structured Bangla-English tutoring and demonstrates transferring this behavior to a sub-1B language model using LoRA, achieving significant improvements in tutoring quality for resource-constrained offline environments.
Introduces CKTN, the first multilingual corpus and benchmark for Cham, Khmer, and Tay-Nung languages, addressing fragmentation in multilingual encoders with a script-aware adaptation recipe.
This paper audits license provenance of over twenty African NLP corpus families, identifies compatibility failures like the JW300 violation and hidden NoDerivs clauses, and provides a due diligence checklist for legally clean dataset creation.
This paper examines the effect of labeled data size on natural language inference performance for 16 African languages using the AfriXNLI benchmark. The results show that scaling behavior is language-sensitive and often non-monotonic, challenging the common assumption of monotonic improvement, and emphasizing the need for language-specific dataset creation and stronger multilingual strategies.