Tag
The paper presents KSE-Web, an analysis of hybrid retrieval and LLM-assisted query expansion for low-resource Khmer semantic search, using a constructed dataset to evaluate BM25, dense, and hybrid methods, finding that BM25 performs best while LLM expansion has limitations for Khmer.
Introduces CKTN, the first multilingual corpus and benchmark for Cham, Khmer, and Tay-Nung languages, addressing fragmentation in multilingual encoders with a script-aware adaptation recipe.
This paper investigates LoRA fine-tuning of the VoxCPM2 TTS model to improve quality for low-resource languages like Khmer, while showing no gain for Korean which the base model already handles well. The adapter yields significant MOS improvement for Khmer with minimal parameter training.
This paper presents a comparative evaluation of embedding models and generator backends for Khmer-language retrieval-augmented question answering in the telecom domain, finding that BGE-M3 performs best for retrieval while generator strengths vary across metrics.