Tag
This paper describes two models for vocabulary difficulty prediction: a black-box LLM fine-tuned with a soft-target loss achieving high accuracy, and an explainable model providing insights into difficulty factors. The models were part of the BEA 2026 Shared Task and achieve strong correlations.
This paper details the RETUYT-INCO team's participation in the BEA 2026 Shared Task 2, introducing a meta-prompting approach for rubric-based scoring of German short answers.
This paper presents a system for the EEUCA 2026 shared task on toxicity detection in gaming chat, achieving 4th place by fine-tuning Llama 3.1 8B with synthetic data augmentation. It highlights a 'validation trap' phenomenon where high validation scores do not correlate with test performance due to dataset distribution shifts.