Tag
This paper introduces PERCEPT, the first large-scale Persian–English code-mixed corpus annotated with Universal Dependencies POS tags, collected from X, Instagram, and Digikala. It presents an LLM-assisted annotation framework and analyzes code-mixing patterns across platforms.
This paper proposes a resource-light algorithm to automatically assign part-of-speech tags to senses in the Al-Mawrid Arabic-English bilingual dictionary by transferring tags from English WordNet after disambiguation, achieving high accuracy with minimal cost.
This paper introduces Approximate Structured Diffusion, a method that combines conditional random fields (CRFs) with discrete diffusion for sequence labelling. It uses a CRF conditioned on noisy label sequences and approximate mean-field inference, achieving a 16.5% error reduction on POS tagging.