Tag
The study investigates whether hallucination detection probes trained on monolingual text transfer to Hinglish, finding robust performance and higher hallucination rates in Hindi and Hinglish compared to English.
This paper introduces PERCEPT, the first large-scale Persian–English code-mixed corpus annotated with Universal Dependencies POS tags, collected from X, Instagram, and Digikala. It presents an LLM-assisted annotation framework and analyzes code-mixing patterns across platforms.
Proposes a unified framework, LCF and LCFEdit, that jointly optimizes construction and injection of code-mixing fingerprints for LLMs using low-resource languages to achieve imperceptible and robust ownership verification.
Introduces Indi-RomCoM, a benchmark for evaluating LLMs on Romanized Code-Mixed (RCM) instructions in four Indic languages, finding that LLMs underperform on RCM tasks and performance degrades with higher code-mixing density.
Developer seeks advice on handling English-Hindi code-mixed text classification without heavy LLMs, as sentence transformers fail on Romanized Hindi.