Tag
A new annotated corpus of persuasion techniques in Bulgarian, Polish, and Russian, covering parliamentary debates and social media, with 25 fine-grained techniques and baseline models for detection and classification.
Analyzes filled pauses across 4 Slavic parliaments using transformer-based detection, replicating known effects (age, speech rate) and finding novel associations with sentiment, political orientation, and power status.
This paper measures tokenizer fertility across 25 European languages on parallel text, revealing a 2.5x spread from English to Greek/Maltese, with Ukrainian paying a 15-18% penalty. It demonstrates domain invariance of fertility rankings, analyzes subword fragmentation, and evaluates cross-lingual few-shot effects.