BOUTEF: A Multilingual Corpus for FakeNews in North Africa -- Language as a Weapon
Summary
This paper introduces BOUTEF, a large-scale multilingual corpus for studying fake news in Algeria and Tunisia, covering Arabic dialects, Arabizi, French, English, and code-switching. It includes empirical analysis of linguistic strategies and engagement dynamics.
View Cached Full Text
Cached at: 06/02/26, 03:36 PM
# BOUTEF: A Multilingual Corpus for FakeNews in North Africa -- Language as a Weapon Source: [https://arxiv.org/abs/2606.00193](https://arxiv.org/abs/2606.00193) [View PDF](https://arxiv.org/pdf/2606.00193) > Abstract:The rapid spread of fake news on social media has become a major challenge, particularly in multilingual and under\-resourced contexts such as North Africa\. In this paper, we introduce BOUTEF, a large\-scale multilingual corpus designed to study the propagation, characteristics, and impact of fake news in Algeria and Tunisia\. The corpus integrates three complementary components: fake narratives, genuine narratives, and associated user\-generated comments, along with verified debunking information\. It covers a wide range of languages and linguistic varieties, including MSA, Algerian and Tunisian dialects, Arabizi, French, English, and code\-switched language\. Building on this resource, we conduct a comprehensive empirical analysis combining quantitative and qualitative approaches\. We examine thematic distributions, linguistic and rhetorical strategies, sentiment patterns, and social engagement dynamics\. Statistical analyses reveal significant associations between thematic categories and message veracity, as well as strong correlations between user engagement and the visibility of fake content\. Our findings show that fake news relies heavily on emotionally charged narratives, sensational framing, and hybrid linguistic practices that enhance virality and audience engagement\. In contrast, debunking content adopts a more factual and verification\-oriented style\. Furthermore, a comparative analysis between Algeria and Tunisia highlights both shared dynamics and country\-specific characteristics shaped by sociopolitical contexts\. The results emphasize the role of informal language practices in the diffusion and reception of misinformation\. By providing a rich, annotated, and publicly available dataset, this work contributes to advancing research on fake news detection, low\-resource language processing, and the understanding of information disorders in complex linguistic environments\. ## Submission history From: Amina Laggoun \[[view email](https://arxiv.org/show-email/39c4ce18/2606.00193)\] **\[v1\]**Fri, 29 May 2026 16:27:47 UTC \(2,062 KB\)
Similar Articles
An End-to-End Hybrid Framework for Rumour Detection in Low-Resources Algerian Dialect
This paper presents an end-to-end hybrid framework for rumour detection in low-resource Algerian dialect social media content, achieving an F1-score of 0.84 by combining transformer embeddings with a classical classifier.
Echoes of Unrest: A Multimodal NLP Framework for Early Warning of Fake News and Violence-Driven Mob Activity
This paper presents a multimodal NLP framework that fuses XLM-RoBERTa and CLIP with geospatial and sarcasm features to detect fake news and predict violence-driven mob activity, achieving 98% test accuracy on a 138,256-sample Bangla/English dataset.
I built an open, from-scratch MT pipeline + parallel corpus for Tunisian Darija (Arabizi) early baseline, and I'm growing it into a curated community corpus [P]
An 18-year-old Tunisian student introduces an open-source machine translation pipeline and parallel corpus for Tunisian Darija in Arabizi script, built from scratch with a small 15.6M-parameter Transformer and an honest baseline BLEU of 3.89, and calls for contributors to ethically expand the corpus.
Dziri Voicebot: An End-to-End Low-Resource Speech-to-Speech Conversational System for Algerian Dialect
This paper presents a modular end-to-end speech-to-speech conversational system for the low-resource Algerian Dialect, integrating ASR, NLU, RAG, and TTS with dedicated datasets and fine-tuned models.
Audience Engagement with Arabic Women's Social Empowerment and Wellbeing: A Decadal Corpus
This paper presents the Arabic Women and Society Corpus, a ten-year collection of over 250,000 Arabic Facebook posts related to women's empowerment and social wellbeing, with engagement metrics for analyzing gender discourse and sentiment.