BOUTEF: A Multilingual Corpus for FakeNews in North Africa -- Language as a Weapon

arXiv cs.CL Papers

Summary

This paper introduces BOUTEF, a large-scale multilingual corpus for studying fake news in Algeria and Tunisia, covering Arabic dialects, Arabizi, French, English, and code-switching. It includes empirical analysis of linguistic strategies and engagement dynamics.

arXiv:2606.00193v1 Announce Type: new Abstract: The rapid spread of fake news on social media has become a major challenge, particularly in multilingual and under-resourced contexts such as North Africa. In this paper, we introduce BOUTEF, a large-scale multilingual corpus designed to study the propagation, characteristics, and impact of fake news in Algeria and Tunisia. The corpus integrates three complementary components: fake narratives, genuine narratives, and associated user-generated comments, along with verified debunking information. It covers a wide range of languages and linguistic varieties, including MSA, Algerian and Tunisian dialects, Arabizi, French, English, and code-switched language. Building on this resource, we conduct a comprehensive empirical analysis combining quantitative and qualitative approaches. We examine thematic distributions, linguistic and rhetorical strategies, sentiment patterns, and social engagement dynamics. Statistical analyses reveal significant associations between thematic categories and message veracity, as well as strong correlations between user engagement and the visibility of fake content. Our findings show that fake news relies heavily on emotionally charged narratives, sensational framing, and hybrid linguistic practices that enhance virality and audience engagement. In contrast, debunking content adopts a more factual and verification-oriented style. Furthermore, a comparative analysis between Algeria and Tunisia highlights both shared dynamics and country-specific characteristics shaped by sociopolitical contexts. The results emphasize the role of informal language practices in the diffusion and reception of misinformation. By providing a rich, annotated, and publicly available dataset, this work contributes to advancing research on fake news detection, low-resource language processing, and the understanding of information disorders in complex linguistic environments.
Original Article
View Cached Full Text

Cached at: 06/02/26, 03:36 PM

# BOUTEF: A Multilingual Corpus for FakeNews in North Africa -- Language as a Weapon
Source: [https://arxiv.org/abs/2606.00193](https://arxiv.org/abs/2606.00193)
[View PDF](https://arxiv.org/pdf/2606.00193)

> Abstract:The rapid spread of fake news on social media has become a major challenge, particularly in multilingual and under\-resourced contexts such as North Africa\. In this paper, we introduce BOUTEF, a large\-scale multilingual corpus designed to study the propagation, characteristics, and impact of fake news in Algeria and Tunisia\. The corpus integrates three complementary components: fake narratives, genuine narratives, and associated user\-generated comments, along with verified debunking information\. It covers a wide range of languages and linguistic varieties, including MSA, Algerian and Tunisian dialects, Arabizi, French, English, and code\-switched language\. Building on this resource, we conduct a comprehensive empirical analysis combining quantitative and qualitative approaches\. We examine thematic distributions, linguistic and rhetorical strategies, sentiment patterns, and social engagement dynamics\. Statistical analyses reveal significant associations between thematic categories and message veracity, as well as strong correlations between user engagement and the visibility of fake content\. Our findings show that fake news relies heavily on emotionally charged narratives, sensational framing, and hybrid linguistic practices that enhance virality and audience engagement\. In contrast, debunking content adopts a more factual and verification\-oriented style\. Furthermore, a comparative analysis between Algeria and Tunisia highlights both shared dynamics and country\-specific characteristics shaped by sociopolitical contexts\. The results emphasize the role of informal language practices in the diffusion and reception of misinformation\. By providing a rich, annotated, and publicly available dataset, this work contributes to advancing research on fake news detection, low\-resource language processing, and the understanding of information disorders in complex linguistic environments\.

## Submission history

From: Amina Laggoun \[[view email](https://arxiv.org/show-email/39c4ce18/2606.00193)\] **\[v1\]**Fri, 29 May 2026 16:27:47 UTC \(2,062 KB\)

Similar Articles