corpus

Tag

Cards List
#corpus

BOUTEF: A Multilingual Corpus for FakeNews in North Africa -- Language as a Weapon

arXiv cs.CL · 2026-06-02 Cached

This paper introduces BOUTEF, a large-scale multilingual corpus for studying fake news in Algeria and Tunisia, covering Arabic dialects, Arabizi, French, English, and code-switching. It includes empirical analysis of linguistic strategies and engagement dynamics.

0 favorites 0 likes
#corpus

Audience Engagement with Arabic Women's Social Empowerment and Wellbeing: A Decadal Corpus

arXiv cs.CL · 2026-05-22 Cached

This paper presents the Arabic Women and Society Corpus, a ten-year collection of over 250,000 Arabic Facebook posts related to women's empowerment and social wellbeing, with engagement metrics for analyzing gender discourse and sentiment.

0 favorites 0 likes
#corpus

ArabDiscrim: A Decade-Long Arabic Facebook Corpus on Racism and Discrimination

arXiv cs.CL · 2026-05-22 Cached

ArabDiscrim is a decade-long lexical resource and corpus of 293K Arabic Facebook posts about racism and discrimination, with engagement signals, morphological regex families, and discrimination axes, supporting fairness-oriented Arabic NLP research.

0 favorites 0 likes
#corpus

Enhancing Scientific Discourse: Machine Translation for the Scientific Domain

arXiv cs.CL · 2026-05-21 Cached

This paper presents the development of parallel and monolingual corpora for scientific machine translation across Spanish-English, French-English, and Portuguese-English, targeting four domains: Cancer Research, Energy Research, Neuroscience, and Transportation. The corpora are used to fine-tune neural machine translation systems, addressing challenges of specialized vocabulary and syntax in scientific text.

0 favorites 0 likes
#corpus

Released a free 9.8M doc Indic multilingual corpus — Hindi, Bengali, Tamil, Telugu + 7 more (CC0, HuggingFace) [P]

Reddit r/MachineLearning · 2026-05-18

Released a free 9.8 million document multilingual Indic corpus (11 languages, CC0 license) on HuggingFace, containing approximately 8.4 billion tokens, built for multilingual research.

0 favorites 0 likes
← Previous
← Back to home

Submit Feedback