hate-speech-detection

Tag

Cards List
#hate-speech-detection

When Explanations Cannot Be Read: Measuring and Correcting SHAP and LIME Rendering for Right-to-Left Languages

arXiv cs.LG ↗ · 2d ago Cached

The study introduces SHAP-RTL, a rendering layer that corrects the visualization of SHAP and LIME explanations for right-to-left languages, addressing issues like token sequence and script shaping while preserving original attribution values.

0 favorites 0 likes
#hate-speech-detection

An Explainable DistilBERT-BiLSTM-Attention Framework for Binary and Multi-Class Hate Speech Detection

arXiv cs.CL ↗ · 2d ago Cached

This paper proposes an explainable hate speech detection framework integrating DistilBERT embeddings, BiLSTM, and an attention mechanism, achieving high F1-scores on benchmark datasets for both binary and multi-class classification.

0 favorites 0 likes
#hate-speech-detection

ParsHate: A Benchmark Dataset for Hate and Target Detection in Persian

arXiv cs.CL ↗ · 2026-09-16 Cached

Introduces ParsHate, a manually annotated dataset of 10,000 Persian tweets for hate speech and target detection, showing moderate performance of current models and emphasizing the need for more advanced methods.

0 favorites 0 likes
#hate-speech-detection

Beyond Accuracy: A Qualitative Analysis of Vision-Language Models for Hate Speech Detection in Memes

arXiv cs.CL ↗ · 2026-08-28 Cached

This paper presents a qualitative analysis of vision-language models for detecting hate speech in memes, evaluating their performance and reasoning under zero-shot and few-shot prompting.

0 favorites 0 likes
#hate-speech-detection

Training-Time Explainability for Multilingual Hate Speech Detection: Aligning Model Reasoning with Human Rationales

arXiv cs.CL ↗ · 2026-08-28 Cached

This paper proposes a training-time explainability framework for multilingual hate speech detection, aligning model reasoning with human rationales to improve classification performance and interpretability, evaluated on English and Hinglish datasets.

0 favorites 0 likes
#hate-speech-detection

From Specialization to Generalization: Instruction-tuned LLMs for Robust Harmful Content Mitigation

arXiv cs.CL ↗ · 2026-08-27 Cached

This paper investigates instruction-tuning general-purpose LLMs for robust harmful content mitigation, specifically hate speech detection, using a unified corpus of 36 datasets, achieving state-of-the-art performance and enhanced cross-domain and cross-lingual generalization.

0 favorites 0 likes
#hate-speech-detection

Efficient Adaptation of LLMs for Hate Speech Detection in Low-Resource Languages: A Comparative Study on Roman Urdu

arXiv cs.AI ↗ · 2026-08-20 Cached

The paper evaluates Large Language Models for hate speech detection in Roman Urdu, a low-resource language, demonstrating that Parameter-Efficient Fine-Tuning with LoRA significantly improves classification performance compared to zero-shot inference.

0 favorites 0 likes
#hate-speech-detection

ZeroR@CHiPSAL 2026: Two-Stage Vision-Language Adaptation with Contrastive Learning for Nepali Meme Classification

arXiv cs.CL ↗ · 2026-08-03 Cached

This paper describes a two-stage vision-language adaptation system for Nepali meme classification, using Qwen3-VL-8B-Instruct with LoRA fine-tuning and contrastive learning. The system achieved 2nd place in hate speech detection and 4th in sentiment analysis at the CHiPSAL 2026 shared task.

0 favorites 0 likes
#hate-speech-detection

Deployment of AIs in the real world

Reddit r/ArtificialInteligence ↗ · 2026-06-08

Reddit has deployed AI/LLMs to analyze all posts and comments in real time for hate speech and harmful content, enabling automatic bans within seconds, contrasting with Instagram and Facebook where such analysis is not applied as rigorously.

0 favorites 0 likes
← Back to home

Submit Feedback