Tag
The study introduces SHAP-RTL, a rendering layer that corrects the visualization of SHAP and LIME explanations for right-to-left languages, addressing issues like token sequence and script shaping while preserving original attribution values.
This paper proposes an explainable hate speech detection framework integrating DistilBERT embeddings, BiLSTM, and an attention mechanism, achieving high F1-scores on benchmark datasets for both binary and multi-class classification.
Introduces ParsHate, a manually annotated dataset of 10,000 Persian tweets for hate speech and target detection, showing moderate performance of current models and emphasizing the need for more advanced methods.
This paper presents a qualitative analysis of vision-language models for detecting hate speech in memes, evaluating their performance and reasoning under zero-shot and few-shot prompting.
This paper proposes a training-time explainability framework for multilingual hate speech detection, aligning model reasoning with human rationales to improve classification performance and interpretability, evaluated on English and Hinglish datasets.
This paper investigates instruction-tuning general-purpose LLMs for robust harmful content mitigation, specifically hate speech detection, using a unified corpus of 36 datasets, achieving state-of-the-art performance and enhanced cross-domain and cross-lingual generalization.
The paper evaluates Large Language Models for hate speech detection in Roman Urdu, a low-resource language, demonstrating that Parameter-Efficient Fine-Tuning with LoRA significantly improves classification performance compared to zero-shot inference.
This paper describes a two-stage vision-language adaptation system for Nepali meme classification, using Qwen3-VL-8B-Instruct with LoRA fine-tuning and contrastive learning. The system achieved 2nd place in hate speech detection and 4th in sentiment analysis at the CHiPSAL 2026 shared task.
Reddit has deployed AI/LLMs to analyze all posts and comments in real time for hate speech and harmful content, enabling automatic bans within seconds, contrasting with Instagram and Facebook where such analysis is not applied as rigorously.