toxicity-detection

Tag

Cards List
#toxicity-detection

A Survey of Toxicity Detection and Mitigation Strategies for Multilingual Language Models

arXiv cs.CL · 2026-06-25 Cached

This survey synthesizes research on toxicity detection and detoxification for multilingual large language models, cataloging threat models, task formulations, detection approaches, and mitigation strategies, while identifying persistent challenges such as uneven language coverage and culturally contingent definitions of harm.

0 favorites 0 likes
#toxicity-detection

On the Limits of LLM Adaptability: Impact of Model-Internalized Priors on Annotation Task Performance

arXiv cs.CL · 2026-06-02 Cached

This paper investigates how LLMs' internal priors affect zero-shot annotation performance, finding that nearly two-thirds of errors resist prompt-based correction and introducing Definition-Specific Familiarity as a better predictor than memorization metrics.

0 favorites 0 likes
#toxicity-detection

Harder to Defend: Towards Chinese Toxicity Attacks via Implicit Enhancement and Obfuscation Rewriting

arXiv cs.CL · 2026-05-22 Cached

The paper introduces CITA, a framework for generating implicit toxicity attacks in Chinese to evaluate and improve LLM toxicity detectors, finding high attack success rates across tested models.

0 favorites 0 likes
#toxicity-detection

Fair and Calibrated Toxicity Detection with Robust Training and Abstention

arXiv cs.LG · 2026-05-15 Cached

This paper studies fairness in toxicity classification across three axes: ranking, calibration, and abstention. It compares ERM, reweighted ERM, and Group DRO methods with post-hoc interventions, finding that calibration disparity is a hidden fairness violation and that abstention itself can be unfair.

0 favorites 0 likes
#toxicity-detection

PSK@EEUCA 2026: Fine-Tuning Large Language Models with Synthetic Data Augmentation for Multi-Class Toxicity Detection in Gaming Chat

arXiv cs.CL · 2026-05-11 Cached

This paper presents a system for the EEUCA 2026 shared task on toxicity detection in gaming chat, achieving 4th place by fine-tuning Llama 3.1 8B with synthetic data augmentation. It highlights a 'validation trap' phenomenon where high validation scores do not correlate with test performance due to dataset distribution shifts.

0 favorites 0 likes
← Back to home

Submit Feedback