harmful-dialogue-detection

Tag

Cards List
#harmful-dialogue-detection

Calling the Bluff: Detecting Ever-Shifting Harmful Chat Dialogue via Ordered Reasoning Chain Regularization

arXiv cs.CL · 2026-08-11 Cached

This paper proposes BRACE, a method that encodes an Ordered Reasoning Chain to detect ever-shifting harmful chat dialogue, achieving high harm-type F1 scores with both encoder and decoder backbones.

0 favorites 0 likes
← Back to home

Submit Feedback