Tag
This paper proposes BRACE, a method that encodes an Ordered Reasoning Chain to detect ever-shifting harmful chat dialogue, achieving high harm-type F1 scores with both encoder and decoder backbones.
Introduction of an AI-powered moderation tool designed to enhance safety in chat experiences by detecting malware and inappropriate content.