Tag
This paper proposes BRACE, a method that encodes an Ordered Reasoning Chain to detect ever-shifting harmful chat dialogue, achieving high harm-type F1 scores with both encoder and decoder backbones.
Adrian is an open-source AI agent runtime security monitoring engine that detects anomalies by jointly analyzing the agent’s behavior logs and reasoning chain, improving accuracy by 35% over behavior-only checks. It supports LangChain integration with a two-line SDK.