Tag
xAI is suing Minnesota over a law that imposes severe penalties for AI-generated non-consensual intimate images, arguing it violates the First Amendment and forces Grok to restrict image editing features.
This paper argues that applying content-safety refusal methods to AI agents is a category error—agentic harm lies in authority misuse rather than output—and proposes action alignment enforced outside the model via least privilege.
Yuvion LLM is a large language model designed for adversarial robustness and content safety, achieving state-of-the-art performance on safety benchmarks and outperforming larger models such as GPT-5.4 and Qwen3-MAX.
NVIDIA releases Nemotron 3.5 Content Safety, a unified multimodal AI safety model that combines multilingual support, custom enterprise policy enforcement, and auditable reasoning (THINK mode) in a single inference call. It builds on the previous Nemotron 3 model by deepening multimodal integration to evaluate text prompts, images, and assistant responses together for more comprehensive safety verdicts.
OpenGuardrails is an open-source platform for AI safety, offering context-aware content-safety and manipulation detection (e.g., prompt injection, jailbreaking) via a unified model, plus a separate NER pipeline for data-leakage identification. It achieves state-of-the-art performance on safety benchmarks and supports private, enterprise-grade deployment.