content-safety

Tag

Cards List
#content-safety

Elon Musk’s xAI is trying to sue its way out of a Grok reckoning

Ars Technica · 3d ago Cached

xAI is suing Minnesota over a law that imposes severe penalties for AI-generated non-consensual intimate images, arguing it violates the First Amendment and forces Grok to restrict image editing features.

0 favorites 0 likes
#content-safety

Agent Safety Is Action Alignment

arXiv cs.AI · 2026-06-30 Cached

This paper argues that applying content-safety refusal methods to AI agents is a category error—agentic harm lies in authority misuse rather than output—and proposes action alignment enforced outside the model via least privilege.

0 favorites 0 likes
#content-safety

Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety

arXiv cs.CL · 2026-06-29 Cached

Yuvion LLM is a large language model designed for adversarial robustness and content safety, achieving state-of-the-art performance on safety benchmarks and outperforming larger models such as GPT-5.4 and Qwen3-MAX.

0 favorites 0 likes
#content-safety

Nemotron 3.5 Content Safety: Customizable Multimodal Safety for Global Enterprise AI

Hugging Face Blog · 2026-06-04 Cached

NVIDIA releases Nemotron 3.5 Content Safety, a unified multimodal AI safety model that combines multilingual support, custom enterprise policy enforcement, and auditable reasoning (THINK mode) in a single inference call. It builds on the previous Nemotron 3 model by deepening multimodal integration to evaluate text prompts, images, and assistant responses together for more comprehensive safety verdicts.

0 favorites 0 likes
#content-safety

OpenGuardrails: An Open-Source Context-Aware AI Guardrails Platform

Papers with Code Trending · 2025-10-22 Cached

OpenGuardrails is an open-source platform for AI safety, offering context-aware content-safety and manipulation detection (e.g., prompt injection, jailbreaking) via a unified model, plus a separate NER pipeline for data-leakage identification. It achieves state-of-the-art performance on safety benchmarks and supports private, enterprise-grade deployment.

0 favorites 0 likes
← Back to home

Submit Feedback