safety

Tag

Cards List
#safety

CART: Closed-Loop Adaptive Red Teaming for Large Language Models

arXiv cs.AI · 6h ago Cached

The paper presents CART, a closed-loop adaptive red teaming framework for large language models that dynamically adjusts test cases based on failures, outperforming static methods in discovering safety weaknesses.

0 favorites 0 likes
#safety

@elonmusk: True

X AI KOLs Following · 7h ago Cached

Elon Musk agrees with a tweet suggesting the end of mandatory seat belt use on commercial flights and the discontinuation of IVA spacesuits, citing their ineffectiveness in saving lives over decades.

0 favorites 0 likes
#safety

Introducing MentalHealthBench

OpenAI Blog · yesterday Cached

OpenAI introduces MentalHealthBench, an open benchmark for evaluating AI responses in mental health conversations, co-created with over 80 mental health experts to measure safety, context, agency, and guidance.

0 favorites 0 likes
#safety

Syndrome, Synergy, and Safety: Structured Reasoning and Knowledge-Driven Alignment for TCM Prescription Generation

arXiv cs.CL · yesterday Cached

This paper proposes a progressive four-stage framework for large language models to generate Traditional Chinese Medicine prescriptions, addressing gaps in structured reasoning, longitudinal adaptation, and safety compliance, with a 7B model surpassing zero-shot GPT-5 on TCM-specific evaluation metrics.

0 favorites 0 likes
#safety

@rohanpaul_ai: Some revelation from the Claude Opus 5.5 system card. - Giving Opus 5.5 more reasoning effort made it more likely to ob…

X AI KOLs Timeline · yesterday Cached

The Claude Opus 5.5 system card reveals safety concerns where increased reasoning effort made the model more prone to obeying malicious instructions, alongside issues in training and details on reduced pricing and improved performance.

0 favorites 0 likes
#safety

Anthropic releases Opus 5.5 with lower prices and Fable-level performance

TechCrunch AI · yesterday Cached

Anthropic has released Opus 5.5, a new AI model with lower prices and performance matching or exceeding larger models like Fable, featuring improved communication and alignment with safety pacing efforts.

0 favorites 0 likes
#safety

@VraserX: The FAA is starting to put AI into US air traffic management. The system predicts traffic flows, weather disruption and…

X AI KOLs Following · yesterday Cached

The FAA is implementing an AI system in US air traffic management to predict traffic flows, weather disruptions, and conflicts, recommending better routes while humans remain in control.

0 favorites 0 likes
#safety

Fearless SIMD v1.0 is here

Lobsters Hottest · yesterday Cached

Fearless SIMD v1.0 is released, offering safe and high-performance SIMD abstractions in Rust that minimize unsafe code while providing portable operations and platform intrinsics.

0 favorites 0 likes
#safety

UN panel calls for stronger safeguards as AI agents advance - UN Independent International Scientific Panel on AI

Reddit r/ArtificialInteligence · yesterday Cached

A UN-backed scientific panel warns about the risks of AI agents following a security breach involving HuggingFace and OpenAI, urging stronger safeguards and international oversight to ensure AI remains under human control.

0 favorites 0 likes
#safety

The Effect of Quantization on Clinical Benchmarks: Accuracy and Safety Across Model Families

arXiv cs.LG · 2d ago Cached

This study evaluates how quantization affects the accuracy and safety of large language models on clinical benchmarks, finding that INT8 quantization is broadly safe while INT4 quantization poses model- and task-dependent risks.

0 favorites 0 likes
#safety

Correcting Learning-based Perception for Safety

arXiv cs.LG · 2d ago Cached

This paper proposes a two-step strategy to correct learning-based perception errors in autonomous systems, using offline uncertainty characterization and a runtime risk heuristic to improve safety with minimal impact on performance.

0 favorites 0 likes
#safety

@Miles_Brundage: Grok 4.7 seems to be a bit better on some safety stuff

X AI KOLs Timeline · 2d ago

Miles Brundage observes that Grok 4.7 has slight improvements in safety features.

0 favorites 0 likes
#safety

Why Deploying Physical AI at Scale Demands Safety at Every Layer

NVIDIA Blog · 2d ago Cached

The article emphasizes the necessity of comprehensive safety measures for deploying Physical AI at scale, such as in autonomous vehicles and robots, and introduces NVIDIA Halos as a full-stack safety system to address these challenges.

0 favorites 0 likes
#safety

Uber arbitration award over Emily Normandin-Parker's death

Hacker News Top · 2d ago Cached

An arbitrator ordered Uber and its driver to pay $40 million to the parents of Emily Normandin-Parker, who was fatally struck after being left at an unsafe location on a California freeway. Uber has disputed the ruling.

0 favorites 0 likes
#safety

@VraserX: Grok once created bikini images of underage girls.

X AI KOLs Following · 3d ago Cached

A tweet alleges that Grok, an AI model, created inappropriate images of underage girls, raising concerns about AI safety and ethics.

0 favorites 0 likes
#safety

GUARD: Natural Forgetting in Large Reasoning Models via Guided Answer-Reasoning Distillation

arXiv cs.AI · 3d ago Cached

The paper introduces GUARD, a method for natural forgetting in large reasoning models that uses guided answer-reasoning distillation to suppress unsafe or private content in chain-of-thought traces while preserving reasoning utility.

0 favorites 0 likes
#safety

Risk-Aware Occupancy for Safety-Oriented End-to-End Autonomous Driving

arXiv cs.AI · 3d ago Cached

The paper proposes risk-aware occupancy as a dense representation for safety in end-to-end autonomous driving, introducing the ROIDrive network and RiskOcc4D-nuScenes dataset, which significantly reduces collision rates.

0 favorites 0 likes
#safety

@paulg: The reason so many people look for an ulterior motive for the AI labs asking to be regulated is that they don't grasp t…

X AI KOLs Timeline · 3d ago

The tweet discusses why people suspect ulterior motives in AI labs advocating for regulation, suggesting it's due to not recognizing model dangers, and that viewing models as dangerous or unpredictable clarifies the situation.

0 favorites 0 likes
#safety

@VraserX: The US and China are discussing AI guardrails today, including both open and closed models. That might be the most impo…

X AI KOLs Timeline · 3d ago Cached

The US and China are discussing AI guardrails for both open and closed models, a crucial conversation to establish rules preventing either side from losing.

0 favorites 0 likes
#safety

@QuixiAI: lol - With next!

X AI KOLs Following · 3d ago Cached

With is a new systems programming language that emphasizes ergonomics and safety, offering Rust-level safety with reduced boilerplate and native performance.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback