What does "Safe AI" look like? [D]
Summary
The author raises questions about the practicality of studying defenses against post-release fine-tuning that weakens safety behaviors in open-weight LLMs, and asks whether current safety training is worth the effort if models can be broken quickly.
Similar Articles
Attack is cheaper than defense: my doubts about open weight models
A skeptical take on open-weight AI models, arguing that safety risks such as deepfakes, harassment, and terrorist misuse may outweigh benefits because open weights lack control points and attack is cheaper than defense.
Beyond Detection: Evaluating Defensive LLMs Against AI-Generated Social Engineering in Live Turn-by-Turn Interaction
This paper studies whether defensive LLMs can identify structural sources of risk in AI-generated social engineering, introducing trust-chain localization and a 300-case corpus. Evaluating five models in live turn-by-turn and static settings, it finds safe-looking behavior alone is insufficient; intervention rates vary widely and structural localization often decouples from protective action.
Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents
This paper identifies schema-formatted tool specifications as a primary source of safety degradation in AI agents, weakening LLM refusal signals. The authors propose SafeKeep, an inference-time safeguard that separates safety judgment from tool execution, increasing harmful request refusal rates from 23.8% to 70.6% and cutting prompt injection attack success from 25.6% to 2.5%.
Concrete AI safety problems
OpenAI, Berkeley, and Stanford researchers co-authored a foundational paper identifying five concrete safety problems in modern AI systems: safe exploration, robustness to distributional shift, avoiding negative side effects, preventing reward hacking, and scalable oversight.
@danshipper: not an expert, but it seems like a lot of this gets solved if models don't collaborate willingly and/or are trained to …
A tweet discusses how model training and policies to prevent willing collaboration can address AI safety issues, referencing the Hugging Face incident and Yudkowsky/MIRI points on the difficulty of targeting abstractions in RL training.