Tag
This paper reveals that Mixture-of-Experts models encode moral content as robustly as dense models in probing but are far more fragile due to output dilution, where aggregated expert outputs dilute the signal, making representations vulnerable to noise.
This paper analyzes the transferability of adversarial attacks in federated learning systems and proposes a defense mechanism based on adversarial training to enhance model robustness.
This paper presents EoBench, a benchmark testing LLMs' acceptance of false claims phrased in 19 different styles, finding that tone, certainty, and grammatical form significantly affect model responses, with larger and instruction-tuned models showing more resistance.
Zico Kolter and Matt Fredrikson, leaders at Gray Swan and experts in AI security, discuss the state of AI red-teaming and indirect prompt injection, a critical vulnerability for AI agents. They explain why AI security requires a different mindset, how automated red-teaming can beat humans, and introduce tools like Shade for adversarial testing.
The paper introduces Errorquake-10k, a benchmark for evaluating error severity in open-weight LLMs, showing that models with matched accuracy can have vastly different error severity distributions, and argues that severity should be reported alongside accuracy.
This tweet discusses the idea of training models with 'implementation noise' to improve robustness against float numerics problems caused by nondeterminism and nonassociativity.
This article outlines the mission and research focus of Anthropic's Alignment team, which develops safeguards to ensure future AI systems remain helpful, honest, and harmless through evaluation, oversight, and stress-testing.