model-robustness

Tag

Cards List
#model-robustness

Output Dilution: Redundant but Fragile Representations in MoE Models

arXiv cs.LG · 2026-08-27 Cached

This paper reveals that Mixture-of-Experts models encode moral content as robustly as dense models in probing but are far more fragile due to output dilution, where aggregated expert outputs dilute the signal, making representations vulnerable to noise.

0 favorites 0 likes
#model-robustness

Rethinking the Transferable Adversarial Attacks and Robust Defense in Federated Learning

arXiv cs.LG · 2026-08-27 Cached

This paper analyzes the transferability of adversarial attacks in federated learning systems and proposes a defense mechanism based on adversarial training to enhance model robustness.

0 favorites 0 likes
#model-robustness

@rohanpaul_ai: LLMs can accept the same false claim differently depending on its tone, certainty, and grammatical form. Small wording …

X AI KOLs Timeline · 2026-07-22 Cached

This paper presents EoBench, a benchmark testing LLMs' acceptance of false claims phrased in 19 different styles, finding that tone, certainty, and grammatical form significantly affect model responses, with larger and instruction-tuned models showing more resistance.

0 favorites 0 likes
#model-robustness

Insights on Indirect Prompt Injection (12 minute read)

TLDR AI · 2026-06-24 Cached

Zico Kolter and Matt Fredrikson, leaders at Gray Swan and experts in AI security, discuss the state of AI red-teaming and indirect prompt injection, a critical vulnerability for AI agents. They explain why AI security requires a different mindset, how automated red-teaming can beat humans, and introduce tools like Shade for adversarial testing.

0 favorites 0 likes
#model-robustness

ERRORQUAKE: Heavy-Tailed Error Severity Distributions in Open-Weight Large Language Models

arXiv cs.LG · 2026-06-05 Cached

The paper introduces Errorquake-10k, a benchmark for evaluating error severity in open-weight LLMs, showing that models with matched accuracy can have vastly different error severity distributions, and argues that severity should be reported alongside accuracy.

0 favorites 0 likes
#model-robustness

@charles_irl: my gut says that to solve float numerics problems from nondeterminism x nonassociativity, we need to think bigger than …

X AI KOLs Following · 2026-05-22 Cached

This tweet discusses the idea of training models with 'implementation noise' to improve robustness against float numerics problems caused by nondeterminism and nonassociativity.

0 favorites 0 likes
#model-robustness

Alignment

Anthropic Research · 2026-05-08 Cached

This article outlines the mission and research focus of Anthropic's Alignment team, which develops safeguards to ensure future AI systems remain helpful, honest, and harmless through evaluation, oversight, and stress-testing.

0 favorites 0 likes
← Back to home

Submit Feedback