Tag
This paper extends missing-information stress-testing to open-ended medical conversation, finding that LLM judge choice materially changes apparent safety and that LLM judges are more permissive than clinicians.
This article teaches how to use Claude with CIA Red Team techniques to stress-test and kill bad ideas before acting on them, saving time and preventing failure.
Patronus AI raises $50M in Series B funding to build simulated digital worlds for stress-testing AI agents, helping ensure they perform reliably in real-world scenarios.
作者构建了一个基于GPT-5.5的自主Codex代理循环运行器,用于测试,目前处于公开测试阶段,提供50次免费运行机会。
This paper introduces AI-MASLD, a stress-audit framework for medical LLMs that reveals how benchmark accuracy can hide serious safety failures, and demonstrates that open-weight models can match or exceed proprietary ones on safety dimensions.
MemFail is a diagnostic benchmark that isolates failure modes of LLM memory systems by formalizing summarization, storage, and retrieval operations, and evaluating them with adversarially designed datasets.
Explains how to use Claude to perform a premortem, a technique by Daniel Kahneman, to stress-test plans by imagining they have already failed.
DetectRL-X is a comprehensive multilingual benchmark for evaluating LLM-generated text detectors across 8 languages and 6 domains, including stress testing with AI-assisted writing operations and perturbations. It reveals strengths and limitations of current detectors in multilingual scenarios.