Tag
The article proposes a new category focused on stress-testing AI agents to improve reliability, using autonomous QA systems that simulate human behavior to detect hidden failures in workflows.
OpenAI and Anthropic are reportedly nearing a deal to stress-test each other's AI systems, as reported by The Information.
The paper introduces a method for controlled outlier generation using diffusion models by manipulating likelihood through Radon-Nikodym derivatives, enabling the creation of low-likelihood samples without retraining.
The Beni robot from Mondo Robotics is an autonomous camera robot designed to track and film dynamic subjects like skateboarders, but it has significant limitations in image quality, tracking response, and control that affect its practicality for content creators.
The developer of GenOS, an open-source framework for multi-agent LLM orchestration using Rust and Git worktrees, is requesting community feedback with worst edge cases to stress-test and improve the system.
The paper introduces PV-SST, a peer-voted social-platform testbed for LLM agents, and finds that peer-ranked feeds increase lexical convergence but do not reliably produce opinion capture or coordination advantages across model families and topics.
The article describes a new arena allocator for the Pony language that addresses memory growth and performance issues found during stress testing, with a design inspired by snmalloc.
This paper formulates PRM stress testing as a quality-diversity search problem using MAP-Elites, characterizing what archive coverage can and cannot certify. Experiments on Qwen2.5-Math-PRM-7B reveal aggregation-dependent vulnerabilities, and a LoRA repair protocol reduces exploit rates.
This paper extends missing-information stress-testing to open-ended medical conversation, finding that LLM judge choice materially changes apparent safety and that LLM judges are more permissive than clinicians.
This article teaches how to use Claude with CIA Red Team techniques to stress-test and kill bad ideas before acting on them, saving time and preventing failure.
Patronus AI raises $50M in Series B funding to build simulated digital worlds for stress-testing AI agents, helping ensure they perform reliably in real-world scenarios.
作者构建了一个基于GPT-5.5的自主Codex代理循环运行器,用于测试,目前处于公开测试阶段,提供50次免费运行机会。
This paper introduces AI-MASLD, a stress-audit framework for medical LLMs that reveals how benchmark accuracy can hide serious safety failures, and demonstrates that open-weight models can match or exceed proprietary ones on safety dimensions.
MemFail is a diagnostic benchmark that isolates failure modes of LLM memory systems by formalizing summarization, storage, and retrieval operations, and evaluating them with adversarially designed datasets.
Explains how to use Claude to perform a premortem, a technique by Daniel Kahneman, to stress-test plans by imagining they have already failed.
DetectRL-X is a comprehensive multilingual benchmark for evaluating LLM-generated text detectors across 8 languages and 6 domains, including stress testing with AI-assisted writing operations and perturbations. It reveals strengths and limitations of current detectors in multilingual scenarios.