Tag
The author built AgentPaySec, a security testing platform for AI agents with financial authority, tested it on a simulated payment agent, found vulnerabilities, and is seeking community feedback on attack scenarios.
The author is introducing a tool for agent regression testing that allows developers to verify tool calls and workflow outcomes after code changes.
The post describes building FetchSandbox, a tool that provides realistic testing environments for AI agents by simulating actual service provider responses to prevent failures in live production.
An analysis of go/no-go criteria for shipping AI agents with real production permissions, proposing a Green/Yellow/Red status system and arguing that serious failures should block release regardless of average success rates.
A fine-tuned model based on GLM-5.2, abliterated and specialized for agent testing and red teaming, achieving 97.5% benign utility on AgentDojo and strong coding benchmarks.
A developer's personal account of the difficulty in objectively evaluating their own AI agent's performance, highlighting the pitfalls of self-testing and the value of unexpected, real-world benchmarks.
Coval announces a Series A funding round to build infrastructure for testing and deploying conversational AI agents in enterprises, inspired by the rigor of autonomous vehicle testing.
A discussion on the challenges of testing non-deterministic AI agents, questioning how developers validate tool usage, behavior, and multi-step workflows without traditional testing patterns.
A discussion on the challenges of testing AI agent harnesses with non-deterministic components, exploring approaches like golden output diffing and using an LLM as a judge, while questioning the validity of such methods.
A developer seeks a model that frequently gets stuck in loops (e.g., GLM Flash) to test loop detection and recovery features for an agent, aiming to develop heuristics that score loop probability and enable backtracking.