Looking for extreme / impossible tasks to properly stress-test my agent.I can’t trust my own judgment anymore
Summary
A developer who built a fully autonomous custom agent architecture that can run for weeks without intervention is asking the community for extreme, adversarial tasks to properly stress-test it, because they can no longer trust their own judgment.
Similar Articles
It's impossible to test your own agent. I tried and failed.
A developer's personal account of the difficulty in objectively evaluating their own AI agent's performance, highlighting the pitfalls of self-testing and the value of unexpected, real-world benchmarks.
[Open-Source] I need your worst edge cases to stress-test GenOS, my new AI agent orchestrator.
The developer of GenOS, an open-source framework for multi-agent LLM orchestration using Rust and Git worktrees, is requesting community feedback with worst edge cases to stress-test and improve the system.
I keep abandoning multi-agent setups because I can't verify the code they ship. How are you handling this?
A developer shares their frustration with multi-agent coding setups where verifying the output of parallel PRs is impractical, and describes building an AI QA agent that uses a real browser (via Browserbase) to automatically click through preview deploys and fail PRs that don't work as expected.
Two weeks ago I asked what breaks between agents and the real world. Two of your replies are now tasks
The author follows up on a previous discussion by creating public benchmark tasks from community-reported incidents to test AI agent failures in real-world scenarios like authentication and booking, and seeks input on invariant checks.
The next big AI category won't build AI agents. It'll try to stress-test them
The article proposes a new category focused on stress-testing AI agents to improve reliability, using autonomous QA systems that simulate human behavior to detect hidden failures in workflows.