Xybernetex is an experimental OpenClaw plugin that monitors AI agent runs to reduce wasted tokens and improve task success through interventions like retries and verification, with the author seeking real-world testers.
I’ve been screwing around with a research project for the last few weeks that has turned into an actual product, and I’m at the point where I need people other than me running it. It’s called Xybernetex. The idea is pretty simple: I don’t actually care how many tokens an individual model call uses. I care how many tokens it takes to successfully finish the task. If an agent uses 20% fewer tokens but fails and has to be run again, we didn’t optimize anything. Xybernetex is an OpenClaw plugin that watches what happens during a run and tries to figure out when the agent needs help. I’ve been experimenting with things like retrying runs that died, adding a verification pass when it makes sense, detecting loops/failures, and putting human approval in front of potentially destructive actions. The results have been interesting enough that I think it’s time to get it out of my lab. So far: Retrying dead runs recovered 54% of them A verification intervention took task success from 67.0% to 78.2% in one of my hard-task experiments There’s a local safety layer for potentially destructive actions I’ve run thousands of graded agent runs while developing the thing The experiment I’m running now is more interesting: when should Xybernetex intervene? Because obviously throwing another model call at every problem is a fantastic way to build a token-saving product that burns more tokens. The goal is to learn which runs benefit from intervention, which intervention works, and when the best thing Xybernetex can do is leave the agent alone. I also deliberately designed it so I don’t need your prompts, files or command contents. The plugin reduces what it sees locally into structural information about the run before talking to the Xybernetex API. Now I need real OpenClaw workloads. I’m looking for a handful of people who use OpenClaw for actual work and are willing to install the plugin, run it, and tell me what happens. Especially if you have workflows that burn a meaningful number of tokens, occasionally get stuck, fail, or require you to babysit them. This is still experimental. I have good results from my own testing, but I have absolutely no intention of pretending that means it’ll magically generalize to everybody else’s workloads. Finding that out is the point. If you’re interested, shoot me a message. I’m also happy to share how the system works and the experimental results.
A developer shares a personal open-source benchmark runner for testing OpenClaw agents on real, messy workflows. The tool allows users to define private evaluation cases, run agents in their actual workspace, and generate reports, aiming to provide more relevant signals than public benchmarks.
OpenClaw is seeking early users to test their open-source model inference plans, sold by concurrency slot with high throughput and no shared pool, in exchange for free access and feedback.
The author is seeking 5-10 beta testers for Behave, an AI agent testing/evaluation tool that goes beyond simple answer checking to catch issues like hallucination, premature conclusions, unsafe advice, and failures to self-correct.
The article critiques current browser AI agents for inefficiency due to repeatedly parsing and reasoning about the same websites, and proposes a model where agents reuse proven interaction paths to reduce token consumption and improve speed.