@geekbb: Auto-optimization tool for Agent harness. It takes over the heavy lifting of harness optimization: you provide a benchmark command and a target repository, and it automatically generates proposals, runs evaluations, records results, keeps the best, discards the rest, and automatically improves the agent's prompts, configurations, and source code. https…
Summary
autoharness is an automated agent harness optimization tool that automatically generates proposals and runs evaluations based on benchmark commands to improve an agent's prompts, configurations, and source code. It supports Codex and Claude.
View Cached Full Text
Cached at: 05/11/26, 12:42 PM
Star this repo if you find it useful!
Built with ❤️ by Kayba and the open-source community.
Similar Articles
Claude Code improved my agent harness by 40% overnight
The author introduces 'Autoharness', a tool that uses Claude Code to autonomously optimize agent harnesses by iterating on prompts and hyperparameters. This resulted in a 40% performance increase on the tau2-airline benchmark.
Automatic Harness Optimization (GitHub Repo)
AutoSaddler is a Microsoft tool that automatically optimizes LLM-agent harnesses by diagnosing execution traces and applying structured updates to prompts, tools, and middleware, demonstrating significant benchmark improvements.
Remote agent harness
A tool for remotely harnessing and managing AI agents.
@GitHub_Daily: Using Claude Code for complex projects, a single agent has limited capabilities. Want multiple agents to collaborate and divide tasks, but manually configuring team structures and skill files is too tedious. Recently found Harness, a Claude Code plugin that automatically generates an entire team architecture from a one-sentence description of your project...
Harness is a Claude Code plugin that automatically generates a multi-agent team architecture based on a one-sentence description. It comes with 6 collaboration modes and 100 ready-made configurations, helping Claude Code transition from solo operation to team collaboration.
@KakaluoteW45042: Building your own agent harness isn't that mysterious: the minimal version has three parts—tool schema validation, structured output parsing, and failure retries. But the real value is in debuggability: when agent behavior is wrong, you can trace it through logs to locate whether it's a prompt, tool return, or retry logic issue. Black-box …
Building your own AI agent toolkit isn't complex; the minimal version includes tool schema validation, structured output parsing, and failure retries, but its true value lies in debuggability, allowing you to pinpoint issues through logs.