A developer built a custom benchmark for coding AI agents and found that a 35B-parameter model outperformed a 120B-parameter one when the harness was optimized, highlighting the importance of tailored evaluation over generic specifications.
I'm building an AI agent from scratch, and one decision I need to make is which open-source model to use. At first, I thought I'd go with GPT-OSS-120B, which is approximately four times the size of Qwen3.6-35B. I assumed the larger model would be substantially better at everything. It wasn't. To test that, I built my custom benchmark. When it comes to your own harness, you cannot trust generic leaderboards or pure specs. Qwen3.6-35B achieved a 95% pass@1 score on the benchmark. GPT-OSS-120B passed the same with just 53%. They were both given the same 19 tasks and evaluated with the same harness. I intentionally picked a random larger model that was easy to spin up just to see how it would perform out of the box. However, during development, the harness was tuned specifically to the Qwen3.6-35B model, hence the huge difference in score. This shows that optimizing your harness for your model matters at least as much as specs alone. That is why I am still skeptical about switching from Claude Code to Pi for my day-to-day work. (similar story to Apple vs. Linux ecosystem) A few details about my benchmark: The 19 tasks are split into 7 easy, 6 medium, and 6 hard tasks in the style of Terminal-Bench. For each task, an instruction is given along with a repository with a seed and a hidden verifier that evaluates the final state and assigns a 0 or 1 reward. The agent is operating in a sandbox and is able to return its work in the form of a git branch. The hidden tests are inserted into the environment only after the agent has run to avoid leaking any leads to the solution. The verifiers themselves are code and not other large language models. For example, a task would fail if the agent modified a file other than the target in the repository, or changed more than 8 lines of code in a file. I used Opik as the observability and evaluation system to track the agent traces, manage dataset versions, and store each run as an experiment. Comparing two models in this case was as simple as comparing two experiments. This setup also revealed a cost analysis. I ran the same model (Qwen3.6-35B) on the same workload using two different payment models: pay-per-token (OpenRouter) and pay-per-GPU-hour (Modal). Paying per token was cheaper, at approximately $0.13 for 1.2M tokens, whereas on Modal the same thing costs $0.46. The GPU is sitting most of the time waiting to process one test at a time, and you're paying for the whole session. Paying per GPU is only cheap if you can batch multiple tasks and keep it running at 100% constantly. Ultimately, model size and leaderboard rankings are only useful up to a point. Until you test your agent in your own use cases, you can never trust generic reports on performance, cost, or latency. What is your best strategy for finding which open models to plug into your harness? For both using agents and building harnesses from scratch. Vibes are not allowed
The author built SmallCode, a coding agent optimized for small local models, achieving 87% benchmark success with a 4B parameter model using techniques like compound tools, improvement loops, and token budgeting.
The AgentScope team introduces PawBench, a benchmark for evaluating the combined performance of models and agent harnesses, analyzing 4,050 test cells to show that harness choice can be as impactful as model upgrades.
Artificial Analysis introduces the Coding Agent Index, a new benchmark suite combining SWE-Bench-Pro-Hard-AA, Terminal-Bench v2, and SWE-Atlas-QnA to evaluate the performance of AI coding agents across diverse tasks.
A comparison showing that an untuned 27B parameter model outperforms a tuned 75B parameter model in agent tasks, highlighting potential inefficiencies in scaling and fine-tuning.
The article benchmarks two open-source coding agents on the same deepseek-v4-flash model, finding similar task success rates but significant differences in performance metrics and a critical bug in one agent's error handling.