@omarsar0: Important discussion. Measuring models against harnesses is completely broken. I prefer to test model quality against m…

X AI KOLs Following News

Summary

The author criticizes the broken system of measuring AI models against test harnesses due to biases and lack of standardization, while speculating that models like Claude may soon generate harnesses dynamically.

Important discussion. Measuring models against harnesses is completely broken. I prefer to test model quality against minimal harnesses like Pi and Hermes Agent. This is not perfect, as there are biases in the harnesses that favor some models and not others. A standard way to do this is missing but important, as harness engineering is where leading AI companies are focusing efforts. Not enough effort here, as things are moving fast and harness optimization is individualistic. On the flip side, I feel like models will eventually have the ability to dynamically generate harnesses on the fly as per task. Claude models do this already to some extent, though pretty inconsistently and remain a mystery. But this could mean that a harness is just a tunable artifact like a system prompt. In that realm, how are we assessing it, and exactly what? Benchmarking will only get murkier from here onwards.
Original Article
View Cached Full Text

Cached at: 08/24/26, 01:47 AM

Important discussion. Measuring models against harnesses is completely broken. I prefer to test model quality against minimal harnesses like Pi and Hermes Agent. This is not perfect, as there are biases in the harnesses that favor some models and not others.

A standard way to do this is missing but important, as harness engineering is where leading AI companies are focusing efforts. Not enough effort here, as things are moving fast and harness optimization is individualistic.

On the flip side, I feel like models will eventually have the ability to dynamically generate harnesses on the fly as per task. Claude models do this already to some extent, though pretty inconsistently and remain a mystery. But this could mean that a harness is just a tunable artifact like a system prompt. In that realm, how are we assessing it, and exactly what? Benchmarking will only get murkier from here onwards.

Onur Solmaz (@onusoz): We need to normalize measuring and judging models against a standardized test harness

“Oh but model X performs best in their own proprietary harness”

I could not care less. When I take exams, I go to the standardized classroom, get the standardized pencil and exam sheet, and

Similar Articles

We NEED a harness benchmark leaderboard

Reddit r/AI_Agents

This article argues for the need of a benchmark leaderboard that compares AI model harnesses (e.g., KimiCode vs OpenCode vs Codex) rather than just models themselves, proposing a repo to test model+harness combinations on cost, runtime, token usage, and score.