@BohuTANG: During the development of Evot, I discovered that to get the best out of Anthropic's Opus series models, the official Claude Code approach is basically the optimal solution, hard to bypass. After in-depth analysis and quantitative verification of the Claude Code prompt, I found that during training they already...

X AI KOLs Timeline News

Summary

During the development of Evot, it was discovered that to maximize the performance of the Anthropic Opus model, the official Claude Code method is the optimal solution, because the Agent Harness behavior pattern is baked into the weights during training, rather than pure prompt engineering; in the future, Agent Harness competition will push behavior down to the model layer.

During the development of Evot, I discovered that to maximize the performance of Anthropic's Opus series models, the official Claude Code approach is essentially the optimal solution and hard to bypass. After deep analysis and quantitative verification of the Claude Code prompt, I found that the Agent Harness behavior pattern has been baked into the weights during training: 1. When outputting long content, the model spontaneously produces <!-- PLACEHOLDER_ --> placeholders for chunked writing — the system prompt never asks for this. 2. Tool names have inherent preference binding inside the model; only the official name yields the best results, and using a different name leads to dramatically different behavior. ... ... This is no longer a matter of prompt engineering — the scaffold behavior pattern was encoded into the weights during model training. For third parties to achieve the same effect, they either align with the official approach or train their own model. I suspect that DeepSeek's Agent Harness (@victor207755822) will have to follow the same path — baking the harness behavior into model training, deeply coupling the model and the scaffold. Cursor Composer 2.5 likely does the same, achieving parity with Opus at 1/10 the cost. The next competitive dimension of Agent Harness: agent behavior sinking from the prompt layer to the model layer. The combination of prompt and weights produces far better results than purely relying on prompt engineering.
Original Article
View Cached Full Text

Cached at: 05/23/26, 08:15 PM

During the development of Evot, it was discovered that to get the best performance out of Anthropic’s Opus series models, the official Claude Code approach is essentially optimal and difficult to circumvent.

After in-depth analysis and quantitative verification of the Claude Code prompt, it was found that the model’s behavior patterns for the Agent Harness were baked in during training:

  1. When generating long content, the model spontaneously produces <!-- PLACEHOLDER_ --> placeholders for chunked writing — the system prompt never explicitly asks for this.
  2. Tool names have inherent preference bindings inside the model; only the official names yield the best results. Changing the name leads to vastly different behavior.

… …

This goes beyond prompt engineering — the scaffold’s behavior patterns are encoded into the model weights during training. For third parties to achieve the same effect, they either need to align with the official approach or train their own models.

I suspect that DeepSeek, when building its Agent Harness (@victor207755822), will have to follow the same path — baking harness behavior into model training, tightly coupling the model and scaffold. Cursor Composer 2.5 likely does the same, achieving parity with Opus at 1/10 the cost.

The next competitive dimension for Agent Harness: agent behavior moves from the prompt layer down to the model layer. When prompt and weights work together in and out, the effect far surpasses relying solely on prompt engineering.

Similar Articles

@shao__meng: Why do Claude Code, Cursor, Codex, Aider, and Cline exhibit different agent behaviors despite potentially sharing the same underlying models? @addyosmani argues: It's due to the "shell" above the model — the Harness, which includes "prompts, ...

X AI KOLs Timeline

The article discusses how Addy Osmani argues that the performance difference between AI coding agents like Claude Code, Cursor, and Cline stems from their 'Harness'—the layer of prompts, tools, and constraints around the model—rather than the underlying model itself. It details best practices for harness engineering, including hooks, sandboxing, and context management, to bridge the gap between model capability and actual agent performance.

@VincentLogic: Anthropic’s talk had some solid insights. Previously, building Agents required manually implementing routing, retry mechanisms, and context compression. Now, the speaker points out that these 'scaffolding' features are already built into the model—stop reinventing the wheel. The most mind-blowing part was the final demo: letting Claude autonomously...

X AI KOLs Timeline

This article comments on Anthropic’s talk regarding Claude, noting that the model now includes built-in Agent scaffolding features such as routing and retries. It highlights a demo showcasing a smooth closed-loop workflow where Claude independently reproduces, fixes, and tests frontend bugs, marking a new era in Agent development.

@NFTCPS: HarnessX is pretty interesting: an agent architecture that can modify itself. Previously, architectural changes relied entirely on manual tuning. When a new model came out, Anthropic removed the planning steps from Claude Code, and Manus refactored its agents five times in six months, each time simplifying. What to change and when to change it — all decided by humans.

X AI KOLs Timeline

HarnessX introduces a framework for self-evolving AI agent harnesses that treats the runtime harness as a first-class object, enabling automatic adaptation via trace-driven reinforcement learning. It achieves average gains of +14.5% across five benchmarks, with larger improvements for weaker models.

@xiaohu: Claude Code's father's own CLAUDE.md is now just two lines... Claude Code team discusses "less is more" sharing how to communicate with models as capabilities increase: "Don't fight the model by adding more, because each generation of models gets stronger. What you painstakingly build today will soon be useless."

X AI KOLs Timeline

Claude Code team shares best practices: CLAUDE.md should be as short as possible and regularly cleared; insists on CLI over GUI because models improve too fast; using AI to fix bugs is already remarkably efficient. Core strategy: subtract, keep configuration light, and trust model capabilities.