agent-improvement

Tag

Cards List
#agent-improvement

@VictorKaiWang1: 95.3% on Terminal Bench 2.1 Deepseek V4 Flash + StateM, 88.8% on TB2.1

X AI KOLs Timeline · 6h ago Cached

Researchers achieve 95.3% accuracy on Terminal-Bench 2.1 using DeepSeek V4 Flash and StateM, matching GPT-5.6 Sol Max performance and exploring agent improvement beyond model scaling.

0 favorites 0 likes
#agent-improvement

@omarsar0: // Adapt the Interface, Not the Model // I am fascinated by the results across my cheap-model-plus-good-harness builds.…

X AI KOLs Following · 2026-05-23 Cached

Proposes Life-Harness, a method that improves frozen LLM agents by adapting the runtime interface instead of model weights, achieving an average 88.5% relative improvement across 126 settings and 18 backbones.

0 favorites 0 likes
← Back to home

Submit Feedback