Tag
A tweet from @reach_vb asking @simonw about methods to evaluate AI models or harnesses, referencing a saturated 'pelican on a bike svg' benchmark.
A new survey from Renmin University reviews nearly 1,000 studies on long-horizon AI agents, arguing that reliable long-horizon intelligence depends on the whole model-harness system, not just larger context windows or stronger models.
A developer argues that the harness (critics, scaffolding) around an AI model is more important than the model itself, sharing an example where a 27B model with good critics became usable for coding work.