Tag
A tweet from @reach_vb asking @simonw about methods to evaluate AI models or harnesses, referencing a saturated 'pelican on a bike svg' benchmark.
A reflection on current practices for verifying AI coding agent output, noting that developers often skim diffs and merge without fully auditing the agent's session activity, raising concerns about code review culture in the age of AI.