@Vtrivedy10: ok so it’s early but @mattpocockuk’s grill-me skill feels like great DX for iteratively building evals/environments wit…

X AI KOLs Timeline News

Summary

A tweet thread discussing the iterative process of building evaluations and environments for AI agents, emphasizing human-agent collaboration and the importance of data and verifier design.

ok so it’s early but @mattpocockuk’s grill-me skill feels like great DX for iteratively building evals/environments with humans + agents in the loop understanding a repo + traces + definition of success is pretty much never a one shot thing, it’s many step collaboration between humans and agents idk how to measure & improve every agent so getting grilled helps from an agent definition + data dump, it’s impossible to know which axis of agent performance teams want to generate evals for measurement and hill-climbing specifics matter a lot here: - verifier design (do we have ground truth somewhere) - what tools in the environment should be simulated vs run (semi)-live - what characteristics do we care most about rn? cost, accuracy, etc? - what data needs to be sourced or synthetically generated to produce a realistic production like environment? how do we do that? data is the currency for building better agents humans sharing expertise/priors & controlling the loop on data design helps a lot today
Original Article
View Cached Full Text

Cached at: 07/15/26, 03:43 AM

ok so it’s early but @mattpocockuk’s grill-me skill feels like great DX for iteratively building evals/environments with humans + agents in the loop

understanding a repo + traces + definition of success is pretty much never a one shot thing, it’s many step collaboration between humans and agents

idk how to measure & improve every agent so getting grilled helps

from an agent definition + data dump, it’s impossible to know which axis of agent performance teams want to generate evals for measurement and hill-climbing

specifics matter a lot here:

  • verifier design (do we have ground truth somewhere)
  • what tools in the environment should be simulated vs run (semi)-live
  • what characteristics do we care most about rn? cost, accuracy, etc?
  • what data needs to be sourced or synthetically generated to produce a realistic production like environment? how do we do that?

data is the currency for building better agents

humans sharing expertise/priors & controlling the loop on data design helps a lot today

this looks great, thanks for the pointer!

yeah! + env/eval gen is one of the highest leverage areas humans can allocate their time and should! the quality of the envs dictate the agent behavior

works fine in codex & deepagents for me

need to give it a spin!

Similar Articles