@jamonholmgren: I'm just going to dump my whole agentic setup out here, because I see too many people missing giant chunks of this and …

X AI KOLs Timeline News

Summary

Jamon Holmgren shares his comprehensive agentic development setup, including workflow docs, self-healing docs, cross-agent review, automated testing, and autonomous agent loops.

I'm just going to dump my whole agentic setup out here, because I see too many people missing giant chunks of this and it's hurting them. Here's what I have and recommend: 0. an AGENTS.md that is a router -- it sends the agent to the right skills, docs, tools 1. a standard workflow doc/skill customized to my needs ... (grab Matt Pocock skills if you don't already have something) ... I tag this in most sessions with `@/AGENT_WORKFLOW.md` and it pulls it in. 2. self-healing docs for every system, and agents are instructed to keep them updated ... I tag the ones I know I need, or let the agent find them through AGENTS.md ... I also provide a more detailed summary in the first 7 lines of every doc, so they're easily greppable to find the right thing, and this is documented in AGENTS.md 3. agents always run the app ... the agent should always actually run the app itself, and test its work and fix issues as it goes, especially if running autonomously / asynchronously 4. end-to-end tests and instructions to write more and keep up to date, and docs on how to write tests, what to avoid, and a list of all the tests and what they test in yet another markdown doc ... write and run targeted tests during implementation, improve and commit with work 5. custom linters at precommit hooks looking for any problems you run across, with `--fix` fixing the problems automatically, OR if that's not feasible, it shells out to a cheaper LLM like Composer 2.5 or Sonnet to fix the problems -- NOT just flagging them, but actually resulting in cleaned code 6. cross-agent review at each major point: research, plan, implementation, and wrap-up. I mean codex, claude, cursor, whatever -- but it shouldn't be the same model reviewing the same code. And specific docs for agent review, what to look for, how to approach it. Also, personas -- looking at the code from different perspectives, such as maintainability, code quality, security, performance, AI smells, domains (e.g. "financial services expert" or whatever) ... and each persona also "owns" a set of system docs too and keeps them up to date 7. agent traces / worksheets that track what the agent is doing each session. if the agent fails partway through, you should be able to hand this worksheet to another agent and it could finish the job. commit this worksheet with the work so it's all connected and easy to reference later (you will reference these later!!), also have the agent apply git tags that correspond to specific worksheet names so they're easy to find 8. automatic agent feedback to you at the end of the session, added to a doc that is also committed with the work, that you periodically ingest into an interactive session and improve your workflows 9. a tools or bin folder that contains python or bash scripts that the agent has skills to make to make its job easier (for example, I have an `agent_review` bash script that lets the agent kick off agent reviews via CLI without knowing each agent's particular incantations) ... docs on how to make scripts effectively, and instructions to constantly build these out more 10. periodic agent sweeps through recent commits, looking for problems / gotchas from a higher level across commits 11. a coding conventions doc that is just for specific coding conventions you want to see in the code base, your review agents use these a lot (but a lot of this should be in linters) 12. an agent loop / night shift skill for autonomous work, that lays out how the agent is to approach this, from an orchestration standpoint 13. a task queue that is accessible to the agent (mine is just a TODOS.md, but yours might be in Linear etc, with a CLI to fetch via API) 14. a periodic false-confidence test audit skill that looks for tests that aren't actually testing what you think they're testing, and that fix those 15. visual regression tests -- take screenshots, compare via tool and with agent visual review, commit with work (git lfs useful here) or at least push into the PR 16. automatic performance benchmark tests that notice when performance degrades 17. performance profiling tools that can be used by agents for targeted benchmarking, trying new techniques, comparing outputs, and comparing profiles 18. end-of-shift full validations, including running all tests, performance, agent reviews, sweeps, everything -- when you return, it's all as pristine as it can be If you have all this, your agentic coding experience is going to be very different than dry prompting and manually guiding it toward the right thing every time.
Original Article
View Cached Full Text

Cached at: 07/12/26, 12:53 PM

I’m just going to dump my whole agentic setup out here, because I see too many people missing giant chunks of this and it’s hurting them.

Here’s what I have and recommend:

  1. an AGENTS.md that is a router – it sends the agent to the right skills, docs, tools

  2. a standard workflow doc/skill customized to my needs … (grab Matt Pocock skills if you don’t already have something) … I tag this in most sessions with @/AGENT_WORKFLOW.md and it pulls it in.

  3. self-healing docs for every system, and agents are instructed to keep them updated … I tag the ones I know I need, or let the agent find them through AGENTS.md … I also provide a more detailed summary in the first 7 lines of every doc, so they’re easily greppable to find the right thing, and this is documented in AGENTS.md

  4. agents always run the app … the agent should always actually run the app itself, and test its work and fix issues as it goes, especially if running autonomously / asynchronously

  5. end-to-end tests and instructions to write more and keep up to date, and docs on how to write tests, what to avoid, and a list of all the tests and what they test in yet another markdown doc … write and run targeted tests during implementation, improve and commit with work

  6. custom linters at precommit hooks looking for any problems you run across, with --fix fixing the problems automatically, OR if that’s not feasible, it shells out to a cheaper LLM like Composer 2.5 or Sonnet to fix the problems – NOT just flagging them, but actually resulting in cleaned code

  7. cross-agent review at each major point: research, plan, implementation, and wrap-up. I mean codex, claude, cursor, whatever – but it shouldn’t be the same model reviewing the same code. And specific docs for agent review, what to look for, how to approach it. Also, personas – looking at the code from different perspectives, such as maintainability, code quality, security, performance, AI smells, domains (e.g. “financial services expert” or whatever) … and each persona also “owns” a set of system docs too and keeps them up to date

  8. agent traces / worksheets that track what the agent is doing each session. if the agent fails partway through, you should be able to hand this worksheet to another agent and it could finish the job. commit this worksheet with the work so it’s all connected and easy to reference later (you will reference these later!!), also have the agent apply git tags that correspond to specific worksheet names so they’re easy to find

  9. automatic agent feedback to you at the end of the session, added to a doc that is also committed with the work, that you periodically ingest into an interactive session and improve your workflows

  10. a tools or bin folder that contains python or bash scripts that the agent has skills to make to make its job easier (for example, I have an agent_review bash script that lets the agent kick off agent reviews via CLI without knowing each agent’s particular incantations) … docs on how to make scripts effectively, and instructions to constantly build these out more

  11. periodic agent sweeps through recent commits, looking for problems / gotchas from a higher level across commits

  12. a coding conventions doc that is just for specific coding conventions you want to see in the code base, your review agents use these a lot (but a lot of this should be in linters)

  13. an agent loop / night shift skill for autonomous work, that lays out how the agent is to approach this, from an orchestration standpoint

  14. a task queue that is accessible to the agent (mine is just a TODOS.md, but yours might be in Linear etc, with a CLI to fetch via API)

  15. a periodic false-confidence test audit skill that looks for tests that aren’t actually testing what you think they’re testing, and that fix those

  16. visual regression tests – take screenshots, compare via tool and with agent visual review, commit with work (git lfs useful here) or at least push into the PR

  17. automatic performance benchmark tests that notice when performance degrades

  18. performance profiling tools that can be used by agents for targeted benchmarking, trying new techniques, comparing outputs, and comparing profiles

  19. end-of-shift full validations, including running all tests, performance, agent reviews, sweeps, everything – when you return, it’s all as pristine as it can be

If you have all this, your agentic coding experience is going to be very different than dry prompting and manually guiding it toward the right thing every time.

Similar Articles

@itsclelia: I have one big problem with agentic engineering: I want agents to operate autonomously, but I also want granular, rever…

X AI KOLs Timeline

I have one big problem with agentic engineering: I want agents to operate autonomously, but I also want granular, reversible control over every change they make. I could solve this by committing every intermediate step to Git, but that would completely pollute my repo history. So I built 𝗮𝗴𝗴𝗶𝘁: a Git-like CLI for local and remote (S3-backed) agent artifact storage, written in Rust . With aggit, my agents can stash intermediate work, create branches safely, restore previous states, and back