@ssh_exe_dev: A Non-Exhaustive Inventory of exe's Software Factory: An agent that looks for security issues systematically. An agent …
Summary
A blog post from exe.dev cataloging their internal software factory: multiple AI agents for security review, alert investigation, log analysis, flaky tests, deploys, plus a custom CMS and self-healing UI tests.
View Cached Full Text
Cached at: 08/08/26, 07:12 PM
A Non-Exhaustive Inventory of exe’s Software Factory:
An agent that looks for security issues systematically. An agent that investigates alerts. An agent that investigates logs. Bots to fix flaky and slow tests. A status page. A system to page our phones (using the excellent and simple PushOver).
Read on:
A Non-Exhaustive Inventory of exe’s Software Factory
Source: https://blog.exe.dev/inventory
- An agent that looks for security issues systematically. Fable refuses to help out, so we systematically look for security issues, with a bias toward recent changes.
- An agent that investigates alerts. Sisyphus keeps track of our alerts. We have ones that page us and ones that merely make noise in Slack. Either way, Sisyphus looks through our logs and metrics and source code, as well as analyzing its own previous investigations, to tease out what’s going on. It gives a great head start when investigating an issue (or just a flaky alert!).
- An agent that investigates logs. Every day, I get an email with interesting trends in our logs.
- Bots to fix flaky and slow tests. A bot is continuously analyzing flaky or slow tests in our CI and suggesting changes.
- A status page. status.exe.devisn’t hosted on exe.dev. We built it ourselves, though.
- A system to page our phones (using the excellent and simple PushOver) When the aforementioned alerts fire, our phones beep very loudly. Traditionally you use PagerDuty for this, but PagerDuty’s durable asset is the entitlement for “Emergency Alerts” from Apple. Turns out PushOver has this as well, and a lovely API.
- An agent that supervises deploys and rollouts. Athena helps do rollouts. Infrastructure deploys are not instantaneous, and even the most patient operators stop paying attention. It checks metrics and logs (and has looked at the source code for what changed in this deploy).
- A blog CMS, with comments, collaborative text editing, embargoes, the whole nine yards If you’re reading this on blog.exe.dev, this ain’t Wordpress. Our blog started out as Markdown files in git, but now there’s a full-featured CMS, with collaborative editing, revision history, comments, embargoes, and a content calendar. A built-in agent (really, Shelley running on the same VM) can import a blog post from whatever you paste in.
- UI tests described as textual paragraphs that lazily materialize into browser instructions but self-heal Who are we kidding? We’re not maintaining Playwright tests by hand anymore. Shelley’s UI tests are increasingly a paragraph of text asking for some behavior. There’s a cache file (checked into git) that makes the test cheap and fast. When it fails, the CI system “heals” it with an LLM, and either fails or checks in the new fixed test. Yes, the build queue modifies the commit on its way through if necessary. More on this in a future post.
- Intrepid reporter bots that report on git commits, our help threads, and so on Every day, we get summaries in Slack about what’s happened in the past day, across git commits and such.
Please note: if you’re writing bots that read untrusted data, understand theLethal Trifecta: private data, untrusted content, and external communication. We happen to think thatexe.devVMs are a great place to isolate these bots, but we also make sure that the tools available to these agentic loops (an agent is just 11 lines of code:https://sketch.dev/blog/agent-loop) are limited in what they can do.
Similar Articles
@0x0SojalSec: Awesome AI Security : Everyone’s racing to deploy AI agents, Almost few peoples is securing them properly. this repo co…
A curated GitHub repository aggregating frameworks, tools, attack matrices, red team guides, policy templates, datasets, and research for securing AI systems, covering topics like prompt injection, jailbreaking, and OWASP/NIST standards.
Devs shipping AI agents what does your security testing look like ?
A developer building security testing tools for AI agents asks the community about their practices for testing against malicious inputs like prompt injection and data exfiltration before shipping.
@josesilesdata: GOODBYE TO CYBERSECURITY! A repository just came out with hundreds of AI security tools in an open-source repository. T…
An open-source repository containing hundreds of AI security tools has been released, featuring techniques for jailbreaking LLMs, prompt injection testing, red team agents, model extraction, and automated pentesting.
@UnTalNixon_exe: THIS IS INSANE An open-source AI that hacks your app BEFORE real attackers do +56,000 stars. Multi-agent. Real working …
Strix is an open-source AI tool that uses multi-agent architecture to autonomously perform security testing, covering OWASP Top 10 vulnerabilities with real-world PoCs and a 30-second setup.
@aarondfrancis: A good way to audit your codebase: • a strong orchestrator inventories every subsystem • it sends fresh read-only agent…
A method using a strong orchestrator and read-only agents with a DSA prompt to audit codebases, finding 93 opportunities across 55 subsystems overnight.