@akshay_pachaar: YC open-sourced the agent harness running YC. Agent frameworks are almost always built for one person. You can stretch …
Summary
Y Combinator open-sources QM, a multiplayer agent harness for running agents across a whole organization, with scoped workspaces, durable sandboxes, and swappable harnesses like Claude Code and Codex.
View Cached Full Text
Cached at: 08/03/26, 11:42 AM
YC open-sourced the agent harness running YC.
Agent frameworks are almost always built for one person. You can stretch a personal assistant to cover a whole company, but the wiring gets complicated fast.
Y Combinator hit that wall directly. They started with a small agent loop in Ruby that could reach internal data, added crons and webhook triggers, then provisioned over 50 Hermes agents as personal assistants. Managing a fleet that size became its own job.
QM is what they built instead, open-sourced it under MIT. The name is short for quartermaster, the person on a ship who keeps things in order belowdecks. It runs inside YC across accounting, legal, events, and engineering, including the work of building QM itself.
The idea that makes it different is the scope.
→ A scope is the unit of isolation, and it applies the same way to a person and to a Slack channel. Each one gets its own memory, files, credential view, permissions, scheduled jobs, and sandbox.
→ That sandbox is durable, so tools installed during one session are still there in the next. Many execution environments get thrown away after every run, which means the agent starts from zero every time.
→ The core is deliberately small. It handles identity, policy, the scheduler, and the agent loop, with Postgres holding sessions, memory, and the queue. The web UI, admin panel, and Slack integration are all plugins over its HTTP API.
→ The harness layer is swappable. Claude Code, Codex, OpenCode, and Pi all drive the same core, so a deployment is not locked to one vendor.
→ Security comes in three levels. Strict pauses every tool call for human approval, Dangerous pauses nothing, and the default Auto runs a classifier over provenance-labelled external data and tool results before they reach the model. Hard denials for things like recursive deletes apply at every level.
That last one is worth sitting with. Screening tool results by where they came from is a concrete answer to prompt injection, and very few open agent systems ship a named default for it.
Repo: http://github.com/yc-software/qm
QM is a bet that the agent loop was never the hard part, and that coordination across an organization is. That is one of several bets every harness makes, thin core versus thick control being the biggest of them.
I wrote the full breakdown of what an agent harness actually is, and where designs like QM land on those tradeoffs.
The article is quoted below.
yc-software/qm
Source: https://github.com/yc-software/qm
qm
A multiplayer agent harness for work. In Slack and on the web.

What is QM?
Most agents are designed like personal assistants. You can make one work for a whole company, but it quickly gets complex. QM is designed for startups. Employees each get their own isolated workspace and work independently without affecting each other, and they can also collaborate with the agent in channels, group messages, and projects.
Each person and each room has its own scoped memory, files, keychain view, permissions, crons, web apps, and durable sandbox.
It’s built with open source in mind. Pick your own harness and model and switch between them — Pi, OpenCode, Codex, and Claude Code all drive the same core, so a deployment isn’t tied to any single vendor.
Features
- Personal and shared scopes. People customize the agent to be theirs, and still work with it collaboratively in Slack channels and projects.
- Slack and web. The same identity and configuration carries between Slack and the web app.
- Admin control. Set org-level configuration, a security posture, and which harnesses and models are available.
- Web apps. Spin up custom internal apps and publish them to the right people.
- Shared skills. Skills are scope-owned and shareable by grant, with admin-gated promotion to the whole org and skill packs imported from git repositories.
- Background work. Crons and watches run work while nobody’s watching.
What you can do with it
- Search internal notes, email, documents, databases, and the web together
- Retrieve information from your company brain
- Build internal apps, publish them to the right people, and keep their data current
- Learn your writing voice from past sends, then triage your inbox on a schedule — labels and reply drafts included
- Work in an existing repository: run tests, open PRs, monitor CI, check system logs
- Track a project in a shared channel and post updates and follow-ups
Architecture
flowchart LR
DB[("Postgres<br/>sessions · memory · queue")]
subgraph CORE["Headless core"]
API["API · identity · policy · scheduler"]
LOOP["Agent loop<br/>(Pi, OpenCode, Claude Code)"]
API <--> LOOP
end
SBX["Per-scope sandbox<br/>files · tools · logged-in services"]
DB <--> API
LOOP <--> SBX
Every turn runs through a central core, which can use a variety of models and harnesses
to generate the response. A Postgres persistence layer holds user data, session history,
and other durable state. The agent has a small, fixed tool surface; one of those tools is
execute, which runs commands in the scope’s own isolated sandbox — its durable computer,
where installed tools stay installed. The web UI, the admin panel, and the public portal
are optional plugins over the core’s HTTP API;
Slack is an optional in-process plugin that core starts
and supervises through a direct service client.
The core runs TypeScript directly on Node and uses Fastify for HTTP. The Slack plugin uses Bolt; the web UI builds with Vite and renders with Lit.
The core itself is generic. Everything specific to one company — org config, custom tools
and skills, sandbox image, infrastructure — lives in a deployment directory that the
qm CLI validates and deploys. Every substrate (harness, session
store, sandbox, memory) sits behind an interface, so production implementations swap in
via one wiring file.
Security and secrets
QM’s approach follows local coding agents like OpenCode, Codex, and Claude Code: the agent acts as the person it’s working for, with their credentials and permissions, and everything it does is audited. An org picks one security posture, which narrower scopes can only tighten:
- Strict — every harness tool call pauses for human approval, except the two no-effect turn enders.
- Auto (default) — a classifier screens provenance-labelled external data and tool results before they reach the model; a deployment can point that at its own screening proxy.
- Dangerous — no content screening, no pauses between tool calls.
The predeclared command policy — approval rules and hard denials for things like recursive deletes or destructive SQL — applies in every posture, Dangerous included.
SECURITY.md has the threat model, the operator assumptions, and the
known limitations.
Deploy it for your org
Create an organization-owned deployment repository that depends on @yc-software/qm:
npm exec --yes --package=@yc-software/qm@latest -- \
qm init . --org <slug> --target <fly-or-aws>
npm install
Initialization materializes a deployment skill for an agent and walks through
infrastructure, web sign-in, connector credentials, optional Slack access, deployment,
and live verification — no source checkout required. Each deployment runs in the
operator’s own cloud account; initialization does not generate or enable deployment CI,
and this repository has no production deployment workflow. See
deployment.md for the details.
Contributing
We take contributions as human-written text, not code — see
CONTRIBUTING.md. Describe the change you’d like informally in a
.txt or .md file in adrs/, and if we’re aligned we’ll handle the
implementation. Report vulnerabilities privately — see SECURITY.md,
not a public issue.
Customize your instance
The deployment repository above carries config and a sandbox layer, and never needs a source checkout. Some organizations want the opposite trade: the whole codebase in one place, so engineers and coding agents read core and customizations together, while the customizations themselves stay private. For that, keep a private fork: a standalone private repository whose history begins as a clone of qm and whose core stays identical to upstream.
Populate it once, then clone it to work in:
gh repo create <org>/qm-private --private
git clone --bare [email protected]:yc-software/qm qm-seed.git
git -C qm-seed.git push --mirror [email protected]:<org>/qm-private
rm -rf qm-seed.git
git clone [email protected]:<org>/qm-private
git -C qm-private remote add upstream [email protected]:yc-software/qm
Create the private fork with a plain clone, as shown above, and never with GitHub’s fork feature. The word “fork” here names the concept — a downstream copy that diverges deliberately and merges from upstream — not GitHub’s Fork button. A GitHub fork inherits the visibility of the repository it came from, so a fork of a public repository cannot be made private. A GitHub fork also shares one object network with the repository it came from, so commits pushed to the fork stay fetchable by SHA from the public side. Many organizations disallow forking private repositories as well. A plain clone has none of these problems, and it costs one thing: the clone is an ordinary repository, so upstream’s CI workflows run live in your own account. Expect to supply the secrets those workflows need, or disable the ones you do not want running.
Everything specific to your organization goes in deploy/layers/<org>/ — config, sandbox
tools and skills, plugin images, infrastructure — in the same shape qm init produces. See
deploy/layers/README.md. Core stays byte-identical to
upstream, which is what keeps merges small.
Two skills maintain the boundary in both directions. update-qm merges upstream qm into
the private fork and opens the sync PR; upstream-pr sends an organization-agnostic fix back to
qm, cutting the branch from upstream/main and checking the outgoing diff, commit
messages, and screenshots for organization identifiers before it pushes. Nothing under
deploy/layers/ ever travels upstream.
Going deeper
docs/getting-started.md— first run, end to endcli/README.md— theqmCLI and the deployment directory contractdocs/deploy-directory.md— the deployment directory in full.env.example— every knob, documented in placeplugins/— the surfaces (Slack, web UI, admin, portal)
License
Except where otherwise noted, QM is available under the MIT License.
Similar Articles
YC open-sourced a multiplayer AI agent
Y Combinator open-sourced QM, the multiplayer AI agent harness they use to run their company. It provides per-team workspaces, per-project memory and scheduled jobs, plus a keychain that shares credentials without exposing raw API keys.
Your agent is only as good as its harness. I open-sourced one with 40 capabilities behind a single function call
An open-source agent harness with 40 capabilities behind a single function call, including persistent memory, Docker sandbox, auto-summarization, stuck-loop detection, budget caps, and live run forking for branching agent execution. Built on Pydantic AI and designed to replace the 2000 lines of glue code every production agent needs.
qm
QM is an open-source multiplayer agent harness for startups, allowing employees to work with AI agents in isolated personal and shared workspaces across Slack and the web. It supports multiple agent harnesses like Pi, OpenCode, Codex, and Claude Code, with scoped memory, files, crons, and durable sandboxes.
@eyad_khrais: https://x.com/eyad_khrais/status/2069552027382980882
A comprehensive guide to building AI agent harnesses, covering tool execution, context management, state/memory, and guardrails, based on lessons from building Claude Code and other harnesses for enterprise.
I open-sourced the "harness" layer for AI agents: run Claude Code/Codex/Gemini with governed MCP tools (browser, editor, secrets)
Open-sourced a desktop workspace that provides a governed runtime for AI coding agents, offering 100+ MCP tools, RBAC, and a self-evolving toolbox.