@eng_khairallah1: Claude Opus 5.5 isn't just for chatting. Builders are using it to run teams of AI agents, write content, review and mak…
Summary
The article provides a guide on building coordinated AI agent teams using Claude Opus 5.5, emphasizing the model's cost-efficiency and architectural advantages for multi-agent systems.
View Cached Full Text
Cached at: 09/29/26, 11:55 AM
Claude Opus 5.5 isn’t just for chatting.
Builders are using it to run teams of AI agents, write content, review and make real money.
Here’s how to build your first team from scratch. https://t.co/0yssII21AE
How to Build Your First Team of AI Agents Using Claude Opus 5.5
Everyone is talking about AI agents.
Save this :)
Build an agent. Deploy an agent. Agent this. Agent that.
Buried in the prompting guide Anthropic published alongside Claude Opus 5.5 is a result that got almost none of the launch coverage.
When Anthropic gave a team of agents an elapsed-time budget, the team kept answer quality comparable to a single agent’s while finishing considerably sooner.
Read that again, because it quietly reframes the question every builder has been asking for two years. The question was always which model. Opus 5.5 answers that one convincingly: it scored 66.4% on Terminal-Bench 4.0 against 55.8% for the far larger Claude Fable 5.1, it costs roughly 40% less than Opus 5 to run on typical workloads, and one early tester completed a 680,000-line code migration with it in less than a day. But the result that matters for what you build next is not about the model at all. It is about what happens when you stop running one agent and start running a coordinated team.
Here is the problem. Most people building “multi-agent systems” right now are building expensive chaos. Five agents with vague roles, full permissions, no spending cap, passing half-finished context to each other and returning confident garbage in parallel. It demos beautifully and fails in production, and the failure usually gets blamed on the model.
The model is not the problem. The architecture is.
This course walks you through building a team that actually works: an orchestrator that plans and delegates, specialists with narrow jobs and narrower permissions, a critic that sends weak work back, and the budgets, caps, and stop conditions that keep the whole thing from spiraling. Everything here uses the Claude Agent SDK and the real Opus 5.5 behaviors documented by Anthropic, not recycled advice from the last model generation.
Let’s build.
Why Teams, and Why Opus 5.5 Changes the Math
Before any code, understand what a team actually buys you. Anthropic’s Agent SDK documentation names four concrete benefits of delegating to subagents, and every one of them is a reason, not a vibe.
Context isolation. Each subagent runs in its own fresh conversation. Its intermediate tool calls and results stay inside it; only its final message returns to the parent. A research subagent can read fifty pages without any of that content bloating the orchestrator’s context. The orchestrator receives a clean summary, not a transcript.
Parallelization. Independent subtasks finish in the time of the slowest one rather than the sum of all of them. Three research angles investigated at once take as long as the hardest one, not three times as long.
Specialized instructions. Each subagent gets a tailored system prompt with its own expertise and constraints, without adding noise to the orchestrator’s instructions.
Tool restrictions. Each subagent can be limited to specific tools. A reviewer that should never edit files simply never receives an edit tool.
Now the Opus 5.5 part. Three documented behaviors make this model unusually well suited to running teams.
First, cost. Opus 5.5 is priced at $4 per million input tokens and $20 per million output, down from $5 and $25 on Opus 5, with cache reads at $0.20 per million. A team multiplies your token spend, so a cheaper frontier model changes which teams are economically sensible to run at all.
Second, effort. Opus 5.5 defaults to medium effort, and Anthropic reports that at medium it matches or exceeds Opus 5 at high on coding and knowledge-work evaluations. Anthropic’s effort documentation explicitly lists subagents as a typical use case for low effort. That means your workers can run lean while your orchestrator keeps the reasoning depth it needs.
Third, time awareness. Anthropic’s prompting guide states that Opus 5.5 “pays close attention to information about elapsed time” and can use it to optimize parallelization in agent teams. That is the finding from the top of this article, and you will use it in Step 10.
The First Rule: Most Tasks Do Not Need a Team
The part the hype skips, and the part that will save you the most money.
A team adds coordination overhead, more API calls, more tokens, and new ways to fail. If your task is one job with one clear way to check it, a single well-prompted agent will beat a team every time, faster and cheaper.
Use a team only when at least one of the four benefits above is doing real work. The job naturally splits into independent pieces that can run at once. Or a subtask would flood the main context with material nobody needs afterward. Or one part of the work genuinely needs different expertise or different permissions than the rest.
If none of those apply, close this tab and build a single agent. If they do apply, keep reading.
The Anatomy of a Team That Works
Every reliable agent team is some combination of three roles.
The orchestrator receives the goal, breaks it into subtasks, assigns each to the right specialist, and assembles the result. It plans and integrates. It does not do the deep work itself. In a well-built team, it is the only agent you talk to.
The specialists each do one narrow job extremely well. A researcher that only gathers and verifies facts. A writer that only drafts from verified material. An analyst that only runs numbers. The narrower the role, the better the output, because a focused instruction beats a vague one every time.
The critic is the role almost everyone skips and the one that separates professional systems from demos. Its only job is to review output against a standard and send it back if it falls short. A team without a critic produces fast, confident garbage. A team with one produces work you can ship.
Get these three right and you have the skeleton of every team worth building.
Step 1: Set Up the Agent SDK
The Claude Agent SDK gives you the same agent loop, tools, and context management that power Claude Code, programmable in Python and TypeScript. It is the right tool for a team you run yourself.
You need Python 3.10 or newer, or Node.js 18 or newer, and an API key from the Claude Console.
bashpython3 -m venv .venv source .venv/bin/activate pip install claude-agent-sdk export ANTHROPIC_API_KEY=your-api-key
For TypeScript:
bashnpm install @anthropic-ai/claude-agent-sdk
The SDK reads the key from the environment of the process that runs your agent. It does not load .env files automatically, so if you keep your key in one, load it yourself before calling the SDK.
One note if you are building a product for other people: Anthropic does not allow third-party developers to offer claude.ai login or rate limits in their products without prior approval. Use API key authentication.
Step 2: Build the Orchestrator Alone First
The most common mistake is designing the full team before a single agent works. Do not do that. Start with the orchestrator running solo, prove the task is achievable, and only then start delegating pieces of it.
pythonimport asyncio from claude_agent_sdk import query, ClaudeAgentOptions, ResultMessage
async def main(): async for message in query( prompt=“Research the three most-discussed AI agent frameworks this month and write a one-page briefing to drafts/briefing.md.”, options=ClaudeAgentOptions( model=“claude-opus-5-5”, allowed_tools=[“WebSearch”, “WebFetch”, “Write”], permission_mode=“acceptEdits”, ), ): if isinstance(message, ResultMessage): print(f“{message.subtype}: ${message.total_cost_usd}“)
asyncio.run(main())
Run it. Read the output. Note where it struggled, where it spent the most tokens, and which parts were genuinely independent of each other. Those notes are your team design. The subtasks that were independent become parallel specialists. The part where quality was shaky becomes the critic’s job.
Step 3: Define Your First Specialist
Specialists are defined with AgentDefinition and passed through the agents option. Two fields are required, and one of them matters far more than people realize.
pythonfrom claude_agent_sdk import AgentDefinition
researcher = AgentDefinition( description=“Gathers and verifies facts on one focused question. Use for any research subtask.”, prompt=“”“You research exactly one question and return verified findings. For every claim, include the source URL. If you cannot verify a claim, mark it UNVERIFIED instead of dropping it or guessing. Return a numbered list: claim, source, confidence (high, medium, low). Do not write prose.”“”, tools=[“WebSearch”, “WebFetch”], effort=“low”, )
The prompt is the specialist’s system prompt. Give it a role, a standard, an output format, and boundaries, exactly as you would brief a contractor.
The description is the field most people underwrite, and it is the one that decides whether your team works. The orchestrator reads each specialist’s description to decide when to delegate to it. A vague description means the orchestrator either never uses the specialist or uses it for the wrong things. Write it as a precise rule for when this agent should be called. If you need to guarantee a specific specialist runs, name it directly in your prompt, for example “Use the researcher agent to check each claim.”
One more thing that trips up everyone: your orchestrator’s allowed_tools must include Agent, because subagents are invoked through the Agent tool.
Step 4: Give Every Role the Least Power It Needs
This is the step that turns a fragile demo into something you can leave running.
A tool you leave out of a specialist’s tools list is not in its session at all. Claude works without it, with no permission prompt and no error. That makes tool restriction the cleanest safety control you have.
Anthropic’s documentation suggests these combinations as a starting point:
-
Read-only analysis: Read, Grep, Glob. Can examine but never modify or execute.
-
Test execution: Bash, Read, Grep. Can run commands and analyze output.
-
Code modification: Read, Edit, Write, Grep, Glob. Full read and write, no command execution.
Apply the principle to every role. Your critic gets read-only tools, because a critic that can edit will “fix” things quietly instead of reporting them. Your researcher gets search and fetch, and nothing that writes. Only the agent whose job is producing the file gets Write.
Opus 5.5 also makes this safer than before. In Anthropic’s containment testing, Opus 5.5 attempted to circumvent boundaries about 85% less often than Opus 5. Stated limits are more load-bearing than they used to be. That is a reason to state them clearly, not a reason to skip them.
Step 5: Solve the Handoff Problem
Here is the detail that silently breaks most first teams.
A subagent starts with a fresh context. It does not see the parent’s conversation history, the parent’s tool results, or the parent’s system prompt. Anthropic’s documentation is explicit: the only content passed from parent to subagent is the Agent tool’s prompt string.
That means if your orchestrator delegates with “now write it up,” the writer has no idea what “it” is. Every file path, every decision, every constraint, every piece of evidence the specialist needs has to be inside the delegation message itself.
Put this directly in your orchestrator’s instructions:
plaintextWhen you delegate, write each subagent a complete brief. Include every file path, finding, decision, and constraint it needs. Assume it has seen nothing of this conversation. A subagent that has to guess is a subagent that will guess wrong.
The flip side also matters. The parent receives the subagent’s final message, but may summarize it. If you need a specialist’s output preserved exactly, instruct the orchestrator to keep it verbatim.
Step 6: Run Independent Work in Parallel
The speed advantage of a team only appears when work genuinely runs at once.
Multiple subagents can run concurrently, and subagents run in the background by default, so the orchestrator can dispatch several and keep working. The practical move is to design your subtasks to be independent. Three researchers each investigating a different framework can run simultaneously. A writer who needs all three reports cannot start until they finish.
Tell the orchestrator explicitly which parts can run together:
plaintextThe three framework investigations are independent. Dispatch all three researchers at once, then wait for every report before briefing the writer.
If you build custom tools for your team, one more lever applies. Marking a tool with readOnlyHint=True tells the SDK it has no side effects, which lets it be called in parallel with other read-only tools. Only mark tools that genuinely modify nothing.
Step 7: Add the Critic
Now the role that makes the output trustworthy.
pythoncritic = AgentDefinition( description=“Reviews a finished draft against the research findings it was written from. Use after every draft, before anything is final.”, prompt=“”“You review a draft against the findings it was written from. Check that every factual claim appears in the findings. Check that nothing marked UNVERIFIED is stated as fact. Check that the structure matches the brief. Return PASS, or FAIL followed by a numbered list of specific fixes. You never edit the draft yourself.”“”, tools=[“Read”], effort=“medium”, )
Then wire the loop into your orchestrator’s instructions: every draft goes to the critic, a FAIL goes back to the writer with the critic’s exact list, and the cycle repeats until PASS.
Cap that cycle. A worker and critic that disagree forever will burn your budget arguing. Tell the orchestrator to stop after three rounds and return the best draft with the critic’s outstanding objections attached. A clear exit is the difference between a system and a loop.
Step 8: Set Effort Per Role
This is one of the biggest cost levers Opus 5.5 gives you, and most people leave it untouched.
Effort has five levels: low, medium, high, xhigh, and max. It controls how much work the model puts into the whole response, including how often and how deeply it thinks. On Opus 5.5, thinking is always on and cannot be disabled, so effort is the primary control for how much the model reasons.
AgentDefinition accepts an effort field, so you can tune every role independently:
-
Researchers and simple workers: low. Anthropic’s effort documentation lists subagents as a typical low use case.
-
Writer and critic: medium. The default, and on Opus 5.5 it matches or beats Opus 5 at high on coding and knowledge work.
-
Orchestrator: medium to start, high if your evals show planning quality improves.
Reserve xhigh for long-running agentic work over roughly 30 minutes, and max for tasks where you have measured a real quality gain. Anthropic’s guidance is to run an effort sweep on your own evals rather than carrying settings over from an earlier model.
Step 9: Cap Depth, Concurrency, and Spend
Claude decides on its own when to spawn subagents and how many. Subagents can spawn their own subagents. Without limits, one prompt can grow into a tree of agents and a bill you did not plan for.
The SDK gives you three caps:
-
Depth: CLAUDE_CODE_MAX_SUBAGENT_SPAWN_DEPTH. Defaults to 3 layers. Set it to 1 so your specialists cannot spawn their own.
-
Concurrency: CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS. Defaults to 20 running at once.
-
Spend: max_budget_usd in Python, maxBudgetUsd in TypeScript. No limit by default.
That last default deserves a moment. Your first team run has no spending ceiling unless you set one. Set one before you press enter, every time.
pythonoptions = ClaudeAgentOptions( env={ “CLAUDE_CODE_MAX_SUBAGENT_SPAWN_DEPTH”: “1”, “CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS”: “5”, }, max_budget_usd=5.0, )
At the cap, the SDK refuses to spawn more subagents, stops background subagents still running, and ends the query with the error_max_budget_usd result subtype.
Step 10: Give the Team a Clock
Now the finding from the top of this article.
Anthropic’s Opus 5.5 prompting guide recommends giving a multi-agent run a sense of time. Append elapsed time to each message in the form elapsed 340s / 1200s, set the budget somewhat above the time you actually want, and tune it on samples. If you cannot predict duration, show elapsed time alone and add this line to the system prompt:
plaintextTime matters here: do not spend time that can be avoided, and the earlier a correct result is obtained, the better.
In Anthropic’s testing, teams given a budget kept answer quality comparable to the single agent’s while finishing considerably sooner.
One critical caveat from the same guide: the budget is advisory. The model uses it to pace itself, not as a hard stop. Enforce the real limit in your own code:
pythontry: await asyncio.wait_for(run_team(), timeout=1200) except asyncio.TimeoutError: print(“Hard stop reached. Check drafts/ for partial output.”)
Step 11: Stop the Team From Stopping Early
One Opus 5.5 behavior will break your team if nobody warns you.
Anthropic’s prompting guide documents that Opus 5.5 tends to end turns with a text update rather than a tool call. In an unattended agent loop, a message with no tool call ends the turn, and the work stops there. Your orchestrator writes a lovely summary announcing what it will do next, and then does not do it.
Anthropic’s guide describes four specific patterns to prevent: a summary that announces the next step with no tool call, an offer to continue unless you object, a list of decisions that do not actually block the work, and deciding a milestone is a good place to report. The model is, in Anthropic’s words, “responsive to instructions that name the specific kinds of early stop you want it to avoid.” So name them in your orchestrator’s system prompt. A condensed version:
plaintextYou are running unattended. A message without a tool call ends your turn and stops all work. Never end a turn with: a summary that announces the next step, an offer to continue, a list of non-blocking decisions, or a report at a milestone. Put status notes in the same message as your next tool call, and keep going until the brief is complete or genuinely blocked.
Anthropic’s guide also recommends keeping the task as a checklist the model updates, and if your harness sends automatic continuations, stopping after two or three so a stuck run cannot loop forever.
The Complete Team
Here is everything assembled into a working research-and-writing team.
pythonimport asyncio from claude_agent_sdk import query, ClaudeAgentOptions, AgentDefinition, ResultMessage
ORCHESTRATOR = “”“You lead a research and writing team. You plan and delegate; you do not research or write yourself.
Process:
- Split the brief into independent research questions. Dispatch one researcher per question, all at once.
- When every report is back, give the writer a complete brief containing all the findings and the output path. Assume it has seen nothing of this conversation.
- Send every draft to the critic. On FAIL, return the critic’s exact fix list to the writer. Stop after three rounds and report outstanding objections.
You are running unattended. A message without a tool call ends your turn and stops all work. Never end a turn with a summary that announces the next step, an offer to continue, or a milestone report. Keep going until the draft passes or you hit the round limit.
Time matters here: do not spend time that can be avoided, and the earlier a correct result is obtained, the better.“”“
AGENTS = { “researcher”: AgentDefinition( description=“Gathers and verifies facts on one focused question. Use for every research subtask.”, prompt=“”“You research exactly one question. For every claim, include the source URL. Mark anything you cannot verify as UNVERIFIED. Return a numbered list: claim, source, confidence. No prose.”“”, tools=[“WebSearch”, “WebFetch”], effort=“low”, ), “writer”: AgentDefinition( description=“Writes a draft strictly from supplied findings. Use only after research is complete.”, prompt=“”“Write the draft using only the findings in your brief. Add no facts of your own. Save it as markdown at the path you are given.”“”, tools=[“Read”, “Write”], effort=“medium”, ), “critic”: AgentDefinition( description=“Reviews a draft against its findings. Use after every draft, before anything is final.”, prompt=“”“Check every factual claim against the findings. Flag anything UNVERIFIED stated as fact. Return PASS, or FAIL with a numbered list of specific fixes. Never edit the draft.”“”, tools=[“Read”], effort=“medium”, ), }
BRIEF = “”“Write a one-page briefing comparing the three most-discussed AI agent frameworks this month: what each does, who it suits, and one honest limitation of each. Save it to drafts/briefing.md.”“”
async def run_team(): async for message in query( prompt=BRIEF, options=ClaudeAgentOptions( model=“claude-opus-5-5”, system_prompt=ORCHESTRATOR, agents=AGENTS, allowed_tools=[“Agent”, “Read”, “Write”, “WebSearch”, “WebFetch”], permission_mode=“acceptEdits”, env={ “CLAUDE_CODE_MAX_SUBAGENT_SPAWN_DEPTH”: “1”, “CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS”: “5”, }, max_budget_usd=5.0, ), ): if isinstance(message, ResultMessage): print(f“{message.subtype}: ${message.total_cost_usd}“)
async def main(): try: await asyncio.wait_for(run_team(), timeout=1200) except asyncio.TimeoutError: print(“Hard stop reached. Check drafts/ for partial output.”) except Exception as error: print(f“Run ended with an error: {error}“)
asyncio.run(main())
Look at what each piece is doing. The orchestrator plans and delegates. Three researchers run in parallel at low effort with search-only tools. The writer works strictly from verified findings. The critic can read but never edit. Depth, concurrency, spend, and wall-clock time are all capped. And the orchestrator knows exactly which early stops to avoid.
The Mistakes That Kill Agent Teams
Building five agents before one works. Earn every new agent. One excellent orchestrator beats five mediocre agents wired together.
Vague descriptions. The orchestrator routes by description. “Helps with research” routes nothing reliably. “Gathers and verifies facts on one focused question” routes correctly.
Thin handoffs. Subagents see only the delegation message. If the brief says “write it up,” the writer has nothing to write from.
No critic. Parallel speed without review just produces wrong answers faster.
Full permissions everywhere. Every tool a role does not need is a mistake it can now make.
No spending cap. The default is unlimited. Set max_budget_usd before your first run.
Trusting the advisory clock. The elapsed-time budget paces the team. Your own timeout stops it.
Letting the team act irreversibly. Drafting an email is delegation. Sending it is your decision. Keep anything that cannot be undone, sending, publishing, spending, deleting, behind a human approval step.
The Honest Truth About Agent Teams
A team of agents will not fix a process you do not understand.
Every specialist needs a clear instruction, and you are the one writing it. If you cannot describe how a task should be done step by step, you cannot delegate it to five agents any better than to one. The code in this course takes an afternoon. The clear thinking about your own process is the real work, and no model release, however strong, does that part for you.
But here is what makes it worth doing. Opus 5.5 made frontier intelligence meaningfully cheaper, made effort a per-role dial, and gave teams a documented way to pace themselves against a clock. The builders who learn to orchestrate now are not competing with AI. They are running a team that researches, writes, and reviews while they do the work only they can do.
Most people will keep typing one prompt into one agent and waiting for one answer.
A small group will build a team this week, cap it, clock it, and give it a critic.
Six weeks from now, those two groups will not be doing the same job.
Start with the orchestrator alone. Then earn the first specialist.
Follow me @eng_khairallah1 for more AI courses, tools, and workflows. New content every week.
hope this was useful for you, Khairallah ❤️
Similar Articles
@0xCodez: https://x.com/0xCodez/status/2058513716509913581
A comprehensive walkthrough on building multi-agent teams with Claude Managed Agents, covering role design, model mixing, and parallel execution to scale from one to 20 agents.
@AnatoliKopadze: Head of Claude Code: "85% of our engineers are running dozens or hundreds of agents. The way you do it is graph enginee…
The Head of Claude Code at Anthropic discusses how 85% of engineers are using AI agents and graph engineering to perform the work of entire teams, highlighting internal practices and future directions.
@eng_khairallah1: https://x.com/eng_khairallah1/status/2058116763372453997
A comprehensive guide teaching non-coders how to build AI agents using Claude and Cowork without writing any code, explaining the core components and providing step-by-step instructions.
I let Codex and Claude Opus work on the same Java AI agent monolith
A developer compares Codex 5.3 and Claude Opus 4.6 on autonomous Java AI agent development, finding that the model with more elegant architecture (Claude) often produced code that never executed, while the more boring and direct Codex improved the working product with practical fixes like timeouts and history recovery.
Introducing Claude Opus 5
Anthropic announces Claude Opus 5, a powerful and cost-effective AI model that approaches the intelligence of Claude Fable 5 at half the price, achieving state-of-the-art results on coding and knowledge work benchmarks.