ProACT: Towards Breakdown-Aware Proactive Agent in Multi-User Collaboration
Summary
ProACT introduces a breakdown-aware agent framework for multi-user collaboration, where the agent proactively detects collaboration breakdowns and decides whether to intervene. The paper also presents the first multi-user collaboration benchmark for evaluating such proactive agents.
View Cached Full Text
Cached at: 07/07/26, 04:38 AM
# Towards Breakdown-Aware Proactive Agent in Multi-User Collaboration
Source: [https://arxiv.org/html/2607.03730](https://arxiv.org/html/2607.03730)
Shu Yang1,Difei Xu1,Jiaxin Pei2,Di Wang1† 1PRADA Lab, King Abdullah University of Science and Technology 2Stanford University
###### Abstract
Conversational agents are increasingly embedded in human collaborative work, yet they remain fundamentallypassiveandreactive: they respond to explicit user requests rather than proactively recognizing moments when a team would benefit from timely intervention as human collaborators often do\. This reactive design substantially limits the use of agents as active participants in multi\-user collaboration, where disagreements, ambiguous goals, forgotten constraints, underspecified plans, discussion loops, and imbalanced participation can gradually undermine group progress\. To move agents from passive assistants toward active participants in multi\-user collaboration, we introduce ProACT, a breakdown\-aware agent framework grounded in theories of common ground, collaborative planning, and coordination work\. ProACT observes the speaker\-attributed conversation history, determines whether the current turn contains a collaboration breakdown requiring intervention, decides whether the agent should stay silent or speak, and, when speaking is needed, routes the case to a targeted collaboration skill\. We further introduce the first multi\-user collaboration benchmark for evaluating proactive agents across project planning, product design, research collaboration, logistics, education, and resource\-constrained decision making\. Across 3,244 turn\-level examples and five LLM backbones, ProACT consistently improves collaborative appropriateness, non\-interruptiveness, conciseness, and judged intervention quality over direct chat\.
ProACT: Towards Breakdown\-Aware Proactive Agent in Multi\-User Collaboration
††footnotetext:Corresponding Author## 1Introduction
Large language model \(LLM\) based agents are increasingly embedded in our collaborative workflows: they can help summarize meetingsKirsteinet al\.\([2024](https://arxiv.org/html/2607.03730#bib.bib35)\), draft plansWeiet al\.\([2025](https://arxiv.org/html/2607.03730#bib.bib29)\), write codeJimenezet al\.\([2024](https://arxiv.org/html/2607.03730#bib.bib36)\), and deliberationsAbdelnabiet al\.\([2024](https://arxiv.org/html/2607.03730#bib.bib38)\); Zhuet al\.\([2025](https://arxiv.org/html/2607.03730#bib.bib40)\)\. Despite this growing role, existing deployed agents remain fundamentallypassiveandreactive\. They wait until the user explicitly asks a question, issues an instruction, or provides a predefined plan and prompts them to implement\. This design misses many moments where an agent could actively participate in collaboration rather than merely assist on demand\.
Figure 1:Illustration of moving agents from passive assistance to proactive participation in multi\-user collaboration\.Specifically, in multi\-user scenarios, breakdowns often arise when participants lose alignment over shared understanding, commitments, responsibilities, or active constraints\(Clark and Brennan,[1991](https://arxiv.org/html/2607.03730#bib.bib2); Grosz and Kraus,[1996](https://arxiv.org/html/2607.03730#bib.bib5)\)\. These failures are rarely abrupt; they accumulate through small coordination gaps\. For example, one participant may optimize for speed while another optimizes for quality, creating a task or priority conflict\(Jehn,[1995](https://arxiv.org/html/2607.03730#bib.bib4)\); a team may repeatedly revisit the same unresolved decision without recognizing that the discussion has entered a loop\.
In such cases, moving agents from passive Assistance to proactive participation in multi\-user collaboration can be quite helpful, as illustrated in Figure[1](https://arxiv.org/html/2607.03730#S1.F1)\. To achieve this, we propose ProACT, a breakdown\-aware proactive framework for agents, which keeps observing the multi\-party conversation and maintains a structured collaboration state\. In contrast to prior work such asLuet al\.\([2025](https://arxiv.org/html/2607.03730#bib.bib37)\), which focuses on predicting user intent from single\-user input, we argue that proactive agents for collaboration require a different formulation from single\-user proactive assistance\. Rather than asking only how to help an individual user, a collaborative proactive agent must ask:*When should I intervene in a group interaction where users may have different, uncertain, or conflicting intents?*This introduces social and normative challenges that are largely absent from single\-user interaction\. So ProACT continues to activately detect emerging breakdowns, such as conflict, uncertainty, underspecified plans, forgotten constraints, and discussion loops\. When a breakdown requires intervention, ProACT selects a targeted collaboration skill and speaks only when a concise, neutral intervention is likely to improve coordination without interrupting productive discussion\.
To validate ProACT, we introduce a benchmark for evaluating proactive agents in multi\-user collaboration\. Each example presents a multi\-party conversation history and asks the agent to decide whether to remain silent or produce a concise group\-facing intervention\. The benchmark spans research collaboration, GitHub issue discussions, project coordination, product design, logistics, education, resource\-constrained decision making, and other collaborative settings, yielding 3,244 turn\-level examples\. We evaluate agents from four perspectives: whether the intervention addresses an evidence\-grounded breakdown, whether it avoids interrupting productive collaboration, whether it remains concise, and its overall intervention quality\. Across five LLM backbones, ProACT consistently improves collaborative appropriateness, non\-interruptiveness, conciseness, and judged intervention quality over baseline, showing that breakdown\-aware skill routing helps agents decide both when to speak and how to intervene\.
## 2Related Work
#### Multi\-user Collaboration
Multi\-user collaboration centers on joint social activity where participants maintain common ground, negotiate shared commitments, and coordinate interdependent actions toward shared or evolving goals\(Bratman,[1992](https://arxiv.org/html/2607.03730#bib.bib6); Clark and Brennan,[1991](https://arxiv.org/html/2607.03730#bib.bib2); Grosz and Kraus,[1996](https://arxiv.org/html/2607.03730#bib.bib5); Malone and Crowston,[1994](https://arxiv.org/html/2607.03730#bib.bib3)\)\. In this setting, progress depends on whether participants remain aligned about what has been understood, what has been decided, who is responsible, which constraints remain active, and whose perspective is missing\. Prior work shows that collaboration can be weakened by task and process conflict\(Jehn,[1995](https://arxiv.org/html/2607.03730#bib.bib4)\), failures of grounding and conversational repair\(Schegloffet al\.,[1977](https://arxiv.org/html/2607.03730#bib.bib14); Clark and Brennan,[1991](https://arxiv.org/html/2607.03730#bib.bib2)\), unmanaged dependencies and articulation work\(Malone and Crowston,[1994](https://arxiv.org/html/2607.03730#bib.bib3); Schmidt and Bannon,[1992](https://arxiv.org/html/2607.03730#bib.bib17)\), inefficient group decision processes\(DeSanctis and Gallupe,[1987](https://arxiv.org/html/2607.03730#bib.bib18); Briggset al\.,[2003](https://arxiv.org/html/2607.03730#bib.bib19)\), and unequal participation or missing information in collective problem solving\(Stasser and Titus,[1985](https://arxiv.org/html/2607.03730#bib.bib1); Woolleyet al\.,[2010](https://arxiv.org/html/2607.03730#bib.bib21)\)\. These studies suggest that effective collaboration requires continuous attention to shared understanding, responsibility, constraints, decision progress, and participation balance\. This motivates proactive agents that can recognize emerging coordination problems and intervene only when a concise, neutral contribution is likely to help the group move forward\.
#### LLM Agent for Coordination
Recent work on LLM agents has shown that language models can be organized as agents that reason, act, communicate, and coordinate through explicit interaction protocols\. Early agent frameworks tightly couple reasoning with tool use, alternating between planning steps and external actions \(e\.g\., calling APIs or using search\)\(Yaoet al\.,[2023](https://arxiv.org/html/2607.03730#bib.bib26)\), while multi\-agent systems such as AutoGen, CAMEL, and MetaGPT coordinate multiple LLM agents through conversation, role assignment, workflow decomposition, or dynamic agent grouping\(Wuet al\.,[2024](https://arxiv.org/html/2607.03730#bib.bib27); Liet al\.,[2023a](https://arxiv.org/html/2607.03730#bib.bib28); Honget al\.,[2024](https://arxiv.org/html/2607.03730#bib.bib30)\)\. Recent systems such as Magentic\-One adopt orchestrator\-style architectures, where a lead agent plans, tracks progress, and delegates subtasks to specialized agents\(Fourneyet al\.,[2024](https://arxiv.org/html/2607.03730#bib.bib33)\)\. Other work evaluates whether LLM agents can coordinate in practice, including coordination games, Theory\-of\-Mind belief tracking, and social\-psychology\-oriented analyses of consensus and debate\(Agasheet al\.,[2025](https://arxiv.org/html/2607.03730#bib.bib34); Liet al\.,[2023b](https://arxiv.org/html/2607.03730#bib.bib41); Zhanget al\.,[2024](https://arxiv.org/html/2607.03730#bib.bib42)\)\. Despite this progress, these settings largely focus on single\-user task assistance or coordination*among agents*, typically under a specified task or workflow, rather than an agent participating in ongoinghumanmulti\-user collaboration\. Our work studies an LLM agent embedded in a multi\-party human discussion, where the key challenge is deciding*when*to intervene proactively and*how*to help without disrupting the group\.
## 3Definitions
### 3\.1Multi\-User Collaboration Environment
A multi\-user collaboration environment is a shared communication space in which a set of participants𝒰=\{u1,…,un\}\\mathcal\{U\}=\\\{u\_\{1\},\\ldots,u\_\{n\}\\\}work toward one or more tasks under evolving goals, constraints, roles, and commitments\. At timesteptt, the observed interaction prefix is represented as a speaker\-attributed message sequence:Ht=\(xi\)i=1tH\_\{t\}=\(x\_\{i\}\)\_\{i=1\}^\{t\}\. For each turni∈\{1,…,t\}i\\in\\\{1,\\ldots,t\\\}, the message is defined asxi=\(si,ci\)x\_\{i\}=\(s\_\{i\},c\_\{i\}\), wheresi∈𝒰s\_\{i\}\\in\\mathcal\{U\}denotes the speaker andcic\_\{i\}denotes the message content\. Compared with single\-user assistance, this environment involves multiple participants who may differ in goals, preferences, knowledge, responsibilities, authority as discribed byYanget al\.\([2026](https://arxiv.org/html/2607.03730#bib.bib44)\)\. The agent observes the shared interaction history and decides whether to remain silent or contribute to the same communication channel\. This setting makes proactive collaboration a timing and coordination problem: the agent must decide when its contribution would help the group move forward and when staying silent would avoid interrupting productive discussion\.
### 3\.2Collaboration Breakdown
Following prior work on human collaboration in §[2](https://arxiv.org/html/2607.03730#S2.SS0.SSS0.Px1), we define a*collaboration breakdown*as an evidence\-grounded state where the group may lose progress, shared understanding, coordination quality, or fair participation without timely repair\. Disagreement, uncertainty, or delay becomes a breakdown when it blocks collaboration by obscuring intent, hiding trade\-offs, violating constraints, underspecifying plans, repeating unresolved issues, or excluding relevant perspectives\.
Table[1](https://arxiv.org/html/2607.03730#S3.T1)summarizes the breakdown categories used in this work\. The taxonomy operationalizes recurring failure modes studied in human collaboration: task and process conflict\(Jehn,[1995](https://arxiv.org/html/2607.03730#bib.bib4)\), grounding and conversational repair\(Schegloffet al\.,[1977](https://arxiv.org/html/2607.03730#bib.bib14); Clark and Brennan,[1991](https://arxiv.org/html/2607.03730#bib.bib2)\), coordination and articulation work\(Schmidt and Bannon,[1992](https://arxiv.org/html/2607.03730#bib.bib17); Malone and Crowston,[1994](https://arxiv.org/html/2607.03730#bib.bib3)\), group decision processes\(DeSanctis and Gallupe,[1987](https://arxiv.org/html/2607.03730#bib.bib18); Briggset al\.,[2003](https://arxiv.org/html/2607.03730#bib.bib19)\), and participation imbalance or missing perspectives in collective problem solving\(Stasser and Titus,[1985](https://arxiv.org/html/2607.03730#bib.bib1); Woolleyet al\.,[2010](https://arxiv.org/html/2607.03730#bib.bib21)\)\. We use these categories to decide whether the current turn contains an intervention\-worthy collaboration issue, while the detailed definitions are given in Table[1](https://arxiv.org/html/2607.03730#S3.T1)\.
Table 1:Collaboration breakdowns observed in multi\-user interactions\.
### 3\.3proactive Collaboration
proactive collaboration refers to an agent’s ability to decide, at each turn of a multi\-user conversation, whether to remain silent or provide a group\-facing intervention before being explicitly asked\. The goal is to support the group process when a contribution can repair an emerging collaboration breakdown, while avoiding unnecessary interruption when participants are already making progress\. Given the conversation historyHtH\_\{t\}, a proactive agent chooses an actionat∈\{∅,ut\},a\_\{t\}\\in\\\{\\varnothing,u\_\{t\}\\\},where∅\\varnothingdenotes silence andutu\_\{t\}denotes a visible group\-facing intervention\. A non\-silence action is appropriate only when the conversation contains evidence of an intervention\-worthy breakdown, the selected intervention matches the collaboration issue, and the response is likely to help the group move forward without disrupting productive discussion\.
This definition connects directly to our evaluation criteria in §[5\.2](https://arxiv.org/html/2607.03730#S5.SS2)\. An intervention should be*appropriate*, meaning that it addresses an evidence\-grounded collaboration issue;*non\-interruptive*, meaning that the agent remains silent when the group is already progressing or self\-repairing; and*concise*, meaning that the response supports the collaboration without dominating the conversation\.
## 4ProACT: Breakdown\-Aware proactive Agent in Multi\-User Collaboration
Figure 2:ProACT trajectory example\. The agent monitors a multi\-user conversation, detects emerging collaboration breakdowns such as discussion loops or conflicts, and applies targeted skills like loop\-breaking or conflict mediation\. The illustration shows how ProACT identifies a looping discussion and intervenes with a concise, evidence\-grounded suggestion to guide the group toward resolution\.ProACT is a skill\-guided framework for LLM\-based proactive participation in multi\-user collaboration\. Building on the definitions in §[3](https://arxiv.org/html/2607.03730#S3), ProACT observes the evolving speaker\-attributed conversation history and decides whether the agent should remainsilentor produce a targeted group\-facingintervention\. The framework first diagnoses whether the current turn contains an intervention\-worthy collaboration breakdown, then routes intervention cases to a relevant collaboration skill\. Each skill is specified as a lightweight operating procedure with a trigger condition, diagnostic cues, reasoning steps, social constraints, and an output contract\. This structure maps proactive participation into an evidence\-grounded decision process where interventions occur only when they are appropriate, non\-interruptive, and concise\. §[4\.1](https://arxiv.org/html/2607.03730#S4.SS1)details the decision pipeline, §[4\.2](https://arxiv.org/html/2607.03730#S4.SS2)presents the skill library, and §[4\.3](https://arxiv.org/html/2607.03730#S4.SS3)describes the structured interface for routing evidence and executing skills\.
### 4\.1ProACT Decision Pipeline
ProACT maps an evolving multi\-user conversation to eitherSilenceor a targeted collaboration intervention\. As illustrated in Figure[2](https://arxiv.org/html/2607.03730#S4.F2), the pipeline has three steps: diagnosing the current collaboration context, deciding whether the agent should speak, and routing intervention cases to an appropriate collaboration skill\.
#### Breakdown\-aware diagnosis\.
At turntt, ProACT observes the conversation historyHtH\_\{t\}, represented as speaker\-attributed messages\. The first step is to assess whether the current interaction contains an intervention\-worthy collaboration breakdown:
\(bt,zt\)=Diagnose\(Ht\),\(b\_\{t\},z\_\{t\}\)=\\mathrm\{Diagnose\}\(H\_\{t\}\),wherebtb\_\{t\}denotes the detected collaboration issue, if any, andzt∈\{0,1\}z\_\{t\}\\in\\\{0,1\\\}indicates whether an intervention is warranted\. The diagnosis covers recurring breakdowns defined in Table[1](https://arxiv.org/html/2607.03730#S3.T1)\. The decision also accounts for local context, such as who is speaking, whether a question is already directed to another participant, whether the group is actively making progress, and whether the evidence is strong enough for a concise intervention\.
#### Skill\-guided proactive action\.
If intervention is unwarranted, ProACT outputsSilence\. If intervention is warranted, ProACT routes the detected issue to a relevant collaboration skill:
kt=r\(bt\)\.k\_\{t\}=r\(b\_\{t\}\)\.The selected skill provides guidance for how the agent should help with that type of collaboration issue\. It specifies the trigger condition, diagnostic cues, social constraints, and expected response format\. The final action is:
at=πkt\(Ht\),at∈\{∅,ut\},a\_\{t\}=\\pi\_\{k\_\{t\}\}\(H\_\{t\}\),\\qquad a\_\{t\}\\in\\\{\\varnothing,u\_\{t\}\\\},where∅\\varnothingdenotes silence andutu\_\{t\}denotes a visible group\-facing intervention\. When ProACT speaks, the response should be grounded in the conversation, address the current collaboration need, preserve human agency, and remain concise\. In this way, ProACT treats proactive participation as a decision about when to help and how to help\.
### 4\.2Collaboration Skill Library
ProACT organizes proactive participation through a library of lightweight collaboration skills, summarized in Table[6](https://arxiv.org/html/2607.03730#A2.T6)\. In agent systems, a skill can be viewed as a reusable action routine: it specifies when the agent should use a capability, what evidence it should check, and how the final action should be formed\. In our setting, these skills are designed for group\-facing collaboration rather than tool execution\. Each skill guides the agent toward a specific repair move, such as asking a clarification question, mediating a conflict, reminding the group of a prior constraint, or breaking a repeated discussion loop\. Each skill is specified by four components: a trigger condition, diagnostic cues, social constraints, and an expected response format\. The trigger condition describes when the skill may be useful\. The diagnostic cues indicate what evidence should appear in the conversation history\. The social constraints guide the agent to remain neutral, avoid unnecessary interruption, and preserve human agency\. The response format keeps the final intervention concise and directly useful to the group\. This design makes ProACT’s interventions more controllable and interpretable: the model selects a skill based on the local conversation context, then produces a short group\-facing utterance grounded in the detected collaboration need\.
### 4\.3Structured Skill Invocation Interface
ProACT uses a structured interface to separate internal skill selection from the visible group\-facing response\. Given the conversation historyHtH\_\{t\}, the model first determines whether the current turn requires intervention\. If no intervention is needed, it returns silence\. If intervention is warranted, it identifies the collaboration issue, selects a relevant skill, and uses the skill instructions to produce the final action:
Ht→\(bt,kt\)→at,at∈\{∅,ut\},H\_\{t\}\\rightarrow\(b\_\{t\},k\_\{t\}\)\\rightarrow a\_\{t\},\\qquad a\_\{t\}\\in\\\{\\varnothing,u\_\{t\}\\\},wherebtb\_\{t\}is the diagnosed collaboration issue,ktk\_\{t\}is the selected skill,∅\\varnothingdenotes silence, andutu\_\{t\}is a visible group\-facing utterance\. The interface constrains the final output format\. For silence, the visible response is empty\. For intervention, the visible response contains only a concise message to the group, without diagnostic labels, skill names, rationale, or meta\-commentary\. This design keeps ProACT’s runtime behavior simple and ensures that the evaluated output is exactly the agent’s group\-facing contribution\.
## 5Evaluation for proactive Collaboration
We construct a benchmark for evaluating proactive collaboration\. Each example contains a multi\-user conversation prefixHtH\_\{t\}and a decision turntt\. The agent must choose between two actions: remain silent or produce a concise, group\-facing intervention\. This setup measures both timing and content: an effective agent should help when the conversation shows a collaboration breakdown and stay silent when human participants are already making progress\.
### 5\.1Dataset Construction
We build the evaluation set from real multi\-user conversations collected from public GitHub issue discussions and opensourced QMSum dataset\(Zhonget al\.,[2021](https://arxiv.org/html/2607.03730#bib.bib25)\)\. These sources cover software collaboration, product design, research meetings, and group decision\-making\. We convert each conversation into a sequence of decision points, where each point asks whether a proactive agent should stay silent or intervene\.
We pre\-label candidate decision points using multiple LLM annotators\. For each point, annotators decide whether a proactive agent should remain silent or intervene\. For intervention cases, they also identify the collaboration issue and provide a short rationale\. We describe the annotation protocol in Appendix[A](https://arxiv.org/html/2607.03730#A1)\. Candidate points with annotator disagreement or ambiguous evidence are sent to human review, yielding 694 human\-reviewed examples\. After annotation, we construct a balanced evaluation set with equal numbers of intervention and silence cases\. To improve coverage of rare collaboration failures, we add 800 synthetic conversations derived from BEAM\(Tavakoliet al\.,[2025](https://arxiv.org/html/2607.03730#bib.bib39)\), which have more interaction turns\. We adapt BEAM scenarios into multi\-user discussions and inject controlled breakdowns from our taxonomy, such as forgotten constraints, discussion loops, participation imbalance, and risky commitments\. These synthetic examples supplement the real conversations; details of the construction procedure are provided in Appendix[B](https://arxiv.org/html/2607.03730#A2)\. The final dataset contains 3,244 examples: 2,444 real examples and 800 synthetic examples, balanced into 1,622 intervention cases and 1,622 silence cases\. Figure[3](https://arxiv.org/html/2607.03730#S5.F3)shows the distribution by topic, collaboration label, and decision position\.
Figure 3:Distribution of the proactive collaboration evaluation dataset by topic, collaboration label, and decision position\. Percentages are computed over all 3,244 examples\.
### 5\.2Evaluation Protocol
We compare ProACT with a Direct Chat baseline under matched input conditions\. For each example, both methods receive the same conversation prefixHtH\_\{t\}\. Direct Chat generates the next assistant response directly, while ProACT uses breakdown diagnosis and skill guidance to decide whether to stay silent or produce a concise group\-facing intervention\. We evaluate outputs with an LLM\. The judge evaluates four criteria\.Appropriatemeasures whether the agent’s action fits the collaboration state: intervening when there is an evidence\-grounded breakdown and staying silent when the group is already making progress\.Non\-interruptivemeasures whether the output avoids disrupting productive human discussion, duplicating an ongoing repair, taking sides prematurely, or dominating the group\.Concisemeasures whether the visible response is short, group\-facing, and free of unnecessary explanation or meta\-commentary\.Qualityis a 1–5 score for visible interventions, reflecting whether the response is useful, grounded in the conversation, neutral, and actionable\. For each exampleii, the judge returns binary indicators of appropriateness, non\-interruption, and conciseness\. We report Appropriate, Non\-int\., and Concise as the average of these indicators over the dataset𝒟\\mathcal\{D\}, which corresponds to the fraction of examples receiving a positive judgment for each criterion\. For Quality, we report the average 1–5 judge score over visible interventions on reference\-intervention examples, so response quality is measured separately from abstention behavior\. Full judging details are provided in Appendix[C](https://arxiv.org/html/2607.03730#A3)\.
## 6Experiments
Figure 4:proactive collaboration evaluation results across five LLM backbones\. Grey markers show Direct Chat and blue markers show our ProACT\.### 6\.1Experimental Setup
We evaluate five agent backbones on the full 3,244\-example test set: GPT\-5\.4, Kimi K2\.5, Claude Sonnet 4\.6, Gemini 3\.1 Pro Preview, and GPT\-OSS\-120B\. For each model, we compare two methods under matched input conditions: Direct Chat and ProACT\. Direct Chat generates the next assistant message in a single pass\. ProACT first diagnoses whether the current turn requires intervention, then either remains silent or invokes a collaboration skill to produce a concise group\-facing intervention\. For all models, we use the same decoding configuration where supported: temperature=1\.0=1\.0and top\-p=0\.98p=0\.98\. All outputs are evaluated by GPT\-5\.4\-Nano as an external judge following the protocol in §[5\.2](https://arxiv.org/html/2607.03730#S5.SS2)\.
### 6\.2Results
Figure 5:GPT\-5\.4 participation timing by reference label\. Each stacked bar decomposes the model behavior into whether the agent responds or stays silent\. ProACT sharply reduces unnecessary responses when the reference label prefers silence\.#### ProACT consistently improves collaborative participation\.
Figure[4](https://arxiv.org/html/2607.03730#S6.F4)compares ProACT with the Direct Chat baseline across five LLM backbones\. ProACT improves every reported metric for every model, including collaborative appropriateness, non\-interruptiveness, conciseness, and intervention quality\. The gains are largest for Kimi K2\.5: appropriateness increases from 0\.222 to 0\.870, non\-interruptiveness from 0\.323 to 0\.942, conciseness from 0\.129 to 0\.964, and intervention quality from 1\.789 to 3\.450\. The same trend holds for Claude Sonnet 4\.6, Gemini 3\.1 Pro Preview, GPT\-OSS\-120B, and GPT\-5\.4\. These consistent improvements show that ProACT helps agents decide both when to participate and how to contribute, producing interventions that are more appropriate, less disruptive, more concise, and higher quality than direct chat responses\.
#### ProACT improves participation timing\.
The gains in Figure[4](https://arxiv.org/html/2607.03730#S6.F4)are partly driven by better participation timing\. Figure[5](https://arxiv.org/html/2607.03730#S6.F5)decomposes GPT\-5\.4 behavior by reference label and judges appropriateness\. On silence\-labeled examples, Direct Chat responds on every turn, with 34\.8% of responses judged inappropriate\. ProACT responds on only 39\.3% of these turns and stays silent in most remaining cases, increasing overall judged appropriateness from 65\.2% to 89\.6%\. On intervention\-labeled examples, ProACT responds less often than Direct Chat, but its visible interventions are of higher quality, improving from 3\.486 to 3\.988\. Inappropriate silence remains low at 4\.7%\. These results show that ProACT improves participation timing: it avoids unnecessary interruptions when the group is already progressing, while still producing useful interventions when collaboration needs support\.
Figure 6:Real examples of ProACT silence on intervention\-labeled turns\. Some apparent misses are acceptable abstentions because participants are already repairing the issue, while true bad misses involve clearer unresolved coordination needs\.Figure 7:GPT\-5\.4 intervention quality gains by topic and decision position\. Values show the absolute improvement of ProACT over Direct Chat\.
#### Quality gains are strongest in social planning tasks and later turns\.
Figure[7](https://arxiv.org/html/2607.03730#S6.F7)analyzes where ProACT improves GPT\-5\.4 intervention quality\. The largest gains appear in topics that require social coordination and planning\. Committee and governance discussions improve by 1\.333 quality points, academic collaboration by 1\.236, and product design by 1\.115\. Programming also improves by 0\.377, but the gain is smaller, likely because many programming examples already contain explicit technical questions, error traces, or reproduction details that make direct responses easier\. The analysis by conversation position shows a similar pattern\. ProACT improves intervention quality across the conversation, with larger gains in later stages\. The quality gain is 0\.226 points within the first 10 turns, 0\.590 points for turns 11 to 50, 1\.009 points for turns 201 to 500, and 1\.400 points after turn 500\. These results suggest that ProACT is especially useful when collaboration accumulates unresolved context, repeated decisions, or coordination failures, where skill\-guided interventions provide more value than direct chat responses\.
#### Most apparent misses reflect conservative abstention\.
A concern is that ProACT may improve non\-interruptiveness by staying silent too often\. We examine intervention\-labeled examples where GPT\-5\.4 ProACT produces no response\. Although these cases appear to be missed interventions under the reference label, most are judged acceptable: 83\.5% of ProACT’s silence decisions in this subset are accepted by the external judge, and rejected silences account for only 4\.7% of all intervention\-labeled examples\.
The qualitative pattern in Figure[6](https://arxiv.org/html/2607.03730#S6.F6)explains this result\. In accepted\-silence cases, participants are already exchanging relevant evidence, answering each other, or moving toward a human\-led resolution\. An additional agent message would likely duplicate the ongoing repair or interrupt productive discussion\. Rejected\-silence cases are different: they contain clearer unresolved coordination needs, such as an unanswered clarification question or an unresolved challenge to issue closure\. This suggests that many apparent misses reflect ambiguity in proactive timing, not simple failures to help\. ProACT tends to be conservative about speaking, which improves group\-facing behavior, while leaving room for stronger detection of cases that require clarification or common\-ground repair\.
## 7Conclusion
We introduced ProACT, a breakdown\-aware framework for proactive agents in multi\-user collaboration\. ProACT observes speaker\-attributed conversation history, detects intervention\-worthy collaboration breakdowns, and chooses between silence and a concise group\-facing intervention through a targeted skill library\. We also built a new benchmark from real multi\-user conversations and BEAM\-derived synthetic cases\. Across five LLM backbones, ProACT improves appropriateness, non\-interruptiveness, conciseness, and intervention quality over Direct Chat\. Further analysis shows that these gains come from better participation timing: ProACT reduces unnecessary interruptions while preserving useful interventions when collaboration needs support\. The strongest gains appear in social planning tasks and later\-stage conversations, where accumulated context and unresolved coordination issues make direct responses less reliable\. We hope this work serves as a starting point for building proactive agents that support human collaboration through timely participation\.
## Limitations
This work evaluates ProACT in a controlled offline, turn\-level setting\. This design enables consistent comparison because each model receives the same conversation history and must make the same silence\-or\-intervention decision, but it does not capture every dynamic of live collaboration, such as trust formation, user adaptation to repeated agent interventions, or downstream reactions after an intervention\. We also use an external LLM judge to scale evaluation of appropriateness, non\-interruption, conciseness, and quality, which supports broad comparison across models but should be complemented by future human studies\. These limitations motivate future work on live multi\-user deployments, richer cultural and organizational settings, and longer\-horizon measures such as trust, perceived fairness, decision quality, coordination cost, and participants’ ability to contest or repair agent behavior\.
Also, although ProACT is designed to support collaboration through concise and neutral interventions, proactive participation also introduces risks that are not fully captured\. In real multi\-user settings, an agent that decides when to speak may gradually change the mode of human collaboration: participants may defer to the agent, treat its interventions as authoritative, or rely on it to structure disagreement and decision making\. More importantly, the same mechanism that helps repair coordination breakdowns could be misused or misaligned to steer group behavior\. For example, a proactive agent might selectively surface evidence, frame options asymmetrically, suppress dissent, or induce participants to cooperate toward goals that they have not explicitly endorsed\. Such risks relate to model deception, persuasion, and manipulation in group decision processes\. Future work should therefore include controlled human\-subject studies and deployment audits to examine how proactive agents affect trust, agency, participation balance, dissent, decision quality, and susceptibility to manipulation\. Practical deployments should also include transparency, logging, contestability, and role or authority constraints so that users can recognize, challenge, or disable agent interventions\.
## References
- Cooperation, competition, and maliciousness: LLM\-stakeholders interactive negotiation\.InThe Thirty\-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track,Cited by:[§1](https://arxiv.org/html/2607.03730#S1.p1.1)\.
- S\. Agashe, Y\. Fan, A\. Reyna, and X\. E\. Wang \(2025\)Llm\-coordination: evaluating and analyzing multi\-agent coordination abilities in large language models\.InFindings of the Association for Computational Linguistics: NAACL 2025,pp\. 8038–8057\.Cited by:[§2](https://arxiv.org/html/2607.03730#S2.SS0.SSS0.Px2.p1.1)\.
- M\. E\. Bratman \(1992\)Shared cooperative activity\.The Philosophical Review101\(2\),pp\. 327–341\.External Links:ISSN 00318108, 15581470,[Link](http://www.jstor.org/stable/2185537)Cited by:[§2](https://arxiv.org/html/2607.03730#S2.SS0.SSS0.Px1.p1.1)\.
- R\. O\. Briggs, G\. De Vreede, and J\. F\. Nunamaker \(2003\)Collaboration engineering with thinklets to pursue sustained success with group support systems\.Journal of Management Information Systems19\(4\),pp\. 31–64\.Cited by:[§2](https://arxiv.org/html/2607.03730#S2.SS0.SSS0.Px1.p1.1),[§3\.2](https://arxiv.org/html/2607.03730#S3.SS2.p2.1)\.
- H\. H\. Clark and S\. E\. Brennan \(1991\)Grounding in communication\.InPerspectives on Socially Shared Cognition,L\. B\. Resnick, J\. M\. Levine, and S\. D\. Teasley \(Eds\.\),pp\. 127–149\.External Links:[Document](https://dx.doi.org/10.1037/10096-006),[Link](https://doi.org/10.1037/10096-006)Cited by:[§1](https://arxiv.org/html/2607.03730#S1.p2.1),[§2](https://arxiv.org/html/2607.03730#S2.SS0.SSS0.Px1.p1.1),[§3\.2](https://arxiv.org/html/2607.03730#S3.SS2.p2.1)\.
- G\. DeSanctis and R\. B\. Gallupe \(1987\)A foundation for the study of group decision support systems\.Management Science33\(5\),pp\. 589–609\.External Links:[Document](https://dx.doi.org/10.1287/mnsc.33.5.589)Cited by:[§2](https://arxiv.org/html/2607.03730#S2.SS0.SSS0.Px1.p1.1),[§3\.2](https://arxiv.org/html/2607.03730#S3.SS2.p2.1)\.
- A\. Fourney, G\. Bansal, H\. Mozannar, C\. Tan, E\. Salinas, F\. Niedtner, G\. Proebsting, G\. Bassman, J\. Gerrits, J\. Alber,et al\.\(2024\)Magentic\-one: a generalist multi\-agent system for solving complex tasks\.arXiv preprint arXiv:2411\.04468\.Cited by:[§2](https://arxiv.org/html/2607.03730#S2.SS0.SSS0.Px2.p1.1)\.
- B\. J\. Grosz and S\. Kraus \(1996\)Collaborative plans for complex group action\.Artif\. Intell\.86\(2\),pp\. 269–357\.External Links:ISSN 0004\-3702,[Link](https://doi.org/10.1016/0004-3702(95)00103-4),[Document](https://dx.doi.org/10.1016/0004-3702%2895%2900103-4)Cited by:[§1](https://arxiv.org/html/2607.03730#S1.p2.1),[§2](https://arxiv.org/html/2607.03730#S2.SS0.SSS0.Px1.p1.1)\.
- S\. Hong, M\. Zhuge, J\. Chen, X\. Zheng, Y\. Cheng, C\. Zhang, J\. Wang, Z\. Wang, S\. K\. S\. Yau, Z\. Lin, L\. Zhou, C\. Ran, L\. Xiao, C\. Wu, and J\. Schmidhuber \(2024\)MetaGPT: meta programming for a multi\-agent collaborative framework\.InInternational Conference on Learning Representations,Cited by:[§2](https://arxiv.org/html/2607.03730#S2.SS0.SSS0.Px2.p1.1)\.
- K\. A\. Jehn \(1995\)A multimethod examination of the benefits and detriments of intragroup conflict\.Administrative Science Quarterly40,pp\. 256\.External Links:[Link](https://api.semanticscholar.org/CorpusID:145774957)Cited by:[§1](https://arxiv.org/html/2607.03730#S1.p2.1),[§2](https://arxiv.org/html/2607.03730#S2.SS0.SSS0.Px1.p1.1),[§3\.2](https://arxiv.org/html/2607.03730#S3.SS2.p2.1)\.
- C\. E\. Jimenez, J\. Yang, A\. Wettig, S\. Yao, K\. Pei, O\. Press, and K\. R\. Narasimhan \(2024\)SWE\-bench: can language models resolve real\-world github issues?\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=VTF8yNQM66)Cited by:[§1](https://arxiv.org/html/2607.03730#S1.p1.1)\.
- F\. Kirstein, T\. Ruas, R\. Kratel, and B\. Gipp \(2024\)Tell me what i need to know: exploring llm\-based \(personalized\) abstractive multi\-source meeting summarization\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track,pp\. 920–939\.Cited by:[§1](https://arxiv.org/html/2607.03730#S1.p1.1)\.
- G\. Li, H\. Hammoud, H\. Itani, D\. Khizbullin, and B\. Ghanem \(2023a\)CAMEL: communicative agents for “mind” exploration of large language model society\.InAdvances in Neural Information Processing Systems,Vol\.36\.Cited by:[§2](https://arxiv.org/html/2607.03730#S2.SS0.SSS0.Px2.p1.1)\.
- H\. Li, Y\. Chong, S\. Stepputtis, J\. Campbell, D\. Hughes, M\. Lewis, and K\. Sycara \(2023b\)Theory of mind for multi\-agent collaboration via large language models\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 180–192\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.13)Cited by:[§2](https://arxiv.org/html/2607.03730#S2.SS0.SSS0.Px2.p1.1)\.
- Y\. Lu, S\. Yang, C\. Qian, G\. Chen, Q\. Luo, Y\. Wu, H\. Wang, X\. Cong, Z\. Zhang, Y\. Lin, W\. Liu, Y\. Wang, Z\. Liu, F\. Liu, and M\. Sun \(2025\)Proactive agent: shifting LLM agents from reactive responses to active assistance\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=sRIU6k2TcU)Cited by:[§1](https://arxiv.org/html/2607.03730#S1.p3.1)\.
- T\. W\. Malone and K\. Crowston \(1994\)The interdisciplinary study of coordination\.ACM Computing Surveys26\(1\),pp\. 87–119\.External Links:[Document](https://dx.doi.org/10.1145/174666.174668)Cited by:[§2](https://arxiv.org/html/2607.03730#S2.SS0.SSS0.Px1.p1.1),[§3\.2](https://arxiv.org/html/2607.03730#S3.SS2.p2.1)\.
- E\. A\. Schegloff, G\. Jefferson, and H\. Sacks \(1977\)The preference for self\-correction in the organization of repair in conversation\.Language53\(2\),pp\. 361–382\.Cited by:[§2](https://arxiv.org/html/2607.03730#S2.SS0.SSS0.Px1.p1.1),[§3\.2](https://arxiv.org/html/2607.03730#S3.SS2.p2.1)\.
- K\. Schmidt and L\. Bannon \(1992\)Taking cscw seriously: supporting articulation work\.Computer Supported Cooperative Work1\(1–2\),pp\. 7–40\.External Links:[Document](https://dx.doi.org/10.1007/BF00752449)Cited by:[§2](https://arxiv.org/html/2607.03730#S2.SS0.SSS0.Px1.p1.1),[§3\.2](https://arxiv.org/html/2607.03730#S3.SS2.p2.1)\.
- G\. Stasser and W\. Titus \(1985\)Pooling of unshared information in group decision making: biased information sampling during discussion\.Journal of Personality and Social Psychology48\(6\),pp\. 1467–1478\.External Links:[Document](https://dx.doi.org/10.1037/0022-3514.48.6.1467)Cited by:[§2](https://arxiv.org/html/2607.03730#S2.SS0.SSS0.Px1.p1.1),[§3\.2](https://arxiv.org/html/2607.03730#S3.SS2.p2.1)\.
- M\. Tavakoli, A\. Salemi, C\. Ye, M\. Abdalla, H\. Zamani, and J\. R\. Mitchell \(2025\)Beyond a million tokens: benchmarking and enhancing long\-term memory in llms\.arXiv preprint arXiv:2510\.27246\.Cited by:[Appendix B](https://arxiv.org/html/2607.03730#A2.p1.1),[§5\.1](https://arxiv.org/html/2607.03730#S5.SS1.p2.1)\.
- H\. Wei, Z\. Zhang, S\. He, T\. Xia, S\. Pan, and F\. Liu \(2025\)Plangenllms: a modern survey of llm planning capabilities\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 19497–19521\.Cited by:[§1](https://arxiv.org/html/2607.03730#S1.p1.1)\.
- A\. W\. Woolley, C\. F\. Chabris, A\. Pentland, N\. Hashmi, and T\. W\. Malone \(2010\)Evidence for a collective intelligence factor in the performance of human groups\.Science330\(6004\),pp\. 686–688\.Cited by:[§2](https://arxiv.org/html/2607.03730#S2.SS0.SSS0.Px1.p1.1),[§3\.2](https://arxiv.org/html/2607.03730#S3.SS2.p2.1)\.
- Q\. Wu, G\. Bansal, J\. Zhang, Y\. Wu, B\. Li, E\. Zhu, L\. Jiang, X\. Zhang, S\. Zhang, J\. Liu, A\. H\. Awadallah, R\. W\. White, D\. Burger, and C\. Wang \(2024\)AutoGen: enabling next\-gen LLM applications via multi\-agent conversations\.InFirst Conference on Language Modeling,Cited by:[§2](https://arxiv.org/html/2607.03730#S2.SS0.SSS0.Px2.p1.1)\.
- S\. Yang, S\. Zhu, H\. Zhu, J\. R\. Enríquez, D\. Wang, A\. Pentland, M\. A\. Bakker, and J\. Pei \(2026\)Multi\-user large language model agents\.arXiv preprint arXiv:2604\.08567\.Cited by:[§3\.1](https://arxiv.org/html/2607.03730#S3.SS1.p1.7)\.
- S\. Yao, J\. Zhao, D\. Yu, N\. Du, I\. Shafran, K\. R\. Narasimhan, and Y\. Cao \(2023\)ReAct: synergizing reasoning and acting in language models\.InInternational Conference on Learning Representations,Cited by:[§2](https://arxiv.org/html/2607.03730#S2.SS0.SSS0.Px2.p1.1)\.
- J\. Zhang, X\. Xu, N\. Zhang, R\. Liu, B\. Hooi, and S\. Deng \(2024\)Exploring collaboration mechanisms for LLM agents: a social psychology view\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics,pp\. 14544–14607\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.782)Cited by:[§2](https://arxiv.org/html/2607.03730#S2.SS0.SSS0.Px2.p1.1)\.
- M\. Zhong, D\. Yin, T\. Yu, A\. Zaidi, M\. Mutuma, R\. Jha, A\. H\. Awadallah, A\. Celikyilmaz, Y\. Liu, X\. Qiu, and D\. Radev \(2021\)QMSum: a new benchmark for query\-based multi\-domain meeting summarization\.InProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,pp\. 5905–5921\.External Links:[Document](https://dx.doi.org/10.18653/v1/2021.naacl-main.472)Cited by:[§5\.1](https://arxiv.org/html/2607.03730#S5.SS1.p1.1)\.
- S\. Zhu, S\. Yang, M\. A\. Bakker, A\. Pentland, and J\. Pei \(2025\)Can ai truly represent your voice in deliberations? a comprehensive study of large\-scale opinion aggregation with llms\.arXiv preprint arXiv:2510\.05154\.Cited by:[§1](https://arxiv.org/html/2607.03730#S1.p1.1)\.
## Appendix ADataset Construction and Annotation Details
We construct the benchmark by converting multi\-user conversations into decision\-point examples for proactive collaboration\. Given a source conversation\(x1,…,xT\)\(x\_\{1\},\\ldots,x\_\{T\}\), a candidate example at turnttcontains the conversation prefixHt=\(xi\)i=1tH\_\{t\}=\(x\_\{i\}\)\_\{i=1\}^\{t\}\. The annotation task asks whether a proactive agent should remain silent or produce a short group\-facing intervention at that point\. Because proactive timing can be context\-dependent, we treat the labels as reference annotations rather than absolute ground truth\.
#### Source artifact provenance and redistribution\.
For each source used in benchmark construction, we retain provenance metadata, including the source type and source identifier when available\. Public collaborative discussions are used as examples of naturally occurring multi\-user coordination, while existing research datasets and scenarios are used only for research evaluation and controlled scenario construction\. We do not treat access to an artifact as permission for unrestricted redistribution\.
LLM annotation prompt template\.System instruction\.You are annotating a multi\-user collaboration conversation for a proactive assistant\. Given the conversation prefix up to the current decision point, decide whether the assistant should stay silent or intervene now\.Task\.Return a structured annotation with three fields\.agent\_actionshould be eithersilenceorintervene\.breakdown\_typeshould be one ofnone,uncertainty,underspecified\_plan,conflict,forgotten\_constraint,drift,looping,imbalanced\_participation, orrisky\_commitment\.rationaleshould be a brief explanation grounded in specific evidence from the visible conversation\.Decision criteria\.Choosesilencewhen participants are already making progress, when a participant has already asked the needed clarification, when the issue is being self\-repaired, when the evidence is weak, or when an assistant message would be duplicative or interruptive\. Chooseinterveneonly when the conversation shows an unresolved collaboration breakdown and a short group\-facing message would likely improve coordination\.Output format\.Return JSON only:\{ "agent\_action": "silence" or "intervene", "breakdown\_type": "\.\.\.", "rationale": "\.\.\." \}
Table 2:LLM annotation prompt used for pre\-labeling proactive collaboration decision\-point examples\.
#### Personally identifying information and offensive content\.
Because our real examples are derived from public collaborative discussions and existing benchmark conversations, they may contain public user handles, names, project\-specific references, URLs, email addresses, or other strings that could identify individuals or organizations\. We therefore screen collected examples before inclusion\. We use automatic pattern checks for common identifiers, including email addresses, phone numbers, URLs, social\-media or GitHub handles, and speaker\-name fields, followed by manual inspection during candidate filtering and human adjudication\. In the benchmark release and in paper examples, speaker names and user handles are replaced with role\-neutral identifiers such asParticipant A,Participant B, orMaintainer\. Direct contact information, private links, and unnecessary project\-specific identifiers are removed or masked\. Reviewer identifiers used during annotation are used only for assignment and adjudication and are not included in the released data\.
#### Candidate extraction\.
We begin with real multi\-user conversations from public collaborative settings, including GitHub issue discussions and QMSum meetings\. These sources cover software collaboration, product design, research meetings, and group decision\-making\. We segment each conversation into candidate decision points and keep cases where the prefix contains enough context for an agent to judge the collaboration state\. We remove examples that are too short, lack multiple participants, contain insufficient context, or do not involve a collaborative process\.
#### LLM pre\-labeling\.
Each candidate decision point is first pre\-labeled by multiple LLM annotators\. The annotators receive the conversation prefix and produce three fields: \(1\) a decision labelyt∈\{Silence,Intervention\}y\_\{t\}\\in\\\{\\textsc\{Silence\},\\textsc\{Intervention\}\\\}; \(2\) a collaboration issue label, selected from the breakdown categories in Table[1](https://arxiv.org/html/2607.03730#S3.T1)orNone; and \(3\) a short rationale grounded in the visible conversation\. Annotators are instructed to chooseSilencewhen participants are already making progress, when the issue is being self\-repaired, when evidence for a breakdown is weak, or when an agent response would mainly duplicate the ongoing discussion\.
#### Human review and adjudication\.
LLM pre\-labeling is used as a screening step, not as the final annotation\. We identify 740 candidate examples with LLM annotator disagreement, unclear evidence, malformed rationales, or ambiguous intervention timing\. Two CS Ph\.D\. students review this subset by inspecting the conversation prefix, the proposed decision label, the breakdown type, and the rationale\. They verify whether intervention cases contain concrete evidence of a collaboration breakdown and whether silence cases reflect productive ongoing human collaboration\. After adjudication, we retain 694 examples with aligned final labels and remove cases that remain underspecified, out of scope, or too ambiguous for reliable evaluation\. Reviewers were instructed to answer the following question: “Given the collaboration history up to the current turn, should a proactive agent intervene now? If yes, what collaboration breakdown is happening, which skill should be used, and what should the agent say?” The core rule was: setshould\_intervene=trueonly when speaking now would likely improve the group process at the current turn\. A collaboration breakdown alone was not sufficient; if the group was already resolving the issue, the agent should remain silent\. For each candidate prefix, reviewers were asked to: \(1\) read the full prefix; \(2\) identify the current collaboration state; \(3\) ask whether a helpful facilitator would speak now; \(4\) if yes, choose the breakdown type and preferred skill; \(5\) mark evidence turns; \(6\) write a brief reference intervention; \(7\) optionally write a bad intervention; and \(8\) write an annotation rationale\. Reviewers usedsilencewhen there was no breakdown, the issue was low\-stakes or already resolved, the evidence was insufficient, or speaking would duplicate or interrupt human facilitation\.
Table 3:Textual reconstruction of the human\-review interface used for validating proactive\-collaboration reference labels\. No separate risk or consent disclaimer was displayed in the review UI; reviewers only labeled existing public/benchmark or synthetic text excerpts and could mark examples asskiporneeds\_discussion\.
## Appendix BSynthetic Data Construction
We use synthetic examples to improve coverage of collaboration failures that are important for proactive agents but appear less frequently in public conversations\. The synthetic data are derived from BEAM\(Tavakoliet al\.,[2025](https://arxiv.org/html/2607.03730#bib.bib39)\), which provides long\-context task scenarios that can be adapted into multi\-user collaborative settings\. We use BEAM as a source of task contexts, constraints, and decision situations, then inject short multi\-user episodes that create controlled collaboration breakdowns\.
#### Scenario adaptation\.
For each selected BEAM context, we identify the task goal, relevant participants, available information, constraints, and possible decision points\. We then adapt the context into a short group\-chat episode in which multiple participants exchange information, propose options, and move toward a decision\. The adapted scenarios cover project planning, product design, logistics, education, resource allocation, research collaboration, and other collaborative tasks\.
#### Synthetic breakdown injection\.
Each synthetic example targets one breakdown category from Table[1](https://arxiv.org/html/2607.03730#S3.T1)\. The injected episode contains setup or evidence turns followed by a final trigger turn\. The setup establishes the relevant context, such as an earlier constraint, an unresolved trade\-off, a missing stakeholder, or an incomplete plan\. The final trigger turn makes the collaboration issue visible at the current decision point\. This design ensures that the proactive intervention is useful at that moment, rather than too early or only later\.
#### Generation prompt\.
Table[4](https://arxiv.org/html/2607.03730#A2.T4)shows the prompt template used to generate BEAM\-derived injected breakdown episodes\. The generator is instructed to produce only natural user/team messages\. It must not reveal annotation labels, breakdown names, preferred skills, ground truth, or any reference to the proactive assistant\. This prevents the synthetic conversation from containing artificial cues that would make the task easier than real collaboration\.
Synthetic breakdown injection prompt template\.System instruction\.You generate realistic multi\-user collaboration chat snippets for synthetic data\. Return JSON only\.User instruction\.Generate only the injected user chat turns for a synthetic multi\-user collaboration dataset\. A long conversation is used as background\. Create a two\-stage group\-chat episode to insert into that conversation\. The episode must contain natural user or team messages, not labels or analysis\. Do not mention annotation, dataset construction, breakdown type, preferred skill, ground truth, labels, or the AI assistant in the injected turns\.Return the setup or evidence messages first, then the final trigger message last\. The setup messages will be inserted earlier in the conversation, and the final trigger message will become the latest visible message after which a proactive collaboration assistant should decide whether to intervene\. Use 3–6 injected turns from at least three human collaborators\. Keep messages concise and realistic\.Target breakdown:\{breakdown\_type\}Definition:\{breakdown\_definition\}Ground\-truth requirements\.The group must not resolve the issue before the final injected turn\. The final injected turn must connect back to the earlier setup or evidence\. The final injected turn must make the intervention useful at the current decision point\. Avoid cues for other labels and instantiate the specified breakdown rather than generic confusion\.Positive pattern:\{positive\_pattern\}Negative pattern to avoid:\{negative\_pattern\}Recent BEAM context before insertion:\{recent\_beam\_context\}Output format\.Return only a JSON array:\[ \{ "speaker": "participant name", "text": "message text" \}, \.\.\. \]
Table 4:Prompt template used to generate synthetic BEAM\-derived collaboration breakdown episodes\.
#### Breakdown\-specific constraints\.
Table[5](https://arxiv.org/html/2607.03730#A2.T5)summarizes the generation constraints used for each breakdown type\. These constraints are designed to create examples where the target breakdown is visible from the conversation prefix while avoiding confounds with other labels\.
Table 5:Breakdown\-specific generation constraints used for synthetic BEAM\-derived examples\.
#### Filtering and labeling\.
Each generated example is checked for three conditions\. First, the target breakdown must be visible from the conversation history without relying on hidden labels\. Second, the final trigger turn must create a clear decision point where a concise group\-facing intervention would be useful\. Third, the participants should not already be resolving the issue by themselves\. We remove examples where the issue is too implicit, the target label is confounded with another breakdown, or the intervention point is unclear\.
Table 6:ProACT skill library and action space\. The agent either remains silent or selects a targeted collaboration skill based on the detected collaboration issue\.
## Appendix CLLM Judge Prompt
Table[7](https://arxiv.org/html/2607.03730#A3.T7)presents the full LLM\-as\-judge prompt used to evaluate candidate outputs in our benchmark\. The prompt asks the judge to assess each action and visible response according to collaborative participation criteria, including appropriateness, non\-interruption, conciseness, grounding, neutrality, and intervention quality\. We format the prompt as a two\-column appendix table to keep the instructions readable\.
Judge prompt used in our benchmark\.System instruction\.You are an expert judge for proactive multi\-user collaboration agents\. Your job is to evaluate a candidate assistant’s response at the current moment in a group conversation\. Judge the candidate by human collaborative participation standards, not by whether it is technically smart in isolation\. Use criteria from human collaboration and facilitation research: maintaining common ground, supporting shared plans and commitments, managing coordination/articulation work, respecting turn\-taking and timing, preserving neutrality and fairness, balancing participation, grounding claims in the visible conversation, being brief, and preserving human agency\.A good proactive response should help move the group conversation forward now with a concrete collaborative next step; improve common ground, coordination, plan completeness, risk awareness, or participation balance; avoid interrupting productive human flow or duplicating what humans are already doing; avoid taking sides, taking over the work, or giving an overlong answer; be grounded in the visible conversation; and be concise enough to fit naturally as one group\-chat contribution\. Silence can be good when speaking would be unnecessary, duplicative, speculative, or interruptive\. Judge whether the candidate’s action itself is collaboratively appropriate: either a useful, well\-timed contribution or appropriate silence\. Return JSON only\.User payload template\.For each evaluated model output, the user message provides the conversation, the candidate response, and the inferred candidate action \(interveneif the response is non\-empty, otherwisesilence\)\. The judge is asked to return:candidate\_action,should\_have\_spoken,collaboratively\_appropriate,helps\_move\_conversation\_forward,non\_interruptive,grounded\_in\_context,concise,not\_overbearing,good\_proactive\_response,quality\_score\_1\_to\_5, and a briefrationale\.Key criteria\.collaboratively\_appropriateis true when the candidate action is the right collaboration move at that moment: useful well\-timed speech, or appropriate silence when speaking would be unnecessary, duplicative, speculative, or interruptive\.non\_interruptiveis true when the action fits naturally at the current turn without interrupting productive human flow, duplicating a human contribution, or prematurely answering a question addressed to someone else\. The 1–5 quality score follows the same rubric: 1 is harmful or disruptive; 3 is mixed but partially useful; 5 is excellent human\-like collaborative participation that is clearly needed, minimally intrusive, grounded, concise, and agency\-preserving\.
Table 7:LLM\-as\-judge prompt used to evaluate candidate actions and responses in proactive collaboration\.
## Appendix DProACT Prompt and Components
The ProACT agent is implemented with a lightweight skill harness\. At inference time, the model receives a collaboration prompt and an explicit list of available ProACT skills, including an applicability check, a breakdown\-diagnosis router, targeted collaboration repair skills, an intervention filter, and a final output contract\. Table[8](https://arxiv.org/html/2607.03730#A5.T8)shows the agent\-facing prompt template\.
## Appendix EThe Use of Large Language Models
For this paper, we leveraged GPT\-5\.4†††[https://openai\.com/](https://openai.com/)and Codex†††[https://openai\.com/codex/](https://openai.com/codex/)to support grammar refinement, LaTeX formatting, and the preparation of figure generation code\. All technical ideas, experimental designs, analyses, conclusions, and writing were developed and carried out entirely by the authors\. The authors have full responsibility for the final text\.
ProACT agent prompt template\.System instruction\.You are a proactive collaboration assistant for multi\-user group conversations\. Your role is to help the group only when a short intervention can improve collaboration\. You should observe the visible conversation, identify whether the current point contains a collaboration breakdown, and decide whether to stay silent or produce a concise group\-facing message\.Input\.The input is a speaker\-attributed conversation history\. Each speaker tag denotes a different participant in the collaboration\.Available ProACT components\.You may use the ProACT components for applicability checking, breakdown diagnosis, skill routing, collaboration repair, intervention filtering, and final response formatting\.Decision rule\.Stay silent when participants are already making progress, self\-repairing the issue, or when an assistant response would be duplicative or interruptive\. Intervene only when the conversation shows an unresolved collaboration issue and a short group\-facing message would likely improve coordination\.Visible output contract\.If the decision is silence, produce no visible group\-chat text\. If the decision is intervention, output only the concise message to the group\. Do not include diagnostic labels, skill names, rationales, or meta\-commentary in the visible response\.
Table 8:Prompt Template used by ProACT\.Similar Articles
Anticipate and Learn: Unleashing Idle-Time Compute in Proactive Agents
ProAct is a proactive agent architecture that leverages idle-time computation to anticipate user needs, improving task completion efficiency and accuracy. It introduces ProActEval, a benchmark spanning 200 scenarios across 40 domains, and achieves significant gains over reactive baselines: 14.8% reduction in required turns, 11.7% decrease in user effort, and 28.1% cut in hallucination rates.
AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators
This paper introduces AgentCollabBench, a diagnostic benchmark for multi-agent systems that evaluates behavioral risks like instruction decay and context leakage across four major LLMs. It argues that communication topology is a critical factor in multi-agent reliability, often overshadowing raw model capability.
What Should Agents Say? Action-state Communication for Efficient Multi-Agent Systems
This paper introduces PACT, a method for structuring agent-to-agent communication in multi-agent LLM systems that uses compact action-state records to reduce token consumption while maintaining or improving task performance, with demonstrated gains on SWE-agent and OpenHands.
AgentCo-op: Retrieval-Based Synthesis of Interoperable Multi-Agent Workflows
AgentCo-op is a retrieval-based synthesis framework for composing interoperable multi-agent workflows from reusable skills, tools, and external agents. It uses typed artifact handoffs and bounded self-guided local repair, achieving strong results on benchmarks and enabling collaborative discovery in open-world genomics tasks.
@adxtyahq: For the past few weeks, me and @iamadityaanjana have been working on a Collaborative Multi-Agent Memory System, and we'…
The authors developed a collaborative multi-agent memory system with shared/private memory scopes, trust-aware retrieval, lineage tracking, and contradiction resolution, and submitted a paper to a conference.