From Task Success to Productive Success: Evaluating Human-AI Collaboration by Quality and Cost

arXiv cs.CL Papers

Summary

This paper introduces a productivity-oriented framework for evaluating human-AI collaboration based on outcome quality relative to interaction cost, showing that identical quality ratings can differ significantly in interaction costs and that subjective user ratings are not reliable for measuring productivity.

arXiv:2609.21117v1 Announce Type: new Abstract: AI productivity is often measured by task completion time, economic value, or improvements in outcome quality. However, these measures usually treat collaboration as a black box where they capture what output was produced, but not the interaction cost required to produce it. Motivated by economics literature, we introduce a productivity-oriented framework for evaluating human-AI collaboration as outcome quality relative to interaction cost. Across two datasets spanning four tasks, we show that: (1) sessions with identical quality ratings can differ by up to 70 times in interaction cost; (2) quality-cost relationships vary by task, with some tasks rewarding extended interaction and others favoring fast convergence; (3) subjective user ratings are not reliable substitutes for productivity; and (4) productive sessions are characterized by agents probing earlier and users spending less effort repairing the interaction. By distinguishing productive success from costly success, our framework makes interactional cost visible and shows how dialogue analysis can inform the evaluation and design of AI systems.
Original Article
View Cached Full Text

Cached at: 09/21/26, 09:03 AM

# From Task Success to Productive Success: Evaluating Human-AI Collaboration by Quality and Cost
Source: [https://arxiv.org/html/2609.21117](https://arxiv.org/html/2609.21117)
Mert İnanMalihe AlikhaniAffiliation:Northeastern University, Boston MAAffiliation:\{imai\.s, inan\.m, m\.alikhani\}@northeastern\.edu

###### Abstract

AI productivity is often measured by task completion time, economic value, or improvements in outcome quality\. However, these measures usually treat collaboration as a black box where they capture what output was produced, but not the interaction cost required to produce it\. Motivated by economics literature, we introduce a productivity\-oriented framework for evaluating human\-AI collaboration as outcome quality relative to interaction cost\. Across two datasets spanning four tasks, we show that: \(1\) sessions with identical quality ratings can differ by up to 70×\\timesin interaction cost; \(2\) quality\-cost relationships vary by task, with some tasks rewarding extended interaction and others favoring fast convergence; \(3\) subjective user ratings are not reliable substitutes for productivity; and \(4\) productive sessions are characterized by agents probing earlier and users spending less effort repairing the interaction\. By distinguishing productive success from costly success, our framework makes interactional cost visible and shows how dialogue analysis can inform the evaluation and design of AI systems\.

## 1Introduction

Two human\-AI collaborations can achieve equally successful outcomes while requiring vastly different amounts of interaction from the user\. In our data, for example, two travel planning sessions with the same quality rating differed by approximately 70 times in interaction cost\. Task success treats these sessions as equivalent, but from the user’s perspective they are not\. Figure[1](https://arxiv.org/html/2609.21117#S1.F1)illustrates this distinction: one is a*productive success*, while the other is a*costly success*\.

As large language models become embedded in work practices, researchers have renewed efforts to measure AI productivity: what work AI systems can perform, how much economic value they may create, and whether they allow people to complete tasks faster or at higher quality\. Recent work estimates the economic exposure of occupations to LLMs[Handa et al\. \(2025\)](https://arxiv.org/html/2609.21117#bib.bib6);[Tamkin and McCrory \(2025\)](https://arxiv.org/html/2609.21117#bib.bib7), measures the time horizon over which agents can complete tasks autonomously[Kwa et al\. \(2026\)](https://arxiv.org/html/2609.21117#bib.bib4);[METR \(2026\)](https://arxiv.org/html/2609.21117#bib.bib5);[Patwardhan et al\. \(2025\)](https://arxiv.org/html/2609.21117#bib.bib15), and evaluates productivity gains in domains such as writing, coding, consulting, and customer support\. These studies show that AI systems can affect productivity, but they often treat the interaction itself as a black box\. They measure the final output, elapsed time, or economic value, without explaininghowthe human and AI system reached that outcome\.

Interaction CostOutput QualityLowHighLowHighProductive SuccessHigh quality, low costCostly SuccessHigh quality, high costCheap FailureLow quality, low costUnproductive FailureLow quality, high cost

Figure 1:Productivity distinguishes*productive success*from*costly success*\. Task success captures output quality, but not the interaction cost required to achieve it\. We define productive collaboration as high quality output achieved with relatively low interaction cost\.We argue that productive human\-AI collaboration should be evaluated as a relationship between outcome quality and interaction cost\. This view builds on the principle of least collaborative effort[Clark and Wilkes\-Gibbs \(1986\)](https://arxiv.org/html/2609.21117#bib.bib18), which holds that collaborators coordinate not to minimize effort for one party alone, but to minimize the joint effort required to establish sufficient common ground[Clark and Brennan \(1991\)](https://arxiv.org/html/2609.21117#bib.bib16)\. In human\-AI settings, this distinction is especially important, where an agent can appear helpful while shifting substantial grounding work onto the user\.

To make this distinction measurable, we introduce a productivity oriented evaluation framework for human\-AI collaboration that defines collaborative productivity as outcome quality relative to interaction cost\. We validate the framework across two datasets, including a newly collected human\-AI collaboration dataset with full interaction trajectories, submitted artifacts, and post task perception measures, for a visualization task\.111Code and data to reproduce our work can be found at[https://github\.com/sakimai/human\-AI\-productivity](https://github.com/sakimai/human-AI-productivity)

Our analysis yields three main findings\. First, sessions with identical quality ratings can differ by one to two orders of magnitude in interaction cost, variation that subjective user ratings do not consistently capture\. Second, the relationship between cost and quality is task dependent: in related work writing, additional interaction tends to accompany better outcomes, while in visualization, higher quality sessions tend to converge with less interaction\. These patterns suggest that productivity must be interpreted within task rather than treated as a universal preference for shorter conversations\. Third, productive collaboration is not simply shorter interaction\. Productive sessions are characterized by agents probing and clarifying earlier, while less productive sessions require users to spend more effort clarifying, redirecting, and repairing\. This suggests that interaction can be productive when agent initiated grounding reduces later clarification, and correction, but costly when the burden of resolving ambiguity falls primarily to the user\.

Our contributions are as follows:

1. 1\.We introduce a productivity oriented evaluation framework for human\-AI collaboration as outcome quality relative to interaction cost \(§[3\.1](https://arxiv.org/html/2609.21117#S3.SS1)\), and show that it captures variation not reflected in user subjective ratings \(§[4\.3](https://arxiv.org/html/2609.21117#S4.SS3)\)\.
2. 2\.We show that collaborations with similar outcome can require dramatically different amounts of user effort \(§[4\.1](https://arxiv.org/html/2609.21117#S4.SS1)\) and that tasks differ in whether additional interaction helps or hurts quality \(§[4\.2](https://arxiv.org/html/2609.21117#S4.SS2)\)\.
3. 3\.We identify dialogue patterns associated with productive collaboration, showing that productive agents reduce user side grounding burden by front loading clarification \(§[4\.4](https://arxiv.org/html/2609.21117#S4.SS4)\)\.

## 2Related Work

#### Evaluating productivity of human\-AI collaboration\.

A growing literature measures whether LLMs make people more productive, by measuring outcomes rather than interactions\. Economic analyses estimate aggregate value and labor market effects\([Handa et al\., 2025](https://arxiv.org/html/2609.21117#bib.bib6);[Tamkin and McCrory, 2025](https://arxiv.org/html/2609.21117#bib.bib7);[Eloundou et al\., 2024](https://arxiv.org/html/2609.21117#bib.bib8);[Webb, 2020](https://arxiv.org/html/2609.21117#bib.bib9);[Gans and Goldfarb, 2026](https://arxiv.org/html/2609.21117#bib.bib10)\), and benchmarks measure autonomous task completion\([Kwa et al\., 2026](https://arxiv.org/html/2609.21117#bib.bib4);[METR, 2026](https://arxiv.org/html/2609.21117#bib.bib5);[Patwardhan et al\., 2025](https://arxiv.org/html/2609.21117#bib.bib15);[Vidgen et al\., 2025](https://arxiv.org/html/2609.21117#bib.bib1)\)\. Empirical experiments with users report substantial speedups for professional writing\([Noy and Zhang, 2023](https://arxiv.org/html/2609.21117#bib.bib27)\), customer support\([Brynjolfsson et al\., 2025](https://arxiv.org/html/2609.21117#bib.bib28)\), code generation\([Peng et al\., 2023](https://arxiv.org/html/2609.21117#bib.bib29);[Imai, 2022](https://arxiv.org/html/2609.21117#bib.bib30);[Qian and Wexler, 2024](https://arxiv.org/html/2609.21117#bib.bib33)\), ad creation\([Ju and Aral, 2026](https://arxiv.org/html/2609.21117#bib.bib34)\), and consulting[Dell’Acqua et al\. \(2026\)](https://arxiv.org/html/2609.21117#bib.bib31)with the largest gains among less experienced workers\. Other studies report null or negative effects, where experienced developers were 19% slower with AI while believing they were 20% faster\([Becker et al\., 2025](https://arxiv.org/html/2609.21117#bib.bib32)\), and consultants who used AI on tasks outside its capabilities performed worse than those without it\([Dell’Acqua et al\., 2026](https://arxiv.org/html/2609.21117#bib.bib31)\)\. These studies measurewhetherAI helps on average, nothowlanguage use during the interaction shapes the outcome\. We close this gap by measuring collaborative productivity at the session level and analyzing the dialogue that produced each outcome\. Earlier work combined task success and dialogue cost into a single score for spoken dialogue systems\([Walker et al\., 1997](https://arxiv.org/html/2609.21117#bib.bib14)\)\. We adapt this framing to human\-AI collaboration, but shift the question from comparing systems to explaining variation between sessions on the same task\.

#### Evaluating human\-AI collaboration\.

Recent work has argued that evaluation should move beyond static, AI alone benchmarks[Lee et al\. \(2023\)](https://arxiv.org/html/2609.21117#bib.bib35);[Fragiadakis et al\. \(2025\)](https://arxiv.org/html/2609.21117#bib.bib13);[Xuan et al\. \(2026\)](https://arxiv.org/html/2609.21117#bib.bib45)\. Interaction centered datasets[Lee et al\. \(2022\)](https://arxiv.org/html/2609.21117#bib.bib36);[Chiang et al\. \(2024\)](https://arxiv.org/html/2609.21117#bib.bib37);[Zheng et al\. \(2024\)](https://arxiv.org/html/2609.21117#bib.bib38);[Zhao et al\. \(2024\)](https://arxiv.org/html/2609.21117#bib.bib17)and benchmarks[Lin et al\. \(2025\)](https://arxiv.org/html/2609.21117#bib.bib39);[Chang et al\. \(2025\)](https://arxiv.org/html/2609.21117#bib.bib40);[Li et al\. \(2025\)](https://arxiv.org/html/2609.21117#bib.bib41);[Zheng et al\. \(2023\)](https://arxiv.org/html/2609.21117#bib.bib42)are aimed to evaluate model behavior in use\. A related line of work evaluates*human\-agent collaboration*, where the AI system can communicate with the user while also taking actions in a shared task environment\([Yao et al\., 2025](https://arxiv.org/html/2609.21117#bib.bib46);[Barres et al\., 2025](https://arxiv.org/html/2609.21117#bib.bib47);[Shao et al\., 2026](https://arxiv.org/html/2609.21117#bib.bib2);[Shen et al\., 2025](https://arxiv.org/html/2609.21117#bib.bib3)\)\. Recent work on LLM training further incorporates efficiency into multiturn objectives, penalizing excessive user reading and writing costs alongside task success and interaction quality\([Wu et al\., 2025](https://arxiv.org/html/2609.21117#bib.bib12)\)\. Most existing evaluations still center on end\-to\-end task success rates[Shridhar et al\. \(2021\)](https://arxiv.org/html/2609.21117#bib.bib48);[Zhou et al\. \(2024\)](https://arxiv.org/html/2609.21117#bib.bib49);[Xie et al\. \(2024\)](https://arxiv.org/html/2609.21117#bib.bib50);[Jimenez et al\. \(2024\)](https://arxiv.org/html/2609.21117#bib.bib51), and the cost of reaching that outcome is often ignored\. By introducing a framework that jointly measures outcome quality and interaction cost, it allows us to distinguish productive and unproductive collaboration\.

#### Grounding acts and positive friction\.

Collaboration requires participants to build a shared conceptual workspace[Roschelle and Teasley \(1995\)](https://arxiv.org/html/2609.21117#bib.bib19), and both speakers share the effort of building it\. Theprinciple of least collaborative effort\([Clark and Wilkes\-Gibbs, 1986](https://arxiv.org/html/2609.21117#bib.bib18)\)holds that participants minimizejointrather than individual effort, and[Clark and Brennan \(1991\)](https://arxiv.org/html/2609.21117#bib.bib16)decompose this into formulation, reception, and repair costs distributed across both speakers\. This framing motivates our framework: a productivity measure for human\-AI collaboration must account for the burden agent output places on users\. Recent work shows that LLMs do not naturally reproduce human grounding behavior\.[Shaikh et al\. \(2024\)](https://arxiv.org/html/2609.21117#bib.bib44)find that LLM generations contain fewer grounding acts than human responses, often presuming common ground rather than actively constructing it\.[Shaikh et al\. \(2025\)](https://arxiv.org/html/2609.21117#bib.bib25)further show that grounding failures in human\-LLM interaction can produce downstream rifts\. In contrast, positive friction argues that some slowdowns improve reliability[İnan et al\. \(2025\)](https://arxiv.org/html/2609.21117#bib.bib26)\. Our framework reconciles these views, that positive friction is productive when it prevents larger downstream repair costs, and unproductive when it shifts repair burden to the user\.

## 3Methods

We evaluate human\-AI collaboration at the session level\. Each session produces a final output and an interaction trajectory\. Our framework asks whether the session achieved high outcome quality relative to the interaction cost required to produce it\. We first define the productivity framework, then describe how we instantiate quality, cost, and dialogue features in our two datasets\.

### 3\.1Collaborative Productivity Framework

We define collaborative productivity as outcome quality relative to interaction cost\. This follows the productivity intuition of output relative to input\([Solow, 1957](https://arxiv.org/html/2609.21117#bib.bib23);[Schreyer and Pilat, 2001](https://arxiv.org/html/2609.21117#bib.bib24)\)in economics, but adapts it to human\-AI collaboration,222We usehuman\-AI collaborationas an umbrella term covering both chat\-only and tool\-using LLM agents\.where the relevant input is the interactional effort required to reach an outcome\.

This formulation is motivated by theories of grounding in communication\. Collaborative work requires participants to establish sufficient common ground for the task at hand\.[Clark and Brennan \(1991\)](https://arxiv.org/html/2609.21117#bib.bib16)distinguish several costs involved in this process, including formulation costs for producing utterances, reception costs for interpreting them, and repair costs for resolving misunderstandings\. In human\-AI collaboration, these costs are distributed across the user and the agent, where users write prompts, read and interpret model outputs, and repair when the interaction goes off track\. We therefore operationalize interaction cost as a property of the full trajectory rather than only the user’s messages\. Our framework also builds on PARADISE, which evaluates spoken dialogue agents by combining task success with dialogue costs\([Walker et al\., 1997](https://arxiv.org/html/2609.21117#bib.bib14)\)\. PARADISE showed that dialogue evaluation should separate what was accomplished from the costs incurred in accomplishing it, and that these quantities can be combined into a performance function\. We adapt this idea from spoken dialogue systems to human\-AI collaboration\.

Let𝒴\\mathcal\{Y\}denote the space of artifacts a collaborative session can produce and𝒯\\mathcal\{T\}the space of interaction trajectories\. We define aquality functionQ:𝒴→ℝQ:\\mathcal\{Y\}\\to\\mathbb\{R\}scoring the artifact and acost functionC:𝒯→ℝ≥0C:\\mathcal\{T\}\\to\\mathbb\{R\}\_\{\\geq 0\}scoring the interaction\. The collaborative productivity of sessioniion a task is

Pz​\(i\)=z⁡\(Qi\)−z⁡\(Ci\)P\_\{z\}\(i\)=z\(Q\_\{i\}\)\-z\(C\_\{i\}\)wherez⁡\(⋅\)z\(\\cdot\)denotes within task z\-score normalization andQi,CiQ\_\{i\},C\_\{i\}are the quality and cost of sessionii\. We standardizeQQandCCbecause tasks differ in both quality scales and expected interaction length, so raw scores are not directly comparable across tasks\. Following multi\-attribute utility approaches\([Keeney and Raiffa, 1993](https://arxiv.org/html/2609.21117#bib.bib11)\), we combine the standardized attributes additively, treating quality as a benefit and interaction cost as a penalty\. Thus,PzP\_\{z\}can be interpreted as an index that shows whether a session achieved higher quality for that task while requiring lower interaction cost\. In Appendix[F](https://arxiv.org/html/2609.21117#A6), we report sensitivity analyses varying the relative weight on quality and cost and show that the main qualitative findings are stable across weights\.

#### Quality as a task dependent design choice\.

The appropriate quality function \(QQ\) depends on the task and evaluation goal\. In code generation, the standard target is functional correctness, usually reported as pass@k[Chen et al\. \(2021\)](https://arxiv.org/html/2609.21117#bib.bib52)or test case pass rate[Hendrycks et al\. \(2021\)](https://arxiv.org/html/2609.21117#bib.bib54)\. For high stakes domains, such as medical[Ayers et al\. \(2023\)](https://arxiv.org/html/2609.21117#bib.bib53);[Tu et al\. \(2025\)](https://arxiv.org/html/2609.21117#bib.bib57)or mental health[Scholich et al\. \(2025\)](https://arxiv.org/html/2609.21117#bib.bib56);[Franke Föyen et al\. \(2025\)](https://arxiv.org/html/2609.21117#bib.bib58)support, expert ratings may be appropriate because correctness and harm require domain judgment\. For service oriented tasks, user ratings may be appropriate when the target outcome is whether the user’s goal was satisfied[Budzianowski et al\. \(2018\)](https://arxiv.org/html/2609.21117#bib.bib55)\. For open\-ended but rubric evaluable tasks, LLM\-as\-a\-judge rubric based scoring provides a consistent way to assess artifact quality across sessions\.

#### Interaction cost as grounding burden\.

Cost measures should reflect the interactional burden imposed by the collaboration\. User tokens approximate formulation effort, agent tokens approximate reception burden, and turns or elapsed time capture coarser interaction overhead\. Repair costs are not added as a separate term in our primary measure because they accumulate through the trajectory itself where repeated clarification and reformulation increase the number of tokens\. In this sense, token and turn based costs operationalize the idea from grounding theory that formulation, reception, and repair effort accumulate over the interaction\.

### 3\.2Data

We instantiate the framework on two datasets spanning four tasks: three structured human\-agent tasks from Collaborative Gym and one newly collected human\-LLM visualization task\.

#### CoGym

We use human\-agent collaboration sessions from the CoGym framework[Shao et al\. \(2026\)](https://arxiv.org/html/2609.21117#bib.bib2)spanning three diverse tasks: travel planning \(n = 112\), tabular analysis \(n = 69\), and related work synthesis \(n = 47\)\. In these tasks, users collaborate with an LLM agent that can act in a shared environment, such as searching databases, editing a shared document, or executing code\. Sessions were collected from real users interacting with LLM agents \(GPT\-4o, Gemini 2\.0 Flash\) through a web interface[Hurst et al\. \(2024\)](https://arxiv.org/html/2609.21117#bib.bib61);[Team et al\. \(2025\)](https://arxiv.org/html/2609.21117#bib.bib62)\. We restrict primary analyses to sessions producing a final artifact \(n=184\), and provide additional details of each task in Appendix[A](https://arxiv.org/html/2609.21117#A1)\.

#### Visualization

We additionally collect a new human\-LLM collaboration dataset of 42 sessions in which participants used GPT\-5\.1[Singh et al\. \(2026\)](https://arxiv.org/html/2609.21117#bib.bib43)to analyze the WildChat\-1M dataset[Zhao et al\. \(2024\)](https://arxiv.org/html/2609.21117#bib.bib17)\. Participants were asked to formulate a hypothesis about the dataset, test it using Python, and submit a visualization with supporting code and a short interpretation\. The task is open\-ended but the final output can still be evaluated with a rubric for relevance, correctness, and evidential support\. We log the full chat trajectory, submitted code, final visualization, and post task survey responses\. After submission, participants rated their confidence in the output, the usefulness of the AI, and perceived interaction productivity\. This dataset complements CoGym by covering a human\-AI collaboration task without environment actions\. We provide additional details on the study setting, custom platform, and procedure in Appendix[B](https://arxiv.org/html/2609.21117#A2)\.

#### Defining Q and C\.

Across both datasets, we use LLM\-as\-judge rubric scored task performance as the primary quality measureQQ\([Chiang et al\., 2024](https://arxiv.org/html/2609.21117#bib.bib37)\), using the task specific rubrics reported in Appendix[C](https://arxiv.org/html/2609.21117#A3)\. This choice reflects our goal of measuring productivity as the quality of the final artifact relative to the interaction cost required to produce it\. On the other hand, user Likert ratings capture how participants perceived the outcome or interaction\. We therefore reserve them for comparison analyses rather than using them to defineQQ\. This separation allows us to ask whether perceived productivity aligns with an independently scored quality\-cost tradeoff, and avoids circularity in later analyses\. Moreover, this avoids using the same user survey both to defineQQand to test whether users perceived the interaction as productive\([Podsakoff et al\., 2003](https://arxiv.org/html/2609.21117#bib.bib22)\)\. In our data, subjective Likert ratings are strongly correlated within each dataset, so we keep them as comparison measures rather than the primary quality measure \(Appendix[D](https://arxiv.org/html/2609.21117#A4)\)\. Our primary cost measure is weighted token cost:

C=user tokens\+0\.5⋅agent tokensC=\\text\{user tokens\}\+0\.5\\cdot\\text\{agent tokens\}User tokens approximate formulation effort, while agent tokens approximate reception burden\([Clark and Brennan, 1991](https://arxiv.org/html/2609.21117#bib.bib16)\)\. We assign agent tokens a lower weight because reading is typically faster than composing text[Brysbaert \(2019\)](https://arxiv.org/html/2609.21117#bib.bib59);[Karat et al\. \(1999\)](https://arxiv.org/html/2609.21117#bib.bib60), while still treating long model outputs as a cost imposed on the user\. Because no token based measure can fully capture effort such as cognitive load and attention, we test whether our conclusions depend on this particular operationalization\. Appendix[F](https://arxiv.org/html/2609.21117#A6)repeats the analyses with user tokens only, total tokens, turn count, and alternative agent token weights; the main conclusions remain qualitatively stable\.

#### Human validation\.

To validate the LLM\-as\-judge quality scores, two authors independently scored a subset of sessions using the same task specific rubrics\. For CoGym, we sample 20 sessions from each task and across the range of LLM\-rubric scores; for the visualization dataset, we score all 42 sessions\. Human scores correlate strongly with LLM\-as\-judge scores \(CoGymρ=\.68\\rho=\.68; visualizationρ=\.74\\rho=\.74\), with high human–human agreement \(ρ=\.81\\rho=\.81\)\. This suggests that the primary quality signal used inPzP\_\{z\}tracks human rubric judgments\.

### 3\.3Grounding and Friction Features

To characterize productive collaboration, we annotate each utterance using two taxonomies: grounding acts\([Shaikh et al\., 2025](https://arxiv.org/html/2609.21117#bib.bib25)\)and positive friction movements\([İnan et al\., 2025](https://arxiv.org/html/2609.21117#bib.bib26)\)\. Grounding acts characterize whether and how participants establish mutual understanding; friction movements characterize deliberate slowdowns that surface assumptions or invite reflection\. The two taxonomies capture different aspects of collaborative discourse\.

#### Grounding acts\.

We adapt the taxonomy of[Shaikh et al\. \(2025\)](https://arxiv.org/html/2609.21117#bib.bib25), which distinguishes acts thatadvancegrounding \(next turn,acknowledge,follow\-up\), acts thatsignal ambiguity\(clarification,overresponse\), and acts thataddressgrounding failures \(repair,reformulation,restart\)\. User and agent acts are labeled separately, since the same surface act carries different functional weight depending on the speaker \(e\.g\., an agent acknowledgment vs\. a user acknowledgment\)\. An utterance may receive multiple labels\.

#### Positive friction movements\.

We adapt the five category positive friction taxonomy from[İnan et al\. \(2025\)](https://arxiv.org/html/2609.21117#bib.bib26):assumption reveal\(surfacing beliefs about the environment, the interlocutor, or the task\),reflective pause\(verbal or behavioral signals of internal deliberation\),reinforcement\(restating a prior utterance for emphasis\),overspecification\(providing more information than requested\), andprobing\(questioning to redirect or clarify\)\. Each utterance receives a single label from these five categories orNot Frictionif none applies\. We retain the single label scheme of the original taxonomy\.

#### Annotation\.

We use GPT\-5\.1 at temperature 0 to label each utterance with grounding acts and positive friction movements, using the annotation guidelines from[Shaikh et al\. \(2025\)](https://arxiv.org/html/2609.21117#bib.bib25)and[İnan et al\. \(2025\)](https://arxiv.org/html/2609.21117#bib.bib26)\. We validate these labels on a subset of 20 sessions, balanced across dataset and productivity quartile\. Two human annotators independently label the same utterances\. Agreement was substantial for positive friction labels \(human–humanκ=\.72\\kappa=\.72; model–humanκ=\.64\\kappa=\.64\) and acceptable for grounding acts \(human–human micro\-F1=\.76=\.76; model–human micro\-F1=\.68=\.68\)\.

## 4Findings

Figure 2:Identical quality ratings can hide large differences in interaction cost\. Each panel shows the distribution of interaction cost within quality buckets for one task\. Cost is shown on a log scale\. Wide horizontal spread within a quality bucket indicates that task success alone cannot distinguish*productive success*from*costly success*\.![Refer to caption](https://arxiv.org/html/2609.21117v1/figures/fig_quality_cost_scatter.png)Figure 3:Task specific relationships between outcome quality and interaction cost\. Each point is a session, with quality and cost standardized within task; dashed lines indicate task means\. The pattern shows that interaction cost is not uniformly beneficial or harmful, where additional interaction accompanies higher quality in some tasks\.We organize the results around four questions\. First, does task success hide interaction costs? \(§[4\.1](https://arxiv.org/html/2609.21117#S4.SS1)\) Second, does the quality\-cost relationship vary by task? \(§[4\.2](https://arxiv.org/html/2609.21117#S4.SS2)\) Third, do subjective ratings capture the same quality\-cost tradeoff? \(§[4\.3](https://arxiv.org/html/2609.21117#S4.SS3)\) Finally, what dialogue patterns characterize productive collaboration? \(§[4\.4](https://arxiv.org/html/2609.21117#S4.SS4)\) These analyses show that productivity is not simply shorter interaction\. It depends on whether interaction cost helps produce quality, and on who bears the grounding work required to reach that quality\.

### 4\.1Cost Variation Within Quality Levels

Identical quality ratings can require very different amounts of interaction\. Figure[2](https://arxiv.org/html/2609.21117#S4.F2)plots interaction cost within each quality bucket separately for all three CoGym tasks and the visualization task\. Across tasks, sessions with the same rating often differ by one to two orders of magnitude in cost\.

The clearest example appears in travel planning\. Among completed sessions in the top quality bucket, interaction costs range from 854 to 60,324 tokens, a 70\.6 times difference\. Similar within quality variation appears in tabular analysis, related work, and visualization\. This motivates measuring productivity as quality relative to cost since task success alone cannot distinguish*productive success*, where high quality is reached with low interaction cost, from*costly success*, where similar quality requires substantially more interaction\. The same pattern holds when quality is measured with human Likert outcome rating rather than LLM rubric scores \(Appendix[E](https://arxiv.org/html/2609.21117#A5)\)\.

### 4\.2Task Structures Shape Quality & Cost

Interaction cost is not uniformly helpful or harmful; its relationship to quality depends on the task\. We compute the within task Spearman correlation between standardized qualityz⁡\(Q\)z\(Q\)and standardized costz⁡\(C\)z\(C\), and Figure[3](https://arxiv.org/html/2609.21117#S4.F3)visualizes the relationship at the session level\.

Related work shows the strongest positive association \(ρQ​C=\+0\.55\\rho\_\{QC\}=\+0\.55\), suggesting that additional interaction often supports more complete synthesis\. Travel planning is also positive but weaker \(ρQ​C=\+0\.27\\rho\_\{QC\}=\+0\.27\), consistent with a task where iteration helps but gains diminish as plans converge\. In contrast, tabular analysis shows little evidence of a positive quality\-cost relationship \(ρQ​C=−0\.13\\rho\_\{QC\}=\-0\.13\) where additional exchanges do not reliably improve quality\. This interpretation is consistent with works showing that data analysis is iterative and often repetitive[Wongsuphasawat et al\. \(2019\)](https://arxiv.org/html/2609.21117#bib.bib20);[Kandel et al\. \(2012\)](https://arxiv.org/html/2609.21117#bib.bib21)\. Visualization shows the most negative association \(ρQ​C=−0\.48\\rho\_\{QC\}=\-0\.48\), suggesting that high quality sessions depend less on iteration and tend to converge quickly on a hypothesis and visualization\.

These differences matter for evaluation\. A universal penalty on interaction length would mischaracterize tasks like related work, where more interaction can reflect useful elaboration\. Conversely, rewarding longer interaction would mischaracterize tasks where additional exchanges reflect debugging, uncertainty, or delayed convergence\. Productivity therefore needs to be interpreted within a task rather than a universal preference for brevity\.

### 4\.3Subjective Ratings Do Not Reliably Capture Productivity

If subjective ratings captured productivity, then users should penalize costly sessions when quality is held fixed\. We test this by computing, for each subjective measureSS, the partial Spearman correlationρ⁡\(C,S∣Q\)\\rho\(C,S\\mid Q\)between interaction costCCand the subjective rating, controlling for qualityQQ\. A productivity sensitive subjective measure should show a negative partial correlation, where among sessions of similar quality, higher cost should correspond to lower ratings\. Across six subjective measures in two datasets, this pattern appears reliably in only one case \(Table[1](https://arxiv.org/html/2609.21117#S4.T1)\)\.

CoGymsatisfactionshows the predicted cost penalty, both unconditionally \(ρ=−0\.22\\rho=\-0\.22,p<0\.01p<0\.01\) and after controlling for quality \(ρ=−0\.24\\rho=\-0\.24,p<0\.01p<0\.01\)\. In contrast, CoGymoutcome qualityandcommunication ratingare cost blind: holding quality fixed, longer or more costly interactions are not rated lower\. But, the visualization study showsSpeedis weakly cost\-blind,effectivenessis cost\-rewarding, andconfidenceis positively associated with cost even after controlling for quality \(ρ=\+0\.41\\rho=\+0\.41,p<0\.01p<0\.01\)\.

These divergences suggest that subjective prompts elicit different constructs\. Quality focused prompts \(outcome quality,communication rating\) anchor users to the final artifact, where interaction cost may be less salient\. Experience focused prompts can split into two interpretations, where satisfaction appears to treat cost as negative experience, while confidence may treat cost as investment or commitment\. Subjective ratings are therefore useful, but they are not interchangeable proxies for productivity\.

Table 1:Partial Spearman correlations of interaction costCCwith subjective measuresSS, controlling for LLM rubric qualityQQ\. CoGym measures are either cost\-blind or cost\-penalizing; Visualization metrics are either cost\-blind or cost\-rewarding\. \*p<0\.05p<0\.05, \*\*p<0\.01p<0\.01\.
### 4\.4Productive Collaboration Shifts Interactional Labor to the Agent

Figure 4:Agent and user probing in top vs\. bottomPzP\_\{z\}quartile \(n=59 each\)\. \(a\) Rate per turn: productive sessions show a labor crossover, with the agent probing more and the user probing less\. \(b\) Counts per session: agent probecountsare identical, but user probe counts more than double in unproductive sessions\. Error bars:±\\pm1 SEM\.We next ask what productive collaboration looks like in the dialogue itself\. We compare the top and bottomPzP\_\{z\}quartiles \(n=59 each\) on grounding acts and positive friction movements\. For each feature, we report Cohen’sddfor the difference in means between quartiles, where positive values indicate features that are more frequent in productive sessions\. Significance is assessed with Mann WhitneyUUtests; full feature results are in Appendix[G](https://arxiv.org/html/2609.21117#A7)\.

#### Friction is not inherently unproductive; its speaker matters\.

The rate of any friction per turn is nearly identical between productive and unproductive sessions \(0\.77 vs\. 0\.79,d=−0\.31d=\-0\.31, n\.s\.\)\. However, unproductive sessions contain more total friction movements because they are longer \(8\.0 vs\. 14\.1,d=−0\.65d=\-0\.65,p<0\.001p<0\.001\)\. The difference between these sessions is therefore not how often friction occurs, butwhoproduces it\.

The clearest example is probing \(Figure[4](https://arxiv.org/html/2609.21117#S4.F4)\)\. Productive sessions devote 44% of agent turns to probing, compared to 27% in unproductive sessions \(d=\+0\.73d=\+0\.73,p<0\.001p<0\.001\)\. Across the full sample, agent probing rate is positively correlated withPzP\_\{z\}\(ρ=\+0\.19\\rho=\+0\.19,p<0\.01p<0\.01, n=233\)\. This effect is concentrated in tasks where iterative clarification is useful, such as related work \(ρ=\+0\.35\\rho=\+0\.35\) and travel planning \(ρ=\+0\.24\\rho=\+0\.24\), consistent with the task level patterns in §[4\.2](https://arxiv.org/html/2609.21117#S4.SS2)\(Table[4](https://arxiv.org/html/2609.21117#A7.T4)\)\. Agent probecountsare identical across quartiles \(n¯=1\.81\\bar\{n\}=1\.81in both; Figure[4](https://arxiv.org/html/2609.21117#S4.F4)b\), which shows that productive sessions reach the same amount of clarification in roughly half the number of turns\. The pattern is stronger early in the session where 48% of agent turns in the first 30% of productive sessions contain probing, vs\. 31% in unproductive ones\. User probing shows the opposite pattern \(d=−0\.52d=\-0\.52,p<0\.01p<0\.01\), suggesting that when the agent does not probe, users must do the clarification work themselves\. Assumption reveal follows the same trend, concentrated on the user side in unproductive sessions \(n=0\.31n=0\.31vs\.1\.371\.37,d=−0\.57d=\-0\.57\)\.

#### Grounding acts show the same asymmetry\.

User repair \(0\.22 vs\. 0\.83,d=−0\.58d=\-0\.58\), clarification \(0\.27 vs\. 0\.63,d=−0\.49d=\-0\.49\), follow\-up \(3\.05 vs\. 6\.02,d=−0\.55d=\-0\.55\) are all elevated in unproductive sessions\. By contrast, user acknowledgment and agreement rates run roughly 3×\\timeshigher in productive sessions \(d=\+0\.58d=\+0\.58for both,p<0\.05p<0\.05\)\. Productive collaboration therefore looks likeagent probes→\\rightarrowuser confirms; unproductive collaboration leaves the grounding work to the user to repair, clarify, and probe\. This matches the principle of least collaborative effort\([Clark and Wilkes\-Gibbs, 1986](https://arxiv.org/html/2609.21117#bib.bib18);[Clark and Brennan, 1991](https://arxiv.org/html/2609.21117#bib.bib16)\)\. Agent side friction can be productive when it surfaces ambiguity early and prevents downstream repair\. The same friction becomes costly when the user must perform it later as repair work\.

## 5Discussion

Our findings suggest that productive human\-AI collaboration is about how interactional work is distributed across the collaboration\.

#### Productive collaboration depends on who bears the grounding work\.

Our dialogue analysis shows that productive collaboration is not simply less interactive or less frictional\. Productive and unproductive sessions contain similar rates of friction; what differs is who performs the grounding work \(§[4\.4](https://arxiv.org/html/2609.21117#S4.SS4)\)\. In productive sessions, agents probe earlier and users confirm; in unproductive sessions, users repair, clarify, and probe the agent\. Calls for seamless interaction often imply that friction should be minimized, while positive friction argues that strategic slowdowns can improve reliability\([İnan et al\., 2025](https://arxiv.org/html/2609.21117#bib.bib26)\)\. Our results suggest that friction is productive when the agent pays a small clarification cost up front, and costly when unresolved ambiguity later becomes user\-side repair\.This implies that agents should be designed to front\-load clarification, surface assumptions, and ask targeted probes before acting, rather than waiting for users to initiate repair\.

#### Productivity is task relative\.

The quality\-cost relationship varies substantially across tasks, from positive in related work synthesis \(ρQ​C=\+0\.55\\rho\_\{QC\}=\+0\.55\) to negative in visualization \(ρQ​C=−0\.48\\rho\_\{QC\}=\-0\.48; §[4\.2](https://arxiv.org/html/2609.21117#S4.SS2)\)\. This means there is no universal effort to outcome curve for human\-AI collaboration\. In some tasks, additional interaction reflects useful elaboration, exploration, or synthesis; in others, it may reflect delayed convergence, debugging, or confusion\. Evaluation protocols that uniformly penalize interaction length may therefore underrate systems on tasks where extended interaction is productive, while protocols that reward engagement may overrate systems on tasks where high quality outcomes should converge quickly\.This implies that productivity should be evaluated and interpreted within a task before comparing across tasks or systems\.

#### Subjective ratings measure different constructs\.

Subjective ratings are valuable, but they are not interchangeable measures of productivity\. In our data, most subjective measures do not penalize interaction cost once quality is held fixed, and some move in the opposite direction \(§[4\.3](https://arxiv.org/html/2609.21117#S4.SS3)\)\. This suggests that quality focused prompts direct users to the artifact, satisfaction may integrate cost as negative experience, and confidence may reflect commitment\. As a result, two studies could reach different conclusions about the same system depending on whether they ask users about satisfaction, confidence, usefulness, or speed\.This has implications in studies of human\-AI collaboration which should specify what each subjective prompt is intended to measure, and should avoid treating perceived quality, usefulness, confidence, and productivity as interchangeable outcomes\.

Productivity is conditional on the evaluation objective\.Our framework does not assume that interaction is inherently costly or that shorter interactions are always preferable\. Interaction cost is meaningful only relative to the outcome the collaboration is intended to achieve, and the quality function \(Q\) must therefore reflect the values relevant to that task\. For example, in an educational setting, immediately giving a student the correct answer may minimize interaction while undermining the objective of learning\. If learning is the intended outcome,QQshould instead capture outcomes such as retention, transfer, or independent problem solving\. We therefore view collaborative productivity as efficiency conditional on task relevant goals and not as a prescription to minimize interaction\.

## 6Conclusion

Task success tells us whether a human\-AI collaboration reached a desired outcome, but not what it cost the user to get there\. We introduced a productivity oriented framework for evaluating collaboration as outcome quality relative to interaction cost, and applied it across four tasks from two datasets\. Our results show that identical quality ratings can hide large differences in interaction cost, that the quality\-cost relationship varies by task, and that subjective ratings do not consistently capture this tradeoff\. Dialogue analysis further shows that productive collaboration shifts grounding work toward the agent\. Productive collaboration is not just successful collaboration made shorter\. It is successful collaboration in which interactional effort is spent where it contributes meaningfully to the outcome rather than being unnecessarily transferred to the user\. These findings motivate the design of agents that surface ambiguity and probe when appropriate to reduce unnecessary user repair\.

## Limitations

#### Interaction cost is an approximation of user effort\.

We use weighted token cost as our primary measure, treating user tokens as a proxy for formulation effort and agent tokens as a proxy for reception burden\. This operationalization captures part of the interactional burden, but it does not directly measure cognitive load, attention, frustration, multitasking, or the effort required to evaluate whether an answer is correct\.

#### Dialogue patterns are correlational\.

Our dialogue analysis identifies features associated with productive sessions, such as agent probing and reduced user repair\. However, these results are observational\. Future experiments should test whether prompting or training agents to front load clarification improves productivity without introducing unnecessary friction\.

#### Generalizability is limited by tasks\.

We evaluate four tasks across two datasets\. These settings provide useful variation, but they do not cover the full range of human\-AI collaboration\. Productivity may behave differently in long horizon work, creative collaboration, or domains where trust, safety, and accountability matter more than speed\.

## Use of AI Assistants

The authors used AI assistants for language polishing after the initial draft\. The authors verified and edited all generated text and remain responsible for the content of the paper\.

## Acknowledgments

This research was supported in part by a grant from the Institute for Information, the Internet, and Democracy \(IIID\) at Northeastern University\. We thank Asteria Kaeberlein for their helpful feedback\.

## References

- Ayerset al\.\(2023\)J\. W\. Ayers, A\. Poliak, M\. Dredze, E\. C\. Leas, Z\. Zhu, J\. B\. Kelley, D\. J\. Faix, A\. M\. Goodman, C\. A\. Longhurst, M\. Hogarth, and D\. M\. SmithComparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum\.JAMA Internal Medicine183\(6\),pp\. 589–596\.External Links:ISSN 2168\-6106,[Document](https://dx.doi.org/10.1001/jamainternmed.2023.1838),[Link](https://doi.org/10.1001/jamainternmed.2023.1838),https://jamanetwork\.com/journals/jamainternalmedicine/articlepdf/2804309/jamainternal\_ayers\_2023\_oi\_230030\_1685974538\.66672\.pdfCited by:[§3\.1](https://arxiv.org/html/2609.21117#S3.SS1.SSS0.Px1.p1.1)\.
- Barreset al\.\(2025\)V\. Barres, H\. Dong, S\. Ray, X\. Si, and K\. Narasimhanτ2\\tau^\{2\}\-Bench: evaluating conversational agents in a dual\-control environment\.External Links:2506\.07982,[Link](https://arxiv.org/abs/2506.07982)Cited by:[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px2.p1.1)\.
- Beckeret al\.\(2025\)J\. Becker, N\. Rush, E\. Barnes, and D\. ReinMeasuring the impact of early\-2025 ai on experienced open\-source developer productivity\.External Links:2507\.09089,[Link](https://arxiv.org/abs/2507.09089)Cited by:[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px1.p1.1)\.
- Brynjolfssonet al\.\(2025\)E\. Brynjolfsson, D\. Li, and L\. RaymondGenerative ai at work\.The Quarterly Journal of Economics140\(2\),pp\. 889–942\.Cited by:[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px1.p1.1)\.
- Brysbaert \(2019\)M\. BrysbaertHow many words do we read per minute? a review and meta\-analysis of reading rate\.Journal of Memory and Language109,pp\. 104047\.External Links:ISSN 0749\-596X,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.jml.2019.104047),[Link](https://www.sciencedirect.com/science/article/pii/S0749596X19300786)Cited by:[§3\.2](https://arxiv.org/html/2609.21117#S3.SS2.SSS0.Px3.p1.2)\.
- Budzianowskiet al\.\(2018\)P\. Budzianowski, T\. Wen, B\. Tseng, I\. Casanueva, S\. Ultes, O\. Ramadan, and M\. GašićMultiWOZ \- a large\-scale multi\-domain Wizard\-of\-Oz dataset for task\-oriented dialogue modelling\.InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing,E\. Riloff, D\. Chiang, J\. Hockenmaier, and J\. Tsujii \(Eds\.\),Brussels, Belgium,pp\. 5016–5026\.External Links:[Link](https://aclanthology.org/D18-1547/),[Document](https://dx.doi.org/10.18653/v1/D18-1547)Cited by:[§3\.1](https://arxiv.org/html/2609.21117#S3.SS1.SSS0.Px1.p1.1)\.
- Changet al\.\(2025\)S\. Chang, A\. Anderson, and J\. M\. HofmanChatBench: from static benchmarks to human\-AI evaluation\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 26009–26038\.External Links:[Link](https://aclanthology.org/2025.acl-long.1262/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.1262),ISBN 979\-8\-89176\-251\-0Cited by:[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px2.p1.1)\.
- Chenet al\.\(2021\)M\. Chen, J\. Tworek, H\. Jun, Q\. Yuan, H\. P\. de Oliveira Pinto, J\. Kaplan, H\. Edwards, Y\. Burda, N\. Joseph, G\. Brockman, A\. Ray, R\. Puri, G\. Krueger, M\. Petrov, H\. Khlaaf, G\. Sastry, P\. Mishkin, B\. Chan, S\. Gray, N\. Ryder, M\. Pavlov, A\. Power, L\. Kaiser, M\. Bavarian, C\. Winter, P\. Tillet, F\. P\. Such, D\. Cummings, M\. Plappert, F\. Chantzis, E\. Barnes, A\. Herbert\-Voss, W\. H\. Guss, A\. Nichol, A\. Paino, N\. Tezak, J\. Tang, I\. Babuschkin, S\. Balaji, S\. Jain, W\. Saunders, C\. Hesse, A\. N\. Carr, J\. Leike, J\. Achiam, V\. Misra, E\. Morikawa, A\. Radford, M\. Knight, M\. Brundage, M\. Murati, K\. Mayer, P\. Welinder, B\. McGrew, D\. Amodei, S\. McCandlish, I\. Sutskever, and W\. ZarembaEvaluating large language models trained on code\.External Links:2107\.03374,[Link](https://arxiv.org/abs/2107.03374)Cited by:[§3\.1](https://arxiv.org/html/2609.21117#S3.SS1.SSS0.Px1.p1.1)\.
- Chianget al\.\(2024\)W\. Chiang, L\. Zheng, Y\. Sheng, A\. N\. Angelopoulos, T\. Li, D\. Li, H\. Zhang, B\. Zhu, M\. Jordan, J\. E\. Gonzalez, and I\. StoicaChatbot arena: an open platform for evaluating llms by human preference\.External Links:2403\.04132,[Link](https://arxiv.org/abs/2403.04132)Cited by:[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px2.p1.1),[§3\.2](https://arxiv.org/html/2609.21117#S3.SS2.SSS0.Px3.p1.1)\.
- Clark and Brennan \(1991\)H\. H\. Clark and S\. E\. BrennanGrounding in communication\.\.Cited by:[§1](https://arxiv.org/html/2609.21117#S1.p3.1),[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px3.p1.1),[§3\.1](https://arxiv.org/html/2609.21117#S3.SS1.p2.1),[§3\.2](https://arxiv.org/html/2609.21117#S3.SS2.SSS0.Px3.p1.2),[§4\.4](https://arxiv.org/html/2609.21117#S4.SS4.SSS0.Px2.p1.1)\.
- Clark and Wilkes\-Gibbs \(1986\)H\. H\. Clark and D\. Wilkes\-GibbsReferring as a collaborative process\.Cognition22\(1\),pp\. 1–39\.Cited by:[§1](https://arxiv.org/html/2609.21117#S1.p3.1),[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px3.p1.1),[§4\.4](https://arxiv.org/html/2609.21117#S4.SS4.SSS0.Px2.p1.1)\.
- Dell’Acquaet al\.\(2026\)F\. Dell’Acqua, E\. McFowland, E\. Mollick, H\. Lifshitz, K\. C\. Kellogg, S\. Rajendran, L\. Krayer, F\. Candelon, and K\. R\. LakhaniNavigating the jagged technological frontier: field experimental evidence of the effects of artificial intelligence on knowledge worker productivity and quality\.Organization Science37\(2\),pp\. 403–423\.External Links:[Document](https://dx.doi.org/10.1287/orsc.2025.21838),[Link](https://doi.org/10.1287/orsc.2025.21838),https://doi\.org/10\.1287/orsc\.2025\.21838Cited by:[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px1.p1.1)\.
- Eloundouet al\.\(2024\)T\. Eloundou, S\. Manning, P\. Mishkin, and D\. RockGPTs are gpts: labor market impact potential of llms\.Science384\(6702\),pp\. 1306–1308\.External Links:[Document](https://dx.doi.org/10.1126/science.adj0998),[Link](https://www.science.org/doi/abs/10.1126/science.adj0998),https://www\.science\.org/doi/pdf/10\.1126/science\.adj0998Cited by:[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px1.p1.1)\.
- Fragiadakiset al\.\(2025\)G\. Fragiadakis, C\. Diou, G\. Kousiouris, and M\. NikolaidouEvaluating human\-ai collaboration: a review and methodological framework\.External Links:2407\.19098,[Link](https://arxiv.org/abs/2407.19098)Cited by:[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px2.p1.1)\.
- Franke Föyenet al\.\(2025\)L\. Franke Föyen, E\. Zapel, M\. Lekander, E\. Hedman\-Lagerlöf, and E\. LindsäterArtificial intelligence vs\. human expert: licensed mental health clinicians’ blinded evaluation of ai\-generated and expert psychological advice on quality, empathy, and perceived authorship\.Internet Interventions41,pp\. 100841\.External Links:ISSN 2214\-7829,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.invent.2025.100841),[Link](https://www.sciencedirect.com/science/article/pii/S2214782925000429)Cited by:[§3\.1](https://arxiv.org/html/2609.21117#S3.SS1.SSS0.Px1.p1.1)\.
- Gans and Goldfarb \(2026\)J\. S\. Gans and A\. GoldfarbO\-ring automation\.Working PaperTechnical Report34639,Working Paper Series,National Bureau of Economic Research\.External Links:[Document](https://dx.doi.org/10.3386/w34639),[Link](http://www.nber.org/papers/w34639)Cited by:[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px1.p1.1)\.
- Handaet al\.\(2025\)K\. Handa, A\. Tamkin, M\. McCain, S\. Huang, E\. Durmus, S\. Heck, J\. Mueller, J\. Hong, S\. Ritchie, T\. Belonax, K\. K\. Troy, D\. Amodei, J\. Kaplan, J\. Clark, and D\. GanguliWhich economic tasks are performed with ai? evidence from millions of claude conversations\.External Links:2503\.04761,[Link](https://arxiv.org/abs/2503.04761)Cited by:[§1](https://arxiv.org/html/2609.21117#S1.p2.1),[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px1.p1.1)\.
- Hendryckset al\.\(2021\)D\. Hendrycks, S\. Basart, S\. Kadavath, M\. Mazeika, A\. Arora, E\. Guo, C\. Burns, S\. Puranik, H\. He, D\. Song, and J\. SteinhardtMeasuring coding challenge competence with apps\.External Links:2105\.09938,[Link](https://arxiv.org/abs/2105.09938)Cited by:[§3\.1](https://arxiv.org/html/2609.21117#S3.SS1.SSS0.Px1.p1.1)\.
- Hurstet al\.\(2024\)A\. Hurst, A\. Lerer, A\. P\. Goucher, A\. Perelman, A\. Ramesh, A\. Clark, A\. Ostrow, A\. Welihinda, A\. Hayes, A\. Radford, A\. Mądry, A\. Baker\-Whitcomb, A\. Beutel, A\. Borzunov, A\. Carney, A\. Chow, A\. Kirillov, A\. Nichol, A\. Paino, A\. Renzin, A\. T\. Passos, A\. Kirillov, A\. Christakis, A\. Conneau, A\. Kamali, A\. Jabri, A\. Moyer, A\. Tam, A\. Crookes, A\. Tootoochian, A\. Tootoonchian, A\. Kumar, A\. Vallone, A\. Karpathy, A\. Braunstein, A\. Cann, A\. Codispoti, A\. Galu, A\. Kondrich, A\. Tulloch, A\. Mishchenko, A\. Baek, A\. Jiang, A\. Pelisse, A\. Woodford, A\. Gosalia, A\. Dhar, A\. Pantuliano, A\. Nayak, A\. Oliver, B\. Zoph, B\. Ghorbani, B\. Leimberger, B\. Rossen, B\. Sokolowsky, B\. Wang, B\. Zweig, B\. Hoover, B\. Samic, B\. McGrew, B\. Spero, B\. Giertler, B\. Cheng, B\. Lightcap, B\. Walkin, B\. Quinn, B\. Guarraci, B\. Hsu, B\. Kellogg, B\. Eastman, C\. Lugaresi, C\. Wainwright, C\. Bassin, C\. Hudson, C\. Chu, C\. Nelson, C\. Li, C\. J\. Shern, C\. Conger, C\. Barette, C\. Voss, C\. Ding, C\. Lu, C\. Zhang, C\. Beaumont, C\. Hallacy, C\. Koch, C\. Gibson, C\. Kim, C\. Choi, C\. McLeavey, C\. Hesse, C\. Fischer, C\. Winter, C\. Czarnecki, C\. Jarvis, C\. Wei, C\. Koumouzelis, D\. Sherburn, D\. Kappler, D\. Levin, D\. Levy, D\. Carr, D\. Farhi, D\. Mely, D\. Robinson, D\. Sasaki, D\. Jin, D\. Valladares, D\. Tsipras, D\. Li, D\. P\. Nguyen, D\. Findlay, E\. Oiwoh, E\. Wong, E\. Asdar, E\. Proehl, E\. Yang, E\. Antonow, E\. Kramer, E\. Peterson, E\. Sigler, E\. Wallace, E\. Brevdo, E\. Mays, F\. Khorasani, F\. P\. Such, F\. Raso, F\. Zhang, F\. von Lohmann, F\. Sulit, G\. Goh, G\. Oden, G\. Salmon, G\. Starace, G\. Brockman, H\. Salman, H\. Bao, H\. Hu, H\. Wong, H\. Wang, H\. Schmidt, H\. Whitney, H\. Jun, H\. Kirchner, H\. P\. de Oliveira Pinto, H\. Ren, H\. Chang, H\. W\. Chung, I\. Kivlichan, I\. O’Connell, I\. O’Connell, I\. Osband, I\. Silber, I\. Sohl, I\. Okuyucu, I\. Lan, I\. Kostrikov, I\. Sutskever, I\. Kanitscheider, I\. Gulrajani, J\. Coxon, J\. Menick, J\. Pachocki, J\. Aung, J\. Betker, J\. Crooks, J\. Lennon, J\. Kiros, J\. Leike, J\. Park, J\. Kwon, J\. Phang, J\. Teplitz, J\. Wei, J\. Wolfe, J\. Chen, J\. Harris, J\. Varavva, J\. G\. Lee, J\. Shieh, J\. Lin, J\. Yu, J\. Weng, J\. Tang, J\. Yu, J\. Jang, J\. Q\. Candela, J\. Beutler, J\. Landers, J\. Parish, J\. Heidecke, J\. Schulman, J\. Lachman, J\. McKay, J\. Uesato, J\. Ward, J\. W\. Kim, J\. Huizinga, J\. Sitkin, J\. Kraaijeveld, J\. Gross, J\. Kaplan, J\. Snyder, J\. Achiam, J\. Jiao, J\. Lee, J\. Zhuang, J\. Harriman, K\. Fricke, K\. Hayashi, K\. Singhal, K\. Shi, K\. Karthik, K\. Wood, K\. Rimbach, K\. Hsu, K\. Nguyen, K\. Gu\-Lemberg, K\. Button, K\. Liu, K\. Howe, K\. Muthukumar, K\. Luther, L\. Ahmad, L\. Kai, L\. Itow, L\. Workman, L\. Pathak, L\. Chen, L\. Jing, L\. Guy, L\. Fedus, L\. Zhou, L\. Mamitsuka, L\. Weng, L\. McCallum, L\. Held, L\. Ouyang, L\. Feuvrier, L\. Zhang, L\. Kondraciuk, L\. Kaiser, L\. Hewitt, L\. Metz, L\. Doshi, M\. Aflak, M\. Simens, M\. Boyd, M\. Thompson, M\. Dukhan, M\. Chen, M\. Gray, M\. Hudnall, M\. Zhang, M\. Aljubeh, M\. Litwin, M\. Zeng, M\. Johnson, M\. Shetty, M\. Gupta, M\. Shah, M\. Yatbaz, M\. J\. Yang, M\. Zhong, M\. Glaese, M\. Chen, M\. Janner, M\. Lampe, M\. Petrov, M\. Wu, M\. Wang, M\. Fradin, M\. Pokrass, M\. Castro, M\. O\. T\. de Castro, M\. Pavlov, M\. Brundage, M\. Wang, M\. Khan, M\. Murati, M\. Bavarian, M\. Lin, M\. Yesildal, N\. Soto, N\. Gimelshein, N\. Cone, N\. Staudacher, N\. Summers, N\. LaFontaine, N\. Chowdhury, N\. Ryder, N\. Stathas, N\. Turley, N\. Tezak, N\. Felix, N\. Kudige, N\. Keskar, N\. Deutsch, N\. Bundick, N\. Puckett, O\. Nachum, O\. Okelola, O\. Boiko, O\. Murk, O\. Jaffe, O\. Watkins, O\. Godement, O\. Campbell\-Moore, P\. Chao, P\. McMillan, P\. Belov, P\. Su, P\. Bak, P\. Bakkum, P\. Deng, P\. Dolan, P\. Hoeschele, P\. Welinder, P\. Tillet, P\. Pronin, P\. Tillet, P\. Dhariwal, Q\. Yuan, R\. Dias, R\. Lim, R\. Arora, R\. Troll, R\. Lin, R\. G\. Lopes, R\. Puri, R\. Miyara, R\. Leike, R\. Gaubert, R\. Zamani, R\. Wang, R\. Donnelly, R\. Honsby, R\. Smith, R\. Sahai, R\. Ramchandani, R\. Huet, R\. Carmichael, R\. Zellers, R\. Chen, R\. Chen, R\. Nigmatullin, R\. Cheu, S\. Jain, S\. Altman, S\. Schoenholz, S\. Toizer, S\. Miserendino, S\. Agarwal, S\. Culver, S\. Ethersmith, S\. Gray, S\. Grove, S\. Metzger, S\. Hermani, S\. Jain, S\. Zhao, S\. Wu, S\. Jomoto, S\. Wu, Shuaiqi, Xia, S\. Phene, S\. Papay, S\. Narayanan, S\. Coffey, S\. Lee, S\. Hall, S\. Balaji, T\. Broda, T\. Stramer, T\. Xu, T\. Gogineni, T\. Christianson, T\. Sanders, T\. Patwardhan, T\. Cunninghman, T\. Degry, T\. Dimson, T\. Raoux, T\. Shadwell, T\. Zheng, T\. Underwood, T\. Markov, T\. Sherbakov, T\. Rubin, T\. Stasi, T\. Kaftan, T\. Heywood, T\. Peterson, T\. Walters, T\. Eloundou, V\. Qi, V\. Moeller, V\. Monaco, V\. Kuo, V\. Fomenko, W\. Chang, W\. Zheng, W\. Zhou, W\. Manassra, W\. Sheu, W\. Zaremba, Y\. Patil, Y\. Qian, Y\. Kim, Y\. Cheng, Y\. Zhang, Y\. He, Y\. Zhang, Y\. Jin, Y\. Dai, and Y\. MalkovGPT\-4o system card\.External Links:2410\.21276,[Link](https://arxiv.org/abs/2410.21276)Cited by:[§3\.2](https://arxiv.org/html/2609.21117#S3.SS2.SSS0.Px1.p1.1)\.
- Imai \(2022\)S\. ImaiIs github copilot a substitute for human pair\-programming? an empirical study\.InProceedings of the ACM/IEEE 44th International Conference on Software Engineering: Companion Proceedings,ICSE ’22,New York, NY, USA,pp\. 319–321\.External Links:ISBN 9781450392235,[Link](https://doi.org/10.1145/3510454.3522684),[Document](https://dx.doi.org/10.1145/3510454.3522684)Cited by:[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px1.p1.1)\.
- İnanet al\.\(2025\)M\. İnan, A\. Sicilia, S\. Dey, V\. Dongre, T\. Srinivasan, J\. Thomason, G\. Tür, D\. Hakkani\-Tür, and M\. AlikhaniBetter slow than sorry: introducing positive friction for reliable dialogue systems\.External Links:2501\.17348,[Link](https://arxiv.org/abs/2501.17348)Cited by:[Table 3](https://arxiv.org/html/2609.21117#A7.T3.2.2.1),[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px3.p1.1),[§3\.3](https://arxiv.org/html/2609.21117#S3.SS3.SSS0.Px2.p1.1),[§3\.3](https://arxiv.org/html/2609.21117#S3.SS3.SSS0.Px3.p1.1),[§3\.3](https://arxiv.org/html/2609.21117#S3.SS3.p1.1),[§5](https://arxiv.org/html/2609.21117#S5.SS0.SSS0.Px1.p1.1)\.
- Jimenezet al\.\(2024\)C\. E\. Jimenez, J\. Yang, A\. Wettig, S\. Yao, K\. Pei, O\. Press, and K\. R\. NarasimhanSWE\-bench: can language models resolve real\-world github issues?\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=VTF8yNQM66)Cited by:[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px2.p1.1)\.
- Ju and Aral \(2026\)H\. Ju and S\. AralCollaborating with ai agents: field experiments on teamwork, productivity, and performance\.External Links:2503\.18238,[Link](https://arxiv.org/abs/2503.18238)Cited by:[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px1.p1.1)\.
- Kandelet al\.\(2012\)S\. Kandel, A\. Paepcke, J\. M\. Hellerstein, and J\. HeerEnterprise data analysis and visualization: an interview study\.IEEE Transactions on Visualization and Computer Graphics18\(12\),pp\. 2917–2926\.External Links:[Document](https://dx.doi.org/10.1109/TVCG.2012.219)Cited by:[§4\.2](https://arxiv.org/html/2609.21117#S4.SS2.p2.1)\.
- Karatet al\.\(1999\)C\. Karat, C\. Halverson, D\. Horn, and J\. KaratPatterns of entry and correction in large vocabulary continuous speech recognition systems\.InProceedings of the SIGCHI Conference on Human Factors in Computing Systems,CHI ’99,New York, NY, USA,pp\. 568–575\.External Links:ISBN 0201485591,[Link](https://doi.org/10.1145/302979.303160),[Document](https://dx.doi.org/10.1145/302979.303160)Cited by:[§3\.2](https://arxiv.org/html/2609.21117#S3.SS2.SSS0.Px3.p1.2)\.
- Keeney and Raiffa \(1993\)R\. L\. Keeney and H\. RaiffaDecisions with multiple objectives: preferences and value trade\-offs\.Cambridge University Press\.Cited by:[§3\.1](https://arxiv.org/html/2609.21117#S3.SS1.p3.2)\.
- Kwaet al\.\(2026\)T\. Kwa, B\. West, J\. Becker, A\. Deng, K\. Garcia, M\. Hasin, S\. Jawhar, M\. Kinniment, N\. Rush, S\. V\. Arx, R\. Bloom, T\. Broadley, H\. Du, B\. Goodrich, N\. Jurkovic, L\. H\. Miles, S\. Nix, T\. Lin, N\. Parikh, D\. Rein, L\. J\. K\. Sato, H\. Wijk, D\. M\. Ziegler, E\. Barnes, and L\. ChanMeasuring ai ability to complete long software tasks\.External Links:2503\.14499,[Link](https://arxiv.org/abs/2503.14499)Cited by:[§1](https://arxiv.org/html/2609.21117#S1.p2.1),[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px1.p1.1)\.
- Leeet al\.\(2022\)M\. Lee, P\. Liang, and Q\. YangCoAuthor: designing a human\-ai collaborative writing dataset for exploring language model capabilities\.InCHI Conference on Human Factors in Computing Systems,CHI ’22,pp\. 1–19\.External Links:[Link](http://dx.doi.org/10.1145/3491102.3502030),[Document](https://dx.doi.org/10.1145/3491102.3502030)Cited by:[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px2.p1.1)\.
- Leeet al\.\(2023\)M\. Lee, M\. Srivastava, A\. Hardy, J\. Thickstun, E\. Durmus, A\. Paranjape, I\. Gerard\-Ursin, X\. L\. Li, F\. Ladhak, F\. Rong, R\. E\. Wang, M\. Kwon, J\. S\. Park, H\. Cao, T\. Lee, R\. Bommasani, M\. S\. Bernstein, and P\. LiangEvaluating human\-language model interaction\.Transactions on Machine Learning Research\.Note:External Links:ISSN 2835\-8856,[Link](https://openreview.net/forum?id=hjDYJUn9l1)Cited by:[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px2.p1.1)\.
- Liet al\.\(2025\)T\. Li, W\. Chiang, E\. Frick, L\. Dunlap, T\. Wu, B\. Zhu, J\. E\. Gonzalez, and I\. StoicaFrom crowdsourced data to high\-quality benchmarks: arena\-hard and benchbuilder pipeline\.InProceedings of the 42nd International Conference on Machine Learning,A\. Singh, M\. Fazel, D\. Hsu, S\. Lacoste\-Julien, F\. Berkenkamp, T\. Maharaj, K\. Wagstaff, and J\. Zhu \(Eds\.\),Proceedings of Machine Learning Research, Vol\.267,pp\. 34209–34231\.External Links:[Link](https://proceedings.mlr.press/v267/li25h.html)Cited by:[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px2.p1.1)\.
- Linet al\.\(2025\)B\. Y\. Lin, Y\. Deng, K\. Chandu, A\. Ravichander, V\. Pyatkin, N\. Dziri, R\. Le Bras, and Y\. ChoiWildBench: benchmarking llms with challenging tasks from real users in the wild\.InInternational Conference on Learning Representations,Y\. Yue, A\. Garg, N\. Peng, F\. Sha, and R\. Yu \(Eds\.\),Vol\.2025,pp\. 47852–47870\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2025/file/771155abaae744e08576f1f3b4b7ac0d-Paper-Conference.pdf)Cited by:[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px2.p1.1)\.
- METR \(2026\)METRTask\-completion time horizons of frontier ai models\.Note:[https://metr\.org/time\-horizons/](https://metr.org/time-horizons/)Cited by:[§1](https://arxiv.org/html/2609.21117#S1.p2.1),[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px1.p1.1)\.
- Noy and Zhang \(2023\)S\. Noy and W\. ZhangExperimental evidence on the productivity effects of generative artificial intelligence\.Science381\(6654\),pp\. 187–192\.External Links:[Document](https://dx.doi.org/10.1126/science.adh2586),[Link](https://www.science.org/doi/abs/10.1126/science.adh2586),https://www\.science\.org/doi/pdf/10\.1126/science\.adh2586Cited by:[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px1.p1.1)\.
- Patwardhanet al\.\(2025\)T\. Patwardhan, R\. Dias, E\. Proehl, G\. Kim, M\. Wang, O\. Watkins, S\. P\. Fishman, M\. Aljubeh, P\. Thacker, L\. Fauconnet, N\. S\. Kim, P\. Chao, S\. Miserendino, G\. Chabot, D\. Li, M\. Sharman, A\. Barr, A\. Glaese, and J\. TworekGDPval: evaluating ai model performance on real\-world economically valuable tasks\.External Links:2510\.04374,[Link](https://arxiv.org/abs/2510.04374)Cited by:[§1](https://arxiv.org/html/2609.21117#S1.p2.1),[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px1.p1.1)\.
- Penget al\.\(2023\)S\. Peng, E\. Kalliamvakou, P\. Cihon, and M\. DemirerThe impact of ai on developer productivity: evidence from github copilot\.External Links:2302\.06590,[Link](https://arxiv.org/abs/2302.06590)Cited by:[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px1.p1.1)\.
- Podsakoffet al\.\(2003\)P\. Podsakoff, S\. MacKenzie, J\. Lee, and N\. PodsakoffCommon method biases in behavioral research: a critical review of the literature and recommended remedies\.Journal of Applied Psychology88,pp\. 879–903\.External Links:[Document](https://dx.doi.org/10.1037/0021-9010.88.5.879)Cited by:[§3\.2](https://arxiv.org/html/2609.21117#S3.SS2.SSS0.Px3.p1.1)\.
- Qian and Wexler \(2024\)C\. Qian and J\. WexlerTake it, leave it, or fix it: measuring productivity and trust in human\-ai collaboration\.InProceedings of the 29th International Conference on Intelligent User Interfaces,IUI ’24,pp\. 370–384\.External Links:[Link](http://dx.doi.org/10.1145/3640543.3645198),[Document](https://dx.doi.org/10.1145/3640543.3645198)Cited by:[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px1.p1.1)\.
- Roschelle and Teasley \(1995\)J\. Roschelle and S\. D\. TeasleyThe construction of shared knowledge in collaborative problem solving\.InComputer supported collaborative learning,pp\. 69–97\.Cited by:[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px3.p1.1)\.
- Scholichet al\.\(2025\)T\. Scholich, M\. Barr, S\. Wiltsey Stirman, and S\. RajA comparison of responses from human therapists and large language model–based chatbots to assess therapeutic communication: mixed methods study\.JMIR Ment Health12,pp\. e69709\.External Links:ISSN 2368\-7959,[Document](https://dx.doi.org/10.2196/69709),[Link](https://mental.jmir.org/2025/1/e69709),[Link](https://doi.org/10.2196/69709)Cited by:[§3\.1](https://arxiv.org/html/2609.21117#S3.SS1.SSS0.Px1.p1.1)\.
- Schreyer and Pilat \(2001\)P\. Schreyer and D\. PilatMeasuring productivity\.OECD Economic studies33\(2\),pp\. 127–170\.Cited by:[§3\.1](https://arxiv.org/html/2609.21117#S3.SS1.p1.1)\.
- Shaikhet al\.\(2024\)O\. Shaikh, K\. Gligorić, A\. Khetan, M\. Gerstgrasser, D\. Yang, and D\. JurafskyGrounding gaps in language model generations\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),K\. Duh, H\. Gomez, and S\. Bethard \(Eds\.\),Mexico City, Mexico,pp\. 6279–6296\.External Links:[Link](https://aclanthology.org/2024.naacl-long.348/),[Document](https://dx.doi.org/10.18653/v1/2024.naacl-long.348)Cited by:[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px3.p1.1)\.
- Shaikhet al\.\(2025\)O\. Shaikh, H\. Mozannar, G\. Bansal, A\. Fourney, and E\. HorvitzNavigating rifts in human\-LLM grounding: study and benchmark\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 20832–20847\.External Links:[Link](https://aclanthology.org/2025.acl-long.1016/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.1016),ISBN 979\-8\-89176\-251\-0Cited by:[Table 3](https://arxiv.org/html/2609.21117#A7.T3.2.19.1),[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px3.p1.1),[§3\.3](https://arxiv.org/html/2609.21117#S3.SS3.SSS0.Px1.p1.1),[§3\.3](https://arxiv.org/html/2609.21117#S3.SS3.SSS0.Px3.p1.1),[§3\.3](https://arxiv.org/html/2609.21117#S3.SS3.p1.1)\.
- Shaoet al\.\(2026\)Y\. Shao, V\. Samuel, Y\. Jiang, J\. Yang, and D\. YangCollaborative gym: a framework for enabling and evaluating human\-agent collaboration\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=GDYueXtKXT)Cited by:[Appendix C](https://arxiv.org/html/2609.21117#A3.p2.1),[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px2.p1.1),[§3\.2](https://arxiv.org/html/2609.21117#S3.SS2.SSS0.Px1.p1.1)\.
- Shenet al\.\(2025\)S\. Z\. Shen, V\. Chen, K\. Gu, A\. Ross, Z\. Ma, J\. Ross, A\. Gu, C\. Si, W\. Chi, A\. Peng, J\. J\. Shen, A\. Talwalkar, T\. Wu, and D\. SontagCompletion≠\\neqcollaboration: scaling collaborative effort with agents\.External Links:2510\.25744,[Link](https://arxiv.org/abs/2510.25744)Cited by:[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px2.p1.1)\.
- Shridharet al\.\(2021\)M\. Shridhar, X\. Yuan, M\. Cote, Y\. Bisk, A\. Trischler, and M\. HausknechtALFWorld: aligning text and embodied environments for interactive learning\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=0IOX0YcCdTn)Cited by:[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px2.p1.1)\.
- Singhet al\.\(2026\)A\. Singh, A\. Fry, A\. Perelman, A\. Tart, A\. Ganesh, A\. El\-Kishky, A\. McLaughlin, A\. Low, A\. Ostrow, A\. Ananthram, A\. Nathan, A\. Luo, A\. Helyar, A\. Madry, A\. Efremov, A\. Spyra, A\. Baker\-Whitcomb, A\. Beutel, A\. Karpenko, A\. Makelov, A\. Neitz, A\. Wei, A\. Barr, A\. Kirchmeyer, A\. Ivanov, A\. Christakis, A\. Gillespie, A\. Tam, A\. Bennett, A\. Wan, A\. Huang, A\. M\. Sandjideh, A\. Yang, A\. Kumar, A\. Saraiva, A\. Vallone, A\. Gheorghe, A\. G\. Garcia, A\. Braunstein, A\. Liu, A\. Schmidt, A\. Mereskin, A\. Mishchenko, A\. Applebaum, A\. Rogerson, A\. Rajan, A\. Wei, A\. Kotha, A\. Srivastava, A\. Agrawal, A\. Vijayvergiya, A\. Tyra, A\. Nair, A\. Nayak, B\. Eggers, B\. Ji, B\. Hoover, B\. Chen, B\. Chen, B\. Barak, B\. Minaiev, B\. Hao, B\. Baker, B\. Lightcap, B\. McKinzie, B\. Wang, B\. Quinn, B\. Fioca, B\. Hsu, B\. Yang, B\. Yu, B\. Zhang, B\. Brenner, C\. R\. Zetino, C\. Raymond, C\. Lugaresi, C\. Paz, C\. Hudson, C\. Whitney, C\. Li, C\. Chen, C\. Cole, C\. Voss, C\. Ding, C\. Shen, C\. Huang, C\. Colby, C\. Hallacy, C\. Koch, C\. Lu, C\. Kaplan, C\. Kim, C\. Minott\-Henriques, C\. Frey, C\. Yu, C\. Czarnecki, C\. Reid, C\. Wei, C\. Decareaux, C\. Scheau, C\. Zhang, C\. Forbes, D\. Tang, D\. Goldberg, D\. Roberts, D\. Palmie, D\. Kappler, D\. Levine, D\. Wright, D\. Leo, D\. Lin, D\. Robinson, D\. Grabb, D\. Chen, D\. Lim, D\. Salama, D\. Bhattacharjee, D\. Tsipras, D\. Li, D\. Yu, D\. Strouse, D\. Williams, D\. Hunn, E\. Bayes, E\. Arbus, E\. Akyurek, E\. Y\. Le, E\. Widmann, E\. Yani, E\. Proehl, E\. Sert, E\. Cheung, E\. Schwartz, E\. Han, E\. Jiang, E\. Mitchell, E\. Sigler, E\. Wallace, E\. Ritter, E\. Kavanaugh, E\. Mays, E\. Nikishin, F\. Li, F\. P\. Such, F\. de Avila Belbute Peres, F\. Raso, F\. Bekerman, F\. Tsimpourlas, F\. Chantzis, F\. Song, F\. Zhang, G\. Raila, G\. McGrath, G\. Briggs, G\. Yang, G\. Parascandolo, G\. Chabot, G\. Kim, G\. Zhao, G\. Valiant, G\. Leclerc, H\. Salman, H\. Wang, H\. Sheng, H\. Jiang, H\. Wang, H\. Jin, H\. Sikchi, H\. Schmidt, H\. Aspegren, H\. Chen, H\. Qiu, H\. Lightman, I\. Covert, I\. Kivlichan, I\. Silber, I\. Sohl, I\. Hammoud, I\. Clavera, I\. Lan, I\. Akkaya, I\. Kostrikov, I\. Kofman, I\. Etinger, I\. Singal, J\. Hehir, J\. Huh, J\. Pan, J\. Wilczynski, J\. Pachocki, J\. Lee, J\. Quinn, J\. Kiros, J\. Kalra, J\. Samaroo, J\. Wang, J\. Wolfe, J\. Chen, J\. Wang, J\. Harb, J\. Han, J\. Wang, J\. Zhao, J\. Chen, J\. Yang, J\. Tworek, J\. Chand, J\. Landon, J\. Liang, J\. Lin, J\. Liu, J\. Wang, J\. Tang, J\. Yin, J\. Jang, J\. Morris, J\. Flynn, J\. Ferstad, J\. Heidecke, J\. Fishbein, J\. Hallman, J\. Grant, J\. Chien, J\. Gordon, J\. Park, J\. Liss, J\. Kraaijeveld, J\. Guay, J\. Mo, J\. Lawson, J\. McGrath, J\. Vendrow, J\. Jiao, J\. Lee, J\. Steele, J\. Wang, J\. Mao, K\. Chen, K\. Hayashi, K\. Xiao, K\. Salahi, K\. Wu, K\. Sekhri, K\. Sharma, K\. Singhal, K\. Li, K\. Nguyen, K\. Gu\-Lemberg, K\. King, K\. Liu, K\. Stone, K\. Yu, K\. Ying, K\. Georgiev, K\. Lim, K\. Tirumala, K\. Miller, L\. Ahmad, L\. Lv, L\. Clare, L\. Fauconnet, L\. Itow, L\. Yang, L\. Romaniuk, L\. Anise, L\. Byron, L\. Pathak, L\. Maksin, L\. Lo, L\. Ho, L\. Jing, L\. Wu, L\. Xiong, L\. Mamitsuka, L\. Yang, L\. McCallum, L\. Held, L\. Bourgeois, L\. Engstrom, L\. Kuhn, L\. Feuvrier, L\. Zhang, L\. Switzer, L\. Kondraciuk, L\. Kaiser, M\. Joglekar, M\. Singh, M\. Shah, M\. Stratta, M\. Williams, M\. Chen, M\. Sun, M\. Cayton, M\. Li, M\. Zhang, M\. Aljubeh, M\. Nichols, M\. Haines, M\. Schwarzer, M\. Gupta, M\. Shah, M\. Y\. Guan, M\. Huang, M\. Dong, M\. Wang, M\. Glaese, M\. Carroll, M\. Lampe, M\. Malek, M\. Sharman, M\. Zhang, M\. Wang, M\. Pokrass, M\. Florian, M\. Pavlov, M\. Wang, M\. Chen, M\. Wang, M\. Feng, M\. Bavarian, M\. Lin, M\. Abdool, M\. Rohaninejad, N\. Soto, N\. Staudacher, N\. LaFontaine, N\. Marwell, N\. Liu, N\. Preston, N\. Turley, N\. Ansman, N\. Blades, N\. Pancha, N\. Mikhaylin, N\. Felix, N\. Handa, N\. Rai, N\. Keskar, N\. Brown, O\. Nachum, O\. Boiko, O\. Murk, O\. Watkins, O\. Gleeson, P\. Mishkin, P\. Lesiewicz, P\. Baltescu, P\. Belov, P\. Zhokhov, P\. Pronin, P\. Guo, P\. Thacker, Q\. Liu, Q\. Yuan, Q\. Liu, R\. Dias, R\. Puckett, R\. Arora, R\. T\. Mullapudi, R\. Gaon, R\. Miyara, R\. Song, R\. Aggarwal, R\. Marsan, R\. Yemiru, R\. Xiong, R\. Kshirsagar, R\. Nuttall, R\. Tsiupa, R\. Eldan, R\. Wang, R\. James, R\. Ziv, R\. Shu, R\. Nigmatullin, S\. Jain, S\. Talaie, S\. Altman, S\. Arnesen, S\. Toizer, S\. Toyer, S\. Miserendino, S\. Agarwal, S\. Yoo, S\. Heon, S\. Ethersmith, S\. Grove, S\. Taylor, S\. Bubeck, S\. Banesiu, S\. Amdo, S\. Zhao, S\. Wu, S\. Santurkar, S\. Zhao, S\. R\. Chaudhuri, S\. Krishnaswamy, Shuaiqi, Xia, S\. Cheng, S\. Anadkat, S\. P\. Fishman, S\. Tobin, S\. Fu, S\. Jain, S\. Mei, S\. Egoian, S\. Kim, S\. Golden, S\. Mah, S\. Lin, S\. Imm, S\. Sharpe, S\. Yadlowsky, S\. Choudhry, S\. Eum, S\. Sanjeev, T\. Khan, T\. Stramer, T\. Wang, T\. Xin, T\. Gogineni, T\. Christianson, T\. Sanders, T\. Patwardhan, T\. Degry, T\. Shadwell, T\. Fu, T\. Gao, T\. Garipov, T\. Sriskandarajah, T\. Sherbakov, T\. Korbak, T\. Kaftan, T\. Hiratsuka, T\. Wang, T\. Song, T\. Zhao, T\. Peterson, V\. Kharitonov, V\. Chernova, V\. Kosaraju, V\. Kuo, V\. Pong, V\. Verma, V\. Petrov, W\. Jiang, W\. Zhang, W\. Zhou, W\. Xie, W\. Zhan, W\. McCabe, W\. DePue, W\. Ellsworth, W\. Bain, W\. Thompson, X\. Chen, X\. Qi, X\. Xiang, X\. Shi, Y\. Dubois, Y\. Yu, Y\. Khakbaz, Y\. Wu, Y\. Qian, Y\. T\. Lee, Y\. Chen, Y\. Zhang, Y\. Xiong, Y\. Tian, Y\. Cha, Y\. Bai, Y\. Yang, Y\. Yuan, Y\. Li, Y\. Zhang, Y\. Yang, Y\. Jin, Y\. Jiang, Y\. Wang, Y\. Wang, Y\. Liu, Z\. Stubenvoll, Z\. Dou, Z\. Wu, and Z\. WangOpenAI gpt\-5 system card\.External Links:2601\.03267,[Link](https://arxiv.org/abs/2601.03267)Cited by:[§3\.2](https://arxiv.org/html/2609.21117#S3.SS2.SSS0.Px2.p1.1)\.
- Solow \(1957\)R\. M\. SolowTechnical change and the aggregate production function\.The Review of Economics and Statistics39\(3\),pp\. 312–320\.External Links:ISSN 00346535, 15309142,[Link](http://www.jstor.org/stable/1926047)Cited by:[§3\.1](https://arxiv.org/html/2609.21117#S3.SS1.p1.1)\.
- Tamkin and McCrory \(2025\)A\. Tamkin and P\. McCroryEstimating ai productivity gains from claude conversations\(Website\)External Links:[Link](https://www.anthropic.com/research/estimating-productivity-gains)Cited by:[§1](https://arxiv.org/html/2609.21117#S1.p2.1),[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px1.p1.1)\.
- Teamet al\.\(2025\)G\. Team, R\. Anil, S\. Borgeaud, J\. Alayrac, J\. Yu, R\. Soricut, J\. Schalkwyk, A\. M\. Dai, A\. Hauth, K\. Millican, D\. Silver, M\. Johnson, I\. Antonoglou, J\. Schrittwieser, A\. Glaese, J\. Chen, E\. Pitler, T\. Lillicrap, A\. Lazaridou, O\. Firat, J\. Molloy, M\. Isard, P\. R\. Barham, T\. Hennigan, B\. Lee, F\. Viola, M\. Reynolds, Y\. Xu, R\. Doherty, E\. Collins, C\. Meyer, E\. Rutherford, E\. Moreira, K\. Ayoub, M\. Goel, J\. Krawczyk, C\. Du, E\. Chi, H\. Cheng, E\. Ni, P\. Shah, P\. Kane, B\. Chan, M\. Faruqui, A\. Severyn, H\. Lin, Y\. Li, Y\. Cheng, A\. Ittycheriah, M\. Mahdieh, M\. Chen, P\. Sun, D\. Tran, S\. Bagri, B\. Lakshminarayanan, J\. Liu, A\. Orban, F\. Güra, H\. Zhou, X\. Song, A\. Boffy, H\. Ganapathy, S\. Zheng, H\. Choe, Á\. Weisz, T\. Zhu, Y\. Lu, S\. Gopal, J\. Kahn, M\. Kula, J\. Pitman, R\. Shah, E\. Taropa, M\. A\. Merey, M\. Baeuml, Z\. Chen, L\. E\. Shafey, Y\. Zhang, O\. Sercinoglu, G\. Tucker, E\. Piqueras, M\. Krikun, I\. Barr, N\. Savinov, I\. Danihelka, B\. Roelofs, A\. White, A\. Andreassen, T\. von Glehn, L\. Yagati, M\. Kazemi, L\. Gonzalez, M\. Khalman, J\. Sygnowski, A\. Frechette, C\. Smith, L\. Culp, L\. Proleev, Y\. Luan, X\. Chen, J\. Lottes, N\. Schucher, F\. Lebron, A\. Rrustemi, N\. Clay, P\. Crone, T\. Kocisky, J\. Zhao, B\. Perz, D\. Yu, H\. Howard, A\. Bloniarz, J\. W\. Rae, H\. Lu, L\. Sifre, M\. Maggioni, F\. Alcober, D\. Garrette, M\. Barnes, S\. Thakoor, J\. Austin, G\. Barth\-Maron, W\. Wong, R\. Joshi, R\. Chaabouni, D\. Fatiha, A\. Ahuja, G\. S\. Tomar, E\. Senter, M\. Chadwick, I\. Kornakov, N\. Attaluri, I\. Iturrate, R\. Liu, Y\. Li, S\. Cogan, J\. Chen, C\. Jia, C\. Gu, Q\. Zhang, J\. Grimstad, A\. J\. Hartman, X\. Garcia, T\. S\. Pillai, J\. Devlin, M\. Laskin, D\. de Las Casas, D\. Valter, C\. Tao, L\. Blanco, A\. P\. Badia, D\. Reitter, M\. Chen, J\. Brennan, C\. Rivera, S\. Brin, S\. Iqbal, G\. Surita, J\. Labanowski, A\. Rao, S\. Winkler, E\. Parisotto, Y\. Gu, K\. Olszewska, R\. Addanki, A\. Miech, A\. Louis, D\. Teplyashin, G\. Brown, E\. Catt, J\. Balaguer, J\. Xiang, P\. Wang, Z\. Ashwood, A\. Briukhov, A\. Webson, S\. Ganapathy, S\. Sanghavi, A\. Kannan, M\. Chang, A\. Stjerngren, J\. Djolonga, Y\. Sun, A\. Bapna, M\. Aitchison, P\. Pejman, H\. Michalewski, T\. Yu, C\. Wang, J\. Love, J\. Ahn, D\. Bloxwich, K\. Han, P\. Humphreys, T\. Sellam, J\. Bradbury, V\. Godbole, S\. Samangooei, B\. Damoc, A\. Kaskasoli, S\. M\. R\. Arnold, V\. Vasudevan, S\. Agrawal, J\. Riesa, D\. Lepikhin, R\. Tanburn, S\. Srinivasan, H\. Lim, S\. Hodkinson, P\. Shyam, J\. Ferret, S\. Hand, A\. Garg, T\. L\. Paine, J\. Li, Y\. Li, M\. Giang, A\. Neitz, Z\. Abbas, S\. York, M\. Reid, E\. Cole, A\. Chowdhery, D\. Das, D\. Rogozińska, V\. Nikolaev, P\. Sprechmann, Z\. Nado, L\. Zilka, F\. Prost, L\. He, M\. Monteiro, G\. Mishra, C\. Welty, J\. Newlan, D\. Jia, M\. Allamanis, C\. H\. Hu, R\. de Liedekerke, J\. Gilmer, C\. Saroufim, S\. Rijhwani, S\. Hou, D\. Shrivastava, A\. Baddepudi, A\. Goldin, A\. Ozturel, A\. Cassirer, Y\. Xu, D\. Sohn, D\. Sachan, R\. K\. Amplayo, C\. Swanson, D\. Petrova, S\. Narayan, A\. Guez, S\. Brahma, J\. Landon, M\. Patel, R\. Zhao, K\. Villela, L\. Wang, W\. Jia, M\. Rahtz, M\. Giménez, L\. Yeung, J\. Keeling, P\. Georgiev, D\. Mincu, B\. Wu, S\. Haykal, R\. Saputro, K\. Vodrahalli, J\. Qin, Z\. Cankara, A\. Sharma, N\. Fernando, W\. Hawkins, B\. Neyshabur, S\. Kim, A\. Hutter, P\. Agrawal, A\. Castro\-Ros, G\. van den Driessche, T\. Wang, F\. Yang, S\. Chang, P\. Komarek, R\. McIlroy, M\. Lučić, G\. Zhang, W\. Farhan, M\. Sharman, P\. Natsev, P\. Michel, Y\. Bansal, S\. Qiao, K\. Cao, S\. Shakeri, C\. Butterfield, J\. Chung, P\. K\. Rubenstein, S\. Agrawal, A\. Mensch, K\. Soparkar, K\. Lenc, T\. Chung, A\. Pope, L\. Maggiore, J\. Kay, P\. Jhakra, S\. Wang, J\. Maynez, M\. Phuong, T\. Tobin, A\. Tacchetti, M\. Trebacz, K\. Robinson, Y\. Katariya, S\. Riedel, P\. Bailey, K\. Xiao, N\. Ghelani, L\. Aroyo, A\. Slone, N\. Houlsby, X\. Xiong, Z\. Yang, E\. Gribovskaya, J\. Adler, M\. Wirth, L\. Lee, M\. Li, T\. Kagohara, J\. Pavagadhi, S\. Bridgers, A\. Bortsova, S\. Ghemawat, Z\. Ahmed, T\. Liu, R\. Powell, V\. Bolina, M\. Iinuma, P\. Zablotskaia, J\. Besley, D\. Chung, T\. Dozat, R\. Comanescu, X\. Si, J\. Greer, G\. Su, M\. Polacek, R\. L\. Kaufman, S\. Tokumine, H\. Hu, E\. Buchatskaya, Y\. Miao, M\. Elhawaty, A\. Siddhant, N\. Tomasev, J\. Xing, C\. Greer, H\. Miller, S\. Ashraf, A\. Roy, Z\. Zhang, A\. Ma, A\. Filos, M\. Besta, R\. Blevins, T\. Klimenko, C\. Yeh, S\. Changpinyo, J\. Mu, O\. Chang, M\. Pajarskas, C\. Muir, V\. Cohen, C\. L\. Lan, K\. Haridasan, A\. Marathe, S\. Hansen, S\. Douglas, R\. Samuel, M\. Wang, S\. Austin, C\. Lan, J\. Jiang, J\. Chiu, J\. A\. Lorenzo, L\. L\. Sjösund, S\. Cevey, Z\. Gleicher, T\. Avrahami, A\. Boral, H\. Srinivasan, V\. Selo, R\. May, K\. Aisopos, L\. Hussenot, L\. B\. Soares, K\. Baumli, M\. B\. Chang, A\. Recasens, B\. Caine, A\. Pritzel, F\. Pavetic, F\. Pardo, A\. Gergely, J\. Frye, V\. Ramasesh, D\. Horgan, K\. Badola, N\. Kassner, S\. Roy, E\. Dyer, V\. C\. Campos, A\. Tomala, Y\. Tang, D\. E\. Badawy, E\. White, B\. Mustafa, O\. Lang, A\. Jindal, S\. Vikram, Z\. Gong, S\. Caelles, R\. Hemsley, G\. Thornton, F\. Feng, W\. Stokowiec, C\. Zheng, P\. Thacker, Ç\. Ünlü, Z\. Zhang, M\. Saleh, J\. Svensson, M\. Bileschi, P\. Patil, A\. Anand, R\. Ring, K\. Tsihlas, A\. Vezer, M\. Selvi, T\. Shevlane, M\. Rodriguez, T\. Kwiatkowski, S\. Daruki, K\. Rong, A\. Dafoe, N\. FitzGerald, K\. Gu\-Lemberg, M\. Khan, L\. A\. Hendricks, M\. Pellat, V\. Feinberg, J\. Cobon\-Kerr, T\. Sainath, M\. Rauh, S\. H\. Hashemi, R\. Ives, Y\. Hasson, E\. Noland, Y\. Cao, N\. Byrd, L\. Hou, Q\. Wang, T\. Sottiaux, M\. Paganini, J\. Lespiau, A\. Moufarek, S\. Hassan, K\. Shivakumar, J\. van Amersfoort, A\. Mandhane, P\. Joshi, A\. Goyal, M\. Tung, A\. Brock, H\. Sheahan, V\. Misra, C\. Li, N\. Rakićević, M\. Dehghani, F\. Liu, S\. Mittal, J\. Oh, S\. Noury, E\. Sezener, F\. Huot, M\. Lamm, N\. D\. Cao, C\. Chen, S\. Mudgal, R\. Stella, K\. Brooks, G\. Vasudevan, C\. Liu, M\. Chain, N\. Melinkeri, A\. Cohen, V\. Wang, K\. Seymore, S\. Zubkov, R\. Goel, S\. Yue, S\. Krishnakumaran, B\. Albert, N\. Hurley, M\. Sano, A\. Mohananey, J\. Joughin, E\. Filonov, T\. Kępa, Y\. Eldawy, J\. Lim, R\. Rishi, S\. Badiezadegan, T\. Bos, J\. Chang, S\. Jain, S\. G\. S\. Padmanabhan, S\. Puttagunta, K\. Krishna, L\. Baker, N\. Kalb, V\. Bedapudi, A\. Kurzrok, S\. Lei, A\. Yu, O\. Litvin, X\. Zhou, Z\. Wu, S\. Sobell, A\. Siciliano, A\. Papir, R\. Neale, J\. Bragagnolo, T\. Toor, T\. Chen, V\. Anklin, F\. Wang, R\. Feng, M\. Gholami, K\. Ling, L\. Liu, J\. Walter, H\. Moghaddam, A\. Kishore, J\. Adamek, T\. Mercado, J\. Mallinson, S\. Wandekar, S\. Cagle, E\. Ofek, G\. Garrido, C\. Lombriser, M\. Mukha, B\. Sun, H\. R\. Mohammad, J\. Matak, Y\. Qian, V\. Peswani, P\. Janus, Q\. Yuan, L\. Schelin, O\. David, A\. Garg, Y\. He, O\. Duzhyi, A\. Älgmyr, T\. Lottaz, Q\. Li, V\. Yadav, L\. Xu, A\. Chinien, R\. Shivanna, A\. Chuklin, J\. Li, C\. Spadine, T\. Wolfe, K\. Mohamed, S\. Das, Z\. Dai, K\. He, D\. von Dincklage, S\. Upadhyay, A\. Maurya, L\. Chi, S\. Krause, K\. Salama, P\. G\. Rabinovitch, P\. K\. R\. M, A\. Selvan, M\. Dektiarev, G\. Ghiasi, E\. Guven, H\. Gupta, B\. Liu, D\. Sharma, I\. H\. Shtacher, S\. Paul, O\. Akerlund, F\. Aubet, T\. Huang, C\. Zhu, E\. Zhu, E\. Teixeira, M\. Fritze, F\. Bertolini, L\. Marinescu, M\. Bölle, D\. Paulus, K\. Gupta, T\. Latkar, M\. Chang, J\. Sanders, R\. Wilson, X\. Wu, Y\. Tan, L\. N\. Thiet, T\. Doshi, S\. Lall, S\. Mishra, W\. Chen, T\. Luong, S\. Benjamin, J\. Lee, E\. Andrejczuk, D\. Rabiej, V\. Ranjan, K\. Styrc, P\. Yin, J\. Simon, M\. R\. Harriott, M\. Bansal, A\. Robsky, G\. Bacon, D\. Greene, D\. Mirylenka, C\. Zhou, O\. Sarvana, A\. Goyal, S\. Andermatt, P\. Siegler, B\. Horn, A\. Israel, F\. Pongetti, C\. "\. Chen, M\. Selvatici, P\. Silva, K\. Wang, J\. Tolins, K\. Guu, R\. Yogev, X\. Cai, A\. Agostini, M\. Shah, H\. Nguyen, N\. Ó\. Donnaile, S\. Pereira, L\. Friso, A\. Stambler, A\. Kurzrok, C\. Kuang, Y\. Romanikhin, M\. Geller, Z\. Yan, K\. Jang, C\. Lee, W\. Fica, E\. Malmi, Q\. Tan, D\. Banica, D\. Balle, R\. Pham, Y\. Huang, D\. Avram, H\. Shi, J\. Singh, C\. Hidey, N\. Ahuja, P\. Saxena, D\. Dooley, S\. P\. Potharaju, E\. O’Neill, A\. Gokulchandran, R\. Foley, K\. Zhao, M\. Dusenberry, Y\. Liu, P\. Mehta, R\. Kotikalapudi, C\. Safranek\-Shrader, A\. Goodman, J\. Kessinger, E\. Globen, P\. Kolhar, C\. Gorgolewski, A\. Ibrahim, Y\. Song, A\. Eichenbaum, T\. Brovelli, S\. Potluri, P\. Lahoti, C\. Baetu, A\. Ghorbani, C\. Chen, A\. Crawford, S\. Pal, M\. Sridhar, P\. Gurita, A\. Mujika, I\. Petrovski, P\. Cedoz, C\. Li, S\. Chen, N\. D\. Santo, S\. Goyal, J\. Punjabi, K\. Kappaganthu, C\. Kwak, P\. LV, S\. Velury, H\. Choudhury, J\. Hall, P\. Shah, R\. Figueira, M\. Thomas, M\. Lu, T\. Zhou, C\. Kumar, T\. Jurdi, S\. Chikkerur, Y\. Ma, A\. Yu, S\. Kwak, V\. Ähdel, S\. Rajayogam, T\. Choma, F\. Liu, A\. Barua, C\. Ji, J\. H\. Park, V\. Hellendoorn, A\. Bailey, T\. Bilal, H\. Zhou, M\. Khatir, C\. Sutton, W\. Rzadkowski, F\. Macintosh, R\. Vij, K\. Shagin, P\. Medina, C\. Liang, J\. Zhou, P\. Shah, Y\. Bi, A\. Dankovics, S\. Banga, S\. Lehmann, M\. Bredesen, Z\. Lin, J\. E\. Hoffmann, J\. Lai, R\. Chung, K\. Yang, N\. Balani, A\. Bražinskas, A\. Sozanschi, M\. Hayes, H\. F\. Alcalde, P\. Makarov, W\. Chen, A\. Stella, L\. Snijders, M\. Mandl, A\. Kärrman, P\. Nowak, X\. Wu, A\. Dyck, K\. Vaidyanathan, R\. R, J\. Mallet, M\. Rudominer, E\. Johnston, S\. Mittal, A\. Udathu, J\. Christensen, V\. Verma, Z\. Irving, A\. Santucci, G\. Elsayed, E\. Davoodi, M\. Georgiev, I\. Tenney, N\. Hua, G\. Cideron, E\. Leurent, M\. Alnahlawi, I\. Georgescu, N\. Wei, I\. Zheng, D\. Scandinaro, H\. Jiang, J\. Snoek, M\. Sundararajan, X\. Wang, Z\. Ontiveros, I\. Karo, J\. Cole, V\. Rajashekhar, L\. Tumeh, E\. Ben\-David, R\. Jain, J\. Uesato, R\. Datta, O\. Bunyan, S\. Wu, J\. Zhang, P\. Stanczyk, Y\. Zhang, D\. Steiner, S\. Naskar, M\. Azzam, M\. Johnson, A\. Paszke, C\. Chiu, J\. S\. Elias, A\. Mohiuddin, F\. Muhammad, J\. Miao, A\. Lee, N\. Vieillard, J\. Park, J\. Zhang, J\. Stanway, D\. Garmon, A\. Karmarkar, Z\. Dong, J\. Lee, A\. Kumar, L\. Zhou, J\. Evens, W\. Isaac, G\. Irving, E\. Loper, M\. Fink, I\. Arkatkar, N\. Chen, I\. Shafran, I\. Petrychenko, Z\. Chen, J\. Jia, A\. Levskaya, Z\. Zhu, P\. Grabowski, Y\. Mao, A\. Magni, K\. Yao, J\. Snaider, N\. Casagrande, E\. Palmer, P\. Suganthan, A\. Castaño, I\. Giannoumis, W\. Kim, M\. Rybiński, A\. Sreevatsa, J\. Prendki, D\. Soergel, A\. Goedeckemeyer, W\. Gierke, M\. Jafari, M\. Gaba, J\. Wiesner, D\. G\. Wright, Y\. Wei, H\. Vashisht, Y\. Kulizhskaya, J\. Hoover, M\. Le, L\. Li, C\. Iwuanyanwu, L\. Liu, K\. Ramirez, A\. Khorlin, A\. Cui, T\. LIN, M\. Wu, R\. Aguilar, K\. Pallo, A\. Chakladar, G\. Perng, E\. A\. Abellan, M\. Zhang, I\. Dasgupta, N\. Kushman, I\. Penchev, A\. Repina, X\. Wu, T\. van der Weide, P\. Ponnapalli, C\. Kaplan, J\. Simsa, S\. Li, O\. Dousse, F\. Yang, J\. Piper, N\. Ie, R\. Pasumarthi, N\. Lintz, A\. Vijayakumar, D\. Andor, P\. Valenzuela, M\. Lui, C\. Paduraru, D\. Peng, K\. Lee, S\. Zhang, S\. Greene, D\. D\. Nguyen, P\. Kurylowicz, C\. Hardin, L\. Dixon, L\. Janzer, K\. Choo, Z\. Feng, B\. Zhang, A\. Singhal, D\. Du, D\. McKinnon, N\. Antropova, T\. Bolukbasi, O\. Keller, D\. Reid, D\. Finchelstein, M\. A\. Raad, R\. Crocker, P\. Hawkins, R\. Dadashi, C\. Gaffney, K\. Franko, A\. Bulanova, R\. Leblond, S\. Chung, H\. Askham, L\. C\. Cobo, K\. Xu, F\. Fischer, J\. Xu, C\. Sorokin, C\. Alberti, C\. Lin, C\. Evans, A\. Dimitriev, H\. Forbes, D\. Banarse, Z\. Tung, M\. Omernick, C\. Bishop, R\. Sterneck, R\. Jain, J\. Xia, E\. Amid, F\. Piccinno, X\. Wang, P\. Banzal, D\. J\. Mankowitz, A\. Polozov, V\. Krakovna, S\. Brown, M\. Bateni, D\. Duan, V\. Firoiu, M\. Thotakuri, T\. Natan, M\. Geist, S\. tan Girgin, H\. Li, J\. Ye, O\. Roval, R\. Tojo, M\. Kwong, J\. Lee\-Thorp, C\. Yew, D\. Sinopalnikov, S\. Ramos, J\. Mellor, A\. Sharma, K\. Wu, D\. Miller, N\. Sonnerat, D\. Vnukov, R\. Greig, J\. Beattie, E\. Caveness, L\. Bai, J\. Eisenschlos, A\. Korchemniy, T\. Tsai, M\. Jasarevic, W\. Kong, P\. Dao, Z\. Zheng, F\. Liu, F\. Yang, R\. Zhu, T\. H\. Teh, J\. Sanmiya, E\. Gladchenko, N\. Trdin, D\. Toyama, E\. Rosen, S\. Tavakkol, L\. Xue, C\. Elkind, O\. Woodman, J\. Carpenter, G\. Papamakarios, R\. Kemp, S\. Kafle, T\. Grunina, R\. Sinha, A\. Talbert, D\. Wu, D\. Owusu\-Afriyie, C\. Du, C\. Thornton, J\. Pont\-Tuset, P\. Narayana, J\. Li, S\. Fatehi, J\. Wieting, O\. Ajmeri, B\. Uria, Y\. Ko, L\. Knight, A\. Héliou, N\. Niu, S\. Gu, C\. Pang, Y\. Li, N\. Levine, A\. Stolovich, R\. Santamaria\-Fernandez, S\. Goenka, W\. Yustalim, R\. Strudel, A\. Elqursh, C\. Deck, H\. Lee, Z\. Li, K\. Levin, R\. Hoffmann, D\. Holtmann\-Rice, O\. Bachem, S\. Arora, C\. Koh, S\. H\. Yeganeh, S\. Põder, M\. Tariq, Y\. Sun, L\. Ionita, M\. Seyedhosseini, P\. Tafti, Z\. Liu, A\. Gulati, J\. Liu, X\. Ye, B\. Chrzaszcz, L\. Wang, N\. Sethi, T\. Li, B\. Brown, S\. Singh, W\. Fan, A\. Parisi, J\. Stanton, V\. Koverkathu, C\. A\. Choquette\-Choo, Y\. Li, T\. Lu, A\. Ittycheriah, P\. Shroff, M\. Varadarajan, S\. Bahargam, R\. Willoughby, D\. Gaddy, G\. Desjardins, M\. Cornero, B\. Robenek, B\. Mittal, B\. Albrecht, A\. Shenoy, F\. Moiseev, H\. Jacobsson, A\. Ghaffarkhah, M\. Rivière, A\. Walton, C\. Crepy, A\. Parrish, Z\. Zhou, C\. Farabet, C\. Radebaugh, P\. Srinivasan, C\. van der Salm, A\. Fidjeland, S\. Scellato, E\. Latorre\-Chimoto, H\. Klimczak\-Plucińska, D\. Bridson, D\. de Cesare, T\. Hudson, P\. Mendolicchio, L\. Walker, A\. Morris, M\. Mauger, A\. Guseynov, A\. Reid, S\. Odoom, L\. Loher, V\. Cotruta, M\. Yenugula, D\. Grewe, A\. Petrushkina, T\. Duerig, A\. Sanchez, S\. Yadlowsky, A\. Shen, A\. Globerson, L\. Webb, S\. Dua, D\. Li, S\. Bhupatiraju, D\. Hurt, H\. Qureshi, A\. Agarwal, T\. Shani, M\. Eyal, A\. Khare, S\. R\. Belle, L\. Wang, C\. Tekur, M\. S\. Kale, J\. Wei, R\. Sang, B\. Saeta, T\. Liechty, Y\. Sun, Y\. Zhao, S\. Lee, P\. Nayak, D\. Fritz, M\. R\. Vuyyuru, J\. Aslanides, N\. Vyas, M\. Wicke, X\. Ma, E\. Eltyshev, N\. Martin, H\. Cate, J\. Manyika, K\. Amiri, Y\. Kim, X\. Xiong, K\. Kang, F\. Luisier, N\. Tripuraneni, D\. Madras, M\. Guo, A\. Waters, O\. Wang, J\. Ainslie, J\. Baldridge, H\. Zhang, G\. Pruthi, J\. Bauer, F\. Yang, R\. Mansour, J\. Gelman, Y\. Xu, G\. Polovets, J\. Liu, H\. Cai, W\. Chen, X\. Sheng, E\. Xue, S\. Ozair, C\. Angermueller, X\. Li, A\. Sinha, W\. Wang, J\. Wiesinger, E\. Koukoumidis, Y\. Tian, A\. Iyer, M\. Gurumurthy, M\. Goldenson, P\. Shah, M\. Blake, H\. Yu, A\. Urbanowicz, J\. Palomaki, C\. Fernando, K\. Durden, H\. Mehta, N\. Momchev, E\. Rahimtoroghi, M\. Georgaki, A\. Raul, S\. Ruder, M\. Redshaw, J\. Lee, D\. Zhou, K\. Jalan, D\. Li, B\. Hechtman, P\. Schuh, M\. Nasr, K\. Milan, V\. Mikulik, J\. Franco, T\. Green, N\. Nguyen, J\. Kelley, A\. Mahendru, A\. Hu, J\. Howland, B\. Vargas, J\. Hui, K\. Bansal, V\. Rao, R\. Ghiya, E\. Wang, K\. Ye, J\. M\. Sarr, M\. M\. Preston, M\. Elish, S\. Li, A\. Kaku, J\. Gupta, I\. Pasupat, D\. Juan, M\. Someswar, T\. M\., X\. Chen, A\. Amini, A\. Fabrikant, E\. Chu, X\. Dong, A\. Muthal, S\. Buthpitiya, S\. Jauhari, N\. Hua, U\. Khandelwal, A\. Hitron, J\. Ren, L\. Rinaldi, S\. Drath, A\. Dabush, N\. Jiang, H\. Godhia, U\. Sachs, A\. Chen, Y\. Fan, H\. Taitelbaum, H\. Noga, Z\. Dai, J\. Wang, C\. Liang, J\. Hamer, C\. Ferng, C\. Elkind, A\. Atias, P\. Lee, V\. Listík, M\. Carlen, J\. van de Kerkhof, M\. Pikus, K\. Zaher, P\. Müller, S\. Zykova, R\. Stefanec, V\. Gatsko, C\. Hirnschall, A\. Sethi, X\. F\. Xu, C\. Ahuja, B\. Tsai, A\. Stefanoiu, B\. Feng, K\. Dhandhania, M\. Katyal, A\. Gupta, A\. Parulekar, D\. Pitta, J\. Zhao, V\. Bhatia, Y\. Bhavnani, O\. Alhadlaq, X\. Li, P\. Danenberg, D\. Tu, A\. Pine, V\. Filippova, A\. Ghosh, B\. Limonchik, B\. Urala, C\. K\. Lanka, D\. Clive, Y\. Sun, E\. Li, H\. Wu, K\. Hongtongsak, I\. Li, K\. Thakkar, K\. Omarov, K\. Majmundar, M\. Alverson, M\. Kucharski, M\. Patel, M\. Jain, M\. Zabelin, P\. Pelagatti, R\. Kohli, S\. Kumar, J\. Kim, S\. Sankar, V\. Shah, L\. Ramachandruni, X\. Zeng, B\. Bariach, L\. Weidinger, T\. Vu, A\. Andreev, A\. He, K\. Hui, S\. Kashem, A\. Subramanya, S\. Hsiao, D\. Hassabis, K\. Kavukcuoglu, A\. Sadovsky, Q\. Le, T\. Strohman, Y\. Wu, S\. Petrov, J\. Dean, and O\. VinyalsGemini: a family of highly capable multimodal models\.External Links:2312\.11805,[Link](https://arxiv.org/abs/2312.11805)Cited by:[§3\.2](https://arxiv.org/html/2609.21117#S3.SS2.SSS0.Px1.p1.1)\.
- Tuet al\.\(2025\)T\. Tu, M\. Schaekermann, A\. Palepu, K\. Saab, J\. Freyberg, R\. Tanno, A\. Wang, B\. Li, M\. Amin, Y\. Cheng, E\. Vedadi, N\. Tomasev, S\. Azizi, K\. Singhal, L\. Hou, A\. Webson, K\. Kulkarni, S\. S\. Mahdavi, C\. Semturs, J\. Gottweis, J\. Barral, K\. Chou, G\. S\. Corrado, Y\. Matias, A\. Karthikesalingam, and V\. NatarajanTowards conversational diagnostic artificial intelligence\.Nature642\(8067\),pp\. 442–450\.External Links:[Link](https://www.nature.com/articles/s41586-025-08866-7),[Document](https://dx.doi.org/10.1038/s41586-025-08866-7)Cited by:[§3\.1](https://arxiv.org/html/2609.21117#S3.SS1.SSS0.Px1.p1.1)\.
- Vidgenet al\.\(2025\)B\. Vidgen, A\. Fennelly, E\. Pinnix, J\. Benchek, D\. Khan, Z\. Richards, A\. Bridges, C\. Huang, K\. Sahu, A\. Kottamasu, B\. Ma, B\. Hunsberger, I\. Robinson, A\. Datta, C\. Mahapatra, D\. Barton, C\. R\. Sunstein, E\. Topol, B\. Foody, and O\. NitskiThe ai productivity index \(apex\)\.External Links:2509\.25721,[Link](https://arxiv.org/abs/2509.25721)Cited by:[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px1.p1.1)\.
- Walkeret al\.\(1997\)M\. A\. Walker, D\. J\. Litman, C\. A\. Kamm, and A\. AbellaPARADISE: a framework for evaluating spoken dialogue agents\.In35th Annual Meeting of the Association for Computational Linguistics and 8th Conference of the European Chapter of the Association for Computational Linguistics,Madrid, Spain,pp\. 271–280\.External Links:[Link](https://aclanthology.org/P97-1035/),[Document](https://dx.doi.org/10.3115/976909.979652)Cited by:[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2609.21117#S3.SS1.p2.1)\.
- Webb \(2020\)M\. WebbThe impact of artificial intelligence on the labor market\.SSRN Electronic Journal\.External Links:[Link](https://web.stanford.edu/~mww/webb_jmp.pdf),[Document](https://dx.doi.org/10.2139/ssrn.3482150)Cited by:[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px1.p1.1)\.
- Wongsuphasawatet al\.\(2019\)K\. Wongsuphasawat, Y\. Liu, and J\. HeerGoals, process, and challenges of exploratory data analysis: an interview study\.External Links:1911\.00568,[Link](https://arxiv.org/abs/1911.00568)Cited by:[§4\.2](https://arxiv.org/html/2609.21117#S4.SS2.p2.1)\.
- Wuet al\.\(2025\)S\. Wu, M\. Galley, B\. Peng, H\. Cheng, G\. Li, Y\. Dou, W\. Cai, J\. Zou, J\. Leskovec, and J\. GaoCollabLLM: from passive responders to active collaborators\.InForty\-second International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=DmH4HHVb3y)Cited by:[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px2.p1.1)\.
- Xieet al\.\(2024\)T\. Xie, D\. Zhang, J\. Chen, X\. Li, S\. Zhao, R\. Cao, T\. J\. Hua, Z\. Cheng, D\. Shin, F\. Lei, Y\. Liu, Y\. Xu, S\. Zhou, S\. Savarese, C\. Xiong, V\. Zhong, and T\. YuOSWORLD: benchmarking multimodal agents for open\-ended tasks in real computer environments\.InProceedings of the 38th International Conference on Neural Information Processing Systems,NIPS ’24,Red Hook, NY, USA\.External Links:ISBN 9798331314385Cited by:[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px2.p1.1)\.
- Xuanet al\.\(2026\)K\. Xuan, P\. Song, P\. Lu, P\. Han, W\. Li, Z\. Zhang, Z\. He, W\. Hua, M\. Li, J\. You, A\. Weller, Y\. Wang, and J\. PeiInteractive evaluation requires a design science\.External Links:2605\.17829,[Link](https://arxiv.org/abs/2605.17829)Cited by:[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px2.p1.1)\.
- Yaoet al\.\(2025\)S\. Yao, N\. Shinn, P\. Razavi, and K\. R\. Narasimhan\{$\\tau$\}\-bench: a benchmark for \\underline\{t\}ool\-\\underline\{a\}gent\-\\underline\{u\}ser interaction in real\-world domains\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=roNSXZpUDN)Cited by:[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px2.p1.1)\.
- Zhaoet al\.\(2024\)W\. Zhao, X\. Ren, J\. Hessel, C\. Cardie, Y\. Choi, and Y\. DengWildChat: 1m chatgpt interaction logs in the wild\.External Links:2405\.01470,[Link](https://arxiv.org/abs/2405.01470)Cited by:[Appendix B](https://arxiv.org/html/2609.21117#A2.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px2.p1.1),[§3\.2](https://arxiv.org/html/2609.21117#S3.SS2.SSS0.Px2.p1.1)\.
- Zhenget al\.\(2024\)L\. Zheng, W\. Chiang, Y\. Sheng, T\. Li, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Li, Z\. Lin, E\. Xing, J\. E\. Gonzalez, I\. Stoica, and H\. ZhangLMSYS\-chat\-1m: a large\-scale real\-world LLM conversation dataset\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=BOfDKxfwt0)Cited by:[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px2.p1.1)\.
- Zhenget al\.\(2023\)L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. Xing, H\. Zhang, J\. Gonzalez, and I\. StoicaJudging llm\-as\-a\-judge with mt\-bench and chatbot arena\.InAdvances in Neural Information Processing Systems,A\. Oh, T\. Naumann, A\. Globerson, K\. Saenko, M\. Hardt, and S\. Levine \(Eds\.\),Vol\.36,pp\. 46595–46623\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2023/file/91f18a1287b398d378ef22505bf41832-Paper-Datasets_and_Benchmarks.pdf)Cited by:[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px2.p1.1)\.
- Zhouet al\.\(2024\)S\. Zhou, F\. F\. Xu, H\. Zhu, X\. Zhou, R\. Lo, A\. Sridhar, X\. Cheng, T\. Ou, Y\. Bisk, D\. Fried, U\. Alon, and G\. NeubigWebArena: a realistic web environment for building autonomous agents\.External Links:2307\.13854,[Link](https://arxiv.org/abs/2307.13854)Cited by:[§2](https://arxiv.org/html/2609.21117#S2.SS0.SSS0.Px2.p1.1)\.

## Appendix ACoGym Dataset Task Details

We use three task environments from Collaborative Gym \(CoGym\): travel planning, related work writing, and tabular analysis\. CoGym is designed for human\-agent collaboration in shared task environments, where both the human and the agent can communicate and interact with task specific tools\.

#### Travel planning\.

In the travel planning task, the human and agent collaborate to produce a detailed itinerary that satisfies a user’s travel goals and constraints\. The task is challenging because the initial request may underspecify important preferences, such as preferred activities, destinations, budget constraints, accommodation needs, or transportation choices\. The agent can use travel\-related search tools, including city, attraction, restaurant, flight, and accommodation search, and can update a shared itinerary editor\. The final artifact is the completed travel plan\.

#### Related work writing\.

In the related work task, the human and agent collaborate to write a related work section for a given research topic\. The agent can search for relevant papers, add or remove papers from a shared library, transfer selected papers into a draft, and update a shared text editor\. The final artifact is the related work section\.

#### Tabular analysis\.

In the tabular analysis task, the human and agent collaborate to derive analytical insights from provided tabular data\. The environment includes shared access to the data, a Jupyter notebook for code execution, and a text editor for documenting findings\. The agent can execute notebook cells and update the written analysis\. The final artifact is an analytical report or finding supported by computations over the table\.

These three tasks provide variation in the role of interaction\. Travel planning emphasizes latent preferences and constraint satisfaction; related work writing emphasizes synthesis, domain expertise, and iterative refinement; and tabular analysis emphasizes tool use, code execution, and evidence generation\. This variation is important for our analysis because productivity should not be interpreted as a universal preference for shorter interactions\. Instead, the relationship between quality and cost depends on the task structure\.

## Appendix BVisualization Dataset Details

#### Task\.

The visualization dataset was collected from an IRB approved study\. Participants completed a structured exploratory analysis task using WildChat\-1M\([Zhao et al\., 2024](https://arxiv.org/html/2609.21117#bib.bib17)\), a large scale dataset of real world ChatGPT interactions\. The task required participants to formulate a testable hypothesis about the dataset, write or adapt Python code to analyze the data, produce a visualization, and submit a short written interpretation of their findings\. This design was intended to capture a realistic human–LLM workflow in which users must move from problem formulation to analysis, visualization, and explanation, rather than simply rate model outputs or complete isolated microtasks\.

#### Participants\.

Participants were university students\. This population is appropriate for the task because the study focuses on how users collaborate with an LLM to complete an open\-ended analytical task\. The study therefore captures collaboration in a setting where participants had a genuine task objective and were expected to produce a usable final artifact\. In total, we collected 42 completed sessions\.

#### Interface and model\.

Participants interacted with the LLM through a custom web platform designed for the study\. The interface presented the task instructions, provided access to the chat based assistant, and collected the final submission\. The assistant was powered by GPT\-5\.1\. The platform logged the full interaction trajectory, including user messages and model responses, so that each final artifact could be analyzed together with the process that produced it\.

#### Collected data\.

For each session, we collect the full chat transcript, timestamps if available, submitted code, final visualization, written interpretation, rubric score, and post task survey responses\.

#### Privacy and release\.

Because the dataset contains natural language interactions with an LLM, chat logs may include incidental personal information or sensitive content entered by participants\. Before any release, we will remove direct identifiers, anonymize participant IDs, and filter or redact potentially identifying content\.

## Appendix CTask Quality Rubrics

We evaluate the final artifact from each collaboration session using task specific grading rubrics\. Each rubric assigns an integer score from 1 to 5, where higher scores indicate higher task quality\. The judge is given the original task request and the final artifact produced by the human\-AI team\. For tabular analysis, the judge is also given excerpts from the Jupyter execution history so that the final report can be evaluated against the computations used to produce it\. The judge is instructed to provide a short justification and then conclude with a final numeric score\.

For the related work writing task, we reuse the rubric provided by Collaborative Gym in Figure 13 of[Shao et al\. \(2026\)](https://arxiv.org/html/2609.21117#bib.bib2)\. We use custom rubrics for travel planning, tabular analysis, and the visualization task, shown below\.

### C\.1Travel Planning Rubric

The travel planning rubric evaluates whether the final itinerary satisfies the user’s request, covers the requested dates, includes concrete travel entities, and is internally realistic\.

> Scoring Rubric for Travel Planning The evaluator is given: \(a\) the user’s travel\-planning request, and \(b\) the final travel plan that the human\-AI team produced, defined as the contents of the shared travel\-plan editor at the end of the session\. Evaluation Criteria Score = 1\.The plan is missing entirely, gibberish, or unrelated to the user’s request, including the origin, destination, or dates\. Score = 2\.The plan exists but only sketches one or two days\. Required slots, such as transportation, accommodation, attraction, breakfast, lunch, and dinner, are missing for most days\. Specifics such as flight numbers, restaurant names, or hotel names are absent or appear fabricated\. Score = 3\.The plan covers all requested days but is partially incomplete\. Some meals or attractions are left blank; flight numbers, accommodations, or restaurants are present but generic; or there are commonsense issues, such as destination mismatches, impossible timing, or missing intercity transportation\. Score = 4\.The plan covers every day and fills most meal, attraction, transportation, and accommodation slots with concrete entities, such as named flights, hotels, restaurants, or attractions\. The itinerary is internally consistent, with transportation between cities scheduled before the day’s activities\. Only minor gaps remain\. Score = 5\.The plan is complete and realistic for every day\. It includes named flights with flight numbers, named hotels, named restaurants for each meal, and named attractions\. It respects the user’s preferences, including budget, interests, and number of travelers; places transportation between cities correctly; and contains no obvious commonsense violations\. The evaluator provides 2–4 sentences of justification and concludes with:Therefore, the final score is N\.

### C\.2Tabular Analysis Rubric

The tabular analysis rubric evaluates whether the final report answers the user’s analytical request, reports concrete numerical findings, and supports its conclusions with appropriate computations\.

> Scoring Rubric for Tabular Data Analysis The evaluator is given: \(a\) the user’s analytical request, \(b\) the final analytical report, defined as the contents of the result editor at the end of the session, and \(c\) excerpts of the Jupyter execution history that produced the report\. The evaluator is instructed to evaluate the final report\. Evaluation Criteria Score = 1\.There is no report, the report is gibberish, or it is unrelated to the question\. Score = 2\.The report addresses the question only superficially\. It contains no specific numerical findings, no statistical justification, unsupported or generic conclusions, or conclusions that contradict the executed code\. Score = 3\.The report answers the question with at least one concrete numerical finding, but lacks rigor\. For example, it omits error bars, significance tests, or sample sizes; examines only one slice of the data; or provides conclusions that are only partially supported\. Score = 4\.The report addresses the question with multiple specific numerical findings and uses appropriate aggregations or simple statistics\. The conclusions follow from the analysis, with only minor gaps, such as missing caveats about data quality or the absence of a formal hypothesis test where one would have helped\. Score = 5\.The report directly answers the user’s question with concrete numbers, uses appropriate statistical testing where relevant, considers alternative explanations or confounders, and presents conclusions with clear caveats\. The findings are consistent with the executed code\. The evaluator provides 2–4 sentences of justification and concludes with:Therefore, the final score is N\.

### C\.3Visualization Task Rubric

The visualization rubric evaluates the quality of the submitted WildChat\-1M analysis, including the hypothesis, code, statistical analysis, visualization, and final interpretation\.

> Scoring Rubric for WildChat\-1M Analysis The evaluator is given: \(a\) the prompt and constraints, \(b\) the final code submitted by the participant, and \(c\) the final output captured at submission\. The evaluator is instructed to evaluate the quality of the submitted analysis\. Evaluation Criteria Score = 1\.No code is submitted, the submission is gibberish, or it is unrelated to the task\. This score is also assigned when the submission violates the data\-handling constraints in an obvious way, such as pasting raw conversations\. Score = 2\.The code attempts to load WildChat but does not actually compute the hypothesized analysis\. There is no plot, only a placeholder plot, no stated hypothesis, or the output is empty or trivially limited to commands such asdf\.head\(\)\. Score = 3\.The code states a clear hypothesis, runs at least one aggregation or group comparison consistent with the hypothesis, and produces some output\. However, statistical testing is missing or weak; the visualization is basic or difficult to interpret; or the conclusions are missing or hand\-waved\. Score = 4\.The hypothesis is clear and testable\. The code uses appropriate aggregations and at least one statistical test or correlation\. The visualization matches the hypothesis and includes proper labels and titles\. The output is consistent with the analysis, and the data\-handling constraints are respected\. Score = 5\.The hypothesis is well motivated and operationalized\. The analysis combines aggregations, an appropriate statistical test with a reported statistic and p\-value or effect size, and a polished, well\-labeled visualization that supports the conclusion\. The code is clean and self\-contained, constraints are fully respected, and the conclusion explicitly follows from the results with appropriate caveats\. The evaluator provides 2–4 sentences of justification and concludes with:Therefore, the final score is N\.

## Appendix DLikert Correlations

In addition to rubric scored task quality, both datasets include subjective Likert ratings collected after the task\. These ratings capture participants’ perceptions of the outcome and interaction, but they are not interchangeable with our primary quality measure\. Table[2](https://arxiv.org/html/2609.21117#A4.T2)reports pairwise Spearman correlations among the available Likert measures within each dataset\.

For CoGym, participants rated the quality of the outcome, the agent, communication with the agent, and overall satisfaction\. These measures are moderately to strongly correlated with one another, suggesting that participants’ subjective impressions of the interaction are not independent constructs\. For example, outcome quality is strongly correlated with both agent rating and communication rating\. For the visualization dataset, perceived usefulness, speed, and confidence are also positively correlated, with the strongest relationship between usefulness and speed\.

These correlations motivate our decision not to use subjective ratings as the primary quality signalQQ\. If subjective ratings were used to define outcome quality, later analyses comparing productivity against users’ perceived experience would become partially circular\. Instead, we use independently rubric scored task performance as the primary quality measure and reserve subjective ratings for comparison analyses\. This allows us to test whether users’ perceptions track the quality\-cost tradeoff captured by our productivity metric\.

Table 2:Inter\-item Spearman correlations among subjective Likert ratings\. Subjective ratings are substantially correlated within each study, motivating our decision to use independently rubric\-scored task performance as the primary quality signalQQand reserve subjective ratings for comparison analyses\.Figure 5:Interaction cost within subjective outcome\-rating levels\. For CoGym, the subjective measure is the participant’s outcome quality rating; for the visualization task, it is confidence in the submitted result, since a direct outcome quality rating was not collected\. Each point represents one collaboration session, and cost is shown on a log scale\. Wide variation within the same rating level shows that subjective outcome ratings can hide large differences in the interaction cost required to produce the final artifact\.
## Appendix ECost Variation Within Human Likert Quality Levels

Section[4\.1](https://arxiv.org/html/2609.21117#S4.SS1)shows that sessions with identical rubric scored quality can require very different amounts of interaction\. Here, we repeat the analysis using the closest available subjective outcome measure in each dataset\. For CoGym, this is the participant’s Likert rating of outcome quality\. For the visualization task, we use confidence in the submitted result as the closest subjective outcome oriented measure\.

Figure[5](https://arxiv.org/html/2609.21117#A4.F5)plots interaction cost within each human outcome rating level\. The same qualitative pattern appears: sessions with the same perceived outcome quality often differ substantially in interaction cost\. Even when users assign the same rating to the final output, the amount of interaction required to reach that output can vary widely\. This reinforces the motivation for our framework: outcome ratings alone, whether produced by a rubric judge or by the user, do not reveal whether the collaboration was productively achieved or costly to produce\. Productivity therefore requires measuring quality and cost jointly rather than treating perceived quality as a complete evaluation signal\.

## Appendix FProductivity Rankings Are Robust to Cost and Weighting Choices

We validate our productivity metric:Pz=z⁡\(Q\)−z⁡\(C\)P\_\{z\}=z\(Q\)\-z\(C\)along two dimensions\. First, we vary the relative weight assigned to agent tokens\. Second, we compare productivity rankings under several alternative cost definitions: total tokens, user tokens only, agent tokens only, total turns, and user turns\.

### F\.1Varying the Agent\-Token Weight

We define a family of weighted token costs:

Cλ=user tokens\+λ⋅agent tokens,C\_\{\\lambda\}=\\text\{user tokens\}\+\\lambda\\cdot\\text\{agent tokens\},whereλ∈\{0\.50,0\.75,1\.00,1\.50,2\.00\}\\lambda\\in\\\{0\.50,0\.75,1\.00,1\.50,2\.00\\\}\. For each value ofλ\\lambda, we recompute productivity as

Pz\(λ\)=z⁡\(Q\)−z⁡\(Cλ\),P\_\{z\}^\{\(\\lambda\)\}=z\(Q\)\-z\(C\_\{\\lambda\}\),withQQandCλC\_\{\\lambda\}standardized within task\. We then compute pairwise Spearman correlations between the resulting productivity rankings\.

Figure[6](https://arxiv.org/html/2609.21117#A6.F6)shows that productivity rankings are stable across agent\-token weights\. For CoGym, pairwise Spearman correlations range from \.77 to 1\.00; for the visualization dataset, they range from \.93 to 1\.00\. As expected, the lowest correlations occur when comparing the most different weighting choices, such asλ=\.50\\lambda=\.50versusλ=2\.00\\lambda=2\.00\. Even in these cases, rankings remain strongly correlated\. This suggests that our conclusions are not driven by the specific choice of assigning agent tokens half the weight of user tokens\.

![Refer to caption](https://arxiv.org/html/2609.21117v1/fig_lambda_heatmap_sidebyside.png)Figure 6:Productivity rankings are stable across alternative agent token weights\. Each cell reports the Spearman correlation between productivity rankings computed usingCλ=user tokens\+λ⋅agent tokensC\_\{\\lambda\}=\\text\{user tokens\}\+\\lambda\\cdot\\text\{agent tokens\}, forλ∈\{0\.50,0\.75,1\.00,1\.50,2\.00\}\\lambda\\in\\\{0\.50,0\.75,1\.00,1\.50,2\.00\\\}\. The primary setting used in the main text isλ=\.50\\lambda=\.50\. Rankings remain strongly correlated across weights for both CoGym and the visualization dataset\.
### F\.2Alternative Cost Measures

We also test whether the productivity rankings are robust to broader changes in the cost measure\. In addition toCdefaultC\_\{\\text\{default\}\}, we compute productivity using five alternative cost definitions:

- •Total tokens: user tokens \+ agent tokens\.
- •User tokens: only tokens written by the user\.
- •Agent tokens: only tokens generated by the agent\.
- •Total turns: number of user and agent turns\.
- •User turns: number of user turns\.

Figure[7](https://arxiv.org/html/2609.21117#A6.F7)reports pairwise Spearman correlations between productivity rankings under these cost definitions\. The rankings are highly stable\. In CoGym, all pairwise correlations are at least \.70, and the primary measure is especially close to total tokens \(ρ=\.99\\rho=\.99\) and agent tokens \(ρ=\.96\\rho=\.96\)\. In the visualization dataset, all pairwise correlations are at least \.91, indicating even stronger agreement across cost definitions\.

The somewhat lower correlations for user only measures in CoGym are expected because CoGym agents can take actions in shared task environments and sometimes produce long tool mediated outputs\. In these settings, considering only user messages can miss part of the reception burden imposed by the agent\. Nevertheless, the overall pattern remains stable that sessions identified as productive under the primary cost definition tend to remain productive under alternative token\- and turn\-based definitions\.

![Refer to caption](https://arxiv.org/html/2609.21117v1/fig_cost_robustness_heatmap_sidebyside.png)Figure 7:Productivity rankings hold across alternative cost definitions\. Each cell reports the Spearman correlation between productivity rankings computed with different cost measures\.CdefaultC\_\{\\text\{default\}\}is the primary weighted token measure used in the main text\. Rankings are strongly correlated across token\-based and turn\-based alternatives, suggesting that the main findings are not an artifact of a single cost operationalization\.

## Appendix GDialogue Feature Analysis: Full Results

Section[4\.4](https://arxiv.org/html/2609.21117#S4.SS4)reports the main dialogue patterns associated with productive collaboration\. Here, we provide the full feature comparison between the top and bottom productivity quartiles\. We compare sessions in the top and bottom quartiles ofPzP\_\{z\}across all tasks \(n=59n=59per quartile\)\. For each feature, we compute the mean rate per turn in each quartile, Cohen’sddfor the difference in means, and a Mann–Whitney U test\. Positive effect sizes indicate features that are more frequent in productive sessions, while negative effect sizes indicate features that are more frequent in unproductive sessions\.

The full results in Table[3](https://arxiv.org/html/2609.21117#A7.T3)support the interpretation in the main text\. Productive sessions are characterized by more agent probing and more user acknowledgments or agreements\. In contrast, unproductive sessions show more user probing, user assumption reveal, user repair, agent overresponse, and agent overspecification\. This pattern suggests that productive collaboration is not defined by the absence of friction\. Rather, productivity depends on how grounding work is distributed: productive sessions show agent side clarification followed by user confirmation, while unproductive sessions leave more repair and clarification work to the user\.

Feature \(rate per turn\)Top quartileBottom quartileCohen’sddppFriction taxonomy\([İnan et al\., 2025](https://arxiv.org/html/2609.21117#bib.bib26)\)agent probing0\.5060\.245\+0\.73\+0\.730\.0010\.001\*\*\*probing\(combined\)0\.4230\.344\+0\.36\+0\.360\.0340\.034\*user reinforcement0\.0470\.032\+0\.14\+0\.140\.7830\.783reinforcement\(combined\)0\.0600\.050\+0\.11\+0\.110\.6500\.650user reflective pause0\.0080\.004\+0\.09\+0\.090\.5800\.580agent reinforcement0\.0640\.061\+0\.02\+0\.020\.4250\.425user overspecification0\.1090\.112−0\.01\-0\.010\.1620\.162reflective pause\(combined\)0\.0090\.017−0\.18\-0\.180\.1550\.155any friction0\.7410\.800−0\.31\-0\.310\.0870\.087overspecification\(combined\)0\.2040\.270−0\.32\-0\.320\.0200\.020\*agent assumption reveal0\.0540\.112−0\.33\-0\.330\.1150\.115agent overspecification0\.3210\.459−0\.38\-0\.380\.0220\.022\*agent reflective pause0\.0070\.033−0\.38\-0\.380\.1870\.187user probing0\.2730\.425−0\.52\-0\.520\.0050\.005\*\*user assumption reveal0\.0380\.120−0\.61\-0\.610\.0000\.000\*\*\*assumption reveal\(combined\)0\.0450\.119−0\.66\-0\.660\.0010\.001\*\*Grounding act taxonomy\([Shaikh et al\., 2025](https://arxiv.org/html/2609.21117#bib.bib25)\)user acknowledgement0\.1650\.049\+0\.58\+0\.580\.0400\.040\*user agreement0\.1560\.045\+0\.58\+0\.580\.0090\.009\*\*user repeat0\.0480\.021\+0\.29\+0\.290\.9220\.922user disagreement0\.0220\.017\+0\.07\+0\.070\.3490\.349user topic switch0\.0290\.026\+0\.04\+0\.040\.3090\.309agent agreement0\.0000\.000\+0\.00\+0\.001\.0001\.000user clarification0\.0540\.067−0\.11\-0\.110\.2200\.220agent clarification0\.0250\.042−0\.12\-0\.120\.8230\.823user grounding friction0\.1410\.167−0\.15\-0\.150\.2570\.257agent display0\.0240\.044−0\.22\-0\.220\.2030\.203agent acknowledgement0\.0320\.077−0\.35\-0\.350\.0380\.038\*user repair0\.0390\.080−0\.36\-0\.360\.0100\.010\*user followup0\.4010\.508−0\.36\-0\.360\.0720\.072agent response0\.6050\.739−0\.41\-0\.410\.0400\.040\*agent overresponse0\.4690\.620−0\.46\-0\.460\.0230\.023\*Table 3:Full rate based feature comparison between top and bottomPzP\_\{z\}quartiles \(n=59 each\)\. Cohen’sddis computed on the difference in means;pp\-values from Mann–WhitneyUUtests\. Features marked “\(combined\)” aggregate user and agent turns\. Significance:\*p<0\.05p<0\.05,\*\*p<0\.01p<0\.01,\*\*\*p<0\.001p<0\.001\.### G\.1Per Task Correlations for Agent Probing

The main text highlights agent probing as the clearest dialogue feature associated with productive collaboration\. Table[4](https://arxiv.org/html/2609.21117#A7.T4)reports the Spearman correlation between agent probing rate andPzP\_\{z\}, both pooled across tasks and separately by task\.

The pooled correlation is positive and significant \(ρ=\.19,p=\.003\\rho=\.19,p=\.003\), but the relationship is task dependent\. Agent probing is most strongly associated with productivity in related work writing \(ρ=\.35\\rho=\.35\) and travel planning \(ρ=\.24\\rho=\.24\), the two tasks where clarification and iterative refinement are often useful\. In contrast, the association is near zero in tabular analysis and visualization\. This supports the broader claim that productive friction is task relative: agent probing appears most useful when the task benefits from eliciting preferences, constraints, or framing decisions, but less useful when success depends on quickly converging on an executable analysis or visualization\.

Table 4:Spearman correlations between agent probing rate \(rate agent probing\) andPzP\_\{z\}, pooled across all sessions and per\-task\. The effect is concentrated in tasks where iterative clarification helps \(related work, travel planning\) and absent in tasks where success depends more on quick convergence \(tabular analysis, visualization\)\.\*p<0\.05p<0\.05,\*\*p<0\.01p<0\.01\.

Similar Articles

The Honest Math of AI Productivity

Reddit r/ArtificialInteligence

A critical analysis of exaggerated AI productivity claims, citing rigorous studies that show modest gains (15-40%) compared to the 5-10x often claimed by vendors, and warns against uncritical adoption of such hype.

The AI productivity numbers don't match what I actually see on my team

Reddit r/artificial

The author, running a small dev team, shares mixed real-world results from using AI coding tools: they speed up boilerplate and onboarding, but produce confident wrong answers on complex problems and increase code review workload, yielding modest net gains far below the often-cited 10x improvement.