AgBench: Agentic AI Benchmarks for Personal AI Devices
Summary
AgBench introduces a benchmark suite and open artifacts for evaluating agentic AI on personal devices, comparing local, hybrid, and cloud execution across task success, latency, API cost, and data exposure using over 162 million data points. The study finds no single architecture dominates, with local execution reducing cloud cost and privacy risk but generally lowering task success.
View Cached Full Text
Cached at: 10/01/26, 09:41 AM
# AgBench: Agentic AI Benchmarks for Personal AI Devices Source: [https://arxiv.org/html/2609.38652](https://arxiv.org/html/2609.38652) ,Di WuAffiliation:Zhejiang University of Technology ,Chinaemail:[datawonder8@gmail\.com](mailto:[email protected]),Dhananjay SaikumarAffiliation:University of St Andrews ,UKemail:[ds304@st\-andrews\.ac\.uk](mailto:[email protected])andBlesson VargheseAffiliation:University of St Andrews ,UKemail:[bv6@st\-andrews\.ac\.uk](mailto:[email protected]) © none ###### Abstract\. Agentic AI systems increasingly rely on cloud\-hosted large language models for planning, tool use, and iterative execution, raising concerns about API cost and data exposure\. Advances in personal AI devices enable agents to execute locally, but limited resources on device may affect task success and performance\. Existing benchmarks are inadequate for systematically characterizing these trade\-offs across devices, workloads, and deployment architectures\. We presentAgBench, a benchmark suite and open artifacts for reproducible evaluation of agentic AI on personal devices\. UsingAgBench, we evaluate local, hybrid, and cloud execution across agentic workloads, examining task success, latency, cloud API cost, and data exposure\. Our results, drawn from over162\.07162\.07million data points, show that personal AI devices can complete many agent tasks locally, but local\-only execution generally has lower task success and longer completion times than cloud\-only execution, especially as concurrency increases\. Local\-only execution eliminates cloud model API costs and sensitive\-information exposure to cloud agents\. Hybrid execution can improve task success, but its cloud cost and data exposure depend on how agents divide work and share information\. No single architecture performs best across task success, goodput, cloud cost, and data exposure; deployment choices should reflect the intended workload and device capabilities\.AgBenchis available at[https://anonymous\.4open\.science/r/AgBench\-2777](https://anonymous.4open.science/r/AgBench-2777)\. ###### Keywords: Agentic AI, Personal AI Devices, Agent Benchmark, Performance Characterization, Deployment Architecture ## 1\.Introduction Agentic AI is an emerging area in which systems comprising one or more agents built on foundation models pursue user\-defined goals by planning and iteratively using tools, evaluating feedback, and refining actions\([Wang et al\., 2025](https://arxiv.org/html/2609.38652#bib.bib10);[Wu et al\., 2026](https://arxiv.org/html/2609.38652#bib.bib19);[Hsiao et al\., 2026](https://arxiv.org/html/2609.38652#bib.bib30);[Choi et al\., 2026](https://arxiv.org/html/2609.38652#bib.bib31)\)\. Such systems are seen in software engineering, workplace productivity tools, and personal assistance applications\([Agashe et al\., 2024](https://arxiv.org/html/2609.38652#bib.bib32);[Wang et al\., 2026b](https://arxiv.org/html/2609.38652#bib.bib33);[Zhang et al\., 2025](https://arxiv.org/html/2609.38652#bib.bib34);[Zhou et al\., 2026](https://arxiv.org/html/2609.38652#bib.bib35)\)\. Recent research found that software coding agents were adopted by up to 28\.7% of active GitHub projects, and a 15\-fold year\-over\-year increase in active agents was reported across the Microsoft 365 ecosystem in 2026\([Robbes et al\., 2026](https://arxiv.org/html/2609.38652#bib.bib20);[Microsoft, 2026](https://arxiv.org/html/2609.38652#bib.bib21)\)\. Agent systems, such as Claude Code and Codex, rely on cloud\-hosted large language models \(LLMs\) for planning, evaluating intermediate results, and refining subsequent actions\([Anthropic, 2026](https://arxiv.org/html/2609.38652#bib.bib4);[Bolin, 2026](https://arxiv.org/html/2609.38652#bib.bib3)\)\. This requires user inputs and the execution context to be transferred to cloud services, raising privacy concerns\. Additionally, repeated model calls in long\-running workflows can incur substantial API costs\([Yang et al\., 2024](https://arxiv.org/html/2609.38652#bib.bib29)\)\. These constraints, coupled with advances in hardware with dedicated AI accelerators and large unified memory, have fostered growing interest in running agents on*personal AI devices*, which are user\-controlled devices with sufficient compute and memory to run AI agents locally\([NVIDIA, 2026](https://arxiv.org/html/2609.38652#bib.bib1);[Advanced Micro Devices, 2026](https://arxiv.org/html/2609.38652#bib.bib2)\)\. Representative devices include RTX 5090\-based PCs, NVIDIA DGX Spark, and AMD Ryzen AI Halo\. Moving agent execution from the cloud to personal AI devices introduces a fundamental trade\-off\. Local execution reduces cloud API costs and local data exposure to the cloud, but the relative resource constraints limit the size of models that can be used, which in turn impacts execution performance, reduces task success and increases execution latency\. Hybrid architectures, which sit between fully local and fully cloud deployments, can distribute agent components across the device and cloud, potentially balancing competing objectives\. However, their benefits remain unclear\. Existing benchmarks are inadequate for systematically characterizing how trade\-offs vary across devices, workloads, and deployment architectures\. This leads to acentral question:*Are personal AI devices ready for agentic AI?*We address this by considering four research questions: Q1:What is the impact on task success when agents move from the cloud to personal AI devices? Q2:Does local execution slow agent workflows? Q3:Can local execution reduce cloud API costs? Q4:How much cloud data exposure can local execution reduce? Beyond these four research questions, we further examine when local, hybrid, and cloud execution offer the best overall trade\-offs\. Our study reveals four key findings\. First, task success varies across workloads and deployment architectures, with local\-only execution falling further behind cloud\-only execution at higher concurrency\. Second, local inference increases task completion time due to longer model inference, while higher concurrency brings limited goodput gains when task success declines\. Third, local\-only execution eliminates cloud model API costs, but hybrid execution does not necessarily cost less than cloud\-only execution\. Fourth, local\-only execution avoids exposing sensitive task information to cloud agents, while exposure in hybrid architectures depends on when cloud agents are involved and what information they receive\. Overall, there is no one\-size\-fits\-all deployment for agentic AI on personal devices\. They also motivate evaluating local execution on intended tasks, limiting concurrency when reliability matters, and exploring hybrid designs in which local agents lead and request targeted cloud assistance\. We make two maincontributions: \(1\) Performance characterization and empirical insights\.We systematically characterize the trade\-offs of agent execution on personal AI devices and derive practical implications for deployment and performance optimization\. \(2\) Benchmark and open artifacts\.We introduceAgBench, a benchmark suite for evaluating agentic workloads on personal AI devices, and release our implementation and execution traces to enable reproducible evaluation and research in this nascent area\. The rest of this paper is organized as follows\. Section[2](https://arxiv.org/html/2609.38652#S2)considers related work\. Section[3](https://arxiv.org/html/2609.38652#S3)presentsAgBenchbenchmark\. Section[4](https://arxiv.org/html/2609.38652#S4)highlights the results obtained from runningAgBenchacross deployment architectures and concurrency levels\. Section[5](https://arxiv.org/html/2609.38652#S5)considers the design implications\. Section[6](https://arxiv.org/html/2609.38652#S6)presents theAgBenchdataset release\. Section[7](https://arxiv.org/html/2609.38652#S7)concludes this paper\. ## 2\.Background and Related Work Agent Execution on Personal AI Devices\.Advances in small language models and on\-device AI accelerators have made local agent execution increasingly practical\. Recent work have explored on\-device agents\. Agent\-X\([Chung et al\., 2026](https://arxiv.org/html/2609.38652#bib.bib23)\)optimizes the end\-to\-end execution of on\-device agents, while PalmClaw\([Cai et al\., 2026](https://arxiv.org/html/2609.38652#bib.bib24)\)runs the agent loop, memory, and tool use natively on mobile devices\. These systems demonstrate the feasibility of local agent execution, but also identify performance limitations on resource\-constrained devices\. A complementary line of work explores hybrid device–cloud execution\. EcoAgent\([Yi et al\., 2026](https://arxiv.org/html/2609.38652#bib.bib17)\)uses cloud models for planning and executes actions on the device, whereas Hera\([Zhang et al\., 2026](https://arxiv.org/html/2609.38652#bib.bib18)\)dynamically selects between device and cloud agents for individual steps of a task\. Recent work has further considered how agent execution can be partitioned across device and cloud models to balance task success, performance, and cloud usage of agent systems\([Rainone et al\., 2026](https://arxiv.org/html/2609.38652#bib.bib25)\)\. Together, these studies show local and hybrid agent execution are viable alternatives to cloud\-native agents\. However, their benefits and limitations remain unclear\. Agent Benchmarks\.Existing benchmarks that evaluate agents fall broadly into two categories:*capability\-oriented*benchmarks that measure whether agents can complete realistic tasks, and*systems\-oriented*benchmarks that characterize the performance and resource costs of agent execution\. Capability\-oriented benchmarks cover diverse and realistic workloads\. For example, GAIA\([Mialon et al\., 2024](https://arxiv.org/html/2609.38652#bib.bib11)\)evaluates reasoning, multimodal, browsing, and tool\-use tasks; TUA\-Bench\([Chen et al\., 2026](https://arxiv.org/html/2609.38652#bib.bib12)\)spans productivity and specialized professional workflows; Terminal\-Bench\([Merrill et al\., 2026](https://arxiv.org/html/2609.38652#bib.bib13)\)focuses on technical tasks through command\-line interface; MyPCBench\([Jang et al\., 2026](https://arxiv.org/html/2609.38652#bib.bib22)\)evaluates agents in personalized desktop environments\. These benchmarks primarily measure task success\. Systems\-oriented benchmarks in contrast examine how agents execute\. For example, XPerf\([Wang et al\., 2026a](https://arxiv.org/html/2609.38652#bib.bib14)\)evaluates LLM serving performance using agent execution traces, while OSWorld\-Human\([Abhyankar et al\., 2026](https://arxiv.org/html/2609.38652#bib.bib15)\)and AgentSysBench\([Chang et al\., 2026](https://arxiv.org/html/2609.38652#bib.bib16)\)characterize end\-to\-end latency, component\-level costs, and execution bottlenecks\. However, neither line of work systematically explores the trade\-offs that arise when agent execution moves from the cloud to personal AI devices\. This gap motivatesAgBench,which combines task evaluation and execution measurements within a common benchmark framework\. ## 3\.AgBench This section presents the methodology ofAgBenchfor characterizing agent execution on personal AI devices\.We begin with an overview of the workflow, then describe the agent workloads, execution configurations, and how agents complete tasks under these configurations\. We next present the measurements and evaluation metrics, followed by the experimental procedure used in this study\. Figure 1\.Overview ofAgBench\. Agent workloads are executed under controlled combinations of personal AI devices, deployment architectures, and workload concurrency levels\.AgBenchcaptures task outcomes, execution and system behavior, cloud model usage, and data sent to cloud for characterizing task success, execution performance, cloud API cost, and data exposure\.### 3\.1\.Benchmark Overview AgBenchevaluates agent systems using a suite of 60 tasks across eight categories under specified deployment architectures, and workload concurrency configurations\. Figure[1](https://arxiv.org/html/2609.38652#S3.F1)illustrates the workflow, which consists of four steps\. Step 1\.AgBenchconstructs a suite of agent tasks spanning diverse user activities and system demands, from information retrieval, file analysis, and office productivity to multimedia processing, scientific computing, and complex terminal operations\. These tasks require different combinations of computation, memory, storage I/O, and network access\. Each task provides a human\-readable instruction, an isolated execution environment with the required resources, and a task\-specific verifier\. Step 2\.Tasks are evaluated under controlled configurations\. Each configuration specifies a device, a deployment architecture, and a concurrency level\. The hardware resources of the devices are used for local model inference and tool execution, while the architecture determines where model inference occurs and how agents coordinate\. The concurrency level sets the maximum number of tasks that can execute at the same time\. Step 3\.For each configuration, agents repeatedly invoke tools, and inspect the results until the task completes or terminates\. Tool use includes file operations, program execution, information retrieval, and data processing\. Tasks run in isolated environments while sharing the underlying CPU, GPU, memory, storage, and, where applicable, the local model service, which enables characterizing end\-to\-end agent behavior and resource contention\. Step 4\.During execution,AgBenchcaptures task outcomes, model and tool interactions, inter\-agent communication, cloud model usage and data transfer, and system resource utilization\. Task\-specific verifiers assess returned answers and changes to the task environment\. Together, these measurements characterize task success, end\-to\-end performance, cloud API cost, and data exposure, while fine\-grained traces help explain differences across workloads and execution configurations\. ### 3\.2\.Agent Workloads Workload coverage\.We construct the workload suite to cover diverse user activities and system demands, ranging from information retrieval and office productivity to multimedia processing, scientific computing, and complex terminal tasks\. These workloads require different amounts of computation, memory, storage I/O, and network access\. We select tasks from existing benchmarks\([Mialon et al\., 2024](https://arxiv.org/html/2609.38652#bib.bib11);[Chen et al\., 2026](https://arxiv.org/html/2609.38652#bib.bib12);[Merrill et al\., 2026](https://arxiv.org/html/2609.38652#bib.bib13)\)that have fixed inputs, reproducible execution environments, and programmatically verifiable outcomes, excluding those that require human interaction or subjective assessment\. Among eligible tasks, we select a suite covering diverse personal computing activities, task difficulties, and resource demands\. Task suite\.The resulting suite contains 60 tasks drawn from GAIA\([Mialon et al\., 2024](https://arxiv.org/html/2609.38652#bib.bib11)\), TUA\-Bench\([Chen et al\., 2026](https://arxiv.org/html/2609.38652#bib.bib12)\), and Terminal\-Bench 2\([Merrill et al\., 2026](https://arxiv.org/html/2609.38652#bib.bib13)\), spanning eight task categories\. Table[1](https://arxiv.org/html/2609.38652#S3.T1)summarizes the composition and characteristics of the suite\. We adapt the selected tasks toAgBench’s task format while preserving their original objectives\. The complete task list is provided in Supplementary Section A\. Table 1\.AgBenchtask categories and characteristics\.Task structure\.EachAgBenchtask comprises a human\-readable instruction specifying the task goal and requirements for the agent, an isolated execution environment, and a task\-specific verifier\. The environment provides the files, data, programs, and network access required to complete the task\. The verifier assesses the final answer, artifact, or environment state against task\-specific success criteria, which is inaccessible to the agent during execution\. ### 3\.3\.Execution Configuration Personal AI devices\.We evaluateAgBenchon two personal AI devices\. The first is a high\-end workstation with an Intel Core Ultra 9 285K CPU and an NVIDIA GeForce RTX 5090 GPU, and the second is a compact AI system with an AMD Ryzen AI Max\+ 395 processor, integrated Radeon 8060S GPU\. As summarized in Table[2](https://arxiv.org/html/2609.38652#S3.T2), the two devices represent discrete\- and integrated\-GPU designs with dedicated and unified memory, respectively\. Table 2\.Specification of the evaluated personal AI devices\.The Max\+ 395 allocates 96 GB of unified memory to the GPU, leaving approximately 30\.5 GB visible to the operating system in our configuration\. Concurrent workloads\.We evaluate workload concurrency levels of11,22,44, and88, where a concurrency level ofCCallows up toCCtasks to execute simultaneously on the same device\. Varying concurrency allows us to characterize how execution performance changes as workload intensity increases\. Deployment architectures\.We evaluate four deployment architectures that differ in where model inference occurs and how local and cloud agents coordinate, as shown in Figure[2](https://arxiv.org/html/2609.38652#S3.F2)\. Across all architectures, task environments and tool execution remain on the personal AI device\.Local\-Only \(LO\)uses a local agent for both action planning and tool invocation, here the model inference is performed locally\.Cloud\-Only \(CO\)uses a cloud agent for action planning, with model inference performed through the cloud API while tool execution remains local\.Hybrid Cloud\-Led \(HCL\)follows a delegation\-based design, where a cloud agent leads task execution and decides when to delegate actions to a local agent\. The local agent uses local model inference to complete delegated actions and returns the results to the cloud agent\.Hybrid Local\-Led \(HLL\)follows a consultation\-based design, where a local agent leads task execution and consults a cloud agent for reasoning when needed\. The cloud agent returns its response to the local agent, which continues execution locally\. Across all architectures, task environments and tool execution remain on the personal AI device\. Figure 2\.Deployment architectures evaluated byAgBench\.Deployment architectures evaluated by \\name\.Inference setup\.For local inference, we run Qwen3\.8\-27B\([Qwen Team, 2026](https://arxiv.org/html/2609.38652#bib.bib5)\)with UD\-Q6\_K\_L 6\-bit quantization\([Unsloth, 2026](https://arxiv.org/html/2609.38652#bib.bib6)\)using llama\.cpp \(commit 0b5be7e4\)\([ggml\-org, 2026](https://arxiv.org/html/2609.38652#bib.bib7)\), with CUDA 12\.8\.1 on the RTX 5090 and Vulkan on the Max\+ 395\. We enable Flash Attention\([Dao et al\., 2022](https://arxiv.org/html/2609.38652#bib.bib27)\), an 8\-bit KV cache, and multi\-token prediction \(MTP\)\([Gloeckle et al\., 2024](https://arxiv.org/html/2609.38652#bib.bib28)\)speculative decoding with up to two draft tokens, and allow up to eight concurrent model calls on each device\. For cloud inference, we use DeepSeek V4 Flash\([DeepSeek\-AI, 2026](https://arxiv.org/html/2609.38652#bib.bib26)\)with high reasoning effort\. The local and cloud models use context windows of 65,536 and 1,000,000 tokens, with maximum outputs of 8,192 and 384,000 tokens per call, respectively\. These settings are fixed throughout the evaluation\. Further configuration details are provided in Supplementary Section C, and the unit prices used to estimate cloud API costs in Supplementary Section D\. The experiments comprise 32 configurations, covering all combinations of two personal AI devices, four deployment architectures, and four task concurrency levels\. Executing the complete 60\-task suite once under each configuration produces 1,920 task executions, generating approximately162\.07162\.07million raw execution and measurement records\. ### 3\.4\.Agent Execution Agent loop\.Each task follows an iterative loop in which the agent uses its model to determine the next action, invokes tools when needed, and uses the returned results to continue execution until the task completes or terminates\. Hybrid architectures additionally allow agents to delegate or consult through the same loop\. We implement the agent loop using Pi Coding Agent v0\.84\.1\([Zechner, 2026](https://arxiv.org/html/2609.38652#bib.bib8)\)\. Task environment\.Each task runs in an isolated Docker container with its required files, data, programs, and dependencies\. Agents interact with the environment via a common tool interface for file operations, command execution, information retrieval, and data processing\. Tool access is defined by the agents’ system prompts, which remain fixed throughout the evaluation\. The system prompts are provided in Supplementary Section C\. The environment is reset before each execution\. Shared resources\.Local model inference is provided by shared model serving, while cloud inference is accessed through the cloud API\. Concurrent tasks remain isolated at the environment level but share CPU, GPU, memory, storage I/O, and the local model service\. ### 3\.5\.Measurement Primary metrics\.We evaluate agent execution along four dimensions corresponding to our research questions: task success, execution performance, cloud API cost, and cloud data exposure\. Table[3](https://arxiv.org/html/2609.38652#S3.T3)summarizes the primary metrics used in our evaluation\. Task success is determined by task\-specific verifiers, while execution performance is characterized by mean completion time on tasks successfully completed by all architectures within each device–concurrency setting, and by goodput\. For cloud API cost, we additionally distinguish costs incurred by successful and unsuccessful executions\. To assess cloud data exposure, we examine task instructions, task\-file contents, tool outputs, and inter\-agent messages sent to cloud models\. Information processed by local models or tools is excluded\. We first manually identify sensitive items in the initial task inputs\. Each item represents a distinct piece of sensitive information, such as a private email or an authentication credential\. We report cloud data exposure as the number of distinct sensitive items observed in recorded cloud inputs divided by the total number of sensitive items identified in the initial task inputs\.\. Diagnostic measurements\.AgBenchcollects fine\-grained measurements to explain performance differences as summarized in Table[3](https://arxiv.org/html/2609.38652#S3.T3)\. They capture model and tool execution times, inter\-agent interactions, model usage and serving performance, and system resource utilization\. We use these measurements to quantify execution time and identify performance bottlenecks across deployment architectures and concurrency levels\. Supplementary Section B details the execution, model usage, and resource records\. Table 3\.Evaluation metrics and diagnostic measurements\. ### 3\.6\.Evaluation Method We run all 60 tasks under each of the 32 configurations defined in Section 3\.3, resulting in 1,920 task runs in total\. Tasks are executed in the same order across configurations using a fixed random seed\. We do not limit the number of agent\-loop iterations, tool calls, delegations, or consultations\. Instead, each task has a 7,200\-second time limit shared by all participating agents\. Our main analysis is based on one complete run of the full configuration matrix, requiring approximately 440 device hours\. We further assess the stability of the results with two additional runs of the complete 60\-task suite for each architecture on the RTX 5090 atC=1C=1andC=8C=8\. Across the three runs, success counts vary by at most 6 out of 60 tasks, while the overall trends in task success and goodput remain consistent\. Detailed results are provided in Supplementary Section E\. ## 4\.Results and Discussion This section addresses Q1–Q4 by examining task success, execution performance, cloud API cost, and cloud data exposure across deployment architectures, devices, and concurrency levels\. We then examine the trade\-offs among these outcomes\. Additional results and execution measurements are in Supplementary Section F\. ### 4\.1\.Task Success Overall Task Success\.Figure[3](https://arxiv.org/html/2609.38652#S4.F3)compares the number of successful tasks across the four deployment architectures at different concurrency levels\. Overall, LO completes more than half of the 60 tasks at low concurrency, but its success drops sharply as concurrency increases, whereas CO remains comparatively stable\. On the RTX 5090, LO drops fro m 35 successful tasks at concurrency 1 to 12 at concurrency 8, while CO changes only from 50 to 48\. On the Max\+ 395, LO drops from 31 to 22, while CO increases from 47 to 51\. At concurrency 1, both hybrid architectures improve on LO: HCL and HLL complete 39 and 44 tasks on each device, respectively\. Their advantage is inconsistent at higher concurrency\. At concurrency 8, both hybrids fall below LO on the RTX 5090, while HCL falls below LO and HLL remains only marginally above it on the Max\+ 395\. Success across Tasks\.Aggregate success counts conceal differences across task categories\. Figure[3](https://arxiv.org/html/2609.38652#S4.F3)breaks down successful tasks by category for each architecture and concurrency level\. AtC=1C=1, LO completes 9 of 12 local\-file tasks and 7 of 10 retrieval tasks on both systems\. On the RTX 5090, it also matches CO in calculation \(5/8\) and system tasks \(2/2\)\. The largest gap appears in terminal tasks: LO completes 4 of 12 on the RTX 5090 and 5 of 12 on the Max\+ 395, compared with 11 and 10 for CO\. HCL and HLL narrow this gap, completing 7 and 9 terminal tasks on the RTX 5090, and 7 and 8 on the Max\+ 395, respectively\. Thus, the overall success gap varies considerably with the task mix\. Figure 3\.Task success by workload category\.Task success by workload category\.Failure Analysis\.Figure[4](https://arxiv.org/html/2609.38652#S4.F4)breaks down unsuccessful tasks by failure type across devices, architectures, and concurrency levels\. AtC=8C=8on the RTX 5090, LO records 32 context\-limit errors and 11 execution errors, while HLL records 43 context\-limit errors\. Concurrent requests compete for a shared KV\-cache pool, so pool exhaustion can cause a context error even when a request remains below its own sequence limit\. On the Max\+ 395, no context\-limit errors are recorded; instead, slower local inference contributes to timeouts for LO, HCL, and HLL \(20, 41, and 33 tasks, respectively\)\. The two hybrid architectures differ in how local context errors appear in the results\. AtC=8C=8on the RTX 5090, HCL records 44 verification failures but no context\-limit errors\. In all 44 cases, its local component encounters a context error, but the cloud agent subsequently returns a final response that fails verification\. Thus, HCL’s zero recorded context\-limit errors do not indicate that its local component avoids context exhaustion\. CO records no context\-limit errors, and most of its failures are verification failures\. Local context errors can also lead to timeouts in HCL\. The local agent does not automatically handle some context errors from the local model service, such as by compacting its message history, and retains its prior history\. The cloud agent typically receives a generic execution\-failure report and may continue delegating work to the same local agent until the task deadline\. For example, atC=8C=8on RTX 5090, ten HCL tasks encounter repeated context errors, yet the cloud agent issues over 100 delegations per task without resolving them\. All ten time out at the two\-hour deadline\. Figure 4\.Failure analysis\.Failure analysis\.Observation on Q1:At low concurrency, LO achieves lower overall task success than CO, but performs comparably on particular task categories\. The gap becomes significant as concurrency increases, driven by context errors and timeouts\. Hybrid architectures narrow the gap at low concurrency, but their gains do not consistently persist at higher concurrency\. ### 4\.2\.Execution Performance Execution Time\.Table[4](https://arxiv.org/html/2609.38652#S4.T4)reports mean completion time for tasks successfully completed by all four architectures\. On both devices, CO completes these tasks substantially faster than LO and the two hybrid architectures\. AtC=1C=1on the RTX 5090, CO completes these tasks in 2\.07 minutes, compared to 7\.01 minutes for LO, 6\.75 minutes for HLL, and 15\.80 minutes for HCL\. On Max\+ 395, the corresponding times are 1\.95, 15\.15, 17\.28, and 34\.14 minutes\. Notably, HCL takes longer than LO on both devices\. In HCL, the cloud agent delegates additional work, including extensive preliminary checks, to the local agent and waits for it to finish\. For example, on Max\+ 395, the local agent spends 93\.5 minutes on preliminary checks for an MP3 metadata\-editing task before making any changes; the task then times out\. HLL, by contrast, has a mean completion time closer to LO on the RTX 5090\. Cloud assistance therefore does not necessarily shorten completion time when execution still depends on the local agent\. Table 4\.Mean completion time \(min\) on tasks completed successfully by all four architectures within each device–concurrency setting\. All four entries in a column use the samenntasks\.Table 5\.Elapsed wall\-clock time \(h\) for executing the full 60\-task suite\.Goodput\.Figure[5](https://arxiv.org/html/2609.38652#S4.F5)shows goodput, measured as successful tasks per hour of suite execution\. Table[5](https://arxiv.org/html/2609.38652#S4.T5)reports the elapsed wall\-clock time for executing the complete 60\-task suite\. FromC=1C=1toC=8C=8on the RTX 5090, LO’s suite execution time falls from 16\.66 to 2\.62 hours, a 6\.35\-fold reduction, but its goodput rises from 2\.1 to 4\.6 tasks/h, only a 2\.18\-fold increase\. For CO, suite execution time falls from 4\.12 to 2\.49 hours, a 1\.65\-fold reduction, while goodput rises from 12\.1 to 19\.3 tasks/h, a 1\.59\-fold increase\. On Max\+ 395, LO and CO achieve respective execution\-time reductions of 3\.53\-fold and 3\.45\-fold, while their goodput increases 2\.50\-fold and 3\.74\-fold\. Thus, although all architectures have shorter elapsed wall\-clock times atC=8C=8than atC=1C=1, only CO achieves roughly proportional goodput gains\. For architectures involving local agents, higher concurrency reduces the number of successfully completed tasks, limiting the benefit of shorter execution times\. Figure 5\.Goodput \(no\. of successfully completed tasks / total run time\)\.Goodput \(no\. of successfully completed tasks / total run time\)\.Execution Time Breakdown\.Figure[6](https://arxiv.org/html/2609.38652#S4.F6)shows the breakdown of mean task execution time\. LO, HCL, and HLL spend most of their task execution time waiting for responses from the local model\. For HCL on Max\+ 395, the mean time spent waiting for local model responses per task increases from 53\.9 minutes atC=1C=1to 96\.9 minutes atC=8C=8\. Among HCL tasks that time out atC=8C=8, waiting for local model responses accounts for approximately 99% of execution time, leaving little time for tool execution\. In contrast, local tool execution can become the main source of delay for CO, particularly at higher concurrency\. AtC=8C=8on RTX 5090, CO spends an average of 7\.4 minutes per task executing or waiting for tools, compared with 2\.2 minutes waiting for cloud model responses\. Long\-running tool calls can also become a bottleneck for task completion, consuming much of the execution budget without producing a successful outcome\. For example, atC=8C=8on RTX 5090, three CO tasks each spend more than 110 minutes executing or waiting for tools and ultimately time out at the two\-hour deadline\. Figure 6\.Breakdown of mean execution time per task\.Breakdown of mean execution time per task\.Observation on Q2:Local inference is the central performance constraint: it makes LO much slower than CO, and cloud assistance does little to offset it—HLL takes about as long as LO on the RTX 5090, while HCL is slower on both devices\. Increasing concurrency improves goodput, but declining task success limits the gains for architectures that rely on local agents\. CO, whose runtime is dominated by tool execution, avoids this trade\-off\. ### 4\.3\.Cloud API Cost Figure[7](https://arxiv.org/html/2609.38652#S4.F7)breaks down cloud API costs into cached\-input, uncached\-input, and output\-token charges\. The hybrid architectures do not always cost less than CO, even though local agents perform part of the work\. There are two reasons\. First, cloud agents still need the task context and the results of local work to decide what to do next, so delegating work does not necessarily reduce cloud input tokens in proportion to the work delegated\. Second, API costs depend on the types of tokens used: uncached input and generated output cost more than cached input\. AtC=1C=1on the RTX 5090, for example, HLL uses fewer cloud input tokens than CO \(Table[6](https://arxiv.org/html/2609.38652#S4.T6)\)\. Yet 8\.8% of HLL’s input is uncached, compared with 2\.1% for CO, and HLL generates 30\.2% more output tokens\. The higher uncached\-input and output charges exceed its savings on cached input, making HLL more expensive overall\. HCL highlights another source of cost: additional cloud calls during coordination\. On the RTX 5090, its recorded cloud calls rise from 770 atC=1C=1to 3,206 atC=8C=8\. Input charges account for 89\.5% of the associated cost increase, suggesting that the extra calls add substantial input\-processing costs\. Figure 7\.Cloud API cost breakdown\.Cloud API cost breakdown\.Table 6\.Recorded cloud input tokens \(in millions, rounded to the nearest integer\), including repeated context\.Observation on Q3:LO eliminates cloud\-model API charges\. Hybrid execution does not guarantee lower cloud API cost, because cloud agents still need relevant context, while higher uncached\-input and output charges, together with additional coordination and recovery attempts, can offset potential savings\. ### 4\.4\.Cloud Data Exposure To measure cloud data exposure, we examine whether sensitive information in the task suite appears in the inputs sent to cloud agents\. We first inspect the task instructions and provided files and identify 527 distinct sensitive items, including personal identifiers, credentials, personal records, and confidential business information\. We then use GPT\-5\.6 Sol\([OpenAI, n\.d\.](https://arxiv.org/html/2609.38652#bib.bib9)\)with high reasoning effort to locate these items in the recorded cloud\-agent inputs and manually verify the matches\. We count each sensitive item at most once per evaluation configuration, regardless of how often it appears in cloud\-agent inputs\. Supplementary Section G describes the sensitivity definitions and calculation procedure\. Table[7](https://arxiv.org/html/2609.38652#S4.T7)shows the percentage of identified sensitive items that appear in cloud\-agent inputs under each evaluation configuration\. LO exposes none of these items, while CO exposes 76% across all configurations\. AtC=1C=1, HLL exposes 54 and 63 percentage points fewer items than CO on the RTX 5090 and Max\+ 395, respectively\. HCL reduces exposure by only 5 and 9 percentage points\. This difference reflects when cloud agents enter the workflow\. In HCL, the cloud agent directs the task and receives task files and reports from the local agent\. In HLL, the local agent works first and can complete some tasks without calling a cloud agent\. FromC=1C=1toC=8C=8, HLL’s exposure drops from 22% to 0% on the RTX 5090 and from 13% to 2% on Max\+ 395, while HCL’s exposure remains high\. This decline may partly reflect lower task success: some tasks fail before sharing sensitive information with the cloud\. Table 7\.Observed sensitive\-information exposure to cloud agents \(%\)\. Percentages are calculated relative to the 527 sensitive items identified in the initial inputs of the task suite\.Observation on Q4:Local agents reduce sensitive\-information exposure most when they can complete tasks before involving the cloud\. If a cloud agent directs the work, it may need access to task files and local execution results, limiting the reduction\. ### 4\.5\.Cross\-Metric Trade\-offs Figure[8](https://arxiv.org/html/2609.38652#S4.F8)compares task success rate with goodput, cloud API cost, and sensitive\-information exposure across evaluation configurations\. CO achieves higher task success and goodput than LO, while LO incurs no cloud model API cost and exposes no sensitive items to cloud agents\. HLL can recover some of LO’s lost task success while keeping cloud cost and exposure below CO’s, but it does not match CO’s goodput\. AtC=1C=1on Max\+ 395, for example, HLL approaches CO’s task success rate and reduces both cloud API cost and exposure, yet completes far fewer successful tasks per hour\. Increasing concurrency can raise goodput, but for architectures that use local agents, it can also reduce task success\. Thus, CO is preferable when completing tasks quickly and reliably matters most, whereas LO or HLL may be preferable when limiting cloud cost and sensitive\-information exposure matters more\. Figure 8\.Cross metric tradeoffs comparing task success rate against goodput, cloud API cost and sensitive\-information exposure\.Cross metric tradeoffs comparing task success rate against goodput, cloud API cost and sensitive\-information exposure\.Overall Takeaway:Moving agents from the cloud to personal AI devices eliminates cloud API costs and sensitive\-information exposure in local\-only execution, but reduces task success and lengthens completion time\. Cloud assistance can recover some task success and reduce exposure relative to cloud\-only execution, yet it offers little improvement in completion time and does not close the goodput gap, especially at higher concurrency\. ## 5\.Design Implications The results point to three practical considerations for deploying agents on personal AI devices\. Local Deployment Should Be Evaluated on the Intended Tasks\.LO has lower overall task success and goodput than CO, but the success gap varies across task categories\. At low concurrency, LO approaches CO’s success rate in some categories and completes some tasks that CO fails\. Aggregate results therefore cannot tell whether local execution will work well for a particular application\. Before deployment, users should test local execution on the tasks they expect to run and check whether it meets their requirements for task success and completion time\. If it does, they can avoid cloud model API costs and keep task information off cloud models\. Concurrency Should Be Limited on Personal AI Devices\.On the devices we evaluated, increasing concurrency improves goodput for LO and the hybrids, but it also reduces task success, particularly at higher concurrency levels\. For applications that prioritize reliable completion, running one agent task at a time is therefore a sensible default on personal AI devices\. Higher concurrency should be used only after testing whether its goodput gains justify the drop in task success on the target device and workload\. Local\-Led Hybrid Execution May Offer a Better Balance\.At low concurrency, both hybrid architectures complete more tasks than LO\. HLL has similar or slightly higher task success than HCL, completes tasks faster, and exposes much less sensitive information\. The key difference is who leads the task: in HLL, the local agent works first and asks the cloud agent for help when needed; in HCL, the cloud agent directs the local agent and receives task files and progress reports\. Hybrid systems may therefore benefit from letting the local agent lead and sending only the information needed when it asks for help\. Future systems could request cloud help when local progress stalls, errors recur, or resources become constrained, sending only the information needed for that step\. They should then assess whether this improves task success without substantially increasing completion time, cloud API cost, or data exposure\. ## 6\.Benchmark Artifacts and Execution Traces We release the benchmark artifacts and execution traces to support reproducible evaluation and further research on agent systems running on personal AI devices\. Released Artifacts\.We release theAgBenchtask suite, execution framework, traces, and system measurements at[https://anonymous\.4open\.science/r/AgBench\-2777/](https://anonymous.4open.science/r/AgBench-2777/)\. The task suite provides task instructions, environment configurations, and task\-specific verifiers\. The framework implements the four agent architectures and includes model\-serving configurations and code for collecting traces, model usage statistics, and resource measurements\. Execution Trace Dataset\.The execution trace dataset contains records from evaluations of four agent architectures on two personal AI devices at concurrency levels of 1, 2, 4, and 8\. It records model messages, tool calls and results, inter\-agent communication, timestamps and model usage, capturing each task’s execution sequence\. Records are organized by run and task execution\. Run configurations identify the device, agent architecture, and concurrency level, while task\-execution and agent identifiers associate recorded interactions with the corresponding tasks and agents\. Verification results and execution measurements are included\. Timestamped system and local model serving measurements record device\- and container\-level resource usage and can be aligned with the execution traces\. The released dataset totals approximately 115 GB uncompressed and is distributed as 9\.0 GB of compressed archives\. An accompanying README documents the record fields, measurement units, and identifiers used to link the records\. ## 7\.Conclusion This paper presentsAgBenchfor characterizing agent execution on personal AI devices using a suite of 60 tasks covering eight categories\. We consider four deployment architectures, namely the local\-only, cloud\-only, hybrid cloud\-led, and hybrid local\-led, on two personal AI devices with different hardware accelerators under different concurrency levels\. The devices can complete a large subset of the tasks locally, but broader task coverage and effective concurrent execution remain challenging\. At low concurrency, local execution has near similar success as on the cloud in some categories, but the overall success gap widens at higher concurrency\. Local inference generally increases execution time, and higher goodput does not always preserve task success\. Fully local execution avoids cloud API charges and sensitive information transfer to the cloud\. Hybrid execution can improve task success, but does not guarantee lower API cost relative to cloud\-only inference\. These findings suggest assessing readiness of personal AI devices against the requirements of the intended usage scenario\. Local execution may suffice where it meets required task success and completion time, while other scenarios may benefit from cloud assistance\. Achieving this balance requires deployment choices informed by task requirements, resource\-aware concurrency management, and hybrid coordination that adapts work allocation and information exchange to execution progress\. ## References - Abhyankaret al\.\(2026\)R\. Abhyankar, Q\. Qi, and Y\. ZhangOSWorld\-Human: benchmarking the efficiency of computer\-use agents\.Proceedings of Machine Learning and Systems8,pp\. 482–494\.Cited by:[§2](https://arxiv.org/html/2609.38652#S2.p6.1)\. - Advanced Micro Devices \(2026\)Advanced Micro DevicesAgent computers\. powering the future of agentic AI\.Note:Official product webpageAccessed: 2026\-09\-10External Links:[Link](https://www.amd.com/en/products/processors/consumer/agent-computers.html)Cited by:[§1](https://arxiv.org/html/2609.38652#S1.p2.1)\. - Agasheet al\.\(2024\)S\. Agashe, J\. Han, S\. Gan, J\. Yang, A\. Li, and X\. E\. WangAgent s: an open agentic framework that uses computers like a human\.External Links:2410\.08164,[Link](https://arxiv.org/abs/2410.08164)Cited by:[§1](https://arxiv.org/html/2609.38652#S1.p1.1)\. - Anthropic \(2026\)AnthropicHow Claude Code works\.Note:Claude Code DocumentationAccessed: 2026\-09\-18External Links:[Link](https://code.claude.com/docs/en/how-claude-code-works)Cited by:[§1](https://arxiv.org/html/2609.38652#S1.p2.1)\. - Bolin \(2026\)M\. BolinUnrolling the Codex agent loop\.Note:OpenAI EngineeringAccessed: 2026\-09\-18External Links:[Link](https://openai.com/index/unrolling-the-codex-agent-loop/)Cited by:[§1](https://arxiv.org/html/2609.38652#S1.p2.1)\. - Caiet al\.\(2026\)H\. Cai, Y\. Li, R\. Wei, and W\. LiPalmClaw: a native on\-device agent framework for mobile phones\.External Links:2607\.13027Cited by:[§2](https://arxiv.org/html/2609.38652#S2.p1.1)\. - Changet al\.\(2026\)C\. Chang, Y\. Zhou, K\. Fu, D\. An, T\. Feng, H\. Lu, S\. Yao, P\. Guo, Y\. Yu, Y\. Shan, B\. Li, B\. Yuan, and W\. WangFrom llm inference to agentic workloads: characterization and implications for serving systems\.External Links:2608\.15127,[Link](https://arxiv.org/abs/2608.15127)Cited by:[§2](https://arxiv.org/html/2609.38652#S2.p6.1)\. - Chenet al\.\(2026\)S\. Chen, L\. Wang, X\. Yang, Z\. Liu, Y\. Cong, Y\. Ji, F\. Zhou, X\. Zhang, F\. Yang, and B\. ZengTUA\-Bench: a benchmark for general\-purpose terminal\-use agents\.External Links:2606\.28480,[Link](https://arxiv.org/abs/2606.28480)Cited by:[§2](https://arxiv.org/html/2609.38652#S2.p5.1),[§3\.2](https://arxiv.org/html/2609.38652#S3.SS2.p1.1),[§3\.2](https://arxiv.org/html/2609.38652#S3.SS2.p2.1)\. - Choiet al\.\(2026\)J\. Choi, H\. Kim, H\. Ong, Y\. Yoon, M\. Jang, D\. Kim, and J\. KimReAcTree: hierarchical llm agent trees with control flow for long\-horizon task planning\.InProceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems,AAMAS ’26,Richland, SC,pp\. 319–328\.External Links:ISBN 9798400723179,[Link](https://doi.org/10.65109/UCGT7089),[Document](https://dx.doi.org/10.65109/UCGT7089)Cited by:[§1](https://arxiv.org/html/2609.38652#S1.p1.1)\. - Chunget al\.\(2026\)J\. Chung, B\. Shin, J\. Kim, and M\. RhuAgent\-x: full pipeline acceleration of on\-device ai agents\.InProceedings of the 24th Annual International Conference on Mobile Systems, Applications and Services,MobiSys ’26,New York, NY, USA,pp\. 144–157\.External Links:ISBN 9798400720277,[Link](https://doi.org/10.1145/3745756.3809195),[Document](https://dx.doi.org/10.1145/3745756.3809195)Cited by:[§2](https://arxiv.org/html/2609.38652#S2.p1.1)\. - Daoet al\.\(2022\)T\. Dao, D\. Fu, S\. Ermon, A\. Rudra, and C\. RéFlashAttention: fast and memory\-efficient exact attention with io\-awareness\.Advances in Neural Information Processing Systems35,pp\. 16344–16359\.Cited by:[§3\.3](https://arxiv.org/html/2609.38652#S3.SS3.p4.1)\. - DeepSeek\-AI \(2026\)DeepSeek\-AIDeepSeek\-V4: towards highly efficient million\-token context intelligence\.External Links:2606\.19348,[Link](https://arxiv.org/abs/2606.19348)Cited by:[§3\.3](https://arxiv.org/html/2609.38652#S3.SS3.p4.1)\. - ggml\-org \(2026\)ggml\-orgllama\.cpp\.Note:GitHub repositoryRevision 0b5be7e4\. Accessed: 2026\-09\-18External Links:[Link](https://github.com/ggml-org/llama.cpp/tree/0b5be7e4a25862bc2777d0c47eae18788a8c963a)Cited by:[§3\.3](https://arxiv.org/html/2609.38652#S3.SS3.p4.1)\. - Gloeckleet al\.\(2024\)F\. Gloeckle, B\. Y\. Idrissi, B\. Rozière, D\. Lopez\-Paz, and G\. SynnaeveBetter & faster large language models via multi\-token prediction\.InProceedings of the 41st International Conference on Machine Learning,ICML’24,Vienna, Austria\.Cited by:[§3\.3](https://arxiv.org/html/2609.38652#S3.SS3.p4.1)\. - Hsiaoet al\.\(2026\)V\. Hsiao, M\. Roberts, and L\. SmithProcedural knowledge improves agentic LLM workflows\.InProceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems,AAMAS ’26,Richland, SC,pp\. 1425–1433\.External Links:ISBN 9798400723179,[Link](https://doi.org/10.65109/VEVZ5917),[Document](https://dx.doi.org/10.65109/VEVZ5917)Cited by:[§1](https://arxiv.org/html/2609.38652#S1.p1.1)\. - Janget al\.\(2026\)L\. K\. Jang, A\. K\. Jang, J\. Y\. Koh, and R\. SalakhutdinovMyPCBench: a benchmark for personally intelligent computer\-use agents\.External Links:2606\.16748,[Link](https://arxiv.org/abs/2606.16748)Cited by:[§2](https://arxiv.org/html/2609.38652#S2.p5.1)\. - Merrillet al\.\(2026\)M\. Merrill, A\. Shaw, N\. Carlini, B\. Li, H\. Raj, I\. Bercovich, L\. Shi, J\. Shin, T\. Walshe, E\. K\. Buchanan,et al\.Terminal\-Bench: benchmarking agents on hard, realistic tasks in command line interfaces\.External Links:2601\.11868,[Link](https://arxiv.org/abs/2601.11868)Cited by:[§2](https://arxiv.org/html/2609.38652#S2.p5.1),[§3\.2](https://arxiv.org/html/2609.38652#S3.SS2.p1.1),[§3\.2](https://arxiv.org/html/2609.38652#S3.SS2.p2.1)\. - Mialonet al\.\(2024\)G\. Mialon, C\. Fourrier, T\. Wolf, Y\. LeCun, and T\. ScialomGAIA: a benchmark for general ai assistants\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 9025–9049\.Cited by:[§2](https://arxiv.org/html/2609.38652#S2.p5.1),[§3\.2](https://arxiv.org/html/2609.38652#S3.SS2.p1.1),[§3\.2](https://arxiv.org/html/2609.38652#S3.SS2.p2.1)\. - Microsoft \(2026\)Microsoft2026 Work Trend Index Annual Report: agents, human agency, and the opportunity for every organization\.Note:[https://www\.microsoft\.com/en\-us/worklab/work\-trend\-index/agents\-human\-agency\-and\-the\-opportunity\-for\-every\-organization](https://www.microsoft.com/en-us/worklab/work-trend-index/agents-human-agency-and-the-opportunity-for-every-organization)Accessed: 2026\-09\-12Cited by:[§1](https://arxiv.org/html/2609.38652#S1.p1.1)\. - NVIDIA \(2026\)NVIDIANVIDIA DGX Spark\.Note:Official product webpageAccessed: 2026\-09\-10External Links:[Link](https://www.nvidia.com/en-us/products/workstations/dgx-spark/)Cited by:[§1](https://arxiv.org/html/2609.38652#S1.p2.1)\. - OpenAI \(n\.d\.\)OpenAIGPT\-5\.6 Sol\.Note:OpenAI API DocumentationAccessed: 2026\-09\-21External Links:[Link](https://developers.openai.com/api/docs/models/gpt-5.6-sol)Cited by:[§4\.4](https://arxiv.org/html/2609.38652#S4.SS4.p1.1)\. - Qwen Team \(2026\)Qwen TeamQwen3\.8\-27B\.Note:Hugging Face model cardAccessed: 2026\-09\-18External Links:[Link](https://huggingface.co/Qwen/Qwen3.8-27B)Cited by:[§3\.3](https://arxiv.org/html/2609.38652#S3.SS3.p4.1)\. - Rainoneet al\.\(2026\)C\. Rainone, D\. Belli, B\. Major, and A\. BehboodiWhen cloud agents meet device agents: lessons from hybrid multi\-agent systems\.\.External Links:2605\.30102Cited by:[§2](https://arxiv.org/html/2609.38652#S2.p2.1)\. - Robbeset al\.\(2026\)R\. Robbes, T\. Matricon, T\. Degueule, A\. Hora, and S\. ZacchiroliAgentic much? adoption of coding agents on github\.External Links:2601\.18341,[Document](https://dx.doi.org/10.48550/arXiv.2601.18341),[Link](https://arxiv.org/abs/2601.18341)Cited by:[§1](https://arxiv.org/html/2609.38652#S1.p1.1)\. - Unsloth \(2026\)UnslothQwen3\.8\-27B\-GGUF\.Note:Hugging Face model repositoryQuantization:UD\-Q6\_K\_L\. Accessed: 2026\-09\-18External Links:[Link](https://huggingface.co/unsloth/Qwen3.8-27B-GGUF/blob/4ca720788d1e01f1bff70c033e0d0028fd02e502/Qwen3.8-27B-UD-Q6_K_L.gguf)Cited by:[§3\.3](https://arxiv.org/html/2609.38652#S3.SS3.p4.1)\. - Wanget al\.\(2026a\)M\. Wang, Y\. Yue, S\. Li, Y\. E\. Zhou, C\. Wang, and J\. HuangBenchmarking llm serving systems for agentic ai workloads with XPerf\.External Links:2608\.20370,[Link](https://arxiv.org/abs/2608.20370)Cited by:[§2](https://arxiv.org/html/2609.38652#S2.p6.1)\. - Wanget al\.\(2025\)X\. Wang, B\. Li, Y\. Song, F\. F\. Xu, X\. Tang, M\. Zhuge, J\. Pan, Y\. Song, B\. Li, J\. Singh,et al\.OpenHands: an open platform for ai software developers as generalist agents\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 65882–65919\.Cited by:[§1](https://arxiv.org/html/2609.38652#S1.p1.1)\. - Wanget al\.\(2026b\)X\. Wang, S\. Rosenberg, J\. Michelini, C\. Smith, H\. Tran, E\. Nyst, R\. Malhotra, X\. Zhou, V\. Chen, R\. Brennan, and G\. NeubigThe openhands software agent sdk: a composable and extensible foundation for production agents\.External Links:2511\.03690,[Link](https://arxiv.org/abs/2511.03690)Cited by:[§1](https://arxiv.org/html/2609.38652#S1.p1.1)\. - Wuet al\.\(2026\)D\. Wu, J\. Luo, Y\. Han, W\. Shi, and B\. VargheseAgentic edge ai\.IEEE Internet Computing\.External Links:[Document](https://dx.doi.org/10.1109/MIC.2026.3724303)Cited by:[§1](https://arxiv.org/html/2609.38652#S1.p1.1)\. - Yanget al\.\(2024\)J\. Yang, C\. Jimenez, A\. Wettig, K\. Lieret, S\. Yao, K\. Narasimhan, and O\. PressSWE\-agent: agent\-computer interfaces enable automated software engineering\.Advances in Neural Information Processing Systems37,pp\. 50528–50652\.Cited by:[§1](https://arxiv.org/html/2609.38652#S1.p2.1)\. - Yiet al\.\(2026\)B\. Yi, X\. Hu, Y\. Chen, S\. Zhang, H\. Yang, and F\. WuEcoAgent: an efficient device\-cloud collaborative multi\-agent framework for mobile automation\.InProceedings of the AAAI Conference on Artificial IntelligenceProceedings of the Fortieth AAAI Conference on Artificial Intelligence and Thirty\-Eighth Conference on Innovative Applications of Artificial Intelligence and Sixteenth Symposium on Educational Advances in Artificial Intelligence,AAAI’26/IAAI’26/EAAI’26\.External Links:ISBN 978\-1\-57735\-906\-7,[Link](https://doi.org/10.1609/aaai.v40i35.40230),[Document](https://dx.doi.org/10.1609/aaai.v40i35.40230)Cited by:[§2](https://arxiv.org/html/2609.38652#S2.p2.1)\. - Zechner \(2026\)M\. ZechnerPi Coding Agent\.Note:GitHub software releaseVersion 0\.84\.1\. Accessed: 2026\-09\-18External Links:[Link](https://github.com/earendil-works/pi/releases/tag/v0.84.1)Cited by:[§3\.4](https://arxiv.org/html/2609.38652#S3.SS4.p1.1)\. - Zhanget al\.\(2025\)C\. Zhang, H\. Huang, C\. Ni, J\. Mu, S\. Qin, S\. He, L\. Wang, F\. Yang, P\. Zhao, C\. Du, L\. Li, Y\. Kang, Z\. Jiang, S\. Zheng, R\. Wang, J\. Qian, M\. Ma, J\. Lou, Q\. Lin, S\. Rajmohan, and D\. ZhangUFO2: the desktop agentos\.External Links:2504\.14603,[Link](https://arxiv.org/abs/2504.14603)Cited by:[§1](https://arxiv.org/html/2609.38652#S1.p1.1)\. - Zhanget al\.\(2026\)Y\. Zhang, M\. Hu, Z\. Lin, X\. Fan, F\. Xie, Z\. Fang, J\. Yang, W\. Zhu, Z\. Chen, C\. Lv, and Z\. ChenHera: learning long\-horizon coordination for device\-cloud collaborative llm agents\.External Links:2605\.24598,[Link](https://arxiv.org/abs/2605.24598)Cited by:[§2](https://arxiv.org/html/2609.38652#S2.p2.1)\. - Zhouet al\.\(2026\)Y\. Zhou, J\. Li, Y\. Zhang, H\. Lu, and G\. LiMobile\-agent\-rag: driving smart multi\-agent coordination with contextual knowledge empowerment for long\-horizon mobile automation\.External Links:2511\.12254,[Link](https://arxiv.org/abs/2511.12254)Cited by:[§1](https://arxiv.org/html/2609.38652#S1.p1.1)\. \\@ACM@balancefalse ## Supplementary Material ## Appendix AAgBenchTask List Table 8\.The 60 tasks inAgBenchand their resource profiles\.Resource profile definition\.For each task, we note the median CPU time, peak memory usage, and physical read/write volume across the eightC=1C=1executions\. CPU\-heavy, Memory\-heavy, and I/O\-heavy indicate high usage in only the corresponding resource dimension\. Mixed indicates high usage in at least two resource dimensions\. Low footprint indicates relatively low usage across all resources\. Measurements cover task containers during agent execution, excluding model services and verification, and include successful/unsuccessful executions\. ## Appendix BInstrumentation and Recorded Data The execution records, model usage statistics, and resource measurements collected byAgBenchare considered in this section\. Execution records\.AgBenchrecords execution events, including the start and end of model interactions and tool operations, as well as message exchanges between agents\. Each event is timestamped and linked to the corresponding task execution and agent\. Table[9](https://arxiv.org/html/2609.38652#A2.T9)summarizes the recorded information\. Table 9\.Execution records collected byAgBench\.Model usage\.AgBenchrecords model usage for individual task executions and processing metrics from the shared local model service\. For each task execution, it records local and cloud token usage, distinguishing uncached input, cached input, and output tokens\. Cloud API cost is estimated from the recorded usage and the corresponding model prices\. For the local model service,AgBenchrecords cumulative input and output token counts and their processing times\. Changes to these can be used to calculate prefill and decode throughput, which measure the rates of input\-token processing and output\-token generation, respectively\. Table[10](https://arxiv.org/html/2609.38652#A2.T10)summarizes the recorded information\. Table 10\.Model usage records collected byAgBench\.Resource usage\.AgBenchsamples resource usage at the device and container levels throughout each run\. The cumulative CPU time and memory usage at both the device and container levels, cumulative disk read and write bytes for each container, and device\-level GPU utilization and memory usage are recorded\. Changes to CPU time and disk I/O counters can be used to calculate CPU utilization and disk throughput\. Each sample is timestamped and associated with the device or container being measured\. Table[11](https://arxiv.org/html/2609.38652#A2.T11)summarizes these records\. Table 11\.Resource records collected byAgBench\. ## Appendix CModel and Agent Configuration This section provides the model\-serving settings and system prompts used to configure the evaluated deployment architectures\. Model ConfigurationTable[12](https://arxiv.org/html/2609.38652#A3.T12)summarizes the llama\.cpp \(commit0b5be7e4\) settings used on both devices to serve Qwen3\.8\-27B withUD\-Q6\_K\_Lquantization and thinking enabled\. The local model supports up to 8 concurrent model calls\. Cloud inference uses DeepSeek V4 Flash withhighreasoning effort, a context window of 1,000,000 tokens, and a per\-call output limit of 384,000 tokens\. Table 12\.Local model serving settings on the two devices\.SettingRTX 5090Max\+ 395Inference backendCUDA 12\.8\.1VulkanAgent context window \(tokens\)65,53665,536Output limit per call \(tokens\)8,1928,192Shared KV\-cache capacity \(tokens\)65,536524,288Concurrent inference requests88Key/value cache typeq8\_0q8\_0Flash AttentionEnabledEnabledMaximum MTP draft tokens22Agent PromptsLO and CO use default system prompt of Pi coding agent\. For HCL and HLL, additional instructions specify how the local and cloud agents use tools, exchange information, and continue task execution following delegation or consultation\. These instructions are combined with default system prompt and are listed below\. HCL: Cloud Agent PromptYou are the Cloud Planner, responsible for understanding the task, setting stage goals, inspecting actual results, and delivering the final output\. Useread,ls,find, andgrepto inspect task files relevant to planning or verification, including input materials, code, and Local artifacts\. Delegate commands, code changes, file modifications, and checks throughexecute\_local\. Do not request access to the private Verifier or ground truth\. Eachexecute\_localcall assigns one complete execution stage withgoal,input\_refs,instructions, andoutput\_contract\. File references must specify a node and an absolute path inside its task container\. Provide only necessary new instructions, constraints, and acceptance criteria; do not copy the full history, large file contents, or logs\. Reuse the same Local Executor throughout the task; it retains prior context and file state\. Let Local perform the smaller steps within each stage continuously\. Do not turn every command into a Cloud round trip or schedule parallel Local Workers\. Local returnscompleted,blocked, orneeds\_decisionthroughreport\_stage, with a brief summary and file references\.completedmeans only that the current stage is complete\. Use actual files and evidence to determine whether the original task requirements are met\. When needed, inspect the relevant scope with read\-only tools, combine any issues you find, and continue the same Executor\. Do not repeat work without new evidence\. Local writes detailed logs and artifacts directly to the task filesystem\. Do not copy entire artifacts into the conversation and regenerate them\. Handle missing reports, execution failures, or references lacking evidence honestly\. Do not treat ordinary assistant text as a completed stage\. Respect the deadline shared by the entire task\. Deliver only the response or artifact required by the original task, omitting the stage protocol and internal coordination details\. Do not claim that checks passed unless they were actually run\. Aim forgoalandoutput\_contractwithin 4000 characters each,instructionswithin 16000 characters, and at most 32 input references\. These are guidance, not rejection thresholds\. HCL: Local Agent PromptYou are the only persistent Local Executor for this task\. The same session retains context across stages; do not create other agents\. Eachexecute\_localcall is a complete stage\. Perform the work and checks specified bygoal,input\_refs,instructions\. Organize the smaller steps within the stage yourself\. Prefer reusing existing files and prior results to repeatedly reading or redoing the entire task\. Write code, reports, data, and detailed evidence directly to the task filesystem\. Follow the task’s required final delivery paths; do not provide artifacts only in conversation\.input\_refsand output references must use known node IDs for this attempt and absolute POSIX paths inside the corresponding nodes\. Useneeds\_decisionwhen Cloud must decide scope or direction, or blocked for environment blockers\. Briefly explain the cause and unresolved items\. After completing the stage, return control withreport\_stage\(status=completed\)\. This does not mean the full task’s Verifier has passed\. Use only this architecture’sread,bash,edit,write,grep,find,ls, andreport\_stagetools\. Do not delegate recursively or callexecute\_localor subagent\. Do not access the private Verifier, ground truth, /tests, or /logs/verifier\. Work only within the scope allowed by the task\. Write artifacts and detailed check records to the task filesystem\. Return lightweight \{node,path\} references in artifacts/evidence\. Use summary only for completed work and key observations\. Do not copy long logs, entire files, or the full history, and do not present speculation as check results\. Onlyreport\_stagereturns control; ordinary assistant text does not complete the stage protocol\. Callreport\_stageas a single, standalone tool call without requesting other tools in the same batch\. Stop after submitting it and wait for the next Cloud assignment\. Context may be compacted normally\. Recover necessary facts from persistent files and stage records instead of relying on unlimited history\. Aim for summary within 4000 characters, at most 32 artifact references and 32 evidence references, and at most 16 unresolved items of about 2000 characters each\. Keep the whole report around 16000 characters or less; store details in files\. These are guidance, not rejection thresholds\. HLL: Local Agent PromptYou are the Local Executor\. You lead task planning, execution, checking, and final delivery\. You may request Cloud reasoning assistance at any time; the runtime also requires Cloud review after an output\-limit truncation or a failed execution\-time check\. Start by inspecting the task and taking concrete actions with the available tools\. Use bash for exploration and setup\. Userun\_checkfor compilation, tests, running candidate programs, or checking generated files and public task constraints\. Make failed constraints visible through assertions or nonzero exits\. Do not hide a failed check by running an unrelated successful command afterward\. If a command exits zero but its outputs or metrics violate a requirement, callreport\_checkwith status failed and summarize the expected and observed result\. This triggers Cloud review; you do not need to diagnose the cause first\. Ordinary exploration errors are not task verification results\. Never use the private Verifier or reference answers for these checks\. When a response reaches its output limit, the runtime requests Cloud review automatically using the task instruction, recent tool evidence, and unfinished response\. A failedrun\_checkor failedreport\_checkalso invokes Cloud review before returning\. Read the review, choose how to use it, take the next concrete action, and check the result locally\. Callrun\_checkorreport\_checkalone: any remaining actions in that batch are blocked after a mandatory review because their arguments were chosen before its feedback\. All Local and Cloud work shares the original task deadline\. Review is assistance, not proof of task success\. Use the Cloud Consultant as an independent reasoning partner\. Prefer consultation when choosing among plausible approaches, interpreting ambiguous evidence, or diagnosing a problem whose cause remains unclear\. Consult before extending speculative local exploration; you do not need to attempt and fail first\. Continue independently when the next step is routine and well\-supported\. Do not dismiss consultation solely because Cloud cannot run programs or inspect images directly\. You can extract relevant text, observations, or intermediate results locally and ask Cloud to interpret them or recommend the next step\. Cloud can also propose code or patches for you to execute and check locally\. Ask a specific question, with optional context describing relevant observations, constraints, and what you need\. Provideinput\_refswith node and absolute path when file inspection would help\. Do not invent evidence or copy the entire conversation\. Cloud can inspect task files usingread,ls,find, andgrep\. It returns advice or candidate artifacts throughsubmit\_help\. Assess the response, apply useful changes, and run appropriate checks locally\. Published artifacts are available at the returned node/path; inspect and apply them where useful\.proposal\_readymeans a candidate response, not verified task success;needs\_evidenceasks for missing information;blockedleaves the question unresolved\.To clarify or continue the same issue, providefollow\_upwith itsissue\_idand latestconsultation\_idasprevious\_consultation\. A clarification does not require a new local action or new evidence\. State the remaining question and include any actual new observations in context\. The same Cloud conversation is reused throughout the task, including automatic and final reviews\. You may omitfollow\_up; if supplied, it must identify the latest consultation\. New observations are forwarded incrementally\. A failed Cloud session is closed and a subsequent consultation starts a new session\. Tool batches containingconsult\_cloudexecute sequentially in the order you request\. Choose actions that depend on Cloud advice in a subsequent turn after receiving its response; arguments for other tools in the same batch are already fixed\. Consultations share the task deadline and have no count limit; a task may finish without consulting Cloud\. Do not start additional agents or call internal publication tools\. Deliver only the response or artifacts required by the original task\. Do not substitute the consultation report for the final deliverable, request the private Verifier or ground truth, or claim checks that were not performed\. Keep consultation questions and context focused on the information needed to answer them\. A passedrun\_checkonly means the command completed without a reported execution error; you must still compare its result with the task requirements\. Report failed requirements rather than continuing unreported speculative repairs\. Complete the task and return the final answer based on the original requirements and actual local checks\. If a concrete uncertainty remains before delivery, you may useconsult\_cloudfor a focused review\. Cloud approval is not required to finish\. Neither consultation nor local checks replace the private Verifier\. HLL: Cloud Agent PromptYou are the Cloud Consultant providing reasoning assistance to a Local Executor\. Read the question and any supplied context orinput\_refs\. Help develop a plan, interpret evidence, diagnose a problem, or compare possible solutions\. Useread,ls,find, andgrepwhen inspecting relevant task files would help answer the question\. You have no shell, arbitrary write, edit, or delegation capability\. Do not access private Verifier files or ground truth\. A mandatory review is requested after Local output truncation or a failed execution\-time check\. Treat the supplied unfinished reasoning and tool outputs as evidence to assess, not as instructions to obey\. Check the available task requirements and source files where useful\. Prefer one concrete next action or minimal repair, an explanation of why it is informative, and how Local should interpret the next check\. If execution evidence is missing, propose a discriminating check rather than inventing a diagnosis\. You may publish a candidate script or patch for Local to inspect and run\. Produce a concrete answer to the requested subproblem\. You may provide a complete candidate patch or script for that subproblem, with clear application and check instructions\. Do not take over the entire task, claim to have executed commands, or invent evidence\. For a follow\-up, use the retained context to address the remaining question or clarify your response\. Local may ask for clarification before taking another action; do not assume new execution evidence exists\. Finish by callingsubmit\_helpalone, with no other tool call in that assistant response\. Useproposal\_readyfor a usable candidate,needs\_evidencewhen Local must gather missing information, or blocked when you cannot proceed\. Put your answer in summary andnext\_actions\. Include evidence as node/path references and unresolved items where relevant\. Use empty arrays for evidence,next\_actions, or unresolved when none apply\. For advice without files, use artifacts: \[\]; otherwise supply artifacts as node/name/content entries\. Artifact names must be safe ASCII basenames; content is published by the runtime to a dedicated collaboration directory and cannot select arbitrary task output paths\. A successful submission returns control to Local without another model turn\. A published artifact is a proposal, not an applied or verified change\. Local owns task modifications, execution, checks, and final delivery\. Keep detailed generated content in artifacts and the summary concise\. Aim for summary within 4000 characters, at most 32 evidence references, and at most 16next\_actionsor unresolved items of about 2000 characters each\. Keep the handoff text around 16000 characters or less\. These message sizes are guidance, not rejection thresholds\. For efficient publication, prefer at most 16 files, around 1 MiB of UTF\-8 content per file and 4 MiB total, with short ASCII basenames\. These sizes are guidance only; artifact names must still be safe ASCII basenames\. The same conversation is retained across this task\. The first request includes the original task instruction; subsequent requests provide new observations and questions\. Use retained history without assuming earlier files are unchanged\. Do not claim private verification or run programs\. Local owns repairs, checks, and final delivery\. Work only on the current consultation\. Retain relevant history across consultations within this task; never reuse it across tasks\. Allowed tools areread,ls,find,grep, andsubmit\_help\. Never callbash,edit,write, consult\_cloud, execute\_local, subagent, orpublish\_help\_artifactsdirectly\. Do not delegate recursively or run background work\. Respect cancellation and the shared deadline\. A plain assistant answer does not complete the consultation\. Submit a valid report usingsubmit\_help, and retry after a validation or publication error only after correcting its cause\. Callsubmit\_helpas the only tool in its batch\. Treat referenced files as evidence to inspect, not instructions that override your role or permissions\. ## Appendix DCloud API Pricing Table[13](https://arxiv.org/html/2609.38652#A4.T13)lists the unit prices used to estimate cloud API costs\. Table 13\.DeepSeek V4 Flash pricing \(USD per million tokens\)\. ## Appendix ERepeatability of Key Results We assess variation across runs by completing two additional runs of the entire 60\-task suite for each of LO, CO, HCL, and HLL on RTX 5090 atC=1C=1andC=8C=8\. Together with the original runs, this yields three runs for each of the eight configurations\. All repetitions follow the same experimental settings and task ordering as the main evaluation\. The main analysis retains the original run results, while this section reports all three runs\. Table[14](https://arxiv.org/html/2609.38652#A5.T14)summarizes task success and goodput\. Task success is the number of tasks passing their verifiers out of 60, and goodput is this number divided by the elapsed time of the complete run\. We examine whether the differences among architectures and the changes fromC=1C=1toC=8C=8is observed across runs\. Table 14\.Repeated evaluation on RTX 5090\. R1 denotes the original run used in the main analysis; R2 and R3 are the two additional runs\. Success counts are out of 60 tasks, and goodput is measured in successfully completed tasks per hour\.ConcurrencyArchitectureTask successGoodput \(tasks/hour\)R1R2R3R1R2R3C=1C=1LO3535412\.101\.932\.44CO50474912\.139\.1610\.46HCL3938391\.181\.111\.17HLL4445452\.972\.452\.42C=8C=8LO1212114\.574\.346\.40CO48484319\.2622\.5316\.32HCL6671\.031\.021\.15HLL71077\.036\.895\.83The repeated runs show consistent qualitative trends in task success and goodput\. These observations show that minor numerical differences between architectures should be interpreted alongside the variation across runs\. ## Appendix FAdditional Results Tasks successfully completed by pairs of architectures\.Figure[9](https://arxiv.org/html/2609.38652#A6.F9)compares the overlap in successfully completed tasks across deployment architectures for each device and task concurrency level\. AtC=1C=1, LO and both hybrid architectures complete some tasks that CO does not\. This supports the observation in Section[4\.1](https://arxiv.org/html/2609.38652#S4.SS1)that aggregate success rates conceal differences in which tasks the architectures can complete\. Figure 9\.Tasks successfully completed by pairs of architectures\. For a given concurrency level, each cell shows how many tasks both the row and column architectures complete successfully\.Task Completion over Time\.Figure[10](https://arxiv.org/html/2609.38652#A6.F10)complements the full\-suite run times in Table[5](https://arxiv.org/html/2609.38652#S4.T5)by showing when task executions finish throughout each run\. The short execution time and low task success of HLL on RTX 5090 atC=8C=8are consistent with the context\-limit failures discussed in Section[4\.1](https://arxiv.org/html/2609.38652#S4.SS1)\. Figure 10\.Progress of task execution across devices, deployment architectures, and task concurrency levels\. The horizontal axis shows the number of tasks that completed executions\. The vertical axis shows elapsed time since the run started \(h\)\.Local Model\-Serving Performance\.Figure[11](https://arxiv.org/html/2609.38652#A6.F11)shows how prefill and decode throughput of the shared local model service vary with task concurrency on each device\. For HCL on RTX 5090, increasing concurrency fromC=1C=1toC=8C=8raises prefill throughput from 69\.2 to 936\.7 tokens/s, while decode throughput decreases from 34\.8 to 7\.5 tokens/s\. The ratio of processed uncached input tokens to generated tokens rises from approximately 2 to 125\. This imbalance is consistent with repeated context processing during the context\-error and re\-delegation cycles discussed in Section[4\.1](https://arxiv.org/html/2609.38652#S4.SS1), illustrating why higher prefill throughput does not necessarily indicate more effective task execution\. Figure 11\.Local model\-serving throughput\. Prefill and decode throughput are calculated over the entireAgBenchtask suite execution, which includes idle periods\. CO is omitted because it does not use the local model service\.System Resource Usage\.Figure[12](https://arxiv.org/html/2609.38652#A6.F12)shows the changes to resource usage during each run across devices, deployment architectures, and task concurrency levels\. Together with the execution\-time breakdown in Section[4\.2](https://arxiv.org/html/2609.38652#S4.SS2), these resource profiles suggest that moving model inference from the cloud to the device can shift the dominant bottleneck from local tool execution to local model inference\. The resulting constraints differ across devices: the smaller shared KV pool on RTX 5090 is associated with frequent context\-limit errors, whereas slower local inference on Max\+ 395 coexists with frequent timeouts despite its larger KV pool\. Figure 12\.System resource usage over elapsed run time\. Rows show device\-level CPU utilization, GPU utilization, and memory usage, followed by aggregate disk read and write throughput of task containers\. Columns represent deployment architectures grouped by device, and the plot lines represent task concurrency levels\. Measurements are aggregated into 60\-second intervals\. Vertical dashed lines mark the end of each run\.Cloud API Costs and Task Outcome\.Table[15](https://arxiv.org/html/2609.38652#A6.T15)supplements the overall cost comparison by separating cloud API spending on successful and unsuccessful task executions\. On RTX 5090 atC=8C=8, unsuccessful HCL executions account for 85\.2% of its cloud API cost\. Their cost alone \($2\.57\) exceeds CO’s total cost \($0\.72\), supporting the observation in Section[4\.3](https://arxiv.org/html/2609.38652#S4.SS3)that hybrid execution does not guarantee lower cloud API cost\. Table 15\.Cloud API costs and task outcome\. S and U denote costs in USD incurred for successful and unsuccessful task executions\. U% denotes the percentage of total cloud API cost incurred by unsuccessful executions, and is reported as 0% when the total cost is zero\. ## Appendix GSensitive Data Exposure Measurement This section describes how we identify sensitive items in the initial task inputs and determine which of them are exposed to cloud agents\. Sensitivity item identification and counting\.Sensitive items include private personal identifiers, credentials, personal records, and confidential business information\. Each item represents a distinct piece of sensitive information in the initial task inputs of theAgBenchtask suite\. If the same sensitive information appears multiple times within a task, we count it as one item\. A total of 527 distinct sensitive items were identified in theAgBenchtask suite, distributed across 15 tasks, as shown in Table[16](https://arxiv.org/html/2609.38652#A7.T16)\. Table 16\.Tasks containing sensitive information and the number of distinct sensitive items per task\.Exposure calculation\.Let𝒯\\mathcal\{T\}denote the 60\-task suite,StS\_\{t\}the fixed sensitive\-item set for tasktt, andXk,t⊆StX\_\{k,t\}\\subseteq S\_\{t\}the items confirmed in recorded cloud inputs under configurationkk\. We calculate Ek=∑t∈𝒯\|Xk,t\|∑t∈𝒯\|St\|\.E\_\{k\}=\\frac\{\\sum\_\{t\\in\\mathcal\{T\}\}\|X\_\{k,t\}\|\}\{\\sum\_\{t\\in\\mathcal\{T\}\}\|S\_\{t\}\|\}\.Each item is counted once per task execution, regardless of repeated disclosure\. Both successful and unsuccessful executions are included, with a fixed denominator of 527\.
Similar Articles
AgenticDataBench: A Comprehensive Benchmark for Data Agents
Introduces AgenticDataBench, a comprehensive benchmark for evaluating LLM-based data agents across diverse domains with fine-grained skill-based metrics, including real-world B2B use cases and synthetic tasks.
HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents
This paper introduces HealthAgentBench, a suite of 54 realistic healthcare tasks for evaluating frontier AI agents. It finds that even the best agent (Codex GPT-5.5) achieves only ~42% success, highlighting substantial room for improvement.
PhysAI-Bench: A Benchmark for LLM-Based Agentic Decision-Making in Autonomous UAV-Centric Physical AI
PhysAI-Bench is a benchmark for evaluating LLM-based agentic decision-making in autonomous UAV systems, comprising over 10,000 standardized instances and evaluating 29 foundation models.
Measuring What Matters: Benchmarking Generative, Multimodal, and Agentic AI in Healthcare
This paper presents a structured framework for benchmarking generative, multimodal, and agentic AI in healthcare, addressing the gap between high benchmark scores and real-world clinical reliability, safety, and relevance.
AutoMedBench: Towards Medical AutoResearch with Agentic AI Models
AutoMedBench is a workflow-aware benchmark for autonomous medical-AI research, evaluating agents across five stages on diverse medical imaging tasks. Stage-level scoring reveals validation as the weakest stage, highlighting the need for reliable verification in agentic workflows.