OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding
Summary
OmegaUse-OfficeVal is a benchmark for evaluating LLM agents on long-horizon office-suite tasks with economic grounding, comparing human costs and LLM inference costs. It includes 100 tasks requiring ~2.3 hours of human labor each, and finds that frontier LLMs are cheaper and faster but still below human quality.
View Cached Full Text
Cached at: 07/30/26, 05:46 AM
Paper page - OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding
Source: https://huggingface.co/papers/2607.27155 Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Largelanguagemodel(LLM)agentsareincreasinglyexpectedtoassistusersincompletingtasks.However,existingbenchmarksprovidelimitedsupportforevaluatingwhetheragentscancarryoutoffice-suiteworkflowsatareasonablecost.WeintroduceOmegaUse-OfficeVal,abenchmarkforevaluatingLLMagentsonlong-horizonoffice-suitetaskswithtask-leveleconomicgrounding.Thebenchmarkcomprises100tasksderivedfromoffice-suiterequestsproposedbypractitionersandadaptedthroughaprivacy-preservingprocess.Onaverage,thesetasksrequire2.32hoursofhumanlabortocomplete.Animportantfeatureofthebenchmarkisthateachtaskispairedwithtwoeconomicsignals:humanlabortimeandtaskpriceproxy.ThesesignalsenabledirectcomparisonsbetweenhumancostsandLLMinferencecosts,aswellasvalue-weightedevaluation.Tosupportstableevaluation,wedevelopcode-basedverifiersfromfine-grainedrubrics.WeevaluateseveralfrontierLLMstogetherwithahumanbaseline.AlthoughallevaluatedLLMsaresubstantiallycheaperandfasterthanhumanworkers,theyhavenotyetapproachedhuman-leveldeliverablequality.Thecodeanddatasetarefullyopen-sourced,andmoreinformationisavailableonourprojectwebsite:https://omegause-officeval.github.io.
View arXiv pageView PDFProject pageGitHub3Add to collection
Get this paper in your agent:
hf papers read 2607\.27155
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.27155 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.27155 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.27155 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents
Introduces PolyWorkBench, a benchmark for evaluating LLM agents on multilingual long-horizon workplace workflows across five domains, demonstrating significant performance degradation compared to monolingual settings.
Mind the Gap: Can Frontier LLMs Pass a Standardized Office Proficiency Exam?
This paper introduces OfficeEval, a benchmark based on China's National Computer Rank Examination (NCRE) to evaluate LLM agents on complex Office automation tasks. Frontier models achieve at best 36.6% in single-turn and 68.8% with agentic systems, far below human-level performance.
CoffeeBench: Benchmarking Long-Horizon LLM Agents in Heterogeneous Multi-Agent Economies
CoffeeBench is a benchmark for evaluating LLM agents in a long-horizon multi-agent economic simulation where firms interact over 90 days to maximize profits, revealing differences in communication patterns and performance among various models.
ORAgentBench: Can LLM Agents Solve Challenging Operations Research Tasks End to End?
本文介绍ORAgentBench,一个用于评估LLM代理在端到端运筹学任务中表现的执行基准,包含107个经过人工审查的任务。实验表明,当前最佳代理仅通过35.51%的任务,揭示了在可靠决策制定方面的重大不足。
OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
OSWorld 2.0 is a new benchmark for evaluating computer-use agents on 108 long-horizon, real-world workflows. Current agents like Claude Opus 4.8 and GPT-5.5 achieve low completion rates, highlighting significant limitations in handling complex, multi-step tasks.