@seclink: Workspace-Bench-Workspaces stores the complete original workspace file system: Various documents, spreadsheets, code, P…
Summary
Workspace-Bench is a benchmark that simulates a real employee's file workspace to evaluate AI agents on tasks involving file discovery and understanding relationships. It includes various file types and role-based scenarios to test agent capabilities.
View Cached Full Text
Cached at: 09/03/26, 12:08 PM
Workspace-Bench-Workspaces stores the complete original workspace file system:
Various documents, spreadsheets, code, PPTs, meeting minutes, dependency files, etc., simulating that pile of numerous and messy, interconnected files on a real employee’s computer.
The Agent has to find files on its own in this workspace, understand file relationships, and complete tasks.
The accompanying tasks are in Workspace-Bench / Workspace-Bench-Lite, with a total of 5 types of role-based workspaces:
Operations Manager Logistics Manager AI Product Manager Backend Developer Researcher
The full version has 388 tasks, about 20,000 files, over 70 formats, with the largest workspace around 20GB. The evaluation looks not only at whether the final answer is correct, but also at whether the Agent found and correctly used the necessary files.
Similar Articles
EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions
EnterpriseClawBench presents a benchmark for enterprise agents based on real-world workplace sessions, offering 852 reproducible tasks and comprehensive evaluation metrics beyond single performance scores.
@stainlu: we built an open-source version of workspace agents - any model, self-hosted - per-session sandbox - credential isolati…
OpenClaw is an open-source, self-hostable workspace-agent platform that sandboxes each session and isolates credentials.
SaaS-Bench: Can Computer-Use Agents Leverage Real-World SaaS to Solve Professional Workflows?
SaaS-Bench is a new benchmark built on 23 deployable SaaS systems across six professional domains, containing 106 long-horizon tasks for evaluating computer-using agents. Experiments show that even the strongest models complete fewer than 4% of tasks end-to-end, highlighting significant limitations in current agent capabilities.
Back to the Future: A workbook time machine for spread sheet creation benchmarks
This paper introduces the workbook time machine, a pipeline that automatically creates benchmarks for evaluating language models on creating derived spreadsheet objects like formulas, charts, pivot tables, and conditional formatting. The authors produce WTM-Corpus and a curated 150-task benchmark WTM-Bench, then evaluate spreadsheet agents across artifact types, step complexity, and instruction granularity.
WorkBench Revisited: Workplace Agents Two Years On
This paper revisits the WorkBench benchmark for workplace agents two years after its initial release, showing that the best agent (Claude Opus 4.8) now completes 89% of tasks with only 2.5% harmful side effects, compared to GPT-4's 43% completion and 26% harm rate in 2024. It finds that capability and safety improve together, open-weight models have drastically lowered costs, and some basic mistakes persist.