@seclink: Workspace-Bench-Workspaces stores the complete original workspace file system: Various documents, spreadsheets, code, P…

X AI KOLs Timeline Tools

Summary

Workspace-Bench is a benchmark that simulates a real employee's file workspace to evaluate AI agents on tasks involving file discovery and understanding relationships. It includes various file types and role-based scenarios to test agent capabilities.

Workspace-Bench-Workspaces stores the complete original workspace file system: Various documents, spreadsheets, code, PPTs, meeting minutes, dependency files, etc., simulating that pile of numerous and messy, interconnected files on a real employee's computer. The Agent has to find files on its own in this workspace, understand file relationships, and complete tasks. The accompanying tasks are in Workspace-Bench / Workspace-Bench-Lite, with a total of 5 types of role-based workspaces: Operations Manager Logistics Manager AI Product Manager Backend Developer Researcher The full version has 388 tasks, about 20,000 files, over 70 formats, with the largest workspace around 20GB. The evaluation looks not only at whether the final answer is correct, but also at whether the Agent found and correctly used the necessary files.
Original Article
View Cached Full Text

Cached at: 09/03/26, 12:08 PM

Workspace-Bench-Workspaces stores the complete original workspace file system:

Various documents, spreadsheets, code, PPTs, meeting minutes, dependency files, etc., simulating that pile of numerous and messy, interconnected files on a real employee’s computer.

The Agent has to find files on its own in this workspace, understand file relationships, and complete tasks.

The accompanying tasks are in Workspace-Bench / Workspace-Bench-Lite, with a total of 5 types of role-based workspaces:

Operations Manager Logistics Manager AI Product Manager Backend Developer Researcher

The full version has 388 tasks, about 20,000 files, over 70 formats, with the largest workspace around 20GB. The evaluation looks not only at whether the final answer is correct, but also at whether the Agent found and correctly used the necessary files.

Similar Articles

Back to the Future: A workbook time machine for spread sheet creation benchmarks

arXiv cs.AI

This paper introduces the workbook time machine, a pipeline that automatically creates benchmarks for evaluating language models on creating derived spreadsheet objects like formulas, charts, pivot tables, and conditional formatting. The authors produce WTM-Corpus and a curated 150-task benchmark WTM-Bench, then evaluate spreadsheet agents across artifact types, step complexity, and instruction granularity.

WorkBench Revisited: Workplace Agents Two Years On

arXiv cs.CL

This paper revisits the WorkBench benchmark for workplace agents two years after its initial release, showing that the best agent (Claude Opus 4.8) now completes 89% of tasks with only 2.5% harmful side effects, compared to GPT-4's 43% completion and 26% harm rate in 2024. It finds that capability and safety improve together, open-weight models have drastically lowered costs, and some basic mistakes persist.