o-net

Tag

Cards List
#o-net

JobBench: Aligning Agent Work With Human Will

arXiv cs.AI · 2026-05-27 Cached

JobBench is a benchmark built from worker surveys to evaluate AI agents on tasks that workers most want automated, covering 130 tasks across 35 professions with detailed rubrics.

0 favorites 0 likes
#o-net

Design and Report Benchmarks for Knowledge Work

arXiv cs.AI · 2026-05-25 Cached

This paper proposes a three-step framework for designing and reporting benchmarks for knowledge work AI, emphasizing alignment between benchmark tasks and real-world work activities. It derives 18 work activities from the O*NET database and analyzes three existing benchmarks (GDPval, OfficeQA Pro, APEX-SWE) to demonstrate gaps between benchmark scores and actual work capability.

0 favorites 0 likes
← Back to home

Submit Feedback