Tag
WorkSurface-Bench benchmarks enterprise agents on the ability to select the correct knowledge surface (document, table, or graph) for a given query, using 1,151 tasks and separate route and answer scores. Experiments show that even gold-constrained agents struggle with answer accuracy, while surface hints improve performance.