Tag
This paper introduces HybridDeepResearch, a benchmark for evaluating AI agents on tasks that require integrating web search and SQL querying, revealing that even state-of-the-art models struggle with maintaining constraints across structured and unstructured data.