π-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows
Summary
π-Bench is a new benchmark comprising 100 multi-turn tasks with hidden user intents across 5 domain-specific user personas, designed to evaluate proactive assistance in long-horizon workflows for personal assistant agents.
View Cached Full Text
Cached at: 05/22/26, 02:24 AM
Paper page - π-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows
Source: https://huggingface.co/papers/2605.14678 Authors:
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Proactive assistance in personal agent systems requires identifying hidden user intents through sustained multi-turn interactions, which current benchmarks fail to adequately evaluate.
The rise ofpersonal assistant agents, e.g., OpenClaw, highlights the growing potential oflarge language modelsto support users across everyday life and work. A core challenge in these settings isproactive assistance, since users often begin with underspecified requests and leave important needs, constraints, or preferences unstated. However, existing benchmarks rarely evaluate whether agents can identify and act on such hidden intents before they are explicitly stated, especially in sustainedmulti-turn interactionswhere user needs emerge gradually. To address this gap, we introduce π-Bench, a benchmark forproactive assistancecomprising 100 multi-turn tasks across 5domain-specific user personas. By incorporating hiddenuser intents, inter-task dependencies, and cross-session continuity, π-Bench evaluates agents’ ability to anticipate and address user needs over extended interactions, jointly measuringproactivityandtask completioninlong-horizon trajectoriesthat better reflect real-world use. Experiments show (1)proactive assistanceremains challenging, (2) a clear distinction betweentask completionandproactivity, and (3) the value of prior interaction for proactive intent resolution in later tasks.
View arXiv pageView PDFProject pageGitHub7Add to collection
Get this paper in your agent:
hf papers read 2605\.14678
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2605.14678 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2605.14678 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2605.14678 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
@tobi: Great idea. Will support this on Shopify docs
Shopify CEO Tobi Lütke endorsed a proposal by Malte Ubl for documentation platforms to read preferred programming languages from the Accept-Language HTTP header, allowing AI agents and developers to receive code examples in languages like Python instead of TypeScript.
"Huge implications - binaries are now basically editable code"
GPT has saturated ValsAI's SRE benchmark, which tests whether AI models can reverse engineer software from binaries, suggesting advanced decompilation capabilities that effectively make compiled code editable.
Our agent found a clean rule across 120 CIKs, published it to four repos, and it was false at 494 companies
An LLM-assisted agent published a false rule about SEC 8-K filing timestamps after confirming it only against the 120-company sample that generated it; the author retracted the claim and argues that agent outputs must be validated against data outside the hypothesis-generation loop.
I Feel about AI
A personal reflection expressing mixed emotions about AI, from surprise at neural network capabilities and fear of existential risk to anger at corporate exploitation and political inaction, concluding that the technology seems promising while the societal outlook is bleak.
@Aniket1836020: ✨ Opportunity for Researcher @OpenAI in London If you want to work on pretraining Astra successors, London training tea…
OpenAI's London training team is hiring researchers to work on pretraining successors to its Astra model and advance the frontier of the company's flagship LLMs.