π-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows

Hugging Face Daily Papers Papers

Summary

π-Bench is a new benchmark comprising 100 multi-turn tasks with hidden user intents across 5 domain-specific user personas, designed to evaluate proactive assistance in long-horizon workflows for personal assistant agents.

The rise of personal assistant agents, e.g., OpenClaw, highlights the growing potential of large language models to support users across everyday life and work. A core challenge in these settings is proactive assistance, since users often begin with underspecified requests and leave important needs, constraints, or preferences unstated. However, existing benchmarks rarely evaluate whether agents can identify and act on such hidden intents before they are explicitly stated, especially in sustained multi-turn interactions where user needs emerge gradually. To address this gap, we introduce π-Bench, a benchmark for proactive assistance comprising 100 multi-turn tasks across 5 domain-specific user personas. By incorporating hidden user intents, inter-task dependencies, and cross-session continuity, π-Bench evaluates agents' ability to anticipate and address user needs over extended interactions, jointly measuring proactivity and task completion in long-horizon trajectories that better reflect real-world use. Experiments show (1) proactive assistance remains challenging, (2) a clear distinction between task completion and proactivity, and (3) the value of prior interaction for proactive intent resolution in later tasks.
Original Article
View Cached Full Text

Cached at: 05/22/26, 02:24 AM

Paper page - π-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows

Source: https://huggingface.co/papers/2605.14678 Authors:

,

,

,

,

,

,

,

,

,

,

,

,

Abstract

Proactive assistance in personal agent systems requires identifying hidden user intents through sustained multi-turn interactions, which current benchmarks fail to adequately evaluate.

The rise ofpersonal assistant agents, e.g., OpenClaw, highlights the growing potential oflarge language modelsto support users across everyday life and work. A core challenge in these settings isproactive assistance, since users often begin with underspecified requests and leave important needs, constraints, or preferences unstated. However, existing benchmarks rarely evaluate whether agents can identify and act on such hidden intents before they are explicitly stated, especially in sustainedmulti-turn interactionswhere user needs emerge gradually. To address this gap, we introduce π-Bench, a benchmark forproactive assistancecomprising 100 multi-turn tasks across 5domain-specific user personas. By incorporating hiddenuser intents, inter-task dependencies, and cross-session continuity, π-Bench evaluates agents’ ability to anticipate and address user needs over extended interactions, jointly measuringproactivityandtask completioninlong-horizon trajectoriesthat better reflect real-world use. Experiments show (1)proactive assistanceremains challenging, (2) a clear distinction betweentask completionandproactivity, and (3) the value of prior interaction for proactive intent resolution in later tasks.

View arXiv pageView PDFProject pageGitHub7Add to collection

Get this paper in your agent:

hf papers read 2605\.14678

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2605.14678 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2605.14678 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2605.14678 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

@tobi: Great idea. Will support this on Shopify docs

X AI KOLs Following

Shopify CEO Tobi Lütke endorsed a proposal by Malte Ubl for documentation platforms to read preferred programming languages from the Accept-Language HTTP header, allowing AI agents and developers to receive code examples in languages like Python instead of TypeScript.

I Feel about AI

Hacker News Top

A personal reflection expressing mixed emotions about AI, from surprise at neural network capabilities and fear of existential risk to anger at corporate exploitation and political inaction, concluding that the technology seems promising while the societal outlook is bleak.