PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say

Hugging Face Daily Papers Papers

Summary

Introduces PrivacyPeek, a benchmark for auditing acquisition-stage privacy leakage in LLM-based agents, showing that agents often gather more sensitive data than needed and that current defenses are insufficient.

LLM-based agents are rapidly advancing, autonomously invoking external tools to complete multi-step tasks for users. However, agents often acquire more sensitive information than the task requires. Existing privacy benchmarks audit what the agent's response or outgoing actions disclose, but overlook the acquisition stage where data first enters the agent's context. The over-acquired information is then one careless action or one attack away from an outright leak. To assess its prevalence, we introduce PrivacyPeek, a benchmark for evaluating acquisition-stage privacy leakage of LLM-based agents, with 1{,}182 cases across 7 acquisition behaviours and 16 application domains. Specifically, Acquisition Inspection examines the agent's tool-call trajectory, both the tools it invokes and the data it receives, to detect when it acquires sensitive information beyond the task scope. Probe Elicitation then issues a follow-up probe and measures how readily an attacker could elicit sensitive information the agent acquired but did not disclose. Our experiments on 10 LLM-based agents across 4 model families show that the unnecessary acquisition of sensitive information is widespread. In addition, we observe a correlation between the task-completion capability and acquisition-stage leakage. Prompt-level defences reduce only a small fraction of acquisition-stage leakage, leaving the majority unmitigated. These results make auditing acquisition-stage privacy both urgent and necessary. Our dataset and code are available at https://github.com/Xuan269/PrivacyPeek-Resource.
Original Article
View Cached Full Text

Cached at: 08/10/26, 06:14 AM

Paper page - PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say

Source: https://huggingface.co/papers/2606.00152

Abstract

LLM-basedagentsarerapidlyadvancing,autonomouslyinvokingexternaltoolstocompletemulti-steptasksforusers.However,agentsoftenacquiremoresensitiveinformationthanthetaskrequires.Existingprivacybenchmarksauditwhattheagent’sresponseoroutgoingactionsdisclose,butoverlooktheacquisitionstagewheredatafirstenterstheagent’scontext.Theover-acquiredinformationisthenonecarelessactionoroneattackawayfromanoutrightleak.Toassessitsprevalence,weintroducePrivacyPeek,abenchmarkforevaluatingacquisition-stageprivacyleakageofLLM-basedagents,with1{,}182casesacross7acquisitionbehavioursand16applicationdomains.Specifically,AcquisitionInspectionexaminestheagent’stool-calltrajectory,boththetoolsitinvokesandthedataitreceives,todetectwhenitacquiressensitiveinformationbeyondthetaskscope.ProbeElicitationthenissuesafollow-upprobeandmeasureshowreadilyanattackercouldelicitsensitiveinformationtheagentacquiredbutdidnotdisclose.Ourexperimentson10LLM-basedagentsacross4modelfamiliesshowthattheunnecessaryacquisitionofsensitiveinformationiswidespread.Inaddition,weobserveacorrelationbetweenthetask-completioncapabilityandacquisition-stageleakage.Prompt-leveldefencesreduceonlyasmallfractionofacquisition-stageleakage,leavingthemajorityunmitigated.Theseresultsmakeauditingacquisition-stageprivacybothurgentandnecessary.Ourdatasetandcodeareavailableathttps://github.com/Xuan269/PrivacyPeek-Resource.

View arXiv pageView PDFGitHub3Add to collection

Get this paper in your agent:

hf papers read 2606\.00152

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2606.00152 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2606.00152 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2606.00152 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

MosaicLeaks: Can your research agent keep a secret?

Hugging Face Blog

MosaicLeaks introduces a new benchmark for measuring privacy leakage in deep-research AI agents, showing that agents often leak private information through external queries and proposing a training method (PA-DR) to reduce leakage while improving task performance.

PrivacyAlign: Contextual Privacy Alignment for LLM Agents

Hugging Face Daily Papers

PrivacyAlign introduces a human-annotated dataset and training framework for aligning LLM agents to respect contextual privacy norms, showing that frontier models still leak sensitive information and that human-grounded evaluation improves alignment.

MosaicLeaks:Privacy Risks in Querying-in-the-Open for Deep Research Agents

arXiv cs.CL

Introduces MosaicLeaks, a benchmark of 1,001 multi-hop deep research tasks that chain private enterprise documents with public web queries to evaluate privacy leakage. Finds that models leak sensitive information at multiple levels, and proposes PA-DR, a reinforcement learning framework that reduces leakage while improving task accuracy.

POLAR-Bench: A Diagnostic Benchmark for Privacy-Utility Trade-offs in LLM Agents

arXiv cs.AI

POLAR-Bench is a diagnostic benchmark that evaluates the privacy-utility trade-off in LLM agents by testing their ability to follow privacy policies while being adversarially probed by third-party models. Results show frontier models protect over 99% of protected attributes but smaller open-weight models leak over half, highlighting gaps in intent-following.