PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say
Summary
Introduces PrivacyPeek, a benchmark for auditing acquisition-stage privacy leakage in LLM-based agents, showing that agents often gather more sensitive data than needed and that current defenses are insufficient.
View Cached Full Text
Cached at: 08/10/26, 06:14 AM
Paper page - PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say
Source: https://huggingface.co/papers/2606.00152
Abstract
LLM-basedagentsarerapidlyadvancing,autonomouslyinvokingexternaltoolstocompletemulti-steptasksforusers.However,agentsoftenacquiremoresensitiveinformationthanthetaskrequires.Existingprivacybenchmarksauditwhattheagent’sresponseoroutgoingactionsdisclose,butoverlooktheacquisitionstagewheredatafirstenterstheagent’scontext.Theover-acquiredinformationisthenonecarelessactionoroneattackawayfromanoutrightleak.Toassessitsprevalence,weintroducePrivacyPeek,abenchmarkforevaluatingacquisition-stageprivacyleakageofLLM-basedagents,with1{,}182casesacross7acquisitionbehavioursand16applicationdomains.Specifically,AcquisitionInspectionexaminestheagent’stool-calltrajectory,boththetoolsitinvokesandthedataitreceives,todetectwhenitacquiressensitiveinformationbeyondthetaskscope.ProbeElicitationthenissuesafollow-upprobeandmeasureshowreadilyanattackercouldelicitsensitiveinformationtheagentacquiredbutdidnotdisclose.Ourexperimentson10LLM-basedagentsacross4modelfamiliesshowthattheunnecessaryacquisitionofsensitiveinformationiswidespread.Inaddition,weobserveacorrelationbetweenthetask-completioncapabilityandacquisition-stageleakage.Prompt-leveldefencesreduceonlyasmallfractionofacquisition-stageleakage,leavingthemajorityunmitigated.Theseresultsmakeauditingacquisition-stageprivacybothurgentandnecessary.Ourdatasetandcodeareavailableathttps://github.com/Xuan269/PrivacyPeek-Resource.
View arXiv pageView PDFGitHub3Add to collection
Get this paper in your agent:
hf papers read 2606\.00152
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2606.00152 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2606.00152 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2606.00152 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
MosaicLeaks: Can your research agent keep a secret?
MosaicLeaks introduces a new benchmark for measuring privacy leakage in deep-research AI agents, showing that agents often leak private information through external queries and proposing a training method (PA-DR) to reduce leakage while improving task performance.
PrivacyAlign: Contextual Privacy Alignment for LLM Agents
PrivacyAlign introduces a human-annotated dataset and training framework for aligning LLM agents to respect contextual privacy norms, showing that frontier models still leak sensitive information and that human-grounded evaluation improves alignment.
MosaicLeaks:Privacy Risks in Querying-in-the-Open for Deep Research Agents
Introduces MosaicLeaks, a benchmark of 1,001 multi-hop deep research tasks that chain private enterprise documents with public web queries to evaluate privacy leakage. Finds that models leak sensitive information at multiple levels, and proposes PA-DR, a reinforcement learning framework that reduces leakage while improving task accuracy.
POLAR-Bench: A Diagnostic Benchmark for Privacy-Utility Trade-offs in LLM Agents
POLAR-Bench is a diagnostic benchmark that evaluates the privacy-utility trade-off in LLM agents by testing their ability to follow privacy policies while being adversarially probed by third-party models. Results show frontier models protect over 99% of protected attributes but smaller open-weight models leak over half, highlighting gaps in intent-following.
When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills
Introduces AntiSkillBench, a benchmark for evaluating privacy leakage, impersonation risk, and defenses in persona-skill pipelines for AI agents, finding that risks persist across backbones and existing defenses are limited.