When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills
Summary
Introduces AntiSkillBench, a benchmark for evaluating privacy leakage, impersonation risk, and defenses in persona-skill pipelines for AI agents, finding that risks persist across backbones and existing defenses are limited.
View Cached Full Text
Cached at: 08/05/26, 09:44 AM
Paper page - When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills
Source: https://huggingface.co/papers/2608.03700
Abstract
Personaskillsdistillpersonalinteractionhistoriesintoportableandexecutableartifactsfordownstreamagents.Whileenablingflexiblepersonalization,thisprocessconcentratesfragmentedpersonalsignals,amplifiestheirimpactthroughreuse,andchallengesdefensesdesignedforindividualrecordsorretrieval-basedmemory.Tosystematicallyinvestigatethesafetyofthepersona-skillpipeline,weintroduceAntiSkillBench,anend-to-endbenchmarkforevaluatingrisksanddefensesacrossthepersona-skillpipeline.Itcomprises:(i)adatasetof7,500persona-groundeddialoguetraces,constructedfrom50behaviorallyrichprofilesspanningdiversetaskscenarios;(ii)anevaluationsuitethatmeasuresskill-levelprivacyleakageandagent-levelattributedisclosureandbehavioralimpersonationacrossthreeskill-distillationstrategies;and(iii)adefenseevaluationcoveringfourconfigurationsacrossonlineandpost-hocinterventions,includingactiverisksuppressionandpassiveprovenanceprotection.Experimentsacrossthreefrontieragentsshowthatpersona-skillriskspersistacrossagentbackbonesanddistillationprotocols,extendingfromexplicitattributestocommunicationstylesandpersonalitytraits.Existingdefensesexhibitlimitedanddistillation-dependenteffectiveness,failingtogeneralizeacrossriskanddistillationstrategies.TheseresultshighlightAntiSkillBenchasachallengingbenchmarkfordevelopingprivacy-preservingandauthenticity-awarepersonaskills.
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2608\.03700
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2608.03700 in a model README.md to link it from this page.
Datasets citing this paper1
#### yonglixiang/AntiSkillBench Viewer• Updatedabout 4 hours ago • 5.05k • 2
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.03700 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
What should a good benchmark for AI agent skill security scanners include?
Discusses the challenges of designing a benchmark for security scanners that evaluate AI agent skills, which introduce new supply-chain risks. It questions whether benchmarks should include real-world malicious samples, synthetic cases, full skill directories, or boundary cases.
MCP-Persona: Benchmarking LLM Agents on Real-World Personal Applications via Environment Simulation
MCP-Persona is a benchmark evaluating LLM agents on personalized tools interacting with individual accounts and local databases. Experiments reveal significant challenges for state-of-the-art agents in personalized tool use.
@Dinosn: A curated list of resources on agent skills security: attacks, defenses, frameworks, and benchmarks for securing AI age…
A curated GitHub awesome list compiling resources on agent skills security, covering attacks (tool poisoning, indirect prompt injection, backdoors), defenses (sandboxing, permissions, formal verification), frameworks like OWASP Agentic Skills Top 10, MITRE ATLAS, NIST AI RMF, and evaluation benchmarks.
Stealth Apart, Harm Together: Skill Cascading Attacks on Skill-Based Agent Systems
Introduces skill cascading attacks on skill-based agent systems, where multiple benign-looking skill modifications collectively cause harm while evading detection, and presents an automated red-teaming framework and benchmark to study this threat.
@dair_ai: Finally, a good paper testing whether Agent Skills actually help. Worth reading if you are maintaining a skill library …
A benchmark study shows that injecting Agent Skills in Web Development tasks often reduces performance and increases token cost, with failure modes like length-distracted and content-misled models, highlighting the need for per-deployment evaluation.