Tag
This paper introduces AgentCIBench, a benchmark to evaluate privacy risks in computer-use agents, finding that 11 of 15 frontier agents leak information in over 50% of scenarios.
RedactionBench is a manually annotated benchmark for evaluating contextual PII redaction in large language models, introducing the R-Score metric and showing that contextual redaction remains an unsolved problem.
This paper introduces Minim, a trusted local broker that performs privacy-aware minimization of UI observations for LLM-powered agents, using contextual integrity to balance task necessity and sensitivity scores. Experiments on WebArena show it reduces irrelevant sensitive leakage while preserving task-critical information.
Proposes Complementary Self-Distillation (SelfCI) to improve contextual integrity in LLMs by balancing utility and privacy. Evaluated on CI-RL and PrivacyLens benchmarks across multiple models.