Using AI Agents to Automate Black-Box Audits of Personalization Algorithms at Scale
Summary
This paper introduces a framework using generative AI agents to automate black-box audits of personalization algorithms, demonstrating with 1,120 agents on X after the 2024 U.S. election that the algorithmic feed amplifies toxic, polarizing, and right-leaning content compared to the chronological feed.
View Cached Full Text
Cached at: 07/01/26, 05:31 AM
# Using AI Agents to Automate Black-Box Audits of Personalization Algorithms at Scale Source: [https://arxiv.org/abs/2606.30801](https://arxiv.org/abs/2606.30801) [View PDF](https://arxiv.org/pdf/2606.30801) > Abstract:Personalization algorithms determine what content users encounter on online platforms\. Auditing these systems is difficult because independent auditors have only black\-box access to the algorithms, while personalization depends on users' attributes, behavior, and evolving interaction histories\. Existing auditing methods face a tradeoff: studies with real users capture realistic behavior but are costly and hard to control, whereas sock\-puppet audits scale more easily but often rely on scripted behavior that limits realism\. Beyond this, both approaches struggle to decouple user attributes from user behavior, limiting our ability to causally understand personalization\. To address this gap, we introduce a framework for black\-box audits of personalization algorithms using generative AI agents as behavioral engines for synthetic accounts\. Each agent is instantiated with a fixed persona, grounded in demographic and political survey data, and interacts with a platform's content by reasoning about it and choosing actions\. Because behavior is fixed within each persona while platform\-visible signals such as age, gender, or location can be experimentally perturbed, our design enables counterfactual auditing of how platforms respond to user attributes\. As a case study, we deploy 1,120 agents on X shortly after the 2024 U\.S\. election, spanning 14 personas and three counterfactual conditions, collecting over 200,000 content exposures\. We find that X's algorithmic feed amplifies toxic, polarizing, political, and right\-leaning content relative to the chronological feed, with amplification varying sharply by user ideology\. Counterfactual analyses show that demographic signals affect content delivery in persona\-dependent ways: pooled effects are largely null, while subgroup\-level effects vary in direction and magnitude\. Our work establishes GenAI\-based agents as a new tool for algorithmic auditing\. ## Submission history From: Alessandro Morosini \[[view email](https://arxiv.org/show-email/9693250a/2606.30801)\] **\[v1\]**Mon, 29 Jun 2026 18:25:09 UTC \(1,395 KB\)
Similar Articles
AI Agent Audits ?
A practitioner shares concerns about an upcoming audit revealing undocumented AI agents in production, highlighting governance gaps and risks with customer PII access.
AI Watchdog: Agent Interfaces for Detecting and Defending Against Manipulative Dark Patterns in AI Conversations
AI Watchdog is a browser-based agent interface that detects manipulative dark patterns in AI conversations, such as sycophancy and brand bias, and alerts users in real-time. Experimental results showed that just-in-time warnings without cognitive forcing significantly reduced compliance with AI-steered recommendations containing dark patterns.
AI agents are fun until they start touching real data
The article discusses the governance challenges that arise when AI agents interact with real company data and tools, highlighting the need for policy enforcement and audit trails, and mentions Trust3 AI as a potential solution.
Adversarial Creation and Detection of AI-Generated Social Bot Content
This paper presents an adversarial methodology for creating and detecting AI-generated social bot content, curating a multilingual, cross-platform dataset of paired human and AI messages. Training on this adversarial data yields detection that significantly outperforms existing content-based bot detection models in real-world settings.
Should AI agents decide audience selection, personalization, and customer journeys?
Explores the potential of AI agents to take over marketing decisions like audience selection and personalization, questioning whether marketers should hand over control to AI.