Benchmarking the Personalization Capabilities of Large Language Models
Summary
This paper introduces SDR-Bench, a benchmark for evaluating the personalization capabilities of large language models in a two-party Bayesian Persuasion framework, finding a consistent plateau across frontier LLMs and validating the framework with a field deployment.
View Cached Full Text
Cached at: 07/24/26, 05:00 AM
# Benchmarking the Personalization Capabilities of Large Language Models
Source: [https://arxiv.org/html/2607.20471](https://arxiv.org/html/2607.20471)
Ashutosh SrivastavaSiddharth YedlapatiVinay AggarwalYaman K Singla Shashwat DixitJitendra AjmeraBalaji Krishnamurthy ![[Uncaptioned image]](https://arxiv.org/html/2607.20471v1/figures/adobe_logo.png)Adobe Media & Data Science Research behavior\-in\-the\-wild@googlegroups\.com
###### Abstract
Personalization, the act of varying a message to induce action from a specific receiver while keeping sender, channel, and time fixed, has a long tradition in psychology and marketing as a two\-party problem in which sender and receiver have independent objectives\. Large language models remove the bounded\-inventory constraint of classical retrieval\-and\-ranking approaches by generating a continuum of message variants conditioned on inferred receiver state, raising the question of how well current models perform personalization in the classical sense\. Existing LLM personalization benchmarks measure sender\-side adaptation, in which the receiver is the same user the model is serving\. The two\-party question, whether a generated message induces its intended action in a third party, has been investigated only through A/B tests and small\-scale human studies that cannot be re\-run against a new model on demand\. We adapt the Bayesian Persuasion framework of Kamenica and Gentzkow \(2011\) to generative agents and instantiate the formulation in sales, where receiver actions are routinely logged against the outreach that induced them\. We releaseSDR\-Bench, a public corpus of 6,279 customer success stories spanning 22 industries and approximately 200 enterprises, served through a temporally constrained simulation that prevents future\-data leakage\. Across frontier LLMs and deep\-research agents, we observe a consistent personalization plateau and on a Fortune 100 tech cohort no model statistically separates successful from unsuccessful outreach\. A field deployment with 12 professional sales representatives validates the framework, with 48 percent of model\-generated content rated immediately useful and senior\-expert agreement at Pearson 0\.82\. We releaseSDR\-ArenaandSDR\-Benchpublicly to support reproducible study of generative personalization at scale\.
Figure 1:Overview of SDR\-Arena showcasing how LLM generated output is compared with artifacts like Sales Emails, Transcripts & Success Stories to benchmark their personalization capability## 1Introduction
Communication as defined by the seminal work of Lasswell \(Lasswell \([1948](https://arxiv.org/html/2607.20471#bib.bib6)\)\) characterizes any communicative act in five variables: who says what to whom, through which channel, with what effect, and at what time\. Within this framework personalization can be located as the act of varying the message, that is, thewhat, while conditioning on \(and holding fixed\) the speaker, receiver, channel, and time\. The act has two parties whose interests need not initially align; a sender, who selects the message with the goal of inducing some action from the receiver, and a receiver, who has independent preferences and chooses whether to act\. The study of how senders shape messages to induce receiver action has been a long tradition in multiple fields, including, psychology, economics, and marketing, beginning with the Yale Communication and Attitude Change program \(Hovlandet al\.\([1953](https://arxiv.org/html/2607.20471#bib.bib3)\)\) and continuing through the Elaboration Likelihood Model \(Petty and Cacioppo \([1986](https://arxiv.org/html/2607.20471#bib.bib7)\)\) and the formal signaling\-game treatment of Bayesian Persuasion \(Kamenica and Gentzkow \([2011](https://arxiv.org/html/2607.20471#bib.bib22)\)\)\.
So far, machine learning research on personalization has approached this problem as learning a policy to retrieve and rank items from a fixed inventory of candidates\. Recommender systems rank items from a catalog against user\-history signals \(Ricciet al\.\([2010](https://arxiv.org/html/2607.20471#bib.bib15)\)\); advertising platforms optimize the selection, and targeting of ads among a fixed pool of pre\-authored creatives \(Choiet al\.\([2020](https://arxiv.org/html/2607.20471#bib.bib13)\)\); persona\-based dialogue systems condition response generation on explicit persona representations while remaining restricted to comparatively narrow conversational domains \(Zhanget al\.\([2018](https://arxiv.org/html/2607.20471#bib.bib26)\)\)\.
Large language models transition this personalization from a retrieval and ranking problem to a generative one\. LLMs can generate a continuum of message variants for a given \(speaker, receiver, channel, time\) tuple, conditioned on whatever attributes of the receiver can be inferred from the available context\. A growing body of work has applied LLMs to generative\-personalization in marketing \(Matzet al\.\([2024](https://arxiv.org/html/2607.20471#bib.bib27)\)\), education \(Tasdelen and Bodemer \([2025](https://arxiv.org/html/2607.20471#bib.bib18)\); Sharmaet al\.\([2025](https://arxiv.org/html/2607.20471#bib.bib17)\)\), and human\-AI interaction \(Chenet al\.\([2024](https://arxiv.org/html/2607.20471#bib.bib20)\)\)\. These applied results raise the question of how well current LLMs perform personalization in the sense the classical literature studies it: as a sender selectingwhatto say in order to induce a specific receiver action\. While some work exists for measuring LLM personalization, however, it measures a very different property compared to the personalization talked about in psychology and economics literature \(Lasswell \([1948](https://arxiv.org/html/2607.20471#bib.bib6)\); Kamenica and Gentzkow \([2011](https://arxiv.org/html/2607.20471#bib.bib22)\)\)\. LaMP \(Salemiet al\.\([2024](https://arxiv.org/html/2607.20471#bib.bib32)\)\) evaluates personalized text generation conditioned on a user’s history of past interactions\. PersoBench \(Afzoonet al\.\([2024](https://arxiv.org/html/2607.20471#bib.bib34)\)\) measures persona consistency in open\-domain dialogue\. PersonaConvBench \(Liet al\.\([2025](https://arxiv.org/html/2607.20471#bib.bib35)\)\) scores persona\-grounded conversational quality\. PersonaLens \(Zhaoet al\.\([2025](https://arxiv.org/html/2607.20471#bib.bib37)\)\) evaluates assistant behavior under declared user preferences\. PersonaMem \(Jianget al\.\([2025](https://arxiv.org/html/2607.20471#bib.bib36)\)\) measures long\-horizon recall of user attributes across sessions\. In these works, the recipient of the LLM generated message is the same user that the model is serving, and the optimization target is one\-party: the alignment between the model’s output and the preferences of the user who issued the prompt, in the same sense that RLHF aligns an assistant to its user\. The two\-party question, in which sender and receiver have independent objectives and the sender’s success is measured bywhether the receiver acts, has primarily been investigated through randomized human\-subject experiments in which participants are exposed to human\- or LLM\-generated persuasive messages and evaluated based on subsequent shifts in attitudes, agreement, or behavioral intentions \(Matzet al\.\([2024](https://arxiv.org/html/2607.20471#bib.bib27)\); Durmuset al\.\([2024](https://arxiv.org/html/2607.20471#bib.bib38)\)\)\. Such studies are tied to a specific methodology, require weeks of execution, and cannot be re\-run against a new model on demand\. Consequently, comparisons between LLMs and human experts for two\-party personalization remain specific to individual studies and are not directly comparable across systems\. Therefore, there is a need for a formal model of personalization which accounts for both the sender and receiver of the message, sufficient to support a reproducible and automated benchmark applicable to arbitrary generative systems\.
This formulation requires an empirical setting where receiver actions are observed and logged against the specific messages that induced them, the messages are authored at the receiver level rather than at the segment level, and human\-authored ground truth at known successful induced actions are available at scale\. Sales outreach artifacts provides a good testbed to measure this as receiver actions \(replying, scheduling a call, closing a deal\) are routinely logged against the specific outreach that induced them \(Terhoet al\.\([2022b](https://arxiv.org/html/2607.20471#bib.bib33)\)\)\. A sales outreach is drafted one\-to\-one by a Sales Development Representative \(SDR\) for a specific prospect\. In an in\-house study conducted with a Fortune 100 enterprise, we observed that personalized SDR outreach achieved approximately seven times the click\-through rate of templated outreach when promoting the same products to comparable prospect groups\.\. Furthermore, the sales funnel produces a layered set of human\-authored artifacts at known successful transitions, namely outreach emails that secured a call, call transcripts that secured deal discussions, and post\-deal customer success stories that document the content that closed the deal\. Together, these properties make sales a natural empirical instantiation of the two\-party formulation where each artifact in the funnel serves as ground truth at a different stage of the same \(seller, prospect, product\) tuple, enabling stage\-specific evaluation of generative personalization\.
We develop our frameworkSDR\-Arena\(illustrated in Fig[1](https://arxiv.org/html/2607.20471#S0.F1)\) on this empirical setting and list our contributions are as follows:
- •We adapt Bayesian Persuasion to generative agents, recovering personalization as informational alignment between an agent’s generated content and the receiver\-specific content implicit in the ground\-truth sales outreach artifact\.
- •We construct SDR\-Bench, a public corpus of 6,279 customer success stories from approximately 200 enterprises across 22 industries, each paired with the seller, prospect, product, and historical timestamp required to evaluate whether an agent can predict the strategic content of the deal\-closing pitch\.
- •We release SDR\-Arena, an evaluation framework that operationalizes the formalization on SDR\-Bench and on proprietary sales artifacts; to prevent future data leakage, where an agent retrieves the very success story it is being asked to predict, SDR\-Arena serves agents a frozen view of the public web at the historical timestamp of each evaluation instance\.
- •We apply the framework to proprietary sales\-email and sales\-transcript corpora from a Fortune 100 tech company and a mid\-sized healthcare firm, comprising approximately 115,000 filtered outreach emails and 5435 outreach calls by 124 SDRs labeled by whether they induced a successful receiver action\.
- •We validate the framework through field deployment with 12 professional sales development representatives across our partner enterprises and a gold\-standard exercise with senior SDRs from five enterprises
Across frontier LLMs and open\-source deep\-research agents, including STORM \(Shaoet al\.\([2024](https://arxiv.org/html/2607.20471#bib.bib25)\)\), ODR \(LangChain \([2025](https://arxiv.org/html/2607.20471#bib.bib39)\)\), GPT\-4o \(OpenAI \([2023](https://arxiv.org/html/2607.20471#bib.bib40)\)\), Claude Sonnet 4\.6 \(Anthropic \([2026](https://arxiv.org/html/2607.20471#bib.bib43)\)\) and Qwen\-2\.5 \(Yanget al\.\([2024](https://arxiv.org/html/2607.20471#bib.bib41)\)\), we observe a consistent personalization plateau\. Alignment scores cluster in the 30 to 43 percent range, and on the tech\-firm cohort no model statistically separates successful from unsuccessful outreach\. Specialized agents such as STORM reach the upper end of the range, but at one to two orders of magnitude greater inference cost; standard LLMs with temporally constrained search occupy a more compute\-efficient frontier\.
Figure 2:Sales Journey: From Prospecting to Outreach to Call and eventual Deal Closure leading to Success Story publication
## 2Problem Formulation
We formalize the empirical sales in the form of generative personalization by first describing the Sales Development Lifecycle, then casting personalization as a Bayesian Persuasion task where the generated outreach is aimed at inducing specific actions from the recipients within the sales funnel\.
### 2\.1Sales Development Lifecycle
A sales journey \(illustrated in Fig\.[2](https://arxiv.org/html/2607.20471#S1.F2)\) begins with Sales Development Representatives \(SDRs\) researching prospective accounts to identify needs and budget signals, then sending tailored outreach emails to schedule an initial call\. The call further develops the prospect’s needs and progresses toward deal closure, with some opportunities materializing into deals and others not\. Following a successful closure, the workflow often culminates in aCustomer Success Story: a publicly documented case study published by the Seller company showcasing how their products helped a customer overcome key challenges\. Released as web articles, these stories validate the partnership by documenting the transition from a ‘pain state’ to a ‘success state’ \(Terhoet al\.\([2022a](https://arxiv.org/html/2607.20471#bib.bib5)\)\)\. Examples include success stories from[Oracle](https://www.oracle.com/cloud/technical-case-studies/careem/),[Salesforce](https://www.salesforce.com/resources/customer-stories/snapology/), and[Adobe](https://business.adobe.com/customer-success-stories/marriott-case-study.html)\.
### 2\.2Bayesian Persuasion Formulation
We adapt the Bayesian Persuasion framework of Kamenica \(Kamenica and Gentzkow \([2011](https://arxiv.org/html/2607.20471#bib.bib22)\)\) to model personalized outreach as a signaling game between a Sender \(the SDR or the LLM agent that replaces the SDR\) and a Receiver \(the prospect\)\. This framework naturally aligns for personalization because it acknowledges that the receiver enters the interaction with a prior belief about their own needs, and the sender’s role is to provide a signal, the personalized message, that updates the receiver’s posterior in favor of action\.
The unobserved receiver stateω\\omegarepresents the latent compatibility between the receiver’s requirements and the sender’s product\. We decomposeω\\omegaat timettas a tupleωt=\{ni,wi\}\\omega\_\{t\}=\\\{n\_\{i\},w\_\{i\}\\\}wherenndenotes the receiver’s explicit needs \(functional requirements, current pain points\) andwwdenotes their latent wants \(strategic goals, avenues for value generation\)\. The receiver chooses an actiona∈\{0,1\}a\\in\\\{0,1\\\}, wherea=1a=1represents a successful transition to the next stage of the sales funnel \(e\.g\., an email leading to a call, or a call leading to deal discussions\), whilea=0a=0represents a failure to progress\.
Following Kamenica \(Kamenica and Gentzkow \([2011](https://arxiv.org/html/2607.20471#bib.bib22)\)\), the Receiver is treated as a rational Bayesian agent with a prior beliefμ\\muoverω\\omega\. The Receiver takes actiona=1a=1if and only if their expected utilityuRu\_\{R\}, conditional on their belief, exceeds a reservation thresholdτ\\tau:
𝔼\[uR\(a=1,ωt,ξ\)∣μ\]≥τ\\mathbb\{E\}\[u\_\{R\}\(a=1,\\omega\_\{t\},\\xi\)\\mid\\mu\]\\geq\\tau
The Sender’s utilityuSu\_\{S\}is aligned with the Receiver actinga=1a=1\. Thus, the Sender’s goal is to deliver a signal that updates the receiver’s posterior belief such that the above condition is satisfied\. Note that the utility function is also subject to exogeneous factorsξ\\xi\(timing, organizational urgency, prior context, noise\) that are independent of the sender’s signal but contribute to the receiver’s utility\. The sender’s signal can shiftμ\\muin favor of acting; it cannot controlξ\\xi\. Consequently, even an optimal signal is not guaranteed to inducea=1a=1, and the receiver’s action is best understood as a probabilistic outcome whose likelihood the sender attempts to maximize\.
The key adaptation for our setting is that the sender does not observeωt\\omega\_\{t\}directly\. The sender \(in our case an LLM agentΦ\\Phi\) operates on an observable contextWtW\_\{t\}\(state of the world at time t\) consisting of public information available at timett, from which the latent state must be inferred\. From this inference, the agent generates an outreachOO, which we represent structurally as a list of pitch points\. Each pitch point is a specific argument linking the seller’s product to one of the receiver’s inferred needs or wants through a particular value proposition\. Formally, the agent’s policy is a mapping:
O=\{pp1,pp2,…,ppk\},O^=\{pp^1,…,pp^n\}=Φ\(Wt,P∣ω^\)O=\\\{pp\_\{1\},pp\_\{2\},\\dots,pp\_\{k\}\\\},\\\\ \\hat\{O\}=\\\{\\hat\{pp\}\_\{1\},\\dots,\\hat\{pp\}\_\{n\}\\\}=\\Phi\(W\_\{t\},P\\mid\\hat\{\\omega\}\)whereω^\\hat\{\\omega\}is the agent’s inferred receiver state,O^\\hat\{O\}is the outreach generated by agentΦ\\PhiandPPis the product\. The generated list serves as the informational signal intended to update the receiver’s posterior such that the probability of the desired stage transition \(a=1a=1\) is maximized\.
### 2\.3Evaluation Methodology and Proxy
The ideal objective of the Sender is to generate an outreachOOthat maximizes the Receiver’s expected utility:O∗=argmaxO𝔼\[uR\(a=1,ω\)∣O\]O^\{\*\}=\\operatorname\*\{arg\\,max\}\_\{O\}\\mathbb\{E\}\[u\_\{R\}\(a=1,\\omega\)\\mid O\]\. Direct evaluation of this objective is intractable as the receiver’s utility functionuRu\_\{R\}and the true stateω\\omegaare unobservable to the agent \(and to the benchmark\)\. We therefore replace the unobservable utility with an observable proxy \- the informational overlap between the agent’s generated content and the content implicit in a human\-authored message that is known to have induced the desired action\.
We utilize a dataset of successful historical outreach attempts across different engagement stages \(Emails, Calls, and Success Stories\)\. Let𝒟=\{\(Ci,ωi,Oi∗\)\}i=1N\\mathcal\{D\}=\\\{\(C\_\{i\},\\omega\_\{i\},O^\{\*\}\_\{i\}\)\\\}\_\{i=1\}^\{N\}be a set of ground truth examples whereOi∗O^\{\*\}\_\{i\}is a human\-authored message that successfully induced actiona=1a=1\(moving to the next funnel stage\)\.Success in email is defined by an outreachOi∗O^\{\*\}\_\{i\}that led to a call; a call success impliesOi∗O^\{\*\}\_\{i\}led to deal discussions; and a success story impliesOi∗O^\{\*\}\_\{i\}led to deal closure\. Because eachOi∗O^\{\*\}\_\{i\}resulted in a positive outcome, it empirically satisfies the receiver’s utility threshold and can be treated as a sample from the set of utility\-maximizing messages for the corresponding \(sender, receiver, product, time\) tuple\. We extract a list of ground truth pitch pointsV∗V^\{\*\}fromOi∗O^\{\*\}\_\{i\}and define the Weighted Coverage Score \(WCS\) as the semantic alignment between the predicted pitch pointsO^\\hat\{O\}and the ground truth pointsV∗V^\{\*\}:
Weighted Coverage Score=𝒮\(O^,V∗\)\\text\{Weighted Coverage Score\}=\\mathcal\{S\}\(\\hat\{O\},V^\{\*\}\)
whereS\(⋅,⋅\)S\(\\cdot,\\cdot\)is a semantic alignment scoring function \(defined in Section[3\.1](https://arxiv.org/html/2607.20471#S3.SS1)\)\.
This allows us to rigorously benchmark Agent performance by measuring theRelevance Alignmentbetween the predicted pitch points \(where the value proposition is embedded\) against those extracted from the successful ground truth\. This serves as a tractable proxy for the Bayesian persuasion objective: a higher matching score implies the Agent has successfully identified the winning strategy that induces the desired action\.
WCS is, by construction, alower boundon personalization quality\. It captures thewhatcomponent of personalization — whether the agent has correctly inferred the receiver’s decision\-relevant needs and wants and identified the value\-generating arguments that historically induced the desired action while abstracting away thehow\(style, tone, formatting\)\. An agent that achieves high WCS has demonstrated that it can recover the receiver\-specific strategic content of a known successful message; an agent that achieves low WCS has not, regardless of how well\-written its output is\.
### 2\.4Dataset Construction
To evaluate generative personalization in real\-world settings, we construct two complementary datasets: \(i\)SDR\-Bench, a large\-scale public benchmark derived from enterprise sales artifacts, and \(ii\) aprivate enterprise datasetcontaining real sales outreach emails and downstream sales outcomes from two organizations\. Together, these datasets provide both reproducible public evaluation and high\-fidelity validation on real\-world personalized communication\.
##### SDR\-Bench Dataset\.
To enable reproducible public benchmarking of generative personalization systems, we construct SDR\-Bench, a large\-scale corpus of publicly available enterprise sales narratives and customer success stories\. We targeted approximately 12,000 global enterprises with revenues exceeding $1B, identifying sitemaps for 8,298 organizations and collecting 117,000 candidate URLs using heuristics tailored to common success\-story paths \(e\.g\.,/customer\-stories,/case\-study\)\.
A multi\-stage filtering pipeline \(Table[3](https://arxiv.org/html/2607.20471#A1.T3)\) removed non\-text formats, generic landing pages, articles lacking verifiable publication dates or identifiable product solutions, and anonymized stories where the customer organization was not explicitly named \(e\.g\., “a large food products company”\)\. This process yielded a final corpus of 6,279 success\-story articles spanning 22 industries\. Distributions of companies and stories by industry are shown in Fig\.[4](https://arxiv.org/html/2607.20471#S2.F4)and Appendix Fig\.[6](https://arxiv.org/html/2607.20471#A1.F6)\. Detailed construction steps are provided in Appendix[A\.1](https://arxiv.org/html/2607.20471#A1.SS1)\.
##### Private Enterprise Dataset\.
To validate our theoretical proxy, we require settings where the ground\-truth messageOi∗O^\{\*\}\_\{i\}and its successful outcome\(a=1\)\(a=1\)are explicitly observed\. We therefore collaborated with two enterprises—a Fortune 100 technology company and a mid\-sized healthcare firm—to collect real human\-authored sales outreach paired with downstream prospect actions\.
For the Fortune 100 company, we collected approximately 100k outreach emails authored by 124 SDRs and 5,435 sales call transcripts over a two\-year period \(2023–2025\), identifying 13,236 instances in which the outreach successfully induced a sales call\. For the healthcare firm, we analyzed 24,506 outreach emails, of which 354 resulted in a scheduled sales call\. These successful outreach instances serve as observed realizations of optimal messages \(O∗O^\{\*\}\) in our relevance\-alignment framework\. Table[1](https://arxiv.org/html/2607.20471#S2.T1)summarizes the dataset construction pipeline\.
##### Human Personalization Strategies\.
To characterize the qualitative structure of expert personalization, we analyzed the strategies employed by SDRs across both datasets \(Fig\.[3](https://arxiv.org/html/2607.20471#S2.F3)\)\. The three most common strategies were: \(i\)industry\-based personalization, tailoring content to sector\-specific trends and pain points; \(ii\)persona\-based personalization, adapting the value proposition to the recipient’s organizational role; and \(iii\)activity\-based personalization, leveraging behavioral signals such as webinar attendance or prior engagement\.
Appendix Fig\.[5](https://arxiv.org/html/2607.20471#A1.F5)provides qualitative examples showing how SDRs adapt the same product positioning differently across recipients with distinct inferred needs \(nin\_\{i\}\) and wants \(wiw\_\{i\}\)\.
Table 1:Processing of Enterprise Sales Email DataMetric / Artifact CategoryHealthcareTechNumber of SDRs3124Total Emails Collected48,150609,191Deduplication31,034186,379Sales Outreach Emails24,50690,809Sales Call scheduled35413,236Golden dataset handpicked400400
Figure 3:Distribution of count of strategies across a random subset of 34,000 emails
Figure 4:Distribution of count of companies by Industry Type
## 3SDR\-Arena
We introduceSDR\-Arena, a scalable framework designed to systematically benchmark LLM\-based agents on generative personalization over sales outreach artifacts\. To ensure a rigorous and valid evaluation, the arena utilizes an isolated environment that provides agents access to aHistorical Internet Simulator\(WtW\_\{t\}\)\. The arena serves as a standardized testbed for comparing diverse agentic workflows, ranging from complex research pipelines to simple tool\-use configurations\. We evaluate two primary configurations on SDR\-Bench and our Enterprise Dataset:
- •LLMs \+ Web Search:A baseline equipping frontier models with standard search tools to measure the marginal utility of agentic workflows against simpler tool\-use capabilities\.
- •Deep Research Agents:Specialized agents that produce comprehensive research via multi\-turn conversation and broad search retrieval over the internet \(Shaoet al\.\([2024](https://arxiv.org/html/2607.20471#bib.bib25)\); LangChain \([2025](https://arxiv.org/html/2607.20471#bib.bib39)\)\)\.
Historical Internet Simulator:This environment prevents “future leakage” by enforcing a strict temporal boundary, ensuring agents only synthesize information that was publicly available at the simulated time of the sales interaction\. The system enforces theWtW\_\{t\}boundary by passing search\_start\_date and search\_end\_date parameters to the BrightData SERP API \(Bright Data \([2026](https://arxiv.org/html/2607.20471#bib.bib42)\)\)\. By restricting results to timett, we ensure that the generated pitch points are constructed solely from context that would have been accessible to a human researcher at the time of the original sales event, preventing the model from ‘cheating’, where an agent might mistakenly find the successful outcome of a deal that hasn’t happened yet in the simulation\.
### 3\.1Evaluation Framework
We define each evaluation instance as a tuple\(S,C,P,t\)\(S,C,P,t\), whereSSis the seller,CCis the prospect,PPrepresents the products, andttis the historical timestamp\. The tuple is extracted from each sales artifact individually\. Please refer Appendix[A\.4](https://arxiv.org/html/2607.20471#A1.SS4)for examples\.
Implementation:The agent is prompted to act as a sales representative forSSpitchingPPtoCCusing the time\-restricted search tool\. The resulting outputO^=\{pp^1,…,pp^n\}\\hat\{O\}=\\\{\\hat\{pp\}\_\{1\},\\dots,\\hat\{pp\}\_\{n\}\\\}is a set of personalized pitch points intended to address the inferred needs and strategic goals of the prospect\. We employ an LLM\-based semantic judge to extract ground truth pitch pointsV∗V^\{\*\}from the historical sales artifact\. We use the raw content of the sales artifact and employ GPT\-4oOpenAI \([2023](https://arxiv.org/html/2607.20471#bib.bib40)\)to perform an ontological extraction of ‘Pitch Points\.’ Each pitch point is required to follow a stricttriad structure:Product/Service→\\rightarrowSpecific Pain Point→\\rightarrowValue Proposition/Mechanism\. To ensure the pitch points are grounded, the extraction model was instructed to provide and validate pitch points with exact ‘evidence quotes’ from the source text for every claim\. An expert study verified the precision and recall of this extraction to be0\.920\.92and0\.970\.97respectively, showing strong alignment with expert judgment\. Refer to App\. Section[A\.8](https://arxiv.org/html/2607.20471#A1.SS8)for more details\.
For eachpp∈V∗pp\\in V^\{\*\}, the judge evaluates whether the agent’s outputO^\\hat\{O\}successfully covered the point\. Performance is measured by theCoverage Score, defined as the fraction of ground truth strategic value propositions successfully recovered by the agent\. We employ aCoverage Judgerelying on a5\-point Likert scalethat grades Sales Effectiveness and Factual Precision, ranging from0 \(Miss / Irrelevant\)through1 \(Marketing Fluff\),2 \(Topic Match\),3 \(Implied / Soft Match\),4 \(Strong Sales Argument\), up to5 \(Strategic Bullseye\)\- A perfect extraction that captures the exact pain point of the recipient and the specific mechanism the product provides to address it\. The full scoring rubric and the judge prompt are in the Appendix\.
Weighted Coverage Score \(WCS\):This metric normalizes the Likert\-scale evaluations into a percentage representing the agent’s*completeness*in capturing the winning sales logic\. For a given success story withNNground truth pitch points, letsi∈\{0,…,5\}s\_\{i\}\\in\\\{0,\\dots,5\\\}be the score assigned by the judge for theii\-th point\. The WCS is calculated as:
WCS=\(∑i=1Nsi5N\)×100%\\text\{WCS\}=\\left\(\\frac\{\\sum\_\{i=1\}^\{N\}s\_\{i\}\}\{5N\}\\right\)\\times 100\\%A score of 100% implies that the agent successfully predicted every critical deal\-winning argument with maximum specificity\. It is important to note that this is aprediction taskrather than a retrieval task: the ground truth serves as a*future artifact*, and agents must predict these winning points using only historical data available at timett\. This metric measures the semantic alignment \(Section[2\.3](https://arxiv.org/html/2607.20471#S2.SS3)\) of agent outputs against thesefuture artifactsresulting in a realistic back testing scenario\.
Table 2:Ground truth artifact evaluation\. Values reported as Coverage\.Enterprise Sales EmailsSuccess StoriesAverage Cost per QueryModelHealthcare \(<1B\)Tech \(\>10B\)OtherTech\.Mfg\.EnergyITAgg\.PromptCompletionCostUnsuccSuccUnsuccSuccSales Call Transcripts\(30\)\(30\)\(30\)\(30\)\(180\)TokensTokens\($\)STORM\-QWEN\-2\.522\.4632\.2743\.1539\.2430\.4343\.4041\.5939\.2444\.6042\.51∼\\sim29k∼\\sim6\.3k∼\\sim0\.135ODR\-QWEN\-2\.515\.4030\.4138\.5139\.8222\.5530\.1630\.9535\.2835\.3333\.53∼\\sim66k∼\\sim8\.6k∼\\sim0\.250QWEN2\.5\-72B \(WEB\)32\.1136\.7239\.5336\.4325\.3432\.0936\.7537\.0438\.2136\.84∼\\sim5\.6k∼\\sim0\.25k∼\\sim0\.002Claude Sonnet 4\.6 \(WEB\)59\.8960\.8964\.0861\.3334\.5252\.660\.55454\.955\.8∼\\sim79\.2k∼\\sim2\.2k∼\\sim0\.270GPT\-4o \(WEB\)35\.7140\.6247\.4445\.17—33\.5038\.9332\.2636\.0135\.42∼\\sim9\.2k∼\\sim0\.4k∼\\sim0\.027GPT\-4o\-mini \(WEB\)36\.1639\.1448\.5145\.63—33\.3038\.5437\.0738\.0537\.46∼\\sim12\.2k∼\\sim0\.6k∼\\sim0\.002GPT\-5\.4\-mini \(WEB\)39\.2843\.6754\.6452\.66—40\.9946\.0241\.1845\.3944\.63∼\\sim7\.6k∼\\sim0\.7k∼\\sim0\.009GPT\-5\.4 \(WEB\)39\.5748\.7153\.8053\.02—38\.2145\.4642\.9046\.6744\.32∼\\sim12\.1k∼\\sim0\.9k∼\\sim0\.044
## 4Results & Analysis
We evaluate models across two categories: frontier LLMs augmented with the temporally\-restrictedSDR\-Arenaweb\-search tool, comprisingClaude Sonnet 4\.6,GPT\-4o,GPT\-4o\-mini,GPT\-5\.4,GPT\-5\.4\-miniandQWEN\-2\.5\-72Band deep research agents,STORMandODR, both built onQWEN\-2\.5\-72B\. These configurations are evaluated across three corpora\. The first is a public corpus of curated customer success stories partitioned by industry:Technology,Manufacturing,Energy, andIT, with an aggregate set of 180 stories\. The second is a corpus of transcripts of sales calls from a company exceeding $10B in revenue\. The third is a corpus of human\-authored enterprise sales emails, divided into two cohorts: aHealthcarecompany with under $1B in revenue and aTechnologycompany exceeding $10B in revenue, each containing 200 successful and 200 unsuccessful emails\.
### 4\.1Discussion of Empirical Findings
We observe several notable trends\. First,Claude Sonnet 4\.6leads all models with an aggregate WCS of 55\.8 on the public success story dataset, representing a meaningful gap above the next best model, GPT\-5\.4\-mini at 44\.63\. Despite this it only recovers roughly half of the strategic content of the human\-authored success story, indicating a clear personalization plateau across all agent families\. Second, frontier LLMs are a more cost\-efficient alternative to deep research agents\. Claude Sonnet 4\.6 achieves the highest WCS at an inference cost comparable to ODR \( $0\.270 vs\. $0\.250\), while surpassing it by more than 20 WCS points\.
TheEnterprise Sales Emailcohorts \(Table[2](https://arxiv.org/html/2607.20471#S3.T2)\) reveal a sector\-dependent pattern\. In theHealthcarecohort, models more consistently assign higher scores to successful outreach than unsuccessful outreach \(e\.g\., STORM:32\.2732\.27vs\.22\.4622\.46\), suggesting they capture personalization cues relevant to specialized, high\-stakes sectors\. In theTechnologycohort, however, this pattern inverts or collapses: several models score unsuccessful emails comparably to or higher than successful ones \(e\.g\., STORM:43\.1543\.15vs\.39\.2439\.24\), indicating that models generate coherent but strategically shallow content insufficient to drive real\-world revenue in competitive markets\.
##### Pre\-training leakage is not driving WCS
A complementary concern is that publicly indexed success stories inSDR\-Benchmay have appeared in LLM pre\-training corpora, inflating WCS through memorization rather than genuine inference\. To probe this, we partition SDR\-Bench by article publication date and re\-evaluate on pre\-2024 vs\. post\-2024 cohorts; since GPT\-4o’s training cutoff sits between Q4\-2023 and early 2024, post\-2024 stories are unlikely to have been seen during pre\-training\. We observe negligible WCS differences across the split \(STORM:0\.420\.42vs\.0\.430\.43; GPT\-4o:0\.360\.36vs\.0\.360\.36\)\. The absence of pre\-training\-era inflation indicates that performance onSDR\-Benchreflects context\-conditioned synthesis, not retrieval of memorized content\.
### 4\.2Human Alignment and Validation
To ensure that our automated metrics reliably reflect real\-world quality, we conducted expert studies calibrating our Coverage Judge against independent human raters, and confirming practical utility through deployment with professional sales representatives\.
To show that the Coverage Judge follows human judgement, we conducted a human study on 20 success stories \(80 model responses across STORM, ODR, GPT, and Qwen\)\. Three independent human annotators, blinded to model identities, scored coverage following the exact protocol of our LLM\-based Coverage Judge\.
The study yields three convergent signals on Judge fidelity\. \(i\) The Judge tracks human scores with strong rank correlation \(Spearman’sρ=0\.7435\\rho=0\.7435,p<0\.0001p<0\.0001\), holding across models \(ODR:0\.77680\.7768; GPT\-4o:0\.75750\.7575; QWEN\-2\.5:0\.73300\.7330; STORM:0\.69630\.6963\)\. \(ii\) The Judge*preserves model ordering*: both human\- and Judge\-graded WCS rank STORM\>\>GPT\-4o\>\>ODR, so absolute\-score differences do not distort comparative conclusions\. \(iii\) The Judge is systematically more conservative than human raters \(STORM:48\.0648\.06vs\.55\.2955\.29; GPT\-4o:37\.9437\.94vs\.46\.0446\.04; ODR:31\.1031\.10vs\.41\.6941\.69\), ruling out score inflation and establishing WCS as a rigorous lower bound that tracks human intuition at scale\.
We show that the WCS\-based ranking transfers to expert SDR judgment in two field studies with 12 senior SDRs from the partner enterprises whose data appears in this paper\. The first measures*per\-pitch usefulness*\- whether any individual model\-generated pitch point would be used verbatim in real outreach, and the second measures*strategy\-level overlap*between model output and SDR\-authored gold standards\. Together they probe the two granularities at which an automated score can mismatch expert judgment: per\-point quality and overall strategic match\.
- •Per\-pitch usefulness:Twelve SDRs used GPT\-4o\+\+SDR\-Arena to generate pitch points for200\+200\+new prospect companies inside their normal outbound pipeline, with each SDR auditing the model output on accounts they were actively working\. For every generated pitch point, the SDR rated, on a binary criterion, whether it both \(a\) reflected genuine understanding of the prospect’s pain points and \(b\) was usable in outreach without rewriting;48\.2%of pitch points met both criteria\.The field rate corresponds to roughly half of agent output being expert\-grade in live deployment, with the remaining points being factually accurate but strategically generic, directly consistent with the personalization plateau identified above\.
- •Gold\-Standard Alignment:We asked senior SDRs \(≥10\\geq 10years of industry experience and≥5\\geq 5years at the firm\) from five enterprises to independently author reference “gold\-standard” strategies to pitch 30 products to 5 prospects each\. The SDRs were not shown any model output during the exercise, so the reference strategies are an independent expert read of what*should*be pitched\. We then computed the overlap between the gold\-standard strategy and the outputs of the four benchmarked models, and correlated this expert\-overlap score against the corresponding automated WCS on SDR\-Bench\. Across models, expert overlap and WCS track at Pearsonr=0\.816r=0\.816\. The WCS\-based ranking therefore transfers to senior\-SDR judgment without re\-tuning the rubric per enterprise, supporting WCS as a calibrated proxy for whether an agent has identified the strategic content a domain expert would pitch\.
Together, these studies establish that our pipeline’s outputs are both factually grounded and meaningful in live sales contexts\.
We also evaluate a closed source deep research agent,GEMINI\-2\.5\-PRO\-DRon a separate 25\-story subset\. Its Deep Research API does not expose temporal\-restriction parameters and its higher inference cost precludes broader evaluation\. On this subset, it achieves a WCS of62\.63\. Notably, the margin between this score and that of Claude Sonnet 4\.6 with web search remains narrow, further underscoring that frontier LLMs with web search constitute a cost\-effective alternative to deep research agents\.
## 5Conclusion
In this work, we introducedSDR\-Arena, the first comprehensive framework for benchmarking the generative personalization capabilities of Large Language Models\. By grounding our evaluation in theBayesian Persuasionframework, we transitioned from subjective assessments of "quality" to a rigorous measure ofRelevance Alignment\. Our experiments utilizeSDR\-Bench—a novel, high\-fidelity corpus of over 6,200 success stories—and a unique enterprise\-scale dataset of successful sales outreach to quantify how effectively LLMs can synthesize winning strategic arguments\.
Our findings reveal a significant “personalization plateau\.”, showing a substantial gap remains between AI\-generated outreach and human\-level strategic proficiency\.
By releasingSDR\-Arena, we provide the research community with the tools necessary to study autonomous personalization while strictly controlling for data leakage\. As LLMs continue to move into high\-stakes business operations, we hope this framework serves as a foundation for developing AI agents that are not only persuasive but verifiably aligned with the nuanced needs of their human recipients\.
## References
- Persobench: benchmarking personalized response generation in large language models\.arXiv preprint arXiv:2410\.03198\.Cited by:[§1](https://arxiv.org/html/2607.20471#S1.p3.1)\.
- Anthropic \(2026\)Claude sonnet 4\.6\.Note:[https://www\.anthropic\.com/claude/sonnet](https://www.anthropic.com/claude/sonnet)Accessed: 2026\-05\-07Cited by:[§1](https://arxiv.org/html/2607.20471#S1.p6.1)\.
- Bright Data \(2026\)Note:Accessed: 2026\-05\-07External Links:[Link](https://brightdata.com/products/serp-api)Cited by:[§3](https://arxiv.org/html/2607.20471#S3.p2.2)\.
- J\. Chen, Z\. Liu, X\. Huang, C\. Wu, Q\. Liu, G\. Jiang, Y\. Pu, Y\. Lei, X\. Chen, X\. Wang,et al\.\(2024\)When large language models meet personalization: perspectives of challenges and opportunities\.World Wide Web27\(4\),pp\. 42\.Cited by:[§1](https://arxiv.org/html/2607.20471#S1.p3.1)\.
- H\. Choi, C\. Mela, S\. Balseiro, and A\. Leary \(2020\)Online display advertising markets: a literature review and future directions\.Information Systems Research31,pp\.\.External Links:[Document](https://dx.doi.org/10.1287/isre.2019.0902)Cited by:[§1](https://arxiv.org/html/2607.20471#S1.p2.1)\.
- E\. Durmus, L\. Lovitt, A\. Tamkin, S\. Ritchie, J\. Clark, and D\. Ganguli \(2024\)External Links:[Link](https://www.anthropic.com/news/measuring-model-persuasiveness)Cited by:[§1](https://arxiv.org/html/2607.20471#S1.p3.1)\.
- C\.I\. Hovland, I\.L\. Janis, and H\.H\. Kelley \(1953\)Communication and persuasion: psychological studies of opinion change\.Yale paperbound,Yale University Press\.External Links:LCCN 53007776,[Link](https://books.google.co.in/books?id=ZYizW6_P-goC)Cited by:[§1](https://arxiv.org/html/2607.20471#S1.p1.1)\.
- B\. Jiang, Y\. Yuan, M\. Shen, Z\. Hao, Z\. Xu, Z\. Chen, Z\. Liu, A\. R\. Vijjini, J\. He, H\. Yu,et al\.\(2025\)Personamem\-v2: towards personalized intelligence via learning implicit user personas and agentic memory\.arXiv preprint arXiv:2512\.06688\.Cited by:[§1](https://arxiv.org/html/2607.20471#S1.p3.1)\.
- E\. Kamenica and M\. Gentzkow \(2011\)Bayesian persuasion\.American Economic Review101\(6\),pp\. 2590–2615\.Cited by:[§1](https://arxiv.org/html/2607.20471#S1.p1.1),[§1](https://arxiv.org/html/2607.20471#S1.p3.1),[§2\.2](https://arxiv.org/html/2607.20471#S2.SS2.p1.1),[§2\.2](https://arxiv.org/html/2607.20471#S2.SS2.p3.5)\.
- LangChain \(2025\)Open deep research\.Note:[https://github\.com/langchain\-ai/open\_deep\_research](https://github.com/langchain-ai/open_deep_research)Accessed: 2026\-05\-07Cited by:[§1](https://arxiv.org/html/2607.20471#S1.p6.1),[2nd item](https://arxiv.org/html/2607.20471#S3.I1.i2.p1.1)\.
- H\. D\. Lasswell \(1948\)The structure and function of communication in society\.InThe communication of ideas,Vol\.37,pp\. 215–228\.Cited by:[§1](https://arxiv.org/html/2607.20471#S1.p1.1),[§1](https://arxiv.org/html/2607.20471#S1.p3.1)\.
- L\. Li, P\. Cai, R\. A\. Rossi, F\. Dernoncourt, B\. Kveton, J\. Wu, T\. Yu, L\. Song, T\. Yang, Y\. Qin,et al\.\(2025\)A personalized conversational benchmark: towards simulating personalized conversations\.arXiv preprint arXiv:2505\.14106\.Cited by:[§1](https://arxiv.org/html/2607.20471#S1.p3.1)\.
- S\. Matz, S\. Vaid, H\. Peters, G\. Harari, and M\. Cerf \(2024\)The potential of generative ai for personalized persuasion at scale\.Scientific Reports14,pp\.\.External Links:[Document](https://dx.doi.org/10.1038/s41598-024-53755-0)Cited by:[§1](https://arxiv.org/html/2607.20471#S1.p3.1)\.
- OpenAI \(2023\)GPT\-4 technical report\.arXiv preprint arXiv:2303\.08774\.External Links:[Link](https://arxiv.org/abs/2303.08774)Cited by:[§1](https://arxiv.org/html/2607.20471#S1.p6.1),[§3\.1](https://arxiv.org/html/2607.20471#S3.SS1.p2.9)\.
- R\. E\. Petty and J\. T\. Cacioppo \(1986\)The elaboration likelihood model of persuasion\.L\. Berkowitz \(Ed\.\),Advances in Experimental Social Psychology, Vol\.19,pp\. 123–205\.External Links:ISSN 0065\-2601,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/S0065-2601%2808%2960214-2),[Link](https://www.sciencedirect.com/science/article/pii/S0065260108602142)Cited by:[§1](https://arxiv.org/html/2607.20471#S1.p1.1)\.
- F\. Ricci, L\. Rokach, and B\. Shapira \(2010\)Introduction to recommender systems handbook\.InRecommender systems handbook,pp\. 1–35\.Cited by:[§1](https://arxiv.org/html/2607.20471#S1.p2.1)\.
- A\. Salemi, S\. Mysore, M\. Bendersky, and H\. Zamani \(2024\)Lamp: when large language models meet personalization\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 7370–7392\.Cited by:[§1](https://arxiv.org/html/2607.20471#S1.p3.1)\.
- Y\. Shao, Y\. Jiang, T\. Kanell, P\. Xu, O\. Khattab, and M\. Lam \(2024\)Assisting in writing Wikipedia\-like articles from scratch with large language models\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),K\. Duh, H\. Gomez, and S\. Bethard \(Eds\.\),Mexico City, Mexico,pp\. 6252–6278\.External Links:[Link](https://aclanthology.org/2024.naacl-long.347/),[Document](https://dx.doi.org/10.18653/v1/2024.naacl-long.347)Cited by:[§1](https://arxiv.org/html/2607.20471#S1.p6.1),[2nd item](https://arxiv.org/html/2607.20471#S3.I1.i2.p1.1)\.
- S\. Sharma, P\. Mittal, M\. Kumar, and V\. Bhardwaj \(2025\)The role of large language models in personalized learning: a systematic review of educational impact\.Discover Sustainability6\(1\),pp\. 1–24\.Cited by:[§1](https://arxiv.org/html/2607.20471#S1.p3.1)\.
- O\. Tasdelen and D\. Bodemer \(2025\)Generative ai in the classroom: effects of context\-personalized learning material and tasks on motivation and performance\.International Journal of Artificial Intelligence in Education,pp\. 1–22\.Cited by:[§1](https://arxiv.org/html/2607.20471#S1.p3.1)\.
- H\. Terho, A\. Salonen, and M\. Yrjänen \(2022a\)Toward a contextualized understanding of inside sales: the role of sales development in effective lead funnel management\.Journal of Business and Industrial Marketing38,pp\.\.External Links:[Document](https://dx.doi.org/10.1108/JBIM-12-2021-0596)Cited by:[§2\.1](https://arxiv.org/html/2607.20471#S2.SS1.p1.1)\.
- H\. Terho, A\. Salonen, and M\. Yrjänen \(2022b\)Toward a contextualized understanding of inside sales: the role of sales development in effective lead funnel management\.Journal of Business & Industrial Marketing38\(2\),pp\. 337–352\.External Links:ISSN 0885\-8624,[Document](https://dx.doi.org/10.1108/JBIM-12-2021-0596),[Link](https://doi.org/10.1108/JBIM-12-2021-0596),https://www\.emerald\.com/jbim/article\-pdf/38/2/337/1388064/jbim\-12\-2021\-0596\.pdfCited by:[§1](https://arxiv.org/html/2607.20471#S1.p4.1)\.
- A\. Yang, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Li, D\. Liu, F\. Huang, H\. Wei,et al\.\(2024\)Qwen2\.5 technical report\.arXiv preprint arXiv:2412\.15115\.External Links:[Link](https://arxiv.org/abs/2412.15115)Cited by:[§1](https://arxiv.org/html/2607.20471#S1.p6.1)\.
- S\. Zhang, E\. Dinan, J\. Urbanek, A\. Szlam, D\. Kiela, and J\. Weston \(2018\)Personalizing dialogue agents: I have a dog, do you have pets too?\.InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),I\. Gurevych and Y\. Miyao \(Eds\.\),Melbourne, Australia,pp\. 2204–2213\.External Links:[Link](https://aclanthology.org/P18-1205/),[Document](https://dx.doi.org/10.18653/v1/P18-1205)Cited by:[§1](https://arxiv.org/html/2607.20471#S1.p2.1)\.
- Z\. Zhao, C\. Vania, S\. Kayal, N\. Khan, S\. B\. Cohen, and E\. Yilmaz \(2025\)Personalens: a benchmark for personalization evaluation in conversational ai assistants\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 18023–18055\.Cited by:[§1](https://arxiv.org/html/2607.20471#S1.p3.1)\.
## Appendix AAppendix
Figure 5:Personalization in actual Sales EmailsFigure 6:Distribution of count of Success Stories by Industry TypeFigure 7:Qualitative Example: Ground truth pitch points scored against pitch points generated by the agent### A\.1SDR\-Bench: Dataset Curation Details
Filtration CriteriaCountDomains Found for Companies with over $1B revenue∼\\sim30kDomains Found for B2B Companies with over $1B revenue12,080Companies whose Sitemap could be found8,298Candidate Success Story URLs based on pattern matching∼\\sim117kCount of Companies covering these 117k URLs1,772Exclude non\-text formats \(videos/pdfs\)∼\\sim79kURLs for which content could be collected∼\\sim31kQwen based filtering using content to exclude listicle, parent, generic pages and pages with no publish date∼\\sim7\.2kFiltering out stories where the customer is not a specific business6279Table 3:Filtration Criteria and Counts for Scraped Public Data
### A\.2Sales Emails
#### A\.2\.1Filtering & Analysis
Letℰ=\{e1,e2,…,eN\}\\mathcal\{E\}=\\\{e\_\{1\},e\_\{2\},\\ldots,e\_\{N\}\\\}denote the raw corpus of sales emails\. We apply a three\-stage filtering pipeline:
- •Language Filtering: We remove all non\-English emails using language detection, yieldingℰen⊂ℰ\\mathcal\{E\}\{\\text\{en\}\}\\subset\\mathcal\{E\}\.
- •Email Deduplication: We identify and remove duplicate email templates using a combination of exact matching and fuzzy string comparision yieldingℰdeduplicated⊂ℰen\\mathcal\{E\}\{\\text\{deduplicated\}\}\\subset\\mathcal\{E\}\{\\text\{en\}\}
- •Intent Classification\. We employ an LLM\-as\-a\-judge paradigm to classify emails into outreach versus non\-outreach categories\. Specifically, we filter out generic conversational emails, administrative correspondence, and non\-sales communications\. Let𝒥:e→\{0,1\}\\mathcal\{J\}:e\\rightarrow\\\{0,1\\\}be the LLM judge function where𝒥\(e\)=1\\mathcal\{J\}\(e\)=1indicates a valid sales outreach email\. Our filtered corpus is thus: ℰfiltered=\{e∈ℰdeduplicated:𝒥\(e\)=1\}\\mathcal\{E\}\{\\text\{filtered\}\}=\\\{e\\in\\mathcal\{E\}\{\\text\{deduplicated\}\}:\\mathcal\{J\}\(e\)=1\\\}
For each emaile∈ℰfilterede\\in\\mathcal\{E\}\{\\text\{filtered\}\}, we use an LLM to extract the set of strategies employed in each email:Strat\(e\)⊆𝒮\\text\{Strat\}\(e\)\\subseteq\\mathcal\{S\}\. This allows us to visualize the following patterns:
- •Strategy Frequency Distribution: The distributionP\(s\)P\(s\)over strategies reveals the current state of human personalization practices\.
- •Product\-Conditional Strategies: The distributionP\(s\|Productk\)P\(s\|\\text\{Product\}k\)identifies product\-specific personalization patterns\.
These distributions provide interpretable insights into how human SDRs currently operationalize personalization\.
Beyond strategy classification, we extract fine\-grained pitch points from each email using an LLM\. For each emailee, we extract:
𝐩𝐩\(e\)=\{pp1,pp2,…,ppk\}\\mathbf\{pp\}\(e\)=\\\{pp\_\{1\},pp\_\{2\},\\ldots,pp\_\{k\}\\\}where eachppipp\_\{i\}represents a discrete pitch point used in the outreach\. These pitch points constitute the ground truth against which DR agent outputs are evaluated\.
The email dataset comprises of annotated outreach emails with the following attributes per sample:
- •Target CompanyTiT\_\{i\}: The recipient organization\.
- •Sender’s CompanySiS\_\{i\}: The sender’s organization\.
- •EmailEiE\_\{i\}: The content of the email
- •Timestamptt: Date when the email was sent
- •ProductPP: The solution being pitched
- •Strategy LabelsStrat\(e\)\\text\{Strat\}\(e\): Personalization strategies used in the email
- •Pitch Pointspp\(e\)\{pp\}\(e\): Pitch points used in the email
#### A\.2\.2Personalization Strategies
In order to systematically characterize the various personalization strategies used in the emails, we employed the following two step pipeline:
First,we asked domain experts to manually annotate a seed set of emails to identify recurring personalization patterns\.Second,we used an LLM to extract and cluster strategies from 500 randomly sampled emails, which were then reconciled with expert annotations to produce a unified taxonomy\.
A personalization strategys∈𝒮s\\in\\mathcal\{S\}is a variable representing the primary information source leveraged to establish relevance between the seller’s value proposition and the buyer’s needs \.We define the Personalization Strategy Space𝒮=\{s1,s2,…,s10\}\\mathcal\{S\}=\\\{s\_\{1\},s\_\{2\},\\ldots,s\_\{10\}\\\}consisting of 10 categories:
- •Industry based: References industry\-specific trends, pain points, competitors, or case studies from the target company’s industry\.
- •Event based: Leverages trigger events \(funding rounds, MA, product launches, earnings reports, news mentions\) to identify timely business needs\.
- •Technology based: References the recipient’s current tech stack to propose replacement, integration, or complementary solutions\.
- •Lead Activity\-based: References direct actions by the specific lead \(whitepaper downloads, webinar attendance, pricing page visits, demo interactions\)\.
- •Buying Group Activity\-based: References collective actions by the lead’s team or buying committee\.
- •Geography\-based: Utilizes physical location or regional regulatory context \(e\.g\., GDPR, CCPA compliance requirements\)\.
- •Lead Persona\-based: Explicitly maps the lead’s role, title, or job responsibilities to role\-specific pain points\.
- •Firmographics\-based: Leverages company\-level metrics \(headcount growth, revenue, department size\) as personalization anchors\.
- •Relationship\-based: References existing customer relationships, cross\-sell or upsell opportunities\.
- •None: Generic outreach lacking recipient\-specific context\.
### A\.3How to measure personalization in an ideal world?
Ideally, one could evaluate personalization by observing how the same Receiver responds to multiple personalized signalssis\_\{i\}, where eachsis\_\{i\}is generated by a different LLM, effectively a multiverse of interventions\. By comparing Receiver actions across these interventions, we could directly quantify the personalization abilities of different LLMs\. Because such a multiverse is unavailable in practice, we construct an empirical benchmark using real\-world sales interactions\.
### A\.4Task Formulation Details
For the success story of[Salesforce](https://www.salesforce.com/resources/customer-stories/snapology/), the tuple would be \(S: Salesforce, C: Snapology of Lehi, P: Salesforce Starter, t: 26\-05\-2023\)\.
### A\.5Alignment of LLM with humans for pitch point extraction
TPFPFNPrecisionRecallF1 Score1381130\.920\.970\.95
Table 4:Comparative analysis between LLM extracted pitch points and human annotations on 30 customer success stories
### A\.6Token Usage vs Performance of Agents
![[Uncaptioned image]](https://arxiv.org/html/2607.20471v1/figures/tokens.png)
Figure 8:Graph of token usage vs performance of various agents
Table[5](https://arxiv.org/html/2607.20471#A1.T5)reports the average per\-outreach token consumption and inference cost for each agent configuration on the SDR\-Bench evaluation set, alongside its WCS\. Costs are computed using public list prices for the corresponding model API at the time of evaluation\. Deep\-research pipelines \(STORM, ODR\) consume one to two orders of magnitude more tokens than standard LLM\-plus\-search baselines, while only marginally improving WCS over the latter\. QWEN\-2\.5\-72B is the most cost\-efficient configuration, achieving WCS within∼\\sim5\.7 points of STORM at∼\\sim67×\\timeslower cost\.
Table 5:Per\-outreach inference cost vs\. WCS on the SDR\-Bench evaluation set\.ModelAvg\. Prompt TokensAvg\. Completion TokensAvg\. Inference CostWCSSTORM∼\\sim29k∼\\sim6\.3k∼\\sim$0\.13542\.51QWEN\-2\.5\-72B∼\\sim5\.6k∼\\sim250∼\\sim$0\.00236\.84GPT\-4o∼\\sim9\.2k∼\\sim427∼\\sim$0\.02735\.42GPT\-4o\-mini∼\\sim12\.2k∼\\sim572∼\\sim$0\.00237\.46GPT\-5\.4\-mini∼\\sim7\.6k∼\\sim662∼\\sim$0\.00944\.63GPT\-5\.4∼\\sim12\.1k∼\\sim904∼\\sim$0\.04444\.32ODR∼\\sim66k∼\\sim8\.6k∼\\sim$0\.25033\.53Claude Sonnet\-4\.6∼\\sim79k∼\\sim2\.0k∼\\sim$0\.27055\.80
### A\.7Prompts Library
LLM as a Judge Evaluation PromptSYSTEM: Role:You are a Senior Sales Enablement Evaluator\. Your goal is to determine if an AI agent has successfully extracted the "Winning Pitch Points" from a customer success story\.INPUT DATA: 1\.GROUND TRUTH \(GT\) WINNING POINTS: \(These are the proven, specific facts that won the deal\) <<<gt\_pitch\_points\_str\>\>\>2\.CANDIDATE PITCH \(Predicted\): \(These are the points generated by the AI agent\) <<<candidate\_pitch\_points\_str\>\>\>TASK: For EACH "Ground Truth Point", determine how well the "Candidate Pitch" covers it\. You are grading onSales EffectivenessandFactual Precision\.SCORING RUBRIC \(0\-5 Scale\):•0 \(Miss / Irrelevant\):The candidate pitch completely misses this concept\. No mention of this feature, benefit, or metric\.•1 \(Marketing Fluff\):Vaguely mentions the topic \(e\.g\., "improved efficiency"\) but lacks ANY specific substance\. Critique: "This is a generic platitude that could apply to any company\."•2 \(Topic Match\):Identifies the correct Product or Pain Point, but misses the specific Solution or Outcome\. \*Example: GT says "Reduced downtime by 40%", Candidate says "Helps with downtime\."\*•3 \(Implied / Soft Match\):Captures core value proposition correctly, but misses "Hero Evidence" \(specific numbers, names, or unique mechanisms\)\. \*Verdict: A solid conversational point, but less persuasive than the Ground Truth\.\*•4 \(Strong Sales Argument\):Captures the core value AND the key mechanism/outcome\. It is a persuasive, accurate representation of the deal\. \*Difference from 5: Might miss a minor detail \(e\.g\., date, exact city\) that doesn’t impact sales persuasion\.\*•5 \(Strategic Bullseye\):A perfect extraction\. Captures theProduct \+ Pain \+ Value \+ Specific Metric/Evidenceessentially verbatim from the Ground Truth\. \*Verdict: "This is exactly why they bought\."\*OUTPUT FORMAT \(JSON Only\):``` { "evaluations": [ { "gt_point_id": <int>, "gt_summary": "<short_summary_of_gt_point>", "best_match_candidate_text": "<text_from_candidate_or_null>", "score": <0-5>, "reasoning": "<concise_sales_analysis>" } ] } ```
Pitch Point Generation PromptI am a BDR at\{seller\}\(\{seller\_website if seller\_website else ""\}\)\. I want to sell\{seller\}Products to\{customer\}\(\{customer\_website if customer\_website else ""\}\)\. How should I pitch\{", "\.join\(products\)\}to\{customer\}? I need to generate 3 targeted pitch points and value propositions that address the customer’s pain point which I can send to\{customer\}or their CEO/Leads/Decision Makers\. Make sure to give me reasoning to each pitch point\.Requirements for each pitch point:•Must be specific to the target company \(use real facts you discover\)•Should connect the sender’s products/services to the target’s needs/pain points•Must be atomic \(one clear value proposition per point\)•Should be compelling and actionableAfter researching, generate 3 pitch points which have a single value proposition per point addressing the customer’s pain point\. Each pitch point should be unique and not repeat the same value proposition\. \(it can repeat pitch points for the same product or/and pain point but the value proposition should always be different\)\.Respond with a JSON array of pitch point strings:``` [ "First pitch point text here", "Second pitch point text here", "Third pitch point text here", ... "Tenth pitch point text here" ] ``` Strictly stick to the above JSON format and structure\.
Pitch Point Extraction PromptYou are an expert sales analyst\. Analyze the following success story content for\{company\_name\}\(the seller\) and their product\{product\_name\}\.Story Content: \{story\_html\}Your goal is to extract specific, causal pitch points that explain WHY the product was successful, and provide EXACT QUOTES from the text as evidence\.STRICT FORMAT REQUIREMENT: Return a JSON object with a single key"pitch\_points"containing a list of objects\. Each object must have the following structure:``` { "summary": "<Product/Service> solved <Specific Pain Point> of <Customer> because of <Specific Value Proposition/Mechanism>[, resulting in <Quantitative Result>]", "evidence": [ "Exact quote from the text supporting the pain point...", "Another exact quote..." ] } ``` GUIDELINES:1\.Product/Service: Use the specific product name if available\.2\.Specific Pain Point: What exact problem was the customer facing? Dig deeper than "inefficiency"\. Differentiate between "slow speed/manual work" and "inaccuracy/errors"\. These are DISTINCT pain points\.3\.Specific Value Proposition/Mechanism: What specific feature or capability solved it?4\.Quantitative Result: IF the text mentions a number\(e\.g\., "70% faster", "saved 10 hours"\), YOU MUST INCLUDE IT in the summary\.5\.Evidence: You MUST quote the original text\. Do not paraphrase in this field\.Example Output:``` { "pitch_points": [ { "summary": "The Aldec G3 solved the high energy costs of T.Z. Osborne because of its PowerTubes technology that recovers kinetic energy, reducing consumption by 20%.", "evidence": ["Facing rising energy costs...", "The PowerTubes technology recovers kinetic energy, reducing consumption by 20%"] }, { "summary": "Oracle Integration solved the issue of manual entry errors at Careem because of its automated GRN matching feature.", "evidence": ["Manual entry led to mismatches and compliance gaps.", "Automating GRN matching has significantly reduced errors."] } ] } ```
### A\.8Human Study to validate pitch point extraction
To validate the LLM’s accuracy and exhaustiveness in extracting pitch points, we conducted a human study on a random sample of 30 customer success stories\. Annotators evaluated each story against the LLM\-extracted pitch points along two dimensions: \(1\)Precision— verifying factual consistency and flagging hallucinations, and \(2\)Recall— identifying any pitch points the LLM missed\. The LLM achieved a precision of0\.92, recall of0\.97, and an F1\-score of0\.95, validating its use as a robust, scalable proxy for ground truth extraction\.
### A\.9Institutional Review Board Approval
The human evaluation studies including the per\-pitch usefulness field deployment with 12 sales development representatives and the gold\-standard alignment exercise with senior SDRs from five enterprises were reviewed and approved by the Institutional Review Board\. All participants were informed of the study’s purpose and provided consent prior to participation\. No sensitive personal data was collected beyond professional judgments on model\-generated sales content, and all responses were anonymized prior to analysis\.
### A\.10Statistical Significance and Confidence Intervals
We acknowledge that the evaluation is conducted on a limited subset due to the high computational cost of deep research agents\. To ensure robustness, we perform bootstrapping \(1,000 iterations\) to compute 95% confidence intervals \(CIs\) for the Weighted Coverage Score \(WCS\) across the SDR\-Bench evaluation set\.
Table 6:Bootstrap Estimates of WCS on SDR\-Bench Evaluation SetModelMean WCS95% CISTORM0\.4246\[0\.4015, 0\.4491\]ODR0\.3358\[0\.3150, 0\.3585\]GPT\-4o0\.3638\[0\.3422, 0\.3860\]Qwen0\.3692\[0\.3496, 0\.3876\]Validation of the Personalization Plateau\.The 95% CIs for GPT\-4o \[0\.3422, 0\.3860\] and Qwen \[0\.3496, 0\.3876\] exhibit substantial overlap, indicating no statistically significant difference in performance\. This supports the existence of apersonalization plateau, where different model architectures converge to a similar performance ceiling under our evaluation framework\.
Significance of STORM\.In contrast, the CI for STORM \[0\.4015, 0\.4491\] does not overlap with those of other models, indicating a statistically significant performance improvement\.
Stability of Estimates\.The relatively narrow width of the confidence intervals suggests stable estimates despite the limited sample size\. The evaluation set comprises approximately 720 agent–environment interactions, providing a sufficiently representative estimate of model performance under the SDR\-Arena setup\.
### A\.11Broader Impacts and Ethical Considerations
Our work raises important societal and ethical considerations, which we address below\.
##### Paradox of Measurement and Dual\-Use Risks\.
Benchmarking personalization creates an inherent tension: quantifying what makes sales outreach effective risks providing a blueprint for scalable, manipulative content\. We argue, however, that the absence of transparent evaluation standards poses a greater risk by allowing opaque commercial systems to operate unchecked\.SDR\-Arenaprovides the transparency needed to distinguish context\-aware assistance from hallucinatory or manipulative outreach\.
##### Privacy and Data Stewardship\.
Our dataset curation followed strict ethical guidelines\. The proprietary email dataset was processed in a secure, access\-controlled environment with all PII anonymized or redacted, and isnotincluded in our public release\. The publicSDR\-Benchis limited to already\-published customer success stories, further filtered to enterprise entities to minimize individual exposure\.
##### Economic Displacement and Human\-AI Collaboration\.
Our findings reveal a “personalization plateau,” suggesting LLMs currently lag behind human experts in identifying nuanced, strategic revenue drivers\. This supports aHuman\-in\-the\-Loopparadigm: our benchmark should guide assistants that reduce research drudgery for humans, not autonomous systems that replace human judgment\.
##### Acceptable Use Policy\.
To mitigate the risks of misuse, the release of our framework and theSDR\-Benchdataset will be accompanied by a restrictive Acceptable Use Policy\. This policy explicitly prohibits the use of our artifacts or fine\-tuned models for:
1. 1\.Unsolicited High\-Volume Outreach:Using the dataset to train agents for mass\-spamming or harassment\.
2. 2\.Deceptive Practices:Generating content that masquerades as human correspondence without disclosure\.
3. 3\.Social Engineering:Leveraging the personalization metrics to craft targeted phishing attacks\.
By bringing scientific rigor to sales agent evaluation, we aim to steer the field toward personalization that respects user context and delivers genuine value, rather than optimizing for engagement at the expense of user trust\.Similar Articles
APeB: Benchmarking Personalization Ability of Large Language Model Agents
Introduces APeB, a benchmark for evaluating personalization in LLM agents, focusing on inferring user intent and preferences from raw queries and interaction histories. Finds that current models struggle with early-stage queries and that history-aware refinement can help.
Benchmarking LLMs
A study or report on benchmarking large language models, likely comparing performance across various tasks.
Large Language Models Should Learn Personalized Rather Than Aggregated Human Preferences
This position paper argues that large language models should learn from personalized rather than aggregated human preferences, highlighting theoretical limitations from social choice theory and practical issues from demographic diversity. It proposes bounded personalization frameworks that respect individual autonomy while maintaining universal safety constraints.
The Evaluation Blind Spot: A Stereological Theory of Benchmark Coverage for Large Language Models
This paper applies stereological theory to LLM benchmarks, revealing that current leaderboards measure only 3–5 independent dimensions, creating geometric blind spots that dominate statistical noise. It provides theoretical bounds on benchmark coverage and a submodular algorithm for efficient benchmark selection.
Ψ-Bench: Evaluating Persona-Sensitive Influencing in Persuasive Dialogues
Introduces Ψ-Bench, a benchmark for evaluating LLMs' ability to influence users through persuasive dialogues with personalized profiles. Tests 10 frontier LLMs and finds significant room for improvement, with profile access boosting performance by 18.24%.