CARRE: Counterfactual Action Retrieval and Reason Evaluation for Explainable Churn Prescription

arXiv cs.CL Papers

Summary

CARRE is a three-stage framework that combines retrieval-augmented generation, counterfactual scoring, and large language model reasoning to provide explainable churn prescriptions for customer retention actions, showing improved risk reduction over baselines.

arXiv:2609.09766v1 Announce Type: new Abstract: Churn models typically identify high-risk customers but do not specify which feasible retention action should be considered or why that action is appropriate. We present CARRE (Counterfactual Action Retrieval and Reason Evaluation), a three-stage framework that combines retrieval-augmented candidate generation, cost-aware counterfactual scoring, and large language model (LLM) reasoning. CARRE retrieves a predefined catalog of retention actions, estimates model-predicted churn-risk changes under explicit feature transformations, and generates a structured churn reason and a profile-grounded explanation for the selected action. On the IBM Telco Customer Churn dataset, CARRE achieves 79.8% greater mean model-predicted risk reduction than the plain SHAP baseline and 80.4% greater reduction than the cost-controlled SHAP+Cost baseline across 313 high-risk test cases; its cost-normalized efficiency is 10.5% higher than that of plain SHAP. On a 136-case reason-stratified evaluation sample, diagnosis-driven prompt refinement increases weak-label agreement from 79.4% to 90.4%, with no auxiliary-plan constraint violations; because the same sample was used for error diagnosis and re-evaluation, the post-refinement result is not an independent estimate of generalization. For 135 explanations generated using the pre-refinement v2 reason outputs, two cross-vendor LLM judges assign mean scores ranging from 4.02 to 5.00 out of 5, although one judge saturates on actionability, and a deterministic audit finds no contradictions among 66 verifiable profile claims. Retrieval ablations show that k=5 provides the best evaluated compromise between high candidate coverage and downstream reasoning agreement in this dataset. These results illustrate how retrieval, model-based counterfactual scoring, and language generation can be separated and jointly evaluated in a prototype churn-prescription pipeline.
Original Article
View Cached Full Text

Cached at: 09/10/26, 08:14 AM

# CARRE: Counterfactual Action Retrieval and Reason Evaluation for Explainable Churn Prescription
Source: [https://arxiv.org/html/2609.09766](https://arxiv.org/html/2609.09766)
Conference:5th Workshop on End\-End Customer Journey Optimization; August 9, 2026; Jeju, Republic of KoreaCCS:Computing methodologies Machine learningCCS:Information systems Decision support systemsCCS:Human\-centered computing Natural language interfacesMinJoo KimAffiliation:Department of Industrial Data Engineering,Hanyang University,Seoul,Republic of Koreaemail:[kmj0921@hanyang\.ac\.kr](mailto:[email protected])SangJin ParkNote:Corresponding authors\.Affiliation:Department of Industrial Data Engineering,Hanyang University,Seoul,Republic of Koreaemail:[psj3493@hanyang\.ac\.kr](mailto:[email protected])andSeungHwan ChoAffiliation:Department of Industrial Data Engineering,Hanyang University,Seoul,Republic of Koreaemail:[shcho95@hanyang\.ac\.kr](mailto:[email protected])

2026

###### Abstract\.

Churn models typically identify high\-risk customers but do not specify which feasible retention action should be considered or why that action is appropriate\. We presentCARRE\(CounterfactualActionRetrieval andReasonEvaluation\), a three\-stage framework that combines retrieval\-augmented candidate generation, cost\-aware counterfactual scoring, and large language model \(LLM\) reasoning\. CARRE retrieves a predefined catalog of retention actions, estimates model\-predicted churn\-risk changes under explicit feature transformations, and generates a structured churn reason and a profile\-grounded explanation for the selected action\. On the IBM Telco Customer Churn dataset, CARRE achieves 79\.8% greater mean model\-predicted risk reduction than the plain SHAP baseline and 80\.4% greater reduction than the cost\-controlled SHAP\+Cost baseline across 313 high\-risk test cases; its cost\-normalized efficiency is 10\.5% higher than that of plain SHAP\. On a 136\-case reason\-stratified evaluation sample, diagnosis\-driven prompt refinement increases weak\-label agreement from 79\.4% to 90\.4%, with no auxiliary\-plan constraint violations; because the same sample was used for error diagnosis and re\-evaluation, the post\-refinement result is not an independent estimate of generalization\. For 135 explanations generated using the pre\-refinement v2 reason outputs, two cross\-vendor LLM judges assign mean scores ranging from 4\.02 to 5\.00 out of 5, although one judge saturates on actionability, and a deterministic audit finds no contradictions among 66 verifiable profile claims\. Retrieval ablations show thatk=5k\{=\}5provides the best evaluated compromise between high candidate coverage and downstream reasoning agreement in this dataset\. These results illustrate how retrieval, model\-based counterfactual scoring, and language generation can be separated and jointly evaluated in a prototype churn\-prescription pipeline\.

###### Keywords:

customer churn prediction, counterfactual explanation, retrieval\-augmented generation, prescriptive analytics, large language model, explainability

## 1\.Introduction

The growth of digital transformation and the platform economy has made subscription\-based revenue models ubiquitous across software, telecommunications, media, and financial services\. Unlike traditional markets centered on one\-time purchases, subscription models allow customers to cancel at any time, making churn management a persistent operational challenge\. In the telecommunications industry, annual churn rates are widely reported in the 15–25% range, and acquiring a new customer is commonly estimated to cost several times more than retaining an existing one\([1](https://arxiv.org/html/2609.09766#bib.bib1)\)\. This economic pressure has spurred extensive machine learning research on churn prediction\. A wide range of approaches have been proposed, from logistic regression and decision trees to gradient\-boosted ensembles, with profit\-driven formulations addressing the operational cost of retention\([21](https://arxiv.org/html/2609.09766#bib.bib21)\)\. Boosted variants of these classifiers consistently outperform their non\-boosted counterparts on public telecommunications churn benchmarks\([1](https://arxiv.org/html/2609.09766#bib.bib1)\)\.

With predictive accuracy alone, operational deployment remains fundamentally limited: a churn probability score of 0\.85 tells a customer retention manager only that a customer is likely to leave, providing no guidance on*why*the customer might churn or*which specific actions*could reduce that risk\. This prediction\-prescription gap goes beyond a mere technical limitation and translates into real business losses\. Repeating uniform retention campaigns without identifying the dominant observable churn\-risk pattern yields poor cost efficiency and fails to deliver appropriate interventions to high\-risk customers in time\. In telecommunications in particular, where thousands of at\-risk customers must be handled simultaneously, individualized prescriptions paired with explanatory rationale are essential for frontline agents to make effective decisions\.

Existing approaches address these requirements only in isolation\. Post\-hoc explanation methods such as SHAP \(SHapley Additive exPlanations\) decompose model predictions into feature contributions, identifying why a customer has a high churn probability\([2](https://arxiv.org/html/2609.09766#bib.bib2)\)\. However, feature importances do not translate directly into actionable recommendations\. Knowing that a customer’s month\-to\-month contract is the most negative feature does not reveal whether to offer a contract upgrade, a price discount, or a service bundle\. Counterfactual explanation methods identify minimal profile changes that would flip the model’s decision, but focus on hypothetical feature transformations rather than business\-feasible interventions with defined costs and eligibility constraints\([3](https://arxiv.org/html/2609.09766#bib.bib3),[4](https://arxiv.org/html/2609.09766#bib.bib4)\)\. LLM\-based systems can generate fluent, personalized text, but without grounding in a predefined action space, they tend to produce generic advice or hallucinated interventions that cannot be operationalized\.

A further challenge is*explanation alignment*: the property that a system’s explanations maintain mutual consistency across churn diagnosis, action recommendation, and frontline delivery\. Specifically, a prescription explanation delivered to a customer\-facing agent must simultaneously satisfy three criteria\. First, it must faithfully reflect the heuristic churn\-risk label derived from the customer profile\. Second, it must be logically coherent with the recommended action\. Third, it must be expressed in language that the frontline agent can act upon immediately\. Prior work has addressed each criterion individually\. Work on generated\-text faithfulness targets the first\([8](https://arxiv.org/html/2609.09766#bib.bib8)\), counterfactual recourse research addresses actionability\([9](https://arxiv.org/html/2609.09766#bib.bib9)\), and LLM\-based explanation systems partially achieve coherence\. To our knowledge, few prior systems jointly evaluate profile\-grounded diagnosis, constrained action selection, and agent\-facing explanation within a single churn\-prescription pipeline\.

This paper proposesCARRE, a unified pipeline connecting prediction, prescription, and explanation through three tightly coupled components:

1. \(1\)RAG\-based candidate retrieval: A FAISS index over six business action documents retrieves the top semantically relevant intervention candidates for each customer profile\.
2. \(2\)Counterfactual \(CF\) optimization: A calibrated logistic regression model simulates each retrieved candidate as a counterfactual intervention and scores it by jointly weighing the reduction in predicted churn probability against the normalized operational cost of the action\. The candidate that best balances these two objectives is selected as the prescription\.
3. \(3\)LLM reasoning with CoT: An LLM equipped with structured classification criteria and chain\-of\-thought \(CoT\) prompting identifies the customer’s churn reason and generates a natural\-language explanation aligned with the retrieved action context\.

By separating semantic search \(Stage 1\), cost\-aware optimization \(Stage 2\), and natural\-language explanation \(Stage 3\), each component operates within its own strengths\. CARRE is built on a core design philosophy in which the LLM’s role is*reasoning*, not retrieval or optimization\.

The main contributions of this paper are as follows:

- •An end\-to\-end RAG\+CF\+LLM framework for personalized churn prescription, evaluated through a multi\-layer protocol combining CF metrics with bootstrap confidence intervals, weak\-label agreement, a cross\-vendor two\-judge automated evaluation, and a deterministic grounding check\.
- •A controlled*SHAP\+Cost*baseline showing that CARRE retains an 80\.4% model\-predicted risk\-reduction advantage under identical cost penalization, together with aλ\\lambdasweep revealing a sharp cost\-penalty regime boundary rather than a smooth trade\-off\.
- •An exploratory four\-model comparison revealing substantial model\-specific variation and category\-specific failure patterns; because model scale, vendor, architecture, and inference settings are confounded, the experiment does not isolate the cause of these differences\.
- •A prompt\-development analysis showing that explicit priority rules and action withholding improve instruction\-following fidelity, followed by diagnosis\-driven v3 refinement that raises agreement from 79\.4% to 90\.4% on the same 136\-case reason\-stratified sample; we report this as a post\-refinement development result rather than an independent generalization estimate\.
- •A two\-step Stage 3 design separating reason classification and auxiliary\-plan generation \(Stage 3\-a,cf\_actionwithheld\) from final explanation generation \(Stage 3\-b, given reason andcf\_action\), targeting the explanation–prescription incoherence that arises when the cost\-optimal action and the rule\-consistent one diverge \(exact auxiliary\-plan–action match is 3/135, or 2%, among explained cases\)\.

The remainder of this paper reviews the relevant literature, formalizes CARRE and its action space, describes the experimental protocol, reports quantitative and qualitative results, and concludes with limitations and future directions\.

## 2\.Related Work

### 2\.1\.Churn Prediction, Probability Calibration, and Prescriptive Analytics

A systematic comparison of classifiers—logistic regression, SVM, decision trees, naive Bayes, artificial neural networks, and their boosted variants—on a telecommunications churn dataset showed that boosting consistently improves over the plain classifiers, with an AdaBoost\-boosted polynomial\-kernel SVM attaining the best accuracy and F\-measure\([1](https://arxiv.org/html/2609.09766#bib.bib1)\)\. Purely statistical evaluation has been critiqued for ignoring operational costs, motivating profit\-driven frameworks that jointly account for retention cost and churn risk\([21](https://arxiv.org/html/2609.09766#bib.bib21)\)\. Since high\-performing classifiers often produce scores that are not directly interpretable as probabilities, calibration via isotonic regression or Platt scaling is applied\([12](https://arxiv.org/html/2609.09766#bib.bib12)\); calibrated probabilities are a critical input to CARRE’s Stage 2 CF scoring\.

The prescriptive analytics literature distinguishes*descriptive*,*predictive*, and*prescriptive*systems\([23](https://arxiv.org/html/2609.09766#bib.bib23)\), and most deployed churn systems remain at the predictive layer\. An engagement\-based retention recommender was proposed for e\-commerce, but without explaining why a recommendation was selected or how it addresses each customer’s churn drivers\([22](https://arxiv.org/html/2609.09766#bib.bib22)\)\. CARRE closes this gap by completing the full loop from prediction to prescription to explanation\.

### 2\.2\.Counterfactual Explanations

A counterfactual explanation presents the minimal profile change that would alter the current prediction\. Counterfactual explanations were formalized as minimal feature transformation, generated without accessing model internals\([3](https://arxiv.org/html/2609.09766#bib.bib3)\); actionability constraints were then introduced to keep proposed changes feasible\([9](https://arxiv.org/html/2609.09766#bib.bib9)\), and*algorithmic recourse*extended this with structural causal models and a difficulty\-proportional cost function\([4](https://arxiv.org/html/2609.09766#bib.bib4)\)\. Diverse counterfactual sets exposing the Pareto frontier can also be generated\([19](https://arxiv.org/html/2609.09766#bib.bib19)\)\. CARRE adopts this cost\-aware formulation but restricts recourse to a predefined business action vocabulary rather than arbitrary feature transformations, ensuring operational feasibility for every prescription\.

### 2\.3\.Retrieval\-Augmented Generation and LLM\-Based Explanation

RAG combines a pre\-trained LLM with an external retriever, reducing hallucination on knowledge\-intensive tasks\([7](https://arxiv.org/html/2609.09766#bib.bib7)\)\. Dense retrieval improves recall over sparse matching\([20](https://arxiv.org/html/2609.09766#bib.bib20)\), multi\-hop variants handle complex questions\([27](https://arxiv.org/html/2609.09766#bib.bib27)\), and Self\-RAG trains the model to decide when to retrieve and to verify its own answers\([25](https://arxiv.org/html/2609.09766#bib.bib25)\)\.

![Three-stage CARRE architecture: semantic retrieval of approved actions, counterfactual risk-cost optimization, and LLM-based reason classification followed by explanation generation.](https://arxiv.org/html/2609.09766v1/figure.png)Figure 1\.Overview of the CARRE pipeline\. Stage 1 \(Candidate Retrieval\) collects relevant action documents via FAISS semantic search\. Stage 2 \(CF Optimization\) selects the optimal prescription \(cf\_action\) by balancing risk reduction and intervention cost\. Stage 3 \(LLM Reasoning & Explanation\) comprises two sub\-steps: Stage 3\-a classifies the churn reason from the customer profile \(cf\_actionwithheld to prevent reverse\-reasoning bias\), and Stage 3\-b generates the natural\-language explanation given the determined reason and thecf\_actionfrom Stage 2\.Three\-stage CARRE architecture: semantic retrieval of approved actions, counterfactual risk\-cost optimization, and LLM\-based reason classification followed by explanation generation\.In the churn prescription setting, RAG constrains the LLM’s plan space to predefined candidate actions to prevent hallucinated interventions, and provides action\-specific details \(target customer profile, heuristic expected\-effect metadata, eligibility conditions\) as context to ground the LLM’s explanations\. The number of retrieved documentskkis a key design decision, since too few risk missing the optimal action, while too many introduce irrelevant context that may confuse LLM reasoning\.

On the reasoning side, TalkToModel queries ML predictions through dialogue\([15](https://arxiv.org/html/2609.09766#bib.bib15)\), CoT prompting improves accuracy on tasks requiring structured rule application\([14](https://arxiv.org/html/2609.09766#bib.bib14)\), and explicitly listing judgment criteria improves rule adherence in constrained generation\([26](https://arxiv.org/html/2609.09766#bib.bib26),[16](https://arxiv.org/html/2609.09766#bib.bib16)\)\. CARRE’s prompt design integrates CoT instructions and explicit classification criteria from these works, motivated by the reverse\-reasoning bias discussed in Section[3](https://arxiv.org/html/2609.09766#S3)and revisited in Section[6](https://arxiv.org/html/2609.09766#S6)\.

## 3\.Problem Formulation and CARRE Framework

CARRE is a three\-stage pipeline that simultaneously outputs*why*a high\-risk customer is likely to churn,*what action*may reduce that risk, and*why that action*is appropriate\. As illustrated in Figure[1](https://arxiv.org/html/2609.09766#acmlabel1), a customer profile is first transformed into a query to retrievekksemantically similar action documents from a FAISS index \(Stage 1\)\. Each retrieved action is then CF\-simulated to select the optimal action by balancing risk reduction and cost \(Stage 2\)\. Finally, the LLM classifies the churn reason from customer features and generates a natural\-language explanation grounded in the retrieved action context \(Stage 3\)\. The core design principle is*separation of concerns*\. Retrieval handles candidate coverage, optimization handles cost efficiency, and the LLM handles only reasoning and language generation, enabling each stage to be independently evaluated and improved\.

### 3\.1\.Action Space and Reason Taxonomy

CARRE’s action space𝒜\\mathcal\{A\}\(target of prescription\) and reason taxonomyℛ\\mathcal\{R\}\(target of classification\) are grounded in prior work that empirically identifies the recurring drivers of telecommunications churn and candidate retention strategies\. Price sensitivity, contract type, payment method, and underutilization of add\-on services have been consistently identified as primary churn factors\([21](https://arxiv.org/html/2609.09766#bib.bib21)\)\. Price discounts, contract switching, auto\-pay enrollment, and technical support or security service additions are correspondingly common retention levers in industry practice, motivating their inclusion in𝒜\\mathcal\{A\}\. Table[1](https://arxiv.org/html/2609.09766#S3.T1)lists the six actions in𝒜\\mathcal\{A\}with their descriptions and relative costs, where cost denotes a researcher\-defined relative weight capturing the operational and financial difficulty of each intervention, ranging fromA\_PAYMENT\_AUTOPAY\(1\.0, lowest\) toA\_CONTRACT\_24M\(3\.5, highest\)\.

Table 1\.Predefined action space𝒜\\mathcal\{A\}with researcher\-defined relative costs\.The six reason labels inℛ\\mathcal\{R\}are derived from domain rules applied in priority order\. Customers withMonthlyCharges≥80\\geq 80—a threshold corresponding to the upper quartile of the charge distribution in this dataset, indicating elevated financial exposure—on a month\-to\-month contract are classified asPRICE\_SENSITIVE\(Step 1\); those who do not meet Step 1 but hold a month\-to\-month contract withtenure≤6\\leq 6months are classified asCONTRACT\_RISK\(Step 2\)\. Customers whosePaymentMethodis electronic check are classified asPAYMENT\_FRICTION\(Step 3\); those using an internet service but subscribed to neitherTechSupportnorOnlineSecurityare classified asSUPPORT\_DEFICIT\(Step 4\)\. Customers withtenure≤3\\leq 3months are classified asLOW\_ENGAGEMENT\(Step 5\), and those satisfying none of the above conditions are classified asOTHER\(Step 6\)\.

The observable profile pattern represented by each heuristic label is interpreted as follows\.PRICE\_SENSITIVEmarks the combination of high monthly charges and a flexible month\-to\-month contract as a study\-specific proxy for potential price sensitivity\.CONTRACT\_RISKmarks short\-tenure customers on month\-to\-month contracts as a proxy for limited contractual commitment\.PAYMENT\_FRICTIONmarks electronic\-check use as a study\-specific proxy for potential payment\-process friction; the label does not establish that the payment method causally increases churn\.SUPPORT\_DEFICITmarks internet\-service customers without technical support or online security as a proxy for limited support coverage\.LOW\_ENGAGEMENTmarks customers within their first three months as a proxy for limited service experience\.OTHERcovers profiles not captured by Steps 1–5\.

These rules serve as*weak labels*for LLM agreement evaluation\. The thresholds, labels, and priority order are study\-specific weak\-label heuristics designed to create a reproducible instruction\-following task; they are not expert\-validated causal diagnoses of why a customer will churn\. The priority order is a researcher\-defined decision rule used to make the weak\-label task deterministic when multiple conditions hold\.

### 3\.2\.Problem Formulation

Let𝐱∈ℝd\\mathbf\{x\}\\in\\mathbb\{R\}^\{d\}be a customer feature vector andf:ℝd→\[0,1\]f:\\mathbb\{R\}^\{d\}\\to\[0,1\]a calibrated churn probability estimator\. For customers withf⁡\(𝐱\)≥θf\(\\mathbf\{x\}\)\\geq\\theta\(thresholdθ=0\.5\\theta=0\.5\), CARRE produces the following structured output tuple\. Below\-threshold customers may receive a reason label for diagnostic evaluation but do not receive a Stage 2 prescription or Stage 3\-b explanation:

\(1\)CARRE​\(𝐱\)=\(r⏟reason,a∗⏟prescription,e⏟explanation\),\\text\{CARRE\}\(\\mathbf\{x\}\)=\(\\underbrace\{r\}\_\{\\text\{reason\}\},\\;\\underbrace\{a^\{\*\}\}\_\{\\text\{prescription\}\},\\;\\underbrace\{e\}\_\{\\text\{explanation\}\}\),wherer∈ℛr\\in\\mathcal\{R\}is a structured churn label,a∗∈𝒜a^\{\*\}\\in\\mathcal\{A\}is a predefined candidate action, andeeis a natural\-language justification\.

The output tuple is designed to satisfy three simultaneous constraints\.rrmust be derivable from𝐱\\mathbf\{x\}under the priority\-ordered rule taxonomy;a∗a^\{\*\}must maximize CF risk reduction net of cost among retrieved candidates; andeemust be logically consistent with bothrranda∗a^\{\*\}\. The last constraint is structurally encouraged by Stage 3\-b, which conditions generation on bothrranda∗a^\{\*\}; consistency is subsequently evaluated rather than guaranteed\.

### 3\.3\.Three\-Stage Pipeline

In Stage 1, each action is represented as a document containing an ID, description, target customer profile, expected effect, and eligibility conditions\. For example, theA\_CONTRACT\_24Mdocument targets “month\-to\-month customers with medium\-to\-high monthly charges showing interest in long\-term commitment\.” The expected\-effect text in these documents \(e\.g\., a stated churn reduction of 0\.25–0\.40\) is heuristic metadata used only for retrieval context; it is not treated as a causal or empirically validated effect estimate, and Stage 2 recomputes predicted\-risk changes directly from the classifier\. Documents are encoded withsentence\-transformers/all\-MiniLM\-L6\-v2and stored in a FAISS flat inner\-product index\([11](https://arxiv.org/html/2609.09766#bib.bib11)\)\. For each customer, a retrieval query is constructed from four key features \(Contract,PaymentMethod,InternetService,MonthlyCharges\), and the top\-kkmost similar documents \(defaultk=5k\{=\}5\) are returned as the candidate action set\.

In Stage 2, for each retrieved actionaia\_\{i\}, its feature modification rules are applied to𝐱\\mathbf\{x\}to construct a CF customer vector𝐱i′\\mathbf\{x\}^\{\\prime\}\_\{i\}; the exact transformations and eligibility conditions are listed in Appendix[A](https://arxiv.org/html/2609.09766#A1)\(Table[10](https://arxiv.org/html/2609.09766#A1.T10)\)\. The CF score is:

\(2\)CF\-score​\(ai\)=\[f⁡\(𝐱\)−f⁡\(𝐱i′\)\]⏟Δ​risk−λ⋅cost​\(ai\)maxj⁡cost​\(aj\),\\text\{CF\-score\}\(a\_\{i\}\)=\\underbrace\{\\bigl\[f\(\\mathbf\{x\}\)\-f\(\\mathbf\{x\}^\{\\prime\}\_\{i\}\)\\bigr\]\}\_\{\\Delta\\text\{risk\}\}\-\\lambda\\cdot\\frac\{\\text\{cost\}\(a\_\{i\}\)\}\{\\max\_\{j\}\\text\{cost\}\(a\_\{j\}\)\},a∗=arg⁡maxai​CF\-score​\(ai\)a^\{\*\}=\\arg\\max\_\{a\_\{i\}\}\\text\{CF\-score\}\(a\_\{i\}\)is selected as thecf\_action\. The hyperparameterλ\\lambdacontrols the risk\-cost trade\-off and is set toλ=0\.25\\lambda\{=\}0\.25based on theλ\\lambdasweep in Table[5](https://arxiv.org/html/2609.09766#S5.T5)\. If no retrieved action achievesΔ​risk\>0\\Delta\\text\{risk\}\>0, Stage 2 returnscf\_action=NONE\.

Stage 3 comprises two sequential sub\-steps that together connect Stage 1’s semantic retrieval and Stage 2’s cost\-optimal prescription\.

Stage 3\-a \(Reason Classification and Auxiliary Plan Generation\)\.Stage 3\-a and Stage 2 operate*in parallel*: both receive the top\-kkaction candidates from Stage 1 independently\. The LLM receives a JSON customer feature profile, the baseline churn probabilityf⁡\(𝐱\)f\(\\mathbf\{x\}\), and the top\-kkretrieved action documents\. Critically,cf\_actionis*not*passed to Stage 3\-a\. Early experiments showed that including the CF\-selected action causes the LLM to anchor its explanation backward from the prescribed action rather than classifying the customer profile against the explicit criteria—a pattern we term*reverse\-reasoning bias*\. Excludingcf\_actionenforces the*direction*of inference: from customer features through the Step 1–6 priority rules to a reason label\. This provides only*procedural*independence from the prescription—the reason label is not back\-derived fromcf\_action—and not*epistemic*independence: because the Step 1–6 criteria are explicitly supplied in the prompt, Stage 3\-a’s task reduces to faithfully applying a fixed decision procedure to the customer profile rather than performing open\-ended domain reasoning about churn drivers \(see Section[5\.2](https://arxiv.org/html/2609.09766#S5.SS2)\)\. This parallel design allows Stage 2 and Stage 3\-a to be independently evaluated and improved\. Stage 3\-a also emits an auxiliaryllm\_planand brief rationale bullets, used only to evaluate constrained instruction following; they are not delivered as the final prescription or explanation\. The deployed prescription is thecf\_actionfrom Stage 2, and the delivered explanation is regenerated in Stage 3\-b from the Stage 3\-a reason and the Stage 2 action\.

The Stage 3\-a prompt incorporates three design choices\. First, explicit Step 1–6 classification criteria are provided, anchoring the LLM to a structured decision procedure\. Second, CoT instructions require the LLM to evaluate each step in priority order before committing to a label\. Third, a fixed plain\-text output schema consisting of aReason:line, aPlan:line, and two rationale bullets is enforced; plan IDs outside the retrieved set are flagged as auxiliary\-plan constraint violations\. A condensed prompt template illustrating these choices is provided in Appendix[B](https://arxiv.org/html/2609.09766#A2)\.

Stage 3\-b \(Explanation Generation\)\.Given the reasonrrfrom Stage 3\-a and the optimal actiona∗a^\{\*\}from Stage 2, Stage 3\-b generates the natural\-language explanationee\. The LLM receives the customer profile,f⁡\(𝐱\)f\(\\mathbf\{x\}\),rr, anda∗a^\{\*\}, and is instructed to provide a profile\-grounded rationale—in 2–3 English sentences—for whya∗a^\{\*\}is appropriate for this customer’s identified churn riskrr, grounding the explanation in specific profile fields\. Because Stage 3\-b conditions generation on bothrranda∗a^\{\*\}, the resulting explanation is intended to align with the actual prescription from Stage 2; alignment is evaluated empirically rather than guaranteed\. This two\-step design targets the incoherence that arises when Stage 2 and Stage 3\-a select different actions, which is the common case at scale \(Section[5\.3](https://arxiv.org/html/2609.09766#S5.SS3)\)\.

Thecf\_action\(Stage 2\) andllm\_plan\(Stage 3\-a\) may differ, as the two stages optimize for distinct objectives: cost\-adjusted risk reduction versus rule\-consistent classification\. When they diverge, the discrepancy serves as a diagnostic signal—indicating either that cost parameters need adjustment or that rule criteria require refinement—while Stage 3\-b conditions the delivered explanation on the actual prescription\. As a result, CARRE simultaneously produces*why*a high\-risk customer is likely to churn \(rr\),*what action*to take \(a∗a^\{\*\}\), and*why that action*is appropriate \(ee\), bridging the prediction\-prescription gap that characterizes most deployed churn systems\.

## 4\.Experimental Setup

We organize the evaluation around three research questions:

RQ1\.:Does direct counterfactual simulation produce more effective and cost\-efficient prescriptions than random, rule\-based, and SHAP\-based alternatives?

RQ2\.:How do retrieval depthkkand the cost penaltyλ\\lambdaaffect candidate coverage, prescription quality, and operational cost?

RQ3\.:How reliably do LLMs apply the reason taxonomy and generate explanations that are coherent, constrained, and grounded in customer evidence?

We use the IBM Telco Customer Churn dataset\([18](https://arxiv.org/html/2609.09766#bib.bib18)\)\. After removing the customer identifier and the churn label, each of the 7,043 customers is described by 19 predictor features covering demographics \(SeniorCitizen, gender\), service subscriptions \(InternetService, TechSupport, OnlineSecurity\), contract conditions \(Contract, tenure, MonthlyCharges\), and payment method\. The binary churn label has a positive class prevalence of 26\.5%, representing moderate class imbalance\. We apply an 80/20 stratified split with seed 42, yielding 5,634 training and 1,409 test instances\.

All experiments are conducted in a Python environment using scikit\-learn, FAISS, sentence\-transformers, and the OpenAI/Groq API; preprocessing, hyperparameters, model identifiers, and seed settings are summarized in Appendix[G](https://arxiv.org/html/2609.09766#A7)\. The CF prescription experiments \(Tables[3](https://arxiv.org/html/2609.09766#S5.T3)–[5](https://arxiv.org/html/2609.09766#S5.T5)\) use all 313 high\-risk test cases \(f⁡\(𝐱\)≥0\.50f\(\\mathbf\{x\}\)\\geq 0\.50\) in the test split, while the LLM reasoning experiments use a 30\-case balanced set \(Table[6](https://arxiv.org/html/2609.09766#S5.T6)\) for model comparison and a scaled 136\-case reason\-stratified sample \(Tables[7](https://arxiv.org/html/2609.09766#S5.T7),[8](https://arxiv.org/html/2609.09766#S5.T8)\) for the adopted model\.

CARRE’s Stage 2 CF scoring requires a reliable churn probability estimatef⁡\(𝐱\)f\(\\mathbf\{x\}\)\. We use aCalibratedClassifierCV\-wrappedLogisticRegression\(C=1\.0, max\_iter=2000\) with balanced class weighting and sigmoid calibration under 3\-fold cross\-validation\. Although tree\-based ensembles such as gradient boosting or random forests can match or slightly exceed logistic regression on aggregate ranking metrics, the isotonic\-calibrated tree ensembles tested in our implementation exhibited coarse and discontinuous probability responses under graded feature perturbations, making it difficult to estimate the subtle risk changes on which CF simulation depends\. Logistic regression instead provides a smooth, continuous probability surface that responds gradually to feature value changes, making it better suited for CF analysis where we directly estimate how much a given action reduces churn probability\. We validate this trade\-off empirically in Appendix[C](https://arxiv.org/html/2609.09766#A3)\(Table[11](https://arxiv.org/html/2609.09766#A3.T11)\)\. Table[2](https://arxiv.org/html/2609.09766#S4.T2)reports three performance metrics on the test set: ROC\-AUC measures discriminative ability under class imbalance, PR\-AUC captures precision\-recall balance on the minority class, and Brier Score quantifies calibration quality\.

Table 2\.Base prediction model performance \(IBM Telco Churn, 1,409 test instances\)\.ROC\-AUC of 0\.842 confirms sufficient discriminative performance for imbalanced binary classification\. PR\-AUC of 0\.633 is substantially above the random baseline \(0\.265\) given a positive class prevalence of 26\.5%\. A Brier score of 0\.138 indicates reasonable probabilistic accuracy on this test split, although a full calibration assessment would also require reliability curves or calibration\-error measures\. This probabilistic accuracy directly affects the Stage 2 CF scores\.

Base classifier choice for CF scoring\.A natural question is why CARRE retains logistic regression rather than a tree ensemble calibrated via isotonic regression, which attains marginally higher ROC\-AUC \(0\.845 for gradient boosting vs\. 0\.842\)\. A graded\-intervention sweep on the 313 high\-risk cases shows that all three tested tree ensembles produce severely non\-smooth CF surfaces: 45–66% of the intervention range is completely flat and single\-step jumps reach 0\.06–0\.10, so small realistic interventions register either as zero risk change or as an over\-large discontinuous drop\. Logistic regression instead yields a fully continuous, gradually responding surface \(0% flat, maximum jump 0\.002\), producing the gradedΔ\\Deltarisk estimates that Stage 2 requires\. Sacrificing this CF reliability for a 0\.003 ROC\-AUC gain is an unfavorable trade, so we retain calibrated logistic regression\. The comparison protocol and full per\-model results are given in Appendix[C](https://arxiv.org/html/2609.09766#A3)\(Table[11](https://arxiv.org/html/2609.09766#A3.T11)\)\.

## 5\.Results

### 5\.1\.Prescription Quality and Cost Efficiency

This section evaluates CARRE’s prescription pipeline in three aspects in sequence: prescription quality against baselines, RAG retrieval depth \(kk\), and the risk\-cost balance hyperparameter \(λ\\lambda\)\. All experiments are conducted on all 313 high\-risk test cases \(f⁡\(𝐱\)≥0\.50f\(\\mathbf\{x\}\)\\geq 0\.50\), the complete high\-risk cohort of the test split rather than a truncated sample\. For the prescription\-method comparison in Table[3](https://arxiv.org/html/2609.09766#S5.T3), per\-case risk\-reduction metrics are reported with 95% bootstrap confidence intervals\.

To compare prescription methods, CARRE is evaluated against four baselines\. The random baseline uniformly samples one eligible action for each customer under a fixed per\-customer seed\. The rule\-based baseline applies a priority\-ordered profile\-to\-action lookup table\. The SHAP\-based baseline maps the highest positive churn\-directed SHAP feature to an eligible candidate action\. The*SHAP\+Cost*baseline selects among SHAP\-implicated eligible actions using the same cost\-penalty structure as CARRE’s Equation[2](https://arxiv.org/html/2609.09766#S3.E2), with normalized SHAP importance replacing the model\-predicted risk\-reduction term\. The exact rule\-based and SHAP\-based mappings are given in Appendix[F](https://arxiv.org/html/2609.09766#A6)\(Table[14](https://arxiv.org/html/2609.09766#A6.T14)\)\. Table[3](https://arxiv.org/html/2609.09766#S5.T3)compares five methods across four metrics: Risk Red\. is the mean model\-predicted churn\-probability reduction per case, Effectiveness is the fraction of cases achievingΔ​risk\>0\\Delta\\text\{risk\}\>0, Avg\. Rel\. Cost is the mean researcher\-defined relative action cost, and Efficiency is the ratio of Risk Red\. to Avg\. Rel\. Cost\.

Table 3\.Prescription method comparison \(all 313 high\-risk cases withf⁡\(𝐱\)≥0\.5f\(\\mathbf\{x\}\)\{\\geq\}0\.5,λ=0\.25\\lambda\{=\}0\.25\)\. Brackets denote 95% bootstrap confidence intervals \(2,000 resamples\)\. SHAP\+Cost applies the identical cost penalty used in CARRE’s CF\-score to SHAP\-selected actions\.CARRE outperforms all baselines on risk reduction and efficiency\. Against the strongest baseline, CARRE achieves 79\.8% higher raw risk reduction than the SHAP\-based approach \(0\.302 vs\. 0\.168\), with non\-overlapping 95% confidence intervals \(\[0\.294, 0\.308\] vs\. \[0\.166, 0\.169\]\)\. Because CARRE also selects higher\-cost actions on average \(Avg\. Rel\. Cost 3\.26 vs\. 2\.00\), we report two complementary controlled comparisons\. First, a cost\-normalized comparison using Efficiency yields a more conservative 10\.5% advantage \(0\.093 vs\. 0\.084\)\. Second, and more directly, the SHAP\+Cost baseline subjects SHAP\-based selection to the identical cost penalty used by CARRE\. SHAP\+Cost yields nearly the same mean predicted\-risk reduction as plain SHAP \(0\.167 vs\. 0\.168\), and their marginal bootstrap intervals overlap; CARRE exceeds SHAP\+Cost by 80\.4% in mean predicted\-risk reduction\. Together these comparisons show that CARRE’s advantage stems from direct CF simulation rather than from cost weighting—penalizing cost does not help feature\-importance\-based selection, because SHAP importance for prediction does not identify the actions that most reduce churn risk\.

Notably, the high effectiveness of the random baseline \(92\.3%\) reflects a design property of the action space rather than a CARRE\-specific achievement: all six actions were deliberately chosen as business interventions targeting high\-risk churn customers, making most of them broadly applicable across this population regardless of selection method\. Effectiveness \(Δ​risk\>0\\Delta\\text\{risk\}\>0\) is therefore a binary coverage metric capturing whether any positive effect is achieved, not the magnitude of that effect\. The more discriminating measures of method quality are Risk Red\. and Efficiency, where CARRE substantially outperforms all baselines: CARRE achieves 120% higher risk reduction than random \(0\.302 vs\. 0\.137\) and 52% higher efficiency \(0\.093 vs\. 0\.061\), confirming that what differentiates methods is not whether they find an effective action, but how much risk reduction they deliver per unit cost\.

CARRE’s advantage on these quality metrics stems from direct CF simulation\. The SHAP\-based approach targets features with high explanation importance, but importance for prediction does not necessarily coincide with the highest CF risk reduction\. CARRE instead directly simulates each action as a CF intervention, measures the resulting churn probability reduction, and selects the optimal action after balancing cost\. The rule\-based baseline scores lowest on both risk reduction \(0\.074\) and effectiveness \(48\.9%\) because the static lookup table handles only one rule at a time, failing to identify the correct action for customers with multiple simultaneous risk factors; its low effectiveness is attributable to rule mismatch, not action space limitations\.

Turning to the effect of RAG retrieval depthkk, Table[4](https://arxiv.org/html/2609.09766#S5.T4)reports four metrics: Recall@kkis the fraction of cases where the globally optimal action is within the top\-kkretrieved set, Risk Red\. is the mean CF risk reduction, Avg\. Rel\. Cost is the mean researcher\-defined relative action cost, and CF\-NONE Rate is the fraction of cases for which Stage 2 returnscf\_action=NONE\.

Table 4\.Effect of retrieval depthkkon CF performance \(313 high\-risk cases\)\.The results reveal a clear retrieval threshold effect\. The sharp recall jump from 1\.3% to 70\.9% askkincreases from 3 to 4 occurs because the retrieval query represents the customer’s current state, so status\-quo\-similar documents tend to rank first through third, while high\-impact actions \(e\.g\.,A\_CONTRACT\_24Mfor month\-to\-month customers\) tend to appear at rank 4 or lower\. Atk=5k\{=\}5, recall reaches 85\.3% and NONE cases are eliminated; the 14\.7% recall gap versus the oracle \(k=6k\{=\}6\) is offset by the benefit of reducing context noise passed to the LLM\. We adoptk=5k\{=\}5as the default\.

Table[5](https://arxiv.org/html/2609.09766#S5.T5)summarizes theλ\\lambdasweep, whereλ\\lambda\(Equation[2](https://arxiv.org/html/2609.09766#S3.E2)\) governs the risk–cost balance\. Risk Red\. \(meanΔ​risk\\Delta\\text\{risk\}per case\) and Avg\. Rel\. Cost \(mean relative intervention cost\) reflect the two objectives of Equation[2](https://arxiv.org/html/2609.09766#S3.E2)\([4](https://arxiv.org/html/2609.09766#bib.bib4),[19](https://arxiv.org/html/2609.09766#bib.bib19)\); Efficiency is their ratio \(a study\-specific composite, not a standard CF metric\); and High\-Cost Rate is the fraction of cases selecting an action with relative cost≥3\.0\\geq 3\.0, included to gauge practical feasibility\.

Table 5\.λ\\lambdasweep: risk\-cost trade\-off \(313 high\-risk cases,k=5k\{=\}5\)\.The sweep reveals a sharp regime boundary rather than a gradual trade\-off\. Forλ∈\[0\.10,0\.25\]\\lambda\{\\in\}\[0\.10,0\.25\], the cost penalty is small relative to the risk\-reduction gap between actions, so the highest\-impact action \(A\_CONTRACT\_24M\) dominates selection and risk reduction is essentially flat \(0\.303 vs\. 0\.302, a 0\.4% difference\)\. Atλ=0\.50\\lambda\{=\}0\.50the penalty overtakes that gap, selection flips almost entirely to low\-cost actions \(A\_PAYMENT\_AUTOPAY\), and risk reduction collapses by 59% \(0\.302→\\to0\.122\)\. Beyondλ=1\.0\\lambda\{=\}1\.0the sweep saturates\.λ=0\.25\\lambda\{=\}0\.25therefore sits at the knee of this transition: it attains near\-maximum risk reduction \(within 0\.4% of theλ=0\.10\\lambda\{=\}0\.10peak\) with Efficiency essentially tied withλ=0\.10\\lambda\{=\}0\.10\(0\.0925 vs\. 0\.0923\), while any further increase crosses the regime boundary and forfeits over half the achievable risk reduction\. We adoptλ=0\.25\\lambda\{=\}0\.25as the default\. The discreteness of this transition follows from the small six\-action space with widely separated costs, and its dataset\-specific boundary is discussed as a limitation in Section[6](https://arxiv.org/html/2609.09766#S6)\. In practice,λ\\lambdais domain\-calibratable: organizations with strict cost ceilings should operate atλ≥0\.5\\lambda\\geq 0\.5, others below the boundary\.

Retrieval\-only action\-document scaling stress test\.A natural deployment question is how Stage 1 retrieval behaves as the action catalog grows from six curated actions to hundreds of dynamic marketing levers\. We ran a retrieval\-only stress test of Stage 1—the downstream CF and LLM stages were not executed—augmenting the six action documents with up to 394 synthetic lever documents \(retrieval documents, not executable actions\) and measuring retrieval runtime and oracle\-action survival over all 313 high\-risk cases\. Runtime is not a bottleneck: FAISS search itself stays below 0\.01 ms per query even at\|𝒜\|=400\|\\mathcal\{A\}\|\{=\}400, and end\-to\-end query latency, dominated by embedding, stays near 1 ms\. Retrieval precision, however, degrades in two stages—exact\-action Recall@5 collapses once near\-duplicate parameter variants appear, and family\-level recall collapses beyond\|𝒜\|≈200\|\\mathcal\{A\}\|\{\\approx\}200\. Scaling to industrial catalogs therefore requires hierarchical or diversity\-aware retrieval rather than a larger flatkk\. The full protocol and results are given in Appendix[D](https://arxiv.org/html/2609.09766#A4)\(Table[12](https://arxiv.org/html/2609.09766#A4.T12)\)\.

### 5\.2\.LLM Reason Classification

This subsection evaluates how reliably LLMs apply the reason taxonomy across models; the two that follow assess explanation quality at scale and the effect of retrieval scope\.Agreementrefers to the rate at which LLM output labels match the weak labels determined by the Step 1–6 priority rules in Section[3](https://arxiv.org/html/2609.09766#S3)\. Because the same rules are supplied in the prompt, this metric measures*instruction\-following fidelity*—how consistently the LLM applies a given decision procedure—rather than independent domain reasoning ability\.

The LLM model comparison was conducted using an earlier model\-comparison prompt with the CF action withheld, explicit classification criteria, and CoT instructions\. This formulation differs from the v2 scaled\-evaluation prompt used for the 136\-case experiments in Sections[5\.3](https://arxiv.org/html/2609.09766#S5.SS3)and[5\.4](https://arxiv.org/html/2609.09766#S5.SS4)\. Table[6](https://arxiv.org/html/2609.09766#S5.T6)compares four LLMs across three metrics: Agreement is the weak\-label agreement rate, F1\-macro is the macro\-averaged F1 across six reason labels, and auxiliary\-plan violations are the fraction of cases whose Stage 3\-allm\_planfalls outside the retrieved set\. A per\-reason breakdown is provided in Appendix[E](https://arxiv.org/html/2609.09766#A5)\(Table[13](https://arxiv.org/html/2609.09766#A5.T13)\)\.

Table 6\.LLM comparison on reason classification \(model\-comparison prompt; 30\-case balanced set, 5 per reason category\)\. Because this formulation differs from the v2 scaled\-evaluation prompt, its 0\.933 agreement is not directly comparable to the 136\-case results in Sections[5\.3](https://arxiv.org/html/2609.09766#S5.SS3)and[5\.4](https://arxiv.org/html/2609.09766#S5.SS4)\.GPT\-4o achieves the highest agreement \(93\.3%\), with two errors out of 30\. One falls in CONTRACT\_RISK \(4/5\), where a customer with tenure≤\\leq3 satisfies both Steps 2 and 5 and GPT\-4o inverts the priority order; the other falls in OTHER \(4/5\)\. GPT\-4o\-mini \(40\.0%\) shows a distinct avoidance pattern, predicting 0% on PRICE\_SENSITIVE and CONTRACT\_RISK while achieving 100% on PAYMENT\_FRICTION and 80% on OTHER, suggesting it anchors heavily on surface payment\-method cues while failing on price\- and contract\-related rules\. Llama\-3\.3\-70B \(50\.0%\) achieves 100% on PRICE\_SENSITIVE but scores zero on CONTRACT\_RISK and OTHER\. Llama\-3\.3\-70B outperforms GPT\-4o\-mini in this small comparison, but the models differ simultaneously in scale, architecture, vendor, and training, so the result does not isolate which factor drives the difference\. Qwen3\-32B \(36\.7%\) shows the most severe bias, scoring zero on CONTRACT\_RISK, PAYMENT\_FRICTION, and OTHER; for CONTRACT\_RISK this reflects systematic*misclassification*rather than abstention, as the model over\-predicts that label yet never applies it correctly, whereas for PAYMENT\_FRICTION and OTHER it abstains from the label entirely\. This pattern is consistent with the known sensitivity of prompted language models to label and contextual biases\([13](https://arxiv.org/html/2609.09766#bib.bib13)\), although the current experiment does not identify its cause\. Qwen3\-32B was evaluated with its built\-in Extended Thinking mode disabled to apply the same external CoT prompt uniformly across all models\. No auxiliary\-plan constraint violations occurred for any of the four models in this 30\-case comparison\. We adopt GPT\-4o as the default LLM for subsequent experiments\. It is then evaluated on a larger 136\-case sample using the v2 scaled\-evaluation prompt\. Because the prompt formulation also changes, this evaluation should not be interpreted as a pure sample\-size comparison with Table[6](https://arxiv.org/html/2609.09766#S5.T6)\.

### 5\.3\.Explanation Quality at Scale

For human evaluation of generated explanations, four independent raters assessed GPT\-4o outputs on a 1–5 Likert scale across four criteria: Faithfulness \(alignment with the customer’s actual data\), Actionability \(perceived operational feasibility of the recommended action\), Coherence \(logical consistency between the stated reason and recommended plan\), and Satisfaction \(overall perceived quality\)\. However, inter\-rater agreement was very low \(Krippendorff’sα=0\.003\\alpha=0\.003–0\.1750\.175\), indicating that the ratings were insufficiently reliable for confirmatory analysis\. We therefore do not treat the per\-criterion human ratings as findings; a rigorous redesign with structured rater calibration, explicit criterion operationalization, andn≥25n\\geq 25cases per reason category would be required before human evaluation can serve as confirmatory evidence\([24](https://arxiv.org/html/2609.09766#bib.bib24)\)\.

In place of the discarded human study, we conducted a scaled automated evaluation on a reason\-stratified sample of 136 cases drawn from the test split \(25 per reason category; 11 forLOW\_ENGAGEMENT\)—a 4\.5×\\timesexpansion over the 30\-case set\. Each case was processed by the full pipeline with the v2 reason prompt, and each explanation was scored 1–5 on the three criteria by two cross\-vendor LLM judges \(GPT\-4o, Llama\-3\.3\-70B\) under a G\-Eval\-style protocol\([10](https://arxiv.org/html/2609.09766#bib.bib10)\)with explicit criterion definitions and mandated field\-by\-field verification \(Table[7](https://arxiv.org/html/2609.09766#S5.T7)\)\. The explanation evaluation therefore reflects the pre\-refinement v2 reason outputs\. Three findings emerge\. First, both judges rate explanation quality high \(faithfulness 4\.02/4\.04, coherence 4\.24/4\.47, actionability 4\.45/5\.00\), with within\-one\-point cross\-vendor agreement of 61–87%\. The Llama judge saturates on actionability \(5\.00\), so its scores on that criterion carry no discriminative signal\. Second, a deterministic grounding check—verifying every categorical and numeric claim against the actual profile with no LLM involvement—found zero contradictions across the 66 verifiable claims present in 55 of the 135 explanations; the other 80 state no checkable claim, and both judges returned valid scores for all 135\. A claim is verifiable when it asserts a contract type, payment method, dollar charge, or tenure checkable against the profile\. Third, under the same v2 scaled\-evaluation prompt, agreement decreases modestly from 83\.3% on the 30\-case ablation set \(Table[8](https://arxiv.org/html/2609.09766#S5.T8)\) to 79\.4% on the 136\-case sample\. On the larger sample, errors concentrate inCONTRACT\_RISKandOTHER:CONTRACT\_RISKreaches 52% because of Step 2 versus Step 5/Step 3 priority inversions, whileOTHERreaches 48% because of over\-application of theSUPPORT\_DEFICITrule\. The 93\.3% result in Table[6](https://arxiv.org/html/2609.09766#S5.T6)used a different model\-comparison prompt and is therefore not directly comparable\. Because both failure patterns were addressable through prompt changes on this sample, we added a targeted enforcement block—lowest\-step\-wins tie\-breaking with worked examples, and an explicit all\-three\-conditions requirement for Step 4\. This*v3*prompt increased agreement from 79\.4% to 90\.4% \(F1\-macro 0\.910, zero violations\) on the same 136 cases:CONTRACT\_RISKrecovers to 100% andOTHERto 88%, at the cost of partial over\-correction onSUPPORT\_DEFICIT\(60%\)\. This post\-refinement result may be optimistic because the sample was used for both error diagnosis and re\-evaluation, and it therefore requires confirmation on an independent held\-out set\. We adopt v3 as CARRE’s final configuration and retain 79\.4% as the pre\-refinement diagnostic estimate\. Zero grounding contradictions coexist with 79\.4% label agreement because the grounding check verifies factual claims against the profile, not the upstream reason label \(Case 3, Section[5\.5](https://arxiv.org/html/2609.09766#S5.SS5)\)\. These LLM judges share training\-distribution biases with the generator and do not replace calibrated human evaluation; they serve as scalable interim evidence pending a redesigned human protocol with rater calibration and per\-category coverage\.

Table 7\.Automated explanation evaluation on a 136\-case reason\-stratified subset of the test split \(135 generated explanations\); explanations use the v2 reason outputs\. Judges score 1–5 under a G\-Eval\-style protocol; grounding is a deterministic claim\-vs\-profile check involving no LLM\.#### Stage 3\-b: Reason–Action Concordance and Explanation Quality\.

Stage 3\-b was applied to the 136\-case reason\-stratified sample to generate explanations aligned with the Stage 2cf\_action\. The Stage 3\-a auxiliaryllm\_planand the Stage 2cf\_actionmatch exactly in 2% of all 135 explained cases \(3/135\)\. Among the 89 cases where both stages select an action, exact match is 3% \(3/89\) and action\-family match is 11% \(10/89\), becauseA\_CONTRACT\_24Mdominates Stage 2 selections \(99/135\) wherever switching to a two\-year contract gives the highest cost\-penalized CF score\. Part of the divergence is structural: 47 of the 136 cases have Stage 3\-aplan=NONE\. Stage 2 nevertheless finds a positive\-Δ\\Deltariskcf\_actionfor 46 of these cases; the remaining case has no CF action, yielding 135 explanations\. Stage 3\-b generates direct profile\-grounded explanations for concordant pairs and bridging rationales for divergent ones \(Case 2, Section[5\.5](https://arxiv.org/html/2609.09766#S5.SS5)\)\. This divergence does not render Stage 3\-a redundant: the reason label is not a prescription input but serves three roles—grounding the Stage 3\-b explanation, framing the agent\-facing conversation, and flagging cost\-dominance forλ\\lambdaor action\-space recalibration\.

### 5\.4\.Effect of Retrieval Scope

To isolate RAG’s contribution to LLM reasoning beyond its established role of constraining the plan space, GPT\-4o was run under three conditions\. Table[8](https://arxiv.org/html/2609.09766#S5.T8)compares three conditions on Agreement \(weak\-label agreement rate\), F1\-macro \(macro\-averaged F1\), and auxiliary\-plan violations \(fraction of Stage 3\-allm\_planselections outside the retrieved set; these do not affect the finalcf\_action\)\.

Table 8\.Effect of retrieval scope on LLM reasoning quality \(GPT\-4o\), at the original 30\-case set and re\-run on the 136\-case sample\. Thek=3k\{=\}3peak observed atn=30n\{=\}30does not replicate at scale; thek=5k\{=\}5default is best atn=136n\{=\}136\. F1\-macro and auxiliary\-plan violations are reported for then=136n\{=\}136run only\. The 136\-case retrieval ablation uses the v2 reason prompt, which is distinct from the model\-comparison formulation in Table[6](https://arxiv.org/html/2609.09766#S5.T6), so itsk=5k\{=\}5/30\-case agreement \(0\.833\) is not directly comparable to that table’s 0\.933\.On the original 30\-case set,k=3k\{=\}3appeared to peak dramatically \(96\.7%\), suggesting that tighter retrieval might outperform thek=5k\{=\}5default\. Re\-running the ablation on the 136\-case sample shows that this peak was a small\-sample artifact: at scale,k=3k\{=\}3\(76\.5%\) is essentially level with No\-RAG \(75\.7%\), and the CARRE defaultk=5k\{=\}5is best on both agreement \(79\.4%\) and F1\-macro \(0\.782\)\. Two conclusions replace the earlier reading\. First, constraining the LLM’s context to retrieved actions still helps—No\-RAG trailsk=5k\{=\}5by 3\.7 percentage points, consistent with irrelevant actions inducing*action\-to\-reason reverse inference*—but the effect is modest rather than dramatic\. Second,k=5k\{=\}5provides the best evaluated compromise: it attains high, though not maximal, oracle\-action coverage—the full\-libraryk=6k\{=\}6condition reaches maximal coverage \(Table[4](https://arxiv.org/html/2609.09766#S5.T4)\)—together with the highest downstream classification agreement, at zero violations in every condition\.

### 5\.5\.Qualitative Case Study

Table[9](https://arxiv.org/html/2609.09766#S5.T9)presents three high\-risk outputs from the CARRE prescription\-evaluation cohort and one low\-risk abstention control \(GPT\-4o,k=5k\{=\}5\)\. The low\-risk control is included only to illustrate threshold\-based non\-intervention and is not part of the 313\-case high\-risk prescription evaluation\.

Table 9\.Case study: concordant success \(top\), Stage 3\-b mismatch bridge \(second\), step priority failure \(third\), and low\-risk abstention control \(bottom\)\. Per\-case scores are from the two\-judge automated evaluation of Section[5\.3](https://arxiv.org/html/2609.09766#S5.SS3), shown as G:xx/L:yyfor the GPT\-4o and Llama\-3\.3\-70B judges \(1–5 scale\); scores from the discarded human study are not reported\. Concordance is defined at the level of action family \(e\.g\., contract\-switch\), not exact action ID\.Case 1 \(concordant success\)\.Customer 2514\-GINMM is correctly classifiedCONTRACT\_RISK; Stage 2 selectsA\_CONTRACT\_24M\(Δ​risk=0\.326\\Delta\\text\{risk\}\{=\}0\.326\) and Stage 3\-b coherently links the contract\-instability reason to the 24\-month prescription\. Both judges rate it at ceiling \(5/5\)\.

Case 2 \(mismatch bridge\)\.Customer 2077\-MPJQO is classifiedPAYMENT\_FRICTION, but Stage 2 selectsA\_CONTRACT\_24M\(Δ​risk=0\.330\\Delta\\text\{risk\}\{=\}0\.330\), a far larger model\-predicted risk reduction than a payment\-method change\. Stage 3\-b bridges the gap via the contract change, but both judges score it below Case 1 \(faithfulness/coherence 3/3\), showing that divergence yields less convincing rationale \(cf\. F5\)\.

Case 3 \(silent priority error\)\.Customer 8258\-GSTJK satisfies both Step 2 and Step 5; the correct label isCONTRACT\_RISK\(Step 2 precedence\), but GPT\-4o applies Step 5 first and misclassifies asLOW\_ENGAGEMENT\. Stage 3\-b still produces a fluent explanation around the wrong reason, rated 5/5 by both judges—the error surfaces only through the weak\-label check, which is why we treat judge scores as surface quality, not end\-to\-end correctness\.

Case 4 \(low\-risk abstention control\)\.Customer 1166\-PQLGG is a long\-tenured \(72\-month\), low\-charge, phone\-only customer on a two\-year contract, with a churn probability of only 0\.002 and no risk pattern matching Steps 1–5, so the diagnostic reason label isOTHER\. Becausef⁡\(𝐱\)f\(\\mathbf\{x\}\)falls below the intervention threshold, Stage 2 returnscf\_action=NONEand Stage 3\-b generates no explanation\. It lies outside the 313\-case evaluation cohort and illustrates threshold\-based abstention only\.

## 6\.Discussion and Limitations

We draw five main findings from the experimental results, alongside four limitations\. Finding F1 addresses RQ1, F2 addresses RQ2, and F3–F5 address RQ3\.

F1: CF optimization outperforms explanation\-based heuristics, and the advantage is not an artifact of cost weighting\.CARRE achieves 79\.8% higher raw risk reduction than the SHAP\-based approach on all 313 high\-risk cases \(Table[3](https://arxiv.org/html/2609.09766#S5.T3)\)\. The controlled*SHAP\+Cost*baseline \(Section[5\.1](https://arxiv.org/html/2609.09766#S5.SS1)\) rules out the built\-in cost penalty as the source of this gap: subjecting SHAP\-based selection to the same penalty leaves its risk reduction unchanged, and CARRE retains an 80\.4% advantage\. The cost\-normalized Efficiency comparison likewise favors CARRE \(10\.5%\)\. The gap therefore stems from direct CF simulation, not cost weighting, and shows that prediction feature importance and score\-selected intervention targets can diverge—questioning the practice of driving retention campaigns directly from SHAP\.

F2: Retrieval scope has a modest but real effect on reasoning quality, and small\-sample ablations overstate it\.Beyond the plan\-space constraint that is RAG’s established benefit, restricting the LLM’s context to retrieved actions improves weak\-label agreement by 3\.7 percentage points over the full action library, with zero auxiliary\-plan violations throughout\. A dramatick=3k\{=\}3peak on the 30\-case set did not replicate at 136 cases\. Thek=5k\{=\}5setting is the best evaluated compromise: high though not maximal oracle\-action coverage and the highest weak\-label agreement, whereas full\-libraryk=6k\{=\}6keeps maximal CF coverage but lower reasoning agreement\.

F3: Prompt engineering is critical for structured LLM classification, and its gains require independent validation\.Withholding the CF\-selected action and adding explicit Step 1–6 criteria with CoT improved rule\-following during prompt development\. Under the same v2 scaled\-evaluation prompt, agreement decreased modestly from 83\.3% on the 30\-case ablation set to 79\.4% on the 136\-case sample\. A targeted v3 enforcement block then increased agreement to 90\.4% on the same 136 cases\. Because the same cases were used for error diagnosis and re\-evaluation, the v3 improvement remains a prompt\-development result requiring confirmation on an independent held\-out sample\. The 93\.3% model\-comparison result in Table[6](https://arxiv.org/html/2609.09766#S5.T6)used a different prompt formulation and is not directly comparable\.

F4: Structured\-reasoning performance varies substantially across models, but the responsible model factors remain confounded\.Llama\-3\.3\-70B reaches 50\.0% and outperforms GPT\-4o\-mini at 40\.0%, but the pair differs in scale, architecture, and vendor at once, so no single factor is isolated\. All three non\-GPT\-4o models fail systematically in tail categories—Qwen3\-32B, for instance, scores zero onCONTRACT\_RISK,PAYMENT\_FRICTION, andOTHER—so the tested open\-source models may need targeted calibration before use in this task, pending broader comparison\.

F5: Cost\-aware optimization systematically overrides intuitive rule\-based prescriptions, and Stage 3\-b bridges the resulting gap\.Across all 135 explained cases, the Stage 3\-a auxiliary plan and the Stage 2 action match exactly in 2% \(3/135\); among the 89 cases where both stages select an action, exact match is 3% \(3/89\) and family\-level match is 11% \(10/89\), becauseA\_CONTRACT\_24Mdominates Stage 2 selections wherever switching to a two\-year contract yields the highest cost\-penalized CF score, even forPAYMENT\_FRICTIONandSUPPORT\_DEFICITprofiles\. The parallel architecture makes this cost\-dominance transparent, and Stage 3\-b anchors explanations to the actual prescription, most effectively when reason and action are related\.

These findings are accompanied by four limitations\. First, weak\-label agreement measures instruction\-following fidelity, not causal correctness\. The taxonomy is a researcher\-defined heuristic over observable profile patterns, not validated against expert\-annotated churn causes, so agreement is not evidence that the labels identify the true cause of churn\. Expert annotation or observed intervention outcomes would give stronger ground truth but were unavailable here\.

Second, the reported risk reductions are model\-predicted changes, not causal or off\-policy validation of real intervention effects—a limitation shared with the algorithmic recourse literature, where recourse without the true structural causal model remains open\([5](https://arxiv.org/html/2609.09766#bib.bib5)\)\.

Third, the six\-action space does not cover all reason types \(no actions targetLOW\_ENGAGEMENTorOTHER\), and its transformations—detailed in Appendix[A](https://arxiv.org/html/2609.09766#A1)—remain study\-specific design choices rather than validated causal interventions\. Future work should expand and externally validate the action vocabulary\([6](https://arxiv.org/html/2609.09766#bib.bib6)\); the scaling test \(Appendix[D](https://arxiv.org/html/2609.09766#A4)\) shows such expansion needs hierarchical or diversity\-aware retrieval rather than flat top\-kk\.

Fourth,λ=0\.25\\lambda\{=\}0\.25andk=5k\{=\}5are empirical defaults chosen on the same data used for evaluation, without a separate validation split, and all reported results were obtained on a single dataset\. Table[5](https://arxiv.org/html/2609.09766#S5.T5)shows model\-predicted risk reduction is flat within the risk\-first regime but collapses by 59% onceλ\\lambdapasses 0\.50, so the qualitative conclusion is robust while the boundary location is specific to the action costs and feature transformations used in this study\. Cross\-dataset robustness of both theλ\\lambdaboundary and the CF advantage remains future work\.

## 7\.Conclusion

We proposed CARRE, a three\-stage pipeline integrating RAG\-based retrieval, cost\-aware CF scoring, and LLM reasoning for churn prescription\. On the IBM Telco dataset, CARRE substantially outperforms SHAP\-based baselines in model\-predicted risk reduction, and the advantage persists under the cost\-controlled SHAP\+Cost comparison, indicating it stems from direct CF simulation rather than cost weighting\. On the 136\-case reason\-stratified sample, diagnosis\-driven prompt refinement raised weak\-label agreement from 79\.4% to 90\.4%; we treat this as a prompt\-development result, not an independent generalization estimate\. For the 135 v2 explanations, two cross\-vendor judges assign mean scores of 4\.02–5\.00, and a deterministic audit finds no contradictions among 66 verifiable profile claims\. Its modular design—retrieval, optimization, LLM justification—lets each component be evaluated and replaced independently\.

Future work includes expanding the action space forLOW\_ENGAGEMENTandOTHER, improving faithfulness via citation\-based generation\([17](https://arxiv.org/html/2609.09766#bib.bib17)\), and extending CARRE to other subscription industries\.

###### Acknowledgements\.

This research was supported by the Korea Institute for Advancement of Technology \(KIAT\) grant funded by the Korean Government \(MOTIE\) \(RS\-2024\-00416131, HRD Program for Industrial Innovation\)\.

## References

- \(1\)Thanasis Vafeiadis, Konstantinos I\. Diamantaras, George Sarigiannidis, and Konstantinos Ch\. Chatzisavvas\. 2015\.A comparison of machine learning techniques for customer churn prediction\.*Simulation Modelling Practice and Theory*, 55:1–9\.[https://doi\.org/10\.1016/j\.simpat\.2015\.03\.003](https://doi.org/10.1016/j.simpat.2015.03.003)\.
- \(2\)Scott M\. Lundberg and Su\-In Lee\. 2017\.A unified approach to interpreting model predictions\.In*Proceedings of the 31st International Conference on Neural Information Processing Systems \(NIPS’17\)*\. Curran Associates Inc\., Red Hook, NY, USA, 4768–4777\.
- \(3\)Sandra Wachter, Brent Mittelstadt, and Chris Russell\. 2018\.Counterfactual explanations without opening the black box: Automated decisions and the GDPR\.*Harvard Journal of Law & Technology*, 31\(2\):841–887\.
- \(4\)Amir\-Hossein Karimi, Bernhard Schölkopf, and Isabel Valera\. 2021\.Algorithmic Recourse: from Counterfactual Explanations to Interventions\.In*Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency \(FAccT ’21\)*\. Association for Computing Machinery, New York, NY, USA, 353–362\.
- \(5\)Amir\-Hossein Karimi, Julius von Kügelgen, Bernhard Schölkopf, and Isabel Valera\. 2020\.Algorithmic recourse under imperfect causal knowledge: a probabilistic approach\.In*Proceedings of the 34th International Conference on Neural Information Processing Systems \(NIPS ’20\)*\. Curran Associates Inc\., Red Hook, NY, USA, Article 23, 265–277\.
- \(6\)Sahil Verma, Varich Boonsanong, Minh Hoang, Keegan Hines, John Dickerson, and Chirag Shah\. 2024\.Counterfactual Explanations and Algorithmic Recourses for Machine Learning: A Review\.*ACM Comput\. Surv\.*56, 12, Article 312 \(December 2024\), 42 pages\.[https://doi\.org/10\.1145/3677119](https://doi.org/10.1145/3677119)\.
- \(7\)Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen\-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela\. 2020\.Retrieval\-augmented generation for knowledge\-intensive NLP tasks\.In*Proceedings of the 34th International Conference on Neural Information Processing Systems \(NIPS ’20\)*\. Curran Associates Inc\., Red Hook, NY, USA, Article 793, 9459–9474\.
- \(8\)Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald\. 2020\.On faithfulness and factuality in abstractive summarization\.In*Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics \(ACL\)*, pages 1906–1919\.
- \(9\)Berk Ustun, Alexander Spangher, and Yang Liu\. 2019\.Actionable Recourse in Linear Classification\.In*Proceedings of the Conference on Fairness, Accountability, and Transparency \(FAT\* ’19\)*\. Association for Computing Machinery, New York, NY, USA, 10–19\.[https://doi\.org/10\.1145/3287560\.3287566](https://doi.org/10.1145/3287560.3287566)\.
- \(10\)Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu\. 2023\.G\-Eval: NLG Evaluation using Gpt\-4 with Better Human Alignment\.In*Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing*, pages 2511–2522, Singapore\. Association for Computational Linguistics\.
- \(11\)Jeff Johnson, Matthijs Douze, and Hervé Jégou\. 2021\.Billion\-scale similarity search with GPUs\.*IEEE Transactions on Big Data*, 7\(3\):535–547\.
- \(12\)Alexandru Niculescu\-Mizil and Rich Caruana\. 2005\.Predicting good probabilities with supervised learning\.In*Proceedings of the 22nd international conference on Machine learning \(ICML ’05\)*\. Association for Computing Machinery, New York, NY, USA, 625–632\.[https://doi\.org/10\.1145/1102351\.1102430](https://doi.org/10.1145/1102351.1102430)\.
- \(13\)Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh\. 2021\.Calibrate before use: Improving few\-shot performance of language models\.In*Proceedings of the 38th International Conference on Machine Learning*, PMLR 139:12697–12706\.
- \(14\)Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H\. Chi, Quoc V\. Le, and Denny Zhou\. 2022\.Chain\-of\-thought prompting elicits reasoning in large language models\.In*Proceedings of the 36th International Conference on Neural Information Processing Systems \(NIPS ’22\)*\. Curran Associates Inc\., Red Hook, NY, USA, Article 1800, 24824–24837\.
- \(15\)Dylan Slack, Satyapriya Krishna, Himabindu Lakkaraju, and Sameer Singh\. 2023\.Explaining machine learning models with interactive natural language conversations using TalkToModel\.*Nat Mach Intell*5, 873–883\.[https://doi\.org/10\.1038/s42256\-023\-00692\-8](https://doi.org/10.1038/s42256-023-00692-8)\.
- \(16\)Jules White, Quchen Fu, Sam Hays, Michael Sandborn, Carlos Olea, Henry Gilbert, Ashraf Elnashar, Jesse Spencer\-Smith, and Douglas C\. Schmidt\. 2023\.A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT\.In*Proceedings of the 30th Conference on Pattern Languages of Programs \(PLoP ’23\)*\. The Hillside Group, USA, Article 5, 1–31\.
- \(17\)Tianyu Gao, Howard Yen, Jiatong Yu, and Danqi Chen\. 2023\.Enabling Large Language Models to Generate Text with Citations\.In*Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing*, pages 6465–6488, Singapore\. Association for Computational Linguistics\.
- \(18\)IBM\. 2019\.Telco customer churn sample data set\.IBM Cognos Analytics Samples\. Retrieved July 28, 2026 from[https://community\.ibm\.com/community/user/blogs/steven\-macko/2019/07/11/telco\-customer\-churn\-1113](https://community.ibm.com/community/user/blogs/steven-macko/2019/07/11/telco-customer-churn-1113)\.
- \(19\)Ramaravind K\. Mothilal, Amit Sharma, and Chenhao Tan\. 2020\.Explaining machine learning classifiers through diverse counterfactual explanations\.In*Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency \(FAT\* ’20\)*\. Association for Computing Machinery, New York, NY, USA, 607–617\.[https://doi\.org/10\.1145/3351095\.3372850](https://doi.org/10.1145/3351095.3372850)\.
- \(20\)Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen\-tau Yih\. 2020\.Dense Passage Retrieval for Open\-Domain Question Answering\.In*Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\)*, pages 6769–6781, Online\. Association for Computational Linguistics\.
- \(21\)Wouter Verbeke, Karel Dejaeger, David Martens, Joon Hur, and Bart Baesens\. 2012\.New insights into churn prediction in the telecommunication sector: A profit driven data mining approach\.*European Journal of Operational Research*, 218\(1\), 211–229\.
- \(22\)Ali Vanderveld, Addhyan Pandey, Angela Han, and Rajesh Parekh\. 2016\.An Engagement\-Based Customer Lifetime Value System for E\-commerce\.In*Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining \(KDD ’16\)*\. Association for Computing Machinery, New York, NY, USA, 293–302\.[https://doi\.org/10\.1145/2939672\.2939693](https://doi.org/10.1145/2939672.2939693)\.
- \(23\)Katerina Lepenioti, Alexandros Bousdekis, Dimitris Apostolou, and Gregoris Mentzas\. 2020\.Prescriptive analytics: Literature review and research challenges\.*Int\. J\. Inf\. Manag\.*50, C \(Feb 2020\), 57–70\.[https://doi\.org/10\.1016/j\.ijinfomgt\.2019\.04\.003](https://doi.org/10.1016/j.ijinfomgt.2019.04.003)\.
- \(24\)Ron Artstein and Massimo Poesio\. 2008\.Inter\-coder agreement for computational linguistics\.*Comput\. Linguist\.*34, 4 \(December 2008\), 555–596\.[https://doi\.org/10\.1162/coli\.07\-034\-R2](https://doi.org/10.1162/coli.07-034-R2)\.
- \(25\)Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi\. 2024\.Self\-RAG: Learning to retrieve, generate, and critique through self\-reflection\.In*Proceedings of the 12th International Conference on Learning Representations \(ICLR\)*\.
- \(26\)Or Honovich, Roee Aharoni, Jonathan Herzig, Hagai Taitelbaum, Doron Kukliansy, Vered Cohen, Thomas Scialom, Idan Szpektor, Avinatan Hassidim, and Yossi Matias\. 2022\.TRUE: Re\-evaluating Factual Consistency Evaluation\.In*Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies*, pages 3905–3920, Seattle, United States\. Association for Computational Linguistics\.
- \(27\)Wenhan Xiong, Xiang Lorraine Li, Srini Iyer, Jingfei Du, Patrick Lewis, William Yang Wang, Yashar Mehdad, Wen\-tau Yih, Sebastian Riedel, Douwe Kiela, and Barlas Oguz\. 2021\.Answering complex open\-domain questions with multi\-hop dense retrieval\.In*Proceedings of the 9th International Conference on Learning Representations \(ICLR\)*\.

These appendices provide the implementation and diagnostic details that complement the main results\. They document the exact action\-to\-feature transformations and eligibility rules, the Stage 3\-a prompt template, the base\-classifier comparison for counterfactual scoring, the action\-space scaling stress test, the per\-reason weak\-label agreement breakdown, the baseline prescription mappings used in the main experiments, and the implementation and reproducibility details for all reported experiments\.

## Appendix AAction\-to\-Feature Transformations and Eligibility Rules

Table[10](https://arxiv.org/html/2609.09766#A1.T10)specifies, for each action, the eligibility condition checked in Stage 2, the exact feature transformation applied to construct the counterfactual profile, and the researcher\-defined relative cost, taken directly from the study code\.A\_CONTRACT\_24Mchanges only theContractfield and applies no monthly\-charge discount; onlyA\_PRICE\_DISCOUNTreduces charges, by 8%\. Eligibility is enforced in Stage 2, which excludes ineligible actions before scoring and additionally restrictsA\_PRICE\_DISCOUNTto customers withMonthlyCharges≥80\{\\geq\}80\.

Table 10\.Action\-to\-feature transformations, eligibility conditions, and relative costs, as implemented in the study code \(apply\_action,eligible\)\.
## Appendix BStage 3\-a Prompt Template

This appendix gives a condensed schematic of the Stage 3\-a prompt\. The actual prompt uses a fixed plain\-text schema: the model returns aReason:line, aPlan:line, and two rationale bullets, which are parsed downstream\. The version shown is the*v2*scaled\-evaluation prompt used for the 79\.4% diagnostic result, the explanations evaluated in Table[7](https://arxiv.org/html/2609.09766#S5.T7), and the 136\-case retrieval ablation in Table[8](https://arxiv.org/html/2609.09766#S5.T8)\. Table[6](https://arxiv.org/html/2609.09766#S5.T6)used an earlier model\-comparison formulation\. The*v3*prompt that yields 90\.4% adds an enforcement block comprising lowest\-step\-wins tie\-breaking with worked examples and an explicit all\-three\-conditions requirement for Step 4; it is otherwise identical to the v2 scaled\-evaluation prompt\.

The design encodes three choices, each addressing a failure observed during development\. First, the retrievedcf\_actionis deliberately withheld, so the model reasons forward from the customer profile to a reason label rather than rationalizing backward from an action it has already been shown\. Second, the six classification rules are listed explicitly and in priority order, turning reason assignment into a deterministic rule\-application task\. Third, the*CoT Instruction*block asks the model to check each rule in turn and state the first one it satisfies before committing to a label; our refinement experiments found this the single most effective ingredient for correcting step\-order errors\. Finally, thePlanoutput is instructed to select from the retrieved action set; out\-of\-set action IDs are counted as auxiliary\-plan violations\.

\[System\]

Youareachurnretentionanalyst\.Givenacustomerprofile

andretrievedinterventiondocuments,classifytheprimary

churnreasonandrecommendoneaction\.

\[Input\]

Customer:\{"Contract":"Month\-to\-month","tenure":3,

"MonthlyCharges":85\.5,"PaymentMethod":

"Electroniccheck","TechSupport":"No",\.\.\.\}

Churnprobability:f\(x\)=0\.82

Retrievedactions:\[A\_PRICE\_DISCOUNT,A\_CONTRACT\_24M,\.\.\.\]

\(cf\_actionisNOTprovided\)

\[ClassificationRules\-\-evaluateinpriorityorder\]

Step1:MonthlyCharges\>=80ANDMonth\-to\-month

\-\>PRICE\_SENSITIVE

Step2:Month\-to\-monthANDtenure<=6\-\>CONTRACT\_RISK

Step3:Electroniccheckpayment\-\>PAYMENT\_FRICTION

Step4:InternetService\!=NoANDTechSupport=No

ANDOnlineSecurity=No\-\>SUPPORT\_DEFICIT

Step5:tenure<=3\-\>LOW\_ENGAGEMENT

Step6:\(otherwise\)\-\>OTHER

\[CoTInstruction\]

Evaluateeachstepinorder\.Statewhichconditionisfirst

satisfied\.Assignthatlabel\.Donotinferreasonfrom

therecommendedaction\.

\[Output\-\-fixedplain\-textschema\]

Reason:<LABEL\>

Plan:<ACTION\_IDfromretrievedsetonlyorNONE\>

\-<profile\-groundedrationalebullet1\>

\-<profile\-groundedrationalebullet2\>

## Appendix CBase Classifier Comparison for CF Scoring

We compare four calibrated base models—LR with sigmoid calibration and three tree ensembles with isotonic calibration \(gradient boosting, random forest, histogram gradient boosting\)—on predictive quality and CF\-surface smoothness\. Smoothness is measured by sweeping a graded price discount \(0–30% in 31 steps\) on the 313 high\-risk cases and recording how the predicted churn probability responds\.*Distinct*is the fraction of distinct probability values along the sweep \(1\.0 = fully continuous\)\.*Flat*is the fraction of adjacent steps with\|Δ​p\|<10−4\|\\Delta p\|<10^\{\-4\}, i\.e\., dead zones where a small feature perturbation registers no change in predicted risk\.*Max jump*is the mean largest single\-step\|Δ​p\|\|\\Delta p\|, capturing discontinuous cliffs\. Table[11](https://arxiv.org/html/2609.09766#A3.T11)below reports the results\. Gradient boosting attains marginally higher ROC\-AUC of 0\.845 against 0\.842 and PR\-AUC of 0\.659 against 0\.633, confirming a marginal ranking advantage for gradient boosting\. All three tree models nevertheless produce coarse, discontinuous CF surfaces, whereas logistic regression yields the fully continuous, graded response that Stage 2 requires\.

The practical takeaway is that counterfactual scoring rewards a smooth probability surface more than raw ranking power\. A simple, well\-calibrated logistic regression therefore serves CARRE better than a stronger tree ensemble: it keeps predicted\-risk changes estimable, it keeps the pipeline easy to interpret and deploy, and it gives up almost nothing in ranking accuracy\. For these reasons we adopt calibrated logistic regression as CARRE’s base estimator\.

Table 11\.Base classifier comparison for CF scoring: predictive quality vs\. CF\-surface smoothness \(313 high\-risk cases; graded 0–30% discount sweep, 31 steps\)\.
## Appendix DAction\-Space Scaling Stress Test

This is a retrieval\-only stress test: the six action documents were augmented with up to 394 synthetic lever documents \(retrieval documents, not executable actions\)—parameter variants of contract terms, discount tiers, segment\-targeted offers, and channel campaigns, written in the same document style as the real actions—and the downstream CF and LLM stages were not run\. For each library size we measure retrieval runtime and the survival of the CF\-oracle action in the top\-kkover all 313 high\-risk cases \(Table[12](https://arxiv.org/html/2609.09766#A4.T12)\)\. Query encoding \(≈\\approx1 ms\) dominates end\-to\-end latency regardless of library size, and index build stays under 7 seconds\. Exact\-action Recall@5 collapses as soon as near\-duplicate parameter variants are added \(\|𝒜\|≥12\|\\mathcal\{A\}\|\{\\geq\}12\), because inner\-product ranking cannot distinguish among semantically near\-identical variants of the same lever\. Family\-level recall—whether*any*document of the oracle action’s lever family survives in the top\-5—remains high \(0\.92–1\.00\) up to\|𝒜\|=100\|\\mathcal\{A\}\|\{=\}100but collapses beyond\|𝒜\|≈200\|\\mathcal\{A\}\|\{\\approx\}200, once a single high\-similarity lever family grows large enough to monopolize the entire top\-kk\. Hierarchical retrieval \(family\-level candidate grouping followed by within\-family CF simulation\) or diversity\-aware retrieval such as maximal marginal relevance addresses this failure mode\.

Table 12\.Action\-space scaling stress test \(313 high\-risk cases; synthetic levers styled after the real action documents\)\. Query latency is encoding \+ search per query\. Exact = CF\-oracle action ID retrieved; Family = any document of the oracle’s lever family retrieved\.
## Appendix EPer\-Reason Weak\-Label Agreement

Table[13](https://arxiv.org/html/2609.09766#A5.T13)breaks down the four\-model comparison of Section[5\.2](https://arxiv.org/html/2609.09766#S5.SS2)by reason category\. The breakdown makes clear that the aggregate agreement gap is not a uniform accuracy difference but is concentrated in a few categories\. GPT\-4o classifies almost every category correctly, whereas the weaker models fail selectively\. GPT\-4o\-mini collapses on the price and contract categories, and both Llama\-3\.3\-70B and Qwen3\-32B score zero onCONTRACT\_RISKandOTHERwhile still handling simpler categories such asPRICE\_SENSITIVE\. This selective failure, rather than an even drop across all categories, is what separates the models, and it suggests that the smaller models fall back on a few surface cues instead of applying the full priority order\.

Table 13\.Per\-reason weak\-label agreement by model \(30\-case set, 5 cases per reason\)\.
## Appendix FBaseline Prescription Mappings

Table[14](https://arxiv.org/html/2609.09766#A6.T14)gives the rule\-based and SHAP\-based baseline mappings, taken directly from the study code\. The rule\-based baseline applies the conditions in priority order and returnsNONEwhen none holds\. The SHAP baseline maps the feature with the highest positive churn\-directed SHAP value to a candidate action; ties and ineligible targets fall through a fixed priority list\. The random baseline draws one eligible action per customer under a fixed per\-customer seed \(no averaging over repetitions\)\.

Table 14\.Rule\-based \(profile condition→\\toaction, in priority order\) and SHAP\-based \(feature→\\toaction\) baseline mappings, as implemented in the study code\.
## Appendix GImplementation and Reproducibility Details

All experiments use seed 42 and an 80/20 stratified split \(5,634 train / 1,409 test\)\. After droppingcustomerIDand theChurnlabel, 19 predictors remain\. The four numeric features \(SeniorCitizen,tenure,MonthlyCharges,TotalCharges\) are median\-imputed and standardized; categorical features are most\-frequent\-imputed and one\-hot encoded \(unknown categories ignored\); all preprocessing is fit on the training split only\. The churn estimator is aCalibratedClassifierCVwrappingLogisticRegression\(C=1\.0C\{=\}1\.0,max\_iter=2000\{=\}2000, class\-balanced\) with sigmoid calibration under 3\-fold cross\-validation\. Bootstrap confidence intervals use 2,000 per\-method \(marginal\) resamples at seed 42\. Retrieval usesall\-MiniLM\-L6\-v2embeddings, L2\-normalized, in a FAISSIndexFlatIP\(cosine similarity\); the scaling stress test generates its synthetic lever documents deterministically at seed 42\. LLM calls usegpt\-4o,gpt\-4o\-mini,llama\-3\.3\-70b\-versatile, andqwen/qwen3\-32bat temperature 0, with the provider\-default top\-pp, no API seed, a 400\-token cap, and up to three API retries on transient errors; malformed outputs are reparsed by a fixed rule and, on repeated failure, counted as a violation\. We used provider aliases rather than pinned model snapshots, which limits exact reproducibility of the LLM evaluations\.

Similar Articles

Explanations-Driven Active Feature Acquisition for Algorithmic Recourse

arXiv cs.LG

The paper proposes an Explanation-Driven Feature Acquisition (EDFA) method that unifies counterfactual, semifactual, and alterfactual explanations to jointly address algorithmic recourse and feature acquisition, aiming for lower-cost and more actionable recourse with validity guarantees.