Uncertainty-Guided LLM Semantic Augmentation for Heterogeneous Treatment Effect Estimation
Summary
This paper proposes CURL, a plug-in adapter that uses estimator uncertainty to allocate pretrained LLM semantic capacity for improving heterogeneous treatment effect (CATE) estimation. It introduces two role-conditioned prompts to construct assignment- and heterogeneity-oriented representations, improving ten host learners on four benchmarks.
View Cached Full Text
Cached at: 07/30/26, 09:59 AM
# Uncertainty-Guided LLM Semantic Augmentation for Heterogeneous Treatment Effect Estimation Source: [https://arxiv.org/html/2607.26599](https://arxiv.org/html/2607.26599) Jialu XuMIIT Key Laboratory of Data and Decision Intelligence, Beihang UniversityBeijingChina[xujialu08@buaa\.edu\.cn](https://arxiv.org/html/2607.26599v1/mailto:[email protected])Mengkun LiangMIIT Key Laboratory of Data and Decision Intelligence, Beihang UniversityBeijingChina[mengkun@buaa\.edu\.cn](https://arxiv.org/html/2607.26599v1/mailto:[email protected]),Guannan LiuMIIT Key Laboratory of Data and Decision Intelligence, Beihang UniversityBeijingChina[liugn@buaa\.edu\.cn](https://arxiv.org/html/2607.26599v1/mailto:[email protected]),Xiaojie MaoSchool of Economics and Management, Tsinghua UniversityBeijingChina[maoxj@sem\.tsinghua\.edu\.cn](https://arxiv.org/html/2607.26599v1/mailto:[email protected])andJunjie WuMIIT Key Laboratory of Data and Decision Intelligence, Beihang UniversityBeijingChina[wujj@buaa\.edu\.cn](https://arxiv.org/html/2607.26599v1/mailto:[email protected]) ###### Abstract\. Estimating heterogeneous treatment effects is central to targeted interventions, such as personalized promotions and precision medicine\. We focus on the conditional average treatment effect \(CATE\), a standard estimand for characterizing such heterogeneity\. Even under standard identification conditions, finite\-sample CATE estimation requires learning the nuisance structure for covariate adjustment and treatment\-effect heterogeneity, often together with an effective representation ofXX\. Raw numerical and categorical encodings can leave semantic relations and higher\-order interactions implicit, making this joint task locally unstable\. A motivating study further shows that this instability appears through partially separable assignment\- and heterogeneity\-side channels\. Building on this observation, we propose CURL \(Causal Uncertainty\-guided Representation Learning\), a plug\-in adapter that uses estimator uncertainty to allocate pretrained semantic capacity to locally unstable units\. CURL queries a frozen LLM through two role\-conditioned prompts, constructs assignment\- and heterogeneity\-oriented representations from the observed covariates, and routes them through separated pathways\. On four benchmarks, CURL improves ten host learners in most settings, while ablation, refinement\-dynamics, route\-reassignment, and probe analyses support the intended design and roles of the two channels\. Heterogeneous Treatment Effects, Causal Machine Learning, Large Language Models, Representation Learning ## 1\.Introduction Estimating heterogeneous treatment effects is central to individualized interventions in precision medicine, targeted marketing, and policy evaluation; we focus on the conditional average treatment effect \(CATE\), which characterizes how expected effects vary with observed covariates\(Shalitet al\.,[2017](https://arxiv.org/html/2607.26599#bib.bib4); Künzelet al\.,[2019](https://arxiv.org/html/2607.26599#bib.bib2); Nie and Wager,[2021](https://arxiv.org/html/2607.26599#bib.bib26)\)\. Reliable CATE estimation requires the covariatesXXto support confounding adjustment and to reveal systematic variation in treatment response\(Imbens and Rubin,[2015](https://arxiv.org/html/2607.26599#bib.bib45)\)\. Estimation can fail for two reasons: relevant individual\-level causal information may be absent fromXX, threatening identification, or, despite identification, the raw encoding ofXXmay make adjustment and heterogeneity structure difficult for finite\-sample learners to exploit\. In this paper, we focus on the second and less\-examined failure mode\. We term the gap between the raw covariate encoding and a task\-effective representation for finite\-sample CATE learning*representational lossiness*\. The notion is relative to the encoding, learner, and sample size, rather than an intrinsic information loss inXXor a failure of identification\. For example, diagnosis codes may omit semantic proximity, while customer categories may obscure behavioral patterns that jointly shape targeting and response\. Meta\-learning, balanced\-representation, and latent\-variable methods provide flexible tools for CATE estimation\(Künzelet al\.,[2019](https://arxiv.org/html/2607.26599#bib.bib2); Shalitet al\.,[2017](https://arxiv.org/html/2607.26599#bib.bib4); Shiet al\.,[2019](https://arxiv.org/html/2607.26599#bib.bib5); Louizoset al\.,[2017](https://arxiv.org/html/2607.26599#bib.bib14)\), but learn representations and associated functions mainly from the observed sample\. When semantic relations and higher\-order interactions remain implicit, limited data cannot effectively share statistical strength across related observations, leaving local estimates unstable\. This motivates an external inductive bias that reorganizes observed information without adding new individual\-level facts\. Figure 1\.Motivating study on IHDP using an S\-Learner with an auxiliary propensity head\. The left panel places samples on the propensity–CATE uncertainty plane\. Q1–Q4 are the four corner groups, and the rest form the middle group\. The right panel reports column\-normalized T\-Brier, Y\-RMSE, and CATE\-PEHE\. Q2 is dominated by assignment error, Q3 by effect\-estimation error, and Q4 by both\.Large language models \(LLMs\) offer such an inductive bias because pretraining captures broad semantic regularities that may be difficult to learn from a single finite observational sample, including in structured tabular prediction\(Wenet al\.,[2024](https://arxiv.org/html/2607.26599#bib.bib57)\)\. Prior work has examined LLMs for causal reasoning and discovery\(Kıcımanet al\.,[2024](https://arxiv.org/html/2607.26599#bib.bib6); Duet al\.,[2025](https://arxiv.org/html/2607.26599#bib.bib8)\), and recent estimators have used LLM predictions with textual confounders or unstructured records\(Veljanovski and Wood\-Doughty,[2024](https://arxiv.org/html/2607.26599#bib.bib9); Dhawanet al\.,[2024](https://arxiv.org/html/2607.26599#bib.bib29); Chenet al\.,[2024](https://arxiv.org/html/2607.26599#bib.bib30)\)\. However, how pretrained semantic knowledge should be incorporated into CATE estimators for structured covariates remains open\. Applying a common embedding to every sample is costly and can introduce irrelevant variation, while an undifferentiated representation may mix information needed for assignment modeling with information needed for heterogeneous response\. A useful design must therefore determine where semantic capacity is needed, construct representations for distinct statistical roles, and route them into the appropriate components of the estimator\. Figure[1](https://arxiv.org/html/2607.26599#S1.F1)provides empirical support for a selective and role\-conditioned design\. We fit an S\-Learner with an auxiliary propensity head on IHDP, derive CATE from its treatment\-conditioned outcome predictions, and estimate both uncertainties via Monte Carlo \(MC\) dropout\(Gal and Ghahramani,[2016](https://arxiv.org/html/2607.26599#bib.bib15)\)\. Q2 has the largest normalized treatment Brier error but low CATE error, whereas Q3 shows the opposite pattern; Q4 is difficult on both dimensions\. These results indicate that local unreliability varies across samples and cannot be captured by a single score\. Propensity uncertainty reflects instability in treatment\-selection modeling relevant to covariate adjustment, whereas CATE uncertainty reflects instability in treatment\-response heterogeneity\. We therefore use both as operational allocation signals rather than evidence of missing causal variables, consistent with prior work on uncertainty\-based identification of unreliable treatment\-effect estimates\(Jessonet al\.,[2020](https://arxiv.org/html/2607.26599#bib.bib47)\)\. Motivated by these empirical patterns, we propose CURL, an uncertainty\-guided representation adapter for CATE estimation underrepresentational lossiness\. CURL uses propensity and CATE uncertainty to prioritize a fixed proportion of locally unstable samples for LLM\-based semantic augmentation, then constructs independent assignment\-oriented and heterogeneity\-oriented representations from their observed profiles\. The former refines the shared covariate representation, while the latter enters the effect component as a separate feature block and is excluded from propensity estimation\. The augmented set is updated as the host estimator evolves\. In this way, CURL expands the effective hypothesis class without changing the observed information set, target estimand, or identifying assumptions\. Our contributions are summarized as follows\. - •We distinguish representational lossiness from causal information insufficiency, characterizing it as a finite\-sample bottleneck in exploiting semantic relations left implicit by raw covariate encodings\. - •We develop CURL, a plug\-in adapter that uses uncertainty to allocate assignment\- and heterogeneity\-oriented semantic representations and integrates them through asymmetric routing, progressive refinement, and training\-calibrated test\-time inference\. - •We evaluate CURL across four benchmarks and ten CATE estimators, with analyses of predictive performance, robustness, refinement dynamics, and semantic\-channel roles\. ## 2\.Related Work ### 2\.1\.Representation Learning for CATE Estimation CATE estimation has developed along three major paradigms\.*Meta\-learning*methods reduce effect estimation to supervised prediction, including the S\-, T\-, and X\-learner\(Künzelet al\.,[2019](https://arxiv.org/html/2607.26599#bib.bib2)\)and the R\-learner\(Nie and Wager,[2021](https://arxiv.org/html/2607.26599#bib.bib26)\)\.*Balanced\-representation*methods learn treatment\-invariant representations, as in TARNet, CFRNet\(Shalitet al\.,[2017](https://arxiv.org/html/2607.26599#bib.bib4)\), and DragonNet\(Shiet al\.,[2019](https://arxiv.org/html/2607.26599#bib.bib5)\)\.*Deep latent\-variable*methods infer hidden factors from proxies inXX, including CEVAE\(Louizoset al\.,[2017](https://arxiv.org/html/2607.26599#bib.bib14)\), TEDVAE\(Zhanget al\.,[2021](https://arxiv.org/html/2607.26599#bib.bib18)\),β\\beta\-Intact\-VAE\(Wu and Fukumizu,[2022](https://arxiv.org/html/2607.26599#bib.bib21)\), CFDiVAE\(Xuet al\.,[2024](https://arxiv.org/html/2607.26599#bib.bib22)\), and iVAE\(Khemakhemet al\.,[2020](https://arxiv.org/html/2607.26599#bib.bib20)\), with further extensions to structure uncertainty, adjustment\-feature selection, and disentangled or robust representations\(Tran and Zheleva,[2022](https://arxiv.org/html/2607.26599#bib.bib52); Wanget al\.,[2023](https://arxiv.org/html/2607.26599#bib.bib55); Zhonget al\.,[2022](https://arxiv.org/html/2607.26599#bib.bib53); Liet al\.,[2024](https://arxiv.org/html/2607.26599#bib.bib44); Wanget al\.,[2024](https://arxiv.org/html/2607.26599#bib.bib56)\)\. Despite these advances, most representations are learned primarily from the study sample and may remain unreliable when the encoding ofXXis poorly aligned with the estimator’s effective hypothesis class\. Sensitivity to model misspecification and proxy construction further motivates using pretrained semantic structure to help finite\-sample estimators exploit information already contained inXX\(Rissanen and Marttinen,[2021](https://arxiv.org/html/2607.26599#bib.bib27); Zhuet al\.,[2025](https://arxiv.org/html/2607.26599#bib.bib23)\)\. ### 2\.2\.LLMs for Causal Augmentation Recent work has explored LLMs for causal analysis along two lines\. On the*qualitative\-structure*side, LLMs have been used for causal discovery and reasoning from text\(Kıcımanet al\.,[2024](https://arxiv.org/html/2607.26599#bib.bib6); Liuet al\.,[2025](https://arxiv.org/html/2607.26599#bib.bib24)\), causal graph construction\(Banet al\.,[2025](https://arxiv.org/html/2607.26599#bib.bib7); Duet al\.,[2025](https://arxiv.org/html/2607.26599#bib.bib8)\), and formal causal\-query benchmarks\(Jinet al\.,[2023](https://arxiv.org/html/2607.26599#bib.bib12)\), although current models may still rely on shallow Rung\-1 associations\(Chiet al\.,[2024](https://arxiv.org/html/2607.26599#bib.bib28)\)\. On the*quantitative\-estimation*side, DoubleLingo uses LLM\-based nuisance models with text confounders\(Veljanovski and Wood\-Doughty,[2024](https://arxiv.org/html/2607.26599#bib.bib9)\); NATURAL leverages LLM\-derived conditional information from text\(Dhawanet al\.,[2024](https://arxiv.org/html/2607.26599#bib.bib29)\); proximal text\-based inference extracts proxies for proximal identification\(Chenet al\.,[2024](https://arxiv.org/html/2607.26599#bib.bib30)\); and GATE generates counterfactual outcomes for small\-sample structured data\(Huynhet al\.,[2025](https://arxiv.org/html/2607.26599#bib.bib43)\)\. These studies primarily use LLMs to process textual inputs, construct auxiliary variables, or directly predict causal quantities\. Less attention has been paid to using pretrained semantic structure as an inductive bias that helps CATE estimators better exploit structured covariates under limited data\. ### 2\.3\.Uncertainty Estimation in Causal Inference Epistemic uncertainty has long been studied in Bayesian machine learning, from the Laplace approximation\(MacKay,[1992](https://arxiv.org/html/2607.26599#bib.bib41)\)to MC dropout\(Gal and Ghahramani,[2016](https://arxiv.org/html/2607.26599#bib.bib15)\), deep ensembles\(Lakshminarayananet al\.,[2017](https://arxiv.org/html/2607.26599#bib.bib31)\), and variational inference\(Blundellet al\.,[2015](https://arxiv.org/html/2607.26599#bib.bib39)\), withKendall and Gal \([2017](https://arxiv.org/html/2607.26599#bib.bib32)\)distinguishing aleatoric and epistemic uncertainty\. In decision\-making, uncertainty supports exploration through UCB\-style bandits and Thompson sampling\(Abbasi\-Yadkoriet al\.,[2011](https://arxiv.org/html/2607.26599#bib.bib33); Zhouet al\.,[2020](https://arxiv.org/html/2607.26599#bib.bib34); Xuet al\.,[2022](https://arxiv.org/html/2607.26599#bib.bib40)\), or conservatism in offline settings\(Anet al\.,[2021](https://arxiv.org/html/2607.26599#bib.bib35); Baiet al\.,[2022](https://arxiv.org/html/2607.26599#bib.bib37); Wuet al\.,[2021](https://arxiv.org/html/2607.26599#bib.bib36)\)\. Closer to causal estimation,Zhanget al\.\([2023](https://arxiv.org/html/2607.26599#bib.bib38)\)account for logging\-policy uncertainty in MSE\-optimal importance weighting, whereas CATE methods mainly use uncertainty for confidence quantification, failure detection, or budgeted sample acquisition\(Alaa and van der Schaar,[2017](https://arxiv.org/html/2607.26599#bib.bib48); Jessonet al\.,[2020](https://arxiv.org/html/2607.26599#bib.bib47); Wenet al\.,[2025](https://arxiv.org/html/2607.26599#bib.bib59)\)\. In contrast, CURL uses uncertainty to allocate semantic augmentation to locally unstable samples\. ## 3\.Problem Formulation Let𝒟n=\{\(Xi,Ti,Yi\)\}i=1n∼Pn\\mathcal\{D\}\_\{n\}=\\\{\(X\_\{i\},T\_\{i\},Y\_\{i\}\)\\\}\_\{i=1\}^\{n\}\\sim P^\{n\}be an i\.i\.d\. observational sample, whereXi∈𝒳⊆ℝdX\_\{i\}\\in\\mathcal\{X\}\\subseteq\\mathbb\{R\}^\{d\}contains pretreatment covariates,Ti∈\{0,1\}T\_\{i\}\\in\\\{0,1\\\}is a binary treatment, andYi=TiYi\(1\)\+\(1−Ti\)Yi\(0\)Y\_\{i\}=T\_\{i\}Y\_\{i\}\(1\)\+\(1\-T\_\{i\}\)Y\_\{i\}\(0\)\. We study the CATE,i\.e\.,τ\(x\)=𝔼\[Y\(1\)−Y\(0\)∣X=x\]\\tau\(x\)=\\mathbb\{E\}\[Y\(1\)\-Y\(0\)\\mid X=x\]\. We assume conditional unconfoundedness\{Y\(0\),Y\(1\)\}⟂T∣X\\\{Y\(0\),Y\(1\)\\\}\\perp T\\mid Xand overlapϵ≤e\(x\)≤1−ϵ\\epsilon\\leq e\(x\)\\leq 1\-\\epsilon, wheree\(x\)=Pr\(T=1∣X=x\)e\(x\)=\\Pr\(T=1\\mid X=x\)\(Imbens and Rubin,[2015](https://arxiv.org/html/2607.26599#bib.bib45)\)\. Because treatment assignment may depend onXXin observational data,e\(x\)e\(x\)summarizes systematic treatment\-selection differences across the covariate space and, depending on the estimator, supports weighting, balancing, or residualization for covariate adjustment\(Rosenbaum and Rubin,[1983](https://arxiv.org/html/2607.26599#bib.bib54); Nie and Wager,[2021](https://arxiv.org/html/2607.26599#bib.bib26)\)\. Together with consistency, these assumptions identifyτ\(x\)=μ1\(x\)−μ0\(x\)\\tau\(x\)=\\mu\_\{1\}\(x\)\-\\mu\_\{0\}\(x\), whereμt\(x\)=𝔼\[Y∣T=t,X=x\]\\mu\_\{t\}\(x\)=\\mathbb\{E\}\[Y\\mid T=t,X=x\]\. Thus,XXis sufficient for causal adjustment, and unmeasured confounding is outside our scope\. Letℋ\\mathcal\{H\}be the function class of a host estimator\. Its best available CATE approximation under the raw encoding ofXXis \(1\)τℋ∗∈argminf∈ℋ𝔼X∼PX\[\{f\(X\)−τ\(X\)\}2\]\.\\tau\_\{\\mathcal\{H\}\}^\{\*\}\\in\\arg\\min\_\{f\\in\\mathcal\{H\}\}\\mathbb\{E\}\_\{X\\sim P\_\{X\}\}\\left\[\\left\\\{f\(X\)\-\\tau\(X\)\\right\\\}^\{2\}\\right\]\. Identification alone does not ensure thatτℋ∗\\tau\_\{\\mathcal\{H\}\}^\{\*\}is easy to express or learn from finite data\. A learner must estimate the assignment and outcome structure needed for adjustment and how treatment effects vary withXX, often while learning an effective representation from the same sample\. When the raw numerical and categorical encoding ofXXleaves semantic relations and higher\-order interactions implicit, these functions may become unnecessarily complex within the host class, limiting statistical sharing across related observations and destabilizing local estimates\. This is precisely the*representational lossiness*introduced in Section 1\. Concretely, althoughXXidentifiesτ\(x\)\\tau\(x\), relevant adjustment and heterogeneity structure remains difficult to recover from its raw encoding\. The gap is thus relative to the encoding, learner, and sample size, rather than an intrinsic information loss inXXor a failure of identification\. ## 4\.Methodology ### 4\.1\.Framework Overview To address therepresentational lossiness,CURL\(CausalUncertainty\-guidedRepresentationLearning\) augments a host estimatorFθF\_\{\\theta\}with a plug\-in representation adapter that makes adjustment\- and heterogeneity\-related structure inXXeasier to exploit under finite data\. CURL preserves the host\-specific prediction structure and main objective, while introducing adapter parametersη\\etafor the semantic projectors, gates, routing components, and any auxiliary propensity head\. We writeΘ=\(θ,η\)\\Theta=\(\\theta,\\eta\)for the joint trainable state andFΘF\_\{\\Theta\}for the wrapped estimator\. Host\-specific instantiations are provided in Appendix[E](https://arxiv.org/html/2607.26599#A5)\. As illustrated in Figure[2](https://arxiv.org/html/2607.26599#S4.F2), CURL operates through four coupled mechanisms\.*Uncertainty\-Guided Semantic Allocation*\(Section[4\.2](https://arxiv.org/html/2607.26599#S4.SS2)\) uses two epistemic\-uncertainty scores, one for the propensity component and one for the CATE component, to prioritize units with high local predictive instability\.*Role\-Conditioned Semantic Augmentation and Routing*\(Section[4\.3](https://arxiv.org/html/2607.26599#S4.SS3)\) queries a frozen LLM independently for each selected unit to build an assignment\-oriented and a heterogeneity\-oriented representation from its observed covariates\.*Role\-Aware Prediction Routing*\(Section[4\.3\.3](https://arxiv.org/html/2607.26599#S4.SS3.SSS3)\) injects the two representations into the host through separated pathways, so that heterogeneity\-oriented information never reaches the propensity component\.*Progressive Refinement and Test\-Time Inference*\(Section[4\.4](https://arxiv.org/html/2607.26599#S4.SS4)\) repeats this diagnosis asΘ\\Thetaevolves, and a calibrated procedure extends the same allocation and routing rules to unseen units at test time\. We index training rounds byr∈\{0,…,R−1\}r\\in\\\{0,\\ldots,R\-1\\\}and writeΘ\(r\)\\Theta^\{\(r\)\}andF\(r\):=FΘ\(r\)F^\{\(r\)\}:=F\_\{\\Theta^\{\(r\)\}\}for the parameter state and wrapped estimator at the start of roundrr; this superscript is used consistently in what follows\. Becauseη\\etaonly augments, and never replaces, the host’s own computation,F\(0\)F^\{\(0\)\}with no cached semantic input reduces exactly to the unaugmented hostFθF\_\{\\theta\}\. Next, we describe the four modules in detail\. Figure 2\.Overall architecture of CURL\. Uncertainty\-guided diagnosis identifies samples that receive semantic augmentation; dual\-channel LLM encoding constructs assignment\- and heterogeneity\-oriented representations; and role\-aware integration routes them into the corresponding components of the host estimator\. The procedure progressively refreshes the uncertainty diagnosis as the wrapped estimator evolves\. ### 4\.2\.Uncertainty\-Guided Semantic Allocation To diagnose the current estimatorF\(r\)F^\{\(r\)\}, we keep its normalization layers in evaluation mode and dropout layers active, and performMMstochastic forward passes using MC dropout\(Gal and Ghahramani,[2016](https://arxiv.org/html/2607.26599#bib.bib15)\)\. Writinge^i\(r,m\)\\widehat\{e\}\_\{i\}^\{\(r,m\)\}andτ^i\(r,m\)\\widehat\{\\tau\}\_\{i\}^\{\(r,m\)\}for the propensity and CATE predictions obtained in themm\-th pass for unitii, we take the standard deviation of each prediction across theMMpasses as the corresponding uncertainty score, \(2\)se,i\(r\)=Stdm=1M\[e^i\(r,m\)\],sτ,i\(r\)=Stdm=1M\[τ^i\(r,m\)\]\.s\_\{e,i\}^\{\(r\)\}=\\operatorname\{Std\}\_\{m=1\}^\{M\}\\\!\\left\[\\widehat\{e\}\_\{i\}^\{\(r,m\)\}\\right\],\\qquad s\_\{\\tau,i\}^\{\(r\)\}=\\operatorname\{Std\}\_\{m=1\}^\{M\}\\\!\\left\[\\widehat\{\\tau\}\_\{i\}^\{\(r,m\)\}\\right\]\.These two scores capture local predictive instability separately for the assignment component and the treatment effect component\. At the initial allocation round, some host estimators do not provide a native propensity output\. To apply the same allocation rule across all hosts and rounds, we setse,i\(0\)=0s\_\{e,i\}^\{\(0\)\}=0for all units in this case\. After wrapping, CURL introduces and trains a propensity head for every host\. Consequently, from the first refinement round onward, bothse,i\(r\)s\_\{e,i\}^\{\(r\)\}andsτ,i\(r\)s\_\{\\tau,i\}^\{\(r\)\}are available and jointly determine semantic allocation\. CURL converts the two scores into a single ranking through their empirical percentile ranks\. LetPRank\(ai\)\\operatorname\{PRank\}\(a\_\{i\}\)denote the percentile rank ofaia\_\{i\}among\{aj\}j=1N\\\{a\_\{j\}\\\}\_\{j=1\}^\{N\}, and letρ∈\(0,1\]\\rho\\in\(0,1\]be the per\-round selection ratio\. The units prioritized for semantic allocation in roundrrare \(3\)𝒮\(r\)=Top⌊ρN⌋\{PRank\(se,i\(r\)\)\+PRank\(sτ,i\(r\)\):i=1,…,N\},\\mathcal\{S\}^\{\(r\)\}=\\operatorname\{Top\}\_\{\\lfloor\\rho N\\rfloor\}\\left\\\{\\operatorname\{PRank\}\\\!\\left\(s\_\{e,i\}^\{\(r\)\}\\right\)\+\\operatorname\{PRank\}\\\!\\left\(s\_\{\\tau,i\}^\{\(r\)\}\\right\):i=1,\\ldots,N\\right\\\},with ties broken deterministically so that the selection budget stays fixed\. Taking percentile ranks places the two signals on a common scale despite their different units and magnitudes, while their sum provides a unified ranking over both uncertainty sources\. Whense,i\(r\)s\_\{e,i\}^\{\(r\)\}is set to zero for all units under the exception above, the ranking reduces to CATE\-side uncertainty alone\. Because selections may overlap across rounds, previously queried units reuse their cached semantic embeddings\. Let𝒮\(r\)\\mathcal\{S\}^\{\(r\)\}denote the units selected in roundrr,𝒞\(r\)\\mathcal\{C\}^\{\(r\)\}those cached by the end of that round, and𝒬\(r\)\\mathcal\{Q\}^\{\(r\)\}the selected but uncached units\. With𝒞\(−1\)=∅\\mathcal\{C\}^\{\(\-1\)\}=\\varnothing, \(4\)𝒞\(r\)=𝒞\(r−1\)∪𝒮\(r\),𝒬\(r\)=𝒮\(r\)∖𝒞\(r−1\)\.\\mathcal\{C\}^\{\(r\)\}=\\mathcal\{C\}^\{\(r\-1\)\}\\cup\\mathcal\{S\}^\{\(r\)\},\\qquad\\mathcal\{Q\}^\{\(r\)\}=\\mathcal\{S\}^\{\(r\)\}\\setminus\\mathcal\{C\}^\{\(r\-1\)\}\. Thus, the LLM is queried only for𝒬\(r\)\\mathcal\{Q\}^\{\(r\)\}, so each training unit is queried at most once\. ### 4\.3\.Role\-Conditioned Semantic Augmentation and Routing After determining which units require additional semantic capacity in Section[4\.2](https://arxiv.org/html/2607.26599#S4.SS2), the next question is*what*semantic information these units should receive and*where*it should act within the estimator\. To answer this, we introduce the notion of a*role*,i\.e\., the function a semantic channel serves in CATE estimation: assignment modeling or treatment\-effect heterogeneity\. Accordingly, this section specifies what role\-conditioned information the selected units should receive, how strongly it is integrated, and which prediction pathways it may influence\. #### 4\.3\.1\.Role\-Conditioned Dual\-Channel Semantic Construction The units in𝒬\(r\)\\mathcal\{Q\}^\{\(r\)\}identified in the previous step are the ones for which CURL queries the LLM\. Each such unit is first rendered into a natural\-language profilepi=ψ\(xi\)p\_\{i\}=\\psi\(x\_\{i\}\)by a deterministic, dataset\-specific constructorψ\\psi\. A dataset\-level hinthinsh\_\{\\mathrm\{ins\}\}is generated once from aggregate covariate differences between the initially selected and full training populations, and is then held fixed for the remainder of training\. Neitherpip\_\{i\}norhinsh\_\{\\mathrm\{ins\}\}contains unit\-level treatment, outcome, counterfactual, ground\-truth effect, or test statistics; full field mappings, aggregation rules, and templates are deferred to Appendix[C\.3](https://arxiv.org/html/2607.26599#A3.SS3)\. From this profile, two role\-specific constructors produce an assignment\-oriented promptPA,i=PA\(pi,hins\)P\_\{A,i\}=P\_\{A\}\(p\_\{i\},h\_\{\\mathrm\{ins\}\}\), eliciting pre\-treatment context for assignment and outcomes, and a heterogeneity\-oriented promptPH,i=PH\(pi,hins\)P\_\{H,i\}=P\_\{H\}\(p\_\{i\},h\_\{\\mathrm\{ins\}\}\), eliciting stable attributes for treatment sensitivity\. These functional roles target distinct estimator pathways rather than alternative descriptions of the unit\. The prompts use independent LLM calls and caches, preventing either from becoming an autoregressive continuation of the other\. For either promptP∈\{PA,i,PH,i\}P\\in\\\{P\_\{A,i\},P\_\{H,i\}\\\}described in Appendix[C\.2](https://arxiv.org/html/2607.26599#A3.SS2), CURL extracts the final\-layer contextual hidden state at the last valid response token, before the output head\. The resulting channel embeddings are \(5\)zA,iLLM=EmbΦ\(PA,i\),zH,iLLM=EmbΦ\(PH,i\),z\_\{A,i\}^\{\\mathrm\{LLM\}\}=\\operatorname\{Emb\}\_\{\\Phi\}\(P\_\{A,i\}\),\\qquad z\_\{H,i\}^\{\\mathrm\{LLM\}\}=\\operatorname\{Emb\}\_\{\\Phi\}\(P\_\{H,i\}\),and are cached permanently once computed\. For either channelq∈\{A,H\}q\\in\\\{A,H\\\}, the sparse input supplied to the adapter in roundrrzeroes out the cached embedding for any unit not yet in𝒞\(r\)\\mathcal\{C\}^\{\(r\)\}, \(6\)zq,i\(r\)=𝕀\{i∈𝒞\(r\)\}zq,iLLM\.z\_\{q,i\}^\{\(r\)\}=\\mathbb\{I\}\\\!\\left\\\{i\\in\\mathcal\{C\}^\{\(r\)\}\\right\\\}z\_\{q,i\}^\{\\mathrm\{LLM\}\}\.Units outside𝒞\(r\)\\mathcal\{C\}^\{\(r\)\}therefore receive zero semantic input and follow the unaugmented host pathway\. #### 4\.3\.2\.Uncertainty\-Gated Asymmetric Fusion The host encoder produces a base representation for each unit,ui\(r\)=fθ\(r\)\(xi\)∈ℝdhu\_\{i\}^\{\(r\)\}=f\_\{\\theta^\{\(r\)\}\}\(x\_\{i\}\)\\in\\mathbb\{R\}^\{d\_\{h\}\}, using only the host parametersθ\(r\)\\theta^\{\(r\)\}and evolving across rounds asθ\(r\)\\theta^\{\(r\)\}is updated\. Two trainable projectorsgAg\_\{A\}andgHg\_\{H\}adapt the sparse semantic inputs intodhd\_\{h\}\-dimensional features used by the wrapped estimator, \(7\)vA,i\(r\)=gA\(zA,i\(r\)\),vH,i\(r\)=gH\(zH,i\(r\)\)\.v\_\{A,i\}^\{\(r\)\}=g\_\{A\}\(z\_\{A,i\}^\{\(r\)\}\),\\qquad v\_\{H,i\}^\{\(r\)\}=g\_\{H\}\(z\_\{H,i\}^\{\(r\)\}\)\.The strength with which each channel is injected is controlled by a scalar gate\. Each gate combines the matching uncertainty score, passed through a small network and the sigmoid functionσ\(a\)=\(1\+e−a\)−1\\sigma\(a\)=\(1\+e^\{\-a\}\)^\{\-1\}, with an indicator restricting injection to units that have actually been queried, \(8\)αi\(r\)\\displaystyle\\alpha\_\{i\}^\{\(r\)\}=𝕀\{i∈𝒞\(r\)\}σ\(MLPA\(se,i\(r\)\)\),\\displaystyle=\\mathbb\{I\}\\\!\\left\\\{i\\in\\mathcal\{C\}^\{\(r\)\}\\right\\\}\\sigma\\\!\\left\(\\operatorname\{MLP\}\_\{A\}\(s\_\{e,i\}^\{\(r\)\}\)\\right\),βi\(r\)\\displaystyle\\beta\_\{i\}^\{\(r\)\}=𝕀\{i∈𝒞\(r\)\}σ\(MLPH\(sτ,i\(r\)\)\)\.\\displaystyle=\\mathbb\{I\}\\\!\\left\\\{i\\in\\mathcal\{C\}^\{\(r\)\}\\right\\\}\\sigma\\\!\\left\(\\operatorname\{MLP\}\_\{H\}\(s\_\{\\tau,i\}^\{\(r\)\}\)\\right\)\.The cache indicator blocks unqueried units, while each gate aligns a role with its matching uncertainty:ses\_\{e\}modulates assignment semantics andsτs\_\{\\tau\}modulates heterogeneity semantics\. WritingConcat\(a,b\)\\operatorname\{Concat\}\(a,b\)for feature\-wise concatenation, CURL forms assignment\-adjusted shared and effect\-sensitive representations as \(9\)uA,i\(r\)\\displaystyle u\_\{A,i\}^\{\(r\)\}=ui\(r\)\+αi\(r\)vA,i\(r\),\\displaystyle=u\_\{i\}^\{\(r\)\}\+\\alpha\_\{i\}^\{\(r\)\}v\_\{A,i\}^\{\(r\)\},\(10\)uH,i\(r\)\\displaystyle u\_\{H,i\}^\{\(r\)\}=Concat\(uA,i\(r\),βi\(r\)vH,i\(r\)\)\.\\displaystyle=\\operatorname\{Concat\}\\\!\\left\(u\_\{A,i\}^\{\(r\)\},\\beta\_\{i\}^\{\(r\)\}v\_\{H,i\}^\{\(r\)\}\\right\)\.The fusion rules serve distinct roles\. Residual addition treats assignment semantics as a correction to the shared state, preserving its dimension and recovering the host representation when inactive\. Concatenation retains heterogeneity semantics as a separate block that the CATE operator can parameterize independently\. This asymmetry is a role\-specific routing bias, not a claim that the embeddings form an identifiable causal decomposition; their redundancy is further discouraged by the orthogonality penalty in Section[4\.4](https://arxiv.org/html/2607.26599#S4.SS4)\. #### 4\.3\.3\.Role\-Aware Prediction Routing The propensity headge,Θ\(r\)g\_\{e,\\Theta^\{\(r\)\}\}of the current wrapped estimator consumes only the assignment\-adjusted shared representation, while its host\-specific CATE operator𝒯Θ\(r\)\\mathcal\{T\}\_\{\\Theta^\{\(r\)\}\}consumesuH,i\(r\)u\_\{H,i\}^\{\(r\)\}, which already containsuA,i\(r\)u\_\{A,i\}^\{\(r\)\}through the concatenation in Eq\. \([10](https://arxiv.org/html/2607.26599#S4.E10)\), \(11\)e^Θ\(r\)\(xi\)=ge,Θ\(r\)\(uA,i\(r\)\),τ^Θ\(r\)\(xi\)=𝒯Θ\(r\)\(uH,i\(r\)\)\.\\widehat\{e\}\_\{\\Theta^\{\(r\)\}\}\(x\_\{i\}\)=g\_\{e,\\Theta^\{\(r\)\}\}\\\!\\left\(u\_\{A,i\}^\{\(r\)\}\\right\),\\quad\\widehat\{\\tau\}\_\{\\Theta^\{\(r\)\}\}\(x\_\{i\}\)=\\mathcal\{T\}\_\{\\Theta^\{\(r\)\}\}\\\!\\left\(u\_\{H,i\}^\{\(r\)\}\\right\)\.The two roles are thus enforced structurally: heterogeneity semantics cannot reach assignment prediction, whereas the CATE operator accesses both the assignment\-adjusted shared state and effect\-specific representation\. This interface covers treatment\-conditioned single\-head, two\-potential\-outcome\-head, direct\-CATE, and latent\-variable hosts; their insertion points and native objectives appear in Appendix[E](https://arxiv.org/html/2607.26599#A5)\. ### 4\.4\.Progressive Refinement and Test\-Time Inference #### 4\.4\.1\.Progressive Training Training begins by fitting the unaugmented host with its native objective, which we write asFitHost\(𝒟\)\\operatorname\{FitHost\}\(\\mathcal\{D\}\), givingF\(0\)=FitHost\(𝒟\)F^\{\(0\)\}=\\operatorname\{FitHost\}\(\\mathcal\{D\}\)and an empty cache𝒞\(−1\)=∅\\mathcal\{C\}^\{\(\-1\)\}=\\varnothing\. Each subsequent roundr=0,…,R−1r=0,\\ldots,R\-1repeats the diagnosis and allocation steps of Eqs\. \([2](https://arxiv.org/html/2607.26599#S4.E2)\)–\([4](https://arxiv.org/html/2607.26599#S4.E4)\), queries the frozen LLM only on the newly surfaced set𝒬\(r\)\\mathcal\{Q\}^\{\(r\)\}, and resumes optimization from the current parameters to obtainF\(r\+1\)F^\{\(r\+1\)\}, without reinitializing either the estimator or its optimizer\. Within such an optimization block, the cached embeddings and the frozen LLMΦ\\Phiare held fixed, while the host, the projectors, the gates, and the prediction heads are trained jointly; the uncertainty scores produced at the end of a block then serve as the gate inputs for the next one, so that training advances through the same cycle of diagnosis, allocation, and integration introduced in Section[4\.1](https://arxiv.org/html/2607.26599#S4.SS1)\. All trainable parameters within a block are updated by minimizing a single objective that adds two auxiliary terms to the host\-side prediction lossℒmain\\mathcal\{L\}\_\{\\mathrm\{main\}\}, which already includes any host\-specific balancing, latent\-variable, or reconstruction term\. WritingBCE\(p,t\)\\operatorname\{BCE\}\(p,t\)for binary cross\-entropy andcos\(a,b\)=⟨a,b⟩/\(∥a∥2∥b∥2\+ϵ\)\\cos\(a,b\)=\\langle a,b\\rangle/\(\\lVert a\\rVert\_\{2\}\\lVert b\\rVert\_\{2\}\+\\epsilon\)for cosine similarity with numerical stabilizerϵ\>0\\epsilon\>0, the two auxiliary terms supervise the propensity head and discourage redundancy between the two semantic channels, \(12\)ℒ\(Θ\)\\displaystyle\\mathcal\{L\}\(\\Theta\)=ℒmain\+λpropℒprop\+λorthoℒortho,\\displaystyle=\\mathcal\{L\}\_\{\\mathrm\{main\}\}\+\\lambda\_\{\\mathrm\{prop\}\}\\mathcal\{L\}\_\{\\mathrm\{prop\}\}\+\\lambda\_\{\\mathrm\{ortho\}\}\\mathcal\{L\}\_\{\\mathrm\{ortho\}\},\(13\)ℒprop\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{prop\}\}=1N∑i=1NBCE\(e^Θ\(r\)\(xi\),ti\),\\displaystyle=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\operatorname\{BCE\}\\\!\\left\(\\widehat\{e\}\_\{\\Theta^\{\(r\)\}\}\(x\_\{i\}\),t\_\{i\}\\right\),\(14\)ℒortho\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{ortho\}\}=1\|𝒞\(r\)\|∑i∈𝒞\(r\)cos2\(vA,i\(r\),vH,i\(r\)\),\\displaystyle=\\frac\{1\}\{\|\\mathcal\{C\}^\{\(r\)\}\|\}\\sum\_\{i\\in\\mathcal\{C\}^\{\(r\)\}\}\\cos^\{2\}\\\!\\left\(v\_\{A,i\}^\{\(r\)\},v\_\{H,i\}^\{\(r\)\}\\right\),withλprop,λortho≥0\\lambda\_\{\\mathrm\{prop\}\},\\lambda\_\{\\mathrm\{ortho\}\}\\geq 0andℒortho=0\\mathcal\{L\}\_\{\\mathrm\{ortho\}\}=0when𝒞\(r\)\\mathcal\{C\}^\{\(r\)\}is empty\. The auxiliary propensity head is present in every evaluated CURL adapter, so settingλprop=0\\lambda\_\{\\mathrm\{prop\}\}=0recovers the same objective form for a host that lacks this interface natively\. Algorithmic details are given in Appendix[D](https://arxiv.org/html/2607.26599#A4)\. #### 4\.4\.2\.Test\-time Inference Test\-time inference follows a calibration\-and\-refinement principle\. CURL converts the relative ranking used during training into a fixed pointwise rule for deciding whether an unseen unit should receive semantic augmentation, then recomputes uncertainty after augmentation to control semantic fusion\. This conversion is necessary because Eq\. \([3](https://arxiv.org/html/2607.26599#S4.E3)\) relies on relative rankings over the full training set, whereas test\-time decisions must be made pointwise without assuming access to other test units or depending on the composition of an arbitrary test batch\. Letse,i\(R\)s\_\{e,i\}^\{\(R\)\}andsτ,i\(R\)s\_\{\\tau,i\}^\{\(R\)\}denote the two uncertainty scores computed at the final training state, following the same definition as Eq\. \([2](https://arxiv.org/html/2607.26599#S4.E2)\)\. Their empirical distribution functions, \(15\)G^etr\(s\)=1N∑i=1N𝕀\(se,i\(R\)≤s\),G^τtr\(s\)=1N∑i=1N𝕀\(sτ,i\(R\)≤s\),\\widehat\{G\}\_\{e\}^\{\\mathrm\{tr\}\}\(s\)=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\mathbb\{I\}\(s\_\{e,i\}^\{\(R\)\}\\leq s\),\\qquad\\widehat\{G\}\_\{\\tau\}^\{\\mathrm\{tr\}\}\(s\)=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\mathbb\{I\}\(s\_\{\\tau,i\}^\{\(R\)\}\\leq s\),map the two scores, which may have different scales, to percentile positions within the final training uncertainty distributions\. Summing these percentiles givesωitr=G^etr\(se,i\(R\)\)\+G^τtr\(sτ,i\(R\)\)\\omega\_\{i\}^\{\\mathrm\{tr\}\}=\\widehat\{G\}\_\{e\}^\{\\mathrm\{tr\}\}\(s\_\{e,i\}^\{\(R\)\}\)\+\\widehat\{G\}\_\{\\tau\}^\{\\mathrm\{tr\}\}\(s\_\{\\tau,i\}^\{\(R\)\}\), and CURL stores their empirical\(1−ρ\)\(1\-\\rho\)\-quantile,κρ=Q1−ρ\(ωitri=1N\)\\kappa\_\{\\rho\}=Q\_\{1\-\\rho\}\(\{\\omega\_\{i\}^\{\\mathrm\{tr\}\}\}\_\{i=1\}^\{N\}\), together with the two distributions above\. Using the sameρ\\rhocalibrates this threshold to the training\-time selection budget and gives a comparable augmentation rate when the test uncertainty distribution is similar\. For an unseen covariate vectorx∗x\_\{\*\}, CURL first sets both semantic inputs to zero and computes pre\-augmentation uncertaintiesse,∗pres\_\{e,\*\}^\{\\mathrm\{pre\}\}andsτ,∗pres\_\{\\tau,\*\}^\{\\mathrm\{pre\}\}through the same Monte Carlo procedure as Eq\. \([2](https://arxiv.org/html/2607.26599#S4.E2)\)\. It queries the LLM if and only if \(16\)a∗=𝕀\(G^etr\(se,∗pre\)\+G^τtr\(sτ,∗pre\)≥κρ\),a\_\{\*\}=\\mathbb\{I\}\\\!\\left\(\\widehat\{G\}\_\{e\}^\{\\mathrm\{tr\}\}\(s\_\{e,\*\}^\{\\mathrm\{pre\}\}\)\+\\widehat\{G\}\_\{\\tau\}^\{\\mathrm\{tr\}\}\(s\_\{\\tau,\*\}^\{\\mathrm\{pre\}\}\)\\geq\\kappa\_\{\\rho\}\\right\),wherea∗∈\{0,1\}a\_\{\*\}\\in\\\{0,1\\\}is the augmentation decision\. A selected unit receives the two embeddings constructed in Section[4\.3](https://arxiv.org/html/2607.26599#S4.SS3), whereas an unselected unit retains zero embeddings\. CURL then recomputes uncertainty under the updated inputs and uses the post\-augmentation scores as the gate inputs of Eq\. \([8](https://arxiv.org/html/2607.26599#S4.E8)\) before applying Eq\. \([11](https://arxiv.org/html/2607.26599#S4.E11)\)\. Thus, the pre\-augmentation scores determine whether the unit is sufficiently unstable relative to the final trained estimator to justify an LLM query, while the post\-augmentation scores determine how strongly the resulting embeddings influence prediction\. No test\-time treatment, outcome, ground\-truth effect, or comparison across other test units enters either step\. ## 5\.Experiments We evaluate CURL on four benchmarks and ten CATE estimators, asking: \(RQ1\) whether its gains are consistent across estimators and robust to backbone and hyperparameter choices; \(RQ2\) which components drive these gains; and \(RQ3\) whether the two semantic channels exhibit their intended assignment\- and heterogeneity\-oriented roles\. ### 5\.1\.Experimental Setup ##### Datasets and metrics\. We use two semi\-synthetic benchmarks with ground\-truth individual effects, IHDP and Adult, evaluated by PEHE,ϵATE\\epsilon\_\{\\mathrm\{ATE\}\}, andϵATT\\epsilon\_\{\\mathrm\{ATT\}\}\(lower is better\)\. Jobs and the randomized Hillstrom email experiment lack individual\-level effect ground truth and are evaluated by AUUC, AUQC, and LIFT@30 \(higher is better\)\. Dataset and metric details are provided in Appendices[A](https://arxiv.org/html/2607.26599#A1)and[B](https://arxiv.org/html/2607.26599#A2)\. ##### Baselines\. We plugCURLinto ten host learners spanning meta\-learners, balanced\-representation methods, and deep latent\-variable methods \(full list in Appendix[A\.2](https://arxiv.org/html/2607.26599#A1.SS2)\)\. We additionally compare with two LLM\-only baselines,LLMandLLM\+MLP, and withGATE\(Huynhet al\.,[2025](https://arxiv.org/html/2607.26599#bib.bib43)\), which directly imputes counterfactual outcomes\. Feature\-disentangled CATE baselines are reported in Appendix[A\.3](https://arxiv.org/html/2607.26599#A1.SS3)\. ##### Implementation\. We select Qwen2\.5\-7B or Llama\-3\-8B per dataset by validation and compare four additional backbones in Section[5\.4](https://arxiv.org/html/2607.26599#S5.SS4)\. We setρ=0\.25\\rho=0\.25,M=30M=30,R=3R=3,λprop=1\.0\\lambda\_\{\\mathrm\{prop\}\}=1\.0, andλortho=0\.1\\lambda\_\{\\mathrm\{ortho\}\}=0\.1, and report means over five random seeds\. Further details are provided in Appendices[A\.2](https://arxiv.org/html/2607.26599#A1.SS2)and[C](https://arxiv.org/html/2607.26599#A3)\. Table 1\.Results on real\-world uplift benchmarks \(AUUC, AUQC, and LIFT@30; higher is better\)\. Mean±\\pmstandard deviation over five runs is reported\. For each host model and metric, the best result amongbase, GATE, andCURLis boldfaced\.Δ\\Deltais the relative gain ofCURLover the host learner \(%\)\.DatasetBase ModelAUUC↑\\uparrowAUQC↑\\uparrowLIFT@30↑\\uparrowbaseGATECURLΔ\\DeltabaseGATECURLΔ\\DeltabaseGATECURLΔ\\DeltaJobsLLM336\.07±\\pm152\.90–––51\.52±\\pm15\.64–––3\.25±\\pm1\.20–––LLM\+MLP810\.99±\\pm130\.44–––69\.11±\\pm22\.21–––2\.56±\\pm1\.23–––S\-Learner506\.40±\\pm210\.85461\.15±\\pm47\.89522\.99±\\pm233\.41\+3\.28%262\.09±\\pm89\.67272\.56±\\pm28\.26352\.08±\\pm64\.15\+34\.34%11\.33±\\pm6\.782\.88±\\pm0\.3011\.53±\\pm4\.21\+1\.76%T\-Learner671\.81±\\pm254\.99522\.02±\\pm157\.29875\.93±\\pm299\.94\+30\.38%351\.35±\\pm113\.76397\.51±\\pm86\.95487\.50±\\pm142\.68\+38\.75%6\.13±\\pm1\.157\.27±\\pm1\.589\.22±\\pm3\.93\+50\.41%R\-Learner770\.80±\\pm332\.56630\.93±\\pm258\.181007\.40±\\pm344\.72\+30\.70%226\.62±\\pm44\.71168\.09±\\pm46\.37240\.12±\\pm88\.37\+5\.96%6\.60±\\pm2\.717\.68±\\pm0\.178\.14±\\pm3\.01\+23\.32%TARNet491\.85±\\pm70\.27489\.47±\\pm170\.58885\.49±\\pm223\.55\+80\.03%218\.54±\\pm45\.24203\.95±\\pm29\.86266\.82±\\pm57\.53\+22\.09%8\.06±\\pm2\.808\.47±\\pm1\.2511\.73±\\pm2\.20\+45\.65%DragonNet918\.04±\\pm318\.45655\.31±\\pm204\.071185\.06±\\pm389\.74\+29\.09%413\.35±\\pm122\.60376\.23±\\pm140\.30454\.69±\\pm64\.98\+10\.00%11\.51±\\pm3\.316\.17±\\pm2\.2712\.24±\\pm8\.31\+6\.36%CFRNet816\.48±\\pm79\.80521\.71±\\pm197\.32819\.61±\\pm123\.81\+0\.38%223\.08±\\pm63\.98237\.80±\\pm89\.07291\.50±\\pm58\.45\+30\.67%7\.32±\\pm1\.177\.67±\\pm1\.898\.65±\\pm1\.12\+18\.18%CEVAE970\.01±\\pm197\.59442\.97±\\pm125\.07975\.08±\\pm180\.33\+0\.52%465\.46±\\pm56\.80472\.25±\\pm52\.95578\.90±\\pm185\.02\+24\.37%9\.10±\\pm1\.795\.06±\\pm0\.3310\.63±\\pm1\.85\+16\.80%TEDVAE862\.90±\\pm436\.75696\.57±\\pm150\.401087\.30±\\pm253\.83\+26\.01%417\.32±\\pm95\.35372\.78±\\pm7\.02420\.50±\\pm81\.05\+0\.76%4\.61±\\pm1\.234\.71±\\pm0\.166\.16±\\pm2\.82\+33\.48%SDD574\.98±\\pm180\.84710\.59±\\pm208\.96743\.96±\\pm165\.72\+29\.39%217\.93±\\pm63\.15208\.53±\\pm60\.82218\.64±\\pm75\.71\+0\.32%7\.85±\\pm3\.168\.99±\\pm2\.079\.47±\\pm3\.26\+20\.70%CiVAE717\.18±\\pm131\.83655\.36±\\pm177\.131023\.04±\\pm278\.45\+42\.65%236\.21±\\pm21\.96179\.78±\\pm112\.48238\.78±\\pm66\.90\+1\.09%4\.10±\\pm2\.194\.85±\\pm0\.575\.13±\\pm1\.77\+25\.10%HillstromLLM0\.016±\\pm0\.009–––0\.011±\\pm0\.001–––0\.858±\\pm0\.153–––LLM\+MLP0\.053±\\pm0\.032–––0\.014±\\pm0\.006–––1\.345±\\pm0\.526–––S\-Learner0\.112±\\pm0\.0440\.115±\\pm0\.0130\.167±\\pm0\.053\+48\.35%0\.044±\\pm0\.0110\.023±\\pm0\.0040\.045±\\pm0\.012\+2\.49%2\.570±\\pm0\.4052\.853±\\pm2\.9653\.239±\\pm0\.977\+26\.04%T\-Learner0\.106±\\pm0\.0350\.131±\\pm0\.0160\.136±\\pm0\.038\+28\.01%0\.020±\\pm0\.0020\.021±\\pm0\.0080\.023±\\pm0\.003\+12\.25%1\.005±\\pm0\.2210\.973±\\pm0\.1571\.109±\\pm0\.239\+10\.38%R\-Learner0\.205±\\pm0\.0490\.132±\\pm0\.0520\.207±\\pm0\.049\+1\.07%0\.046±\\pm0\.0030\.044±\\pm0\.0420\.049±\\pm0\.005\+6\.75%4\.273±\\pm0\.9632\.841±\\pm0\.8155\.195±\\pm1\.563\+21\.57%TARNet0\.063±\\pm0\.0200\.062±\\pm0\.0160\.065±\\pm0\.022\+2\.22%0\.022±\\pm0\.0020\.020±\\pm0\.0050\.026±\\pm0\.005\+14\.67%0\.991±\\pm0\.1920\.858±\\pm0\.1001\.014±\\pm0\.186\+2\.36%DragonNet0\.109±\\pm0\.0210\.122±\\pm0\.0580\.124±\\pm0\.024\+13\.98%0\.021±\\pm0\.0010\.024±\\pm0\.0080\.026±\\pm0\.004\+21\.40%0\.995±\\pm0\.1630\.865±\\pm0\.3941\.099±\\pm0\.203\+10\.49%CFRNet0\.130±\\pm0\.0410\.107±\\pm0\.0120\.162±\\pm0\.049\+24\.27%0\.037±\\pm0\.0050\.032±\\pm0\.0230\.044±\\pm0\.007\+19\.07%2\.674±\\pm0\.5382\.497±\\pm2\.5463\.734±\\pm1\.399\+39\.61%CEVAE0\.063±\\pm0\.0210\.053±\\pm0\.0160\.068±\\pm0\.015\+6\.97%0\.019±\\pm0\.0020\.014±\\pm0\.0020\.021±\\pm0\.002\+11\.46%0\.973±\\pm0\.1510\.864±\\pm0\.3711\.168±\\pm0\.189\+20\.03%TEDVAE0\.111±\\pm0\.0350\.113±\\pm0\.0160\.116±\\pm0\.037\+4\.23%0\.021±\\pm0\.0010\.021±\\pm0\.0050\.022±\\pm0\.001\+2\.83%0\.689±\\pm0\.2990\.875±\\pm0\.1871\.126±\\pm0\.034\+63\.37%SDD0\.115±\\pm0\.0180\.120±\\pm0\.0160\.128±\\pm0\.019\+10\.49%0\.021±\\pm0\.0010\.029±\\pm0\.0150\.023±\\pm0\.003\+9\.35%1\.022±\\pm0\.1111\.709±\\pm0\.5401\.028±\\pm0\.116\+0\.60%CiVAE0\.066±\\pm0\.0220\.062±\\pm0\.0140\.073±\\pm0\.027\+11\.13%0\.030±\\pm0\.0090\.020±\\pm0\.0030\.033±\\pm0\.011\+11\.04%1\.425±\\pm0\.3641\.407±\\pm0\.5981\.511±\\pm0\.367\+6\.04% Table 2\.Results on semi\-synthetic datasets \(PEHE,ϵATE\\epsilon\_\{\\mathrm\{ATE\}\}, andϵATT\\epsilon\_\{\\mathrm\{ATT\}\}; lower is better\)\. Mean±\\pmstandard deviation over five runs is reported\. For each host model and metric, the best result amongbase, GATE, andCURLis boldfaced\.Δ\\Deltais the relative error reduction ofCURLover the host learner \(%\)\.DatasetBase ModelPEHE↓\\downarrowϵATE\\epsilon\_\{\\mathrm\{ATE\}\}↓\\downarrowϵATT\\epsilon\_\{\\mathrm\{ATT\}\}↓\\downarrowbaseGATECURLΔ\\DeltabaseGATECURLΔ\\DeltabaseGATECURLΔ\\DeltaIHDPLLM44\.117±\\pm26\.698–––54\.308±\\pm41\.443–––53\.266±\\pm40\.708–––LLM\+MLP2\.487±\\pm0\.935–––1\.398±\\pm0\.389–––1\.231±\\pm0\.420–––S\-Learner4\.616±\\pm0\.8383\.499±\\pm1\.1062\.926±\\pm0\.790\+36\.62%3\.050±\\pm0\.2902\.492±\\pm0\.3501\.950±\\pm0\.221\+36\.08%2\.771±\\pm0\.2892\.394±\\pm0\.3911\.437±\\pm0\.045\+48\.15%T\-Learner1\.643±\\pm0\.4161\.450±\\pm0\.5471\.438±\\pm0\.354\+12\.45%0\.614±\\pm0\.1590\.595±\\pm0\.0650\.511±\\pm0\.089\+16\.88%0\.406±\\pm0\.1440\.526±\\pm0\.1360\.275±\\pm0\.051\+32\.21%R\-Learner2\.629±\\pm0\.7252\.625±\\pm0\.8902\.297±\\pm0\.284\+12\.64%1\.179±\\pm0\.3231\.184±\\pm0\.0661\.163±\\pm0\.270\+1\.40%0\.956±\\pm0\.2970\.986±\\pm0\.3150\.936±\\pm0\.243\+2\.04%TARNet1\.712±\\pm0\.5091\.370±\\pm0\.3851\.324±\\pm0\.254\+22\.67%0\.609±\\pm0\.2270\.436±\\pm0\.1210\.200±\\pm0\.047\+67\.26%0\.478±\\pm0\.2000\.664±\\pm0\.2770\.343±\\pm0\.097\+28\.30%DragonNet1\.654±\\pm0\.4731\.531±\\pm0\.4191\.342±\\pm0\.217\+18\.83%0\.513±\\pm0\.1580\.505±\\pm0\.0770\.315±\\pm0\.055\+38\.60%0\.332±\\pm0\.1200\.550±\\pm0\.1950\.225±\\pm0\.072\+32\.26%CFRNet1\.767±\\pm0\.6321\.797±\\pm0\.6111\.405±\\pm0\.342\+20\.47%0\.412±\\pm0\.0850\.459±\\pm0\.0340\.179±\\pm0\.068\+56\.64%0\.389±\\pm0\.2010\.617±\\pm0\.1270\.307±\\pm0\.056\+21\.06%CEVAE1\.879±\\pm0\.7301\.548±\\pm0\.4991\.508±\\pm0\.402\+19\.77%0\.953±\\pm0\.4770\.937±\\pm0\.1850\.328±\\pm0\.152\+65\.61%0\.841±\\pm0\.4290\.910±\\pm0\.4550\.258±\\pm0\.122\+69\.32%TEDVAE2\.150±\\pm0\.3142\.409±\\pm0\.7051\.990±\\pm0\.592\+7\.46%1\.410±\\pm0\.5471\.955±\\pm0\.4561\.303±\\pm1\.032\+7\.54%0\.531±\\pm0\.0301\.929±\\pm0\.7060\.455±\\pm0\.039\+14\.36%SDD1\.875±\\pm0\.6421\.737±\\pm0\.3371\.571±\\pm0\.396\+16\.20%0\.500±\\pm0\.1840\.543±\\pm0\.0910\.497±\\pm0\.085\+0\.42%0\.369±\\pm0\.0690\.544±\\pm0\.4240\.336±\\pm0\.045\+8\.89%CiVAE3\.099±\\pm0\.9243\.096±\\pm1\.0013\.011±\\pm0\.846\+2\.85%0\.231±\\pm0\.0240\.412±\\pm0\.1720\.226±\\pm0\.058\+2\.38%0\.690±\\pm0\.1480\.657±\\pm0\.1660\.598±\\pm0\.184\+13\.42%AdultLLM57\.312±\\pm20\.847–––1\.569±\\pm0\.702–––0\.708±\\pm0\.295–––LLM\+MLP0\.429±\\pm0\.063–––0\.191±\\pm0\.065–––0\.170±\\pm0\.060–––S\-Learner0\.260±\\pm0\.0360\.471±\\pm0\.1120\.213±\\pm0\.041\+18\.05%0\.158±\\pm0\.0480\.454±\\pm0\.1150\.095±\\pm0\.050\+39\.91%0\.135±\\pm0\.0610\.454±\\pm0\.1350\.113±\\pm0\.031\+16\.06%T\-Learner0\.993±\\pm0\.2071\.067±\\pm0\.3490\.894±\\pm0\.270\+9\.97%0\.354±\\pm0\.0700\.917±\\pm0\.2930\.161±\\pm0\.103\+54\.55%0\.396±\\pm0\.0370\.877±\\pm0\.2800\.281±\\pm0\.075\+29\.17%R\-Learner0\.065±\\pm0\.0120\.153±\\pm0\.0170\.062±\\pm0\.021\+5\.35%0\.005±\\pm0\.0040\.101±\\pm0\.0130\.004±\\pm0\.004\+20\.75%0\.007±\\pm0\.0040\.098±\\pm0\.0210\.002±\\pm0\.001\+72\.97%TARNet0\.658±\\pm0\.0881\.049±\\pm0\.3140\.649±\\pm0\.125\+1\.32%0\.331±\\pm0\.0690\.889±\\pm0\.2590\.175±\\pm0\.096\+46\.93%0\.355±\\pm0\.0440\.855±\\pm0\.2480\.275±\\pm0\.072\+22\.76%DragonNet0\.607±\\pm0\.0471\.032±\\pm0\.3340\.596±\\pm0\.059\+1\.73%0\.337±\\pm0\.0490\.867±\\pm0\.2830\.146±\\pm0\.088\+56\.72%0\.346±\\pm0\.0260\.828±\\pm0\.2780\.240±\\pm0\.065\+30\.62%CFRNet0\.947±\\pm0\.2700\.986±\\pm0\.3340\.941±\\pm0\.128\+0\.67%0\.772±\\pm0\.2860\.837±\\pm0\.2840\.452±\\pm0\.260\+41\.46%0\.780±\\pm0\.3080\.804±\\pm0\.2770\.491±\\pm0\.238\+36\.98%CEVAE0\.408±\\pm0\.0620\.653±\\pm0\.2860\.386±\\pm0\.145\+5\.41%0\.203±\\pm0\.1010\.537±\\pm0\.2910\.080±\\pm0\.064\+60\.60%0\.177±\\pm0\.1030\.536±\\pm0\.3150\.113±\\pm0\.067\+36\.51%TEDVAE0\.443±\\pm0\.1920\.621±\\pm0\.7330\.436±\\pm0\.085\+1\.63%0\.212±\\pm0\.0890\.457±\\pm0\.6790\.211±\\pm0\.052\+0\.61%0\.223±\\pm0\.1230\.435±\\pm0\.6410\.221±\\pm0\.083\+0\.90%SDD0\.642±\\pm0\.1191\.034±\\pm0\.3170\.717±\\pm0\.216\-11\.75%0\.345±\\pm0\.0800\.857±\\pm0\.2690\.334±\\pm0\.089\+3\.16%0\.339±\\pm0\.0950\.813±\\pm0\.2660\.329±\\pm0\.082\+2\.98%CiVAE0\.326±\\pm0\.2600\.302±\\pm0\.1230\.114±\\pm0\.053\+65\.06%0\.195±\\pm0\.2310\.191±\\pm0\.1320\.037±\\pm0\.025\+81\.12%0\.210±\\pm0\.2180\.194±\\pm0\.1370\.041±\\pm0\.021\+80\.64% ### 5\.2\.Main Results Tables[1](https://arxiv.org/html/2607.26599#S5.T1)and[2](https://arxiv.org/html/2607.26599#S5.T2)report the main results on the four benchmarks, respectively\. We summarize three observations\. First,CURLimproves its host learner across the vast majority of settings, with one degradation on Adult\. The improvement appears on both benchmarks with ground truth and benchmarks evaluated by uplift metrics, indicating that it is not specific to a particular evaluation protocol\. These patterns are consistent withCURLhelping finite\-sample estimators exploit semantic structure inXXthat is difficult to learn from its raw encoding\. Second, the improvement spans meta\-learners, balanced\-representation methods, and latent\-variable methods, supporting the use ofCURLas a plug\-in adapter rather than an architecture\-specific estimator\. The gains for latent\-variable baselines further suggest that pretrained semantic structure can complement representations learned primarily from the study sample, even when the host learner is highly flexible\. Third, the gains cannot be explained by generic LLM features alone\. BothLLMandLLM\+MLPare substantially weaker thanCURL, and in many cases weaker than conventional CATE learners, showing that directly using LLM outputs or embeddings is insufficient\. Under the same LLM backbone,CURLalso outperformsGATEin most settings, indicating that its benefit comes from uncertainty\-guided allocation, role\-conditioned semantic construction, and separated routing rather than from raw LLM capacity\. ### 5\.3\.Ablation Studies Figure[3](https://arxiv.org/html/2607.26599#S5.F3)compares the fullCURLwith its ablated variants across four representative host learners\. The complete model achieves the best result in all eight host–dataset configurations, indicating that its gains do not depend on a particular estimator architecture\. Figure 3\.Ablation results across four host learners on IHDP \(PEHE, lower is better\) and Jobs \(AUQC, higher is better\)\. Each cell reports the mean over five runs\. Colors are normalized separately within each host\-learner column, with greener cells indicating better performance\.Semantic construction and routing\. Removing either semantic channel \(w/ozHLLMz\_\{H\}^\{\\mathrm\{LLM\}\}orw/ozALLMz\_\{A\}^\{\\mathrm\{LLM\}\}\) consistently underperforms the full model, showing that the two role\-conditioned representations provide complementary information\. Replacing them with a shared representation \(MergedzMz\_\{M\}\) also reduces performance\. The declines caused by removingℒortho\\mathcal\{L\}\_\{\\mathrm\{ortho\}\}or role\-aware routing further support channel separation and the asymmetric prediction pathways\. Allocation and refinement\. Using only CATE uncertainty \(CATE\-onlysτs\_\{\\tau\}\) remains competitive for several IHDP hosts but loses substantial performance on Jobs, suggesting that assignment\-side uncertainty provides an important complementary allocation signal\. Random allocation is less reliable and never matches the full uncertainty\-guided strategy\. Finally,w/o Refinementis consistently inferior toCURL, supporting iterative re\-diagnosis and semantic reallocation as the host estimator evolves\. ### 5\.4\.Sensitivity Analysis We assess the robustness ofCURLwith respect to the LLM backbone, the semantic\-allocation ratioρ\\rho, and the number of refinement roundsRR\. Figure 4\.Sensitivity ofCURLto different LLM backbones on Hillstrom \(AUQC\) and IHDP \(PEHE\) across multiple host learners\.Figure[4](https://arxiv.org/html/2607.26599#S5.F4)reports results under six backbones: Phi\-3\.8B, Qwen2\.5\-7B, Llama2\-7B, Llama3\-8B, Llama2\-13B, and Qwen2\.5\-14B\. Performance varies only modestly across backbones, and larger models are not uniformly better\. This suggests thatCURLis not strongly dependent on a particular backbone or model scale\. Figure[5](https://arxiv.org/html/2607.26599#S5.F5)examines the allocation ratioρ\\rhoand the number of refinement roundsRRon Jobs\. Performance remains stable over broad ranges, with moderate settings generally performing best\. Increasingρ\\rhobeyond this range raises the LLM cost without consistent gains, while excessive refinement rounds yield diminishing returns\. Overall,CURLis not highly sensitive to either hyperparameter around the selected configuration\. \(a\)Effect of the allocation ratio\. \(b\)Effect of the refinement rounds\. Figure 5\.Sensitivity ofCURLon Jobs \(AUQC\) across multiple host learners: \(a\) semantic\-allocation ratioρ\\rhoand \(b\) number of refinement roundsRR\. ### 5\.5\.Progressive Refinement Dynamics We further examine how semantic allocation evolves across refinement rounds\. SettingR=5R=5, we track the cumulative semantic\-cache ratio\|𝒞\(r\)\|/N\|\\mathcal\{C\}^\{\(r\)\}\|/Nand the mean uncertainty scoresse\(r\)s\_\{e\}^\{\(r\)\}andsτ\(r\)s\_\{\\tau\}^\{\(r\)\}on IHDP and Jobs across four representative host learners\. Figure[6](https://arxiv.org/html/2607.26599#S5.F6)summarizes the Jobs results in the main text, while the full results are provided in Appendix[F](https://arxiv.org/html/2607.26599#A6)\. Figure 6\.Progressive refinement dynamics on Jobs across four representative host learners\. We report the cumulative semantic\-cache ratio\|𝒞\(r\)\|/N\|\\mathcal\{C\}^\{\(r\)\}\|/Nand the mean uncertainty scoresse\(r\)s\_\{e\}^\{\(r\)\}andsτ\(r\)s\_\{\\tau\}^\{\(r\)\}over refinement rounds\.As shown in Figure[6](https://arxiv.org/html/2607.26599#S5.F6), the semantic cache expands rapidly in the early rounds and then gradually stabilizes\. Meanwhile, both uncertainty scores exhibit an overall downward trend across most settings\. This pattern is consistent with progressive refinement reducing local predictive instability while the set of units requiring semantic augmentation gradually stabilizes\. ### 5\.6\.Role Validation of the Semantic Channels We evaluate whether the two LLM\-derived channels align with their intended roles through route reassignment and probing, and test whether their gains reflect representational content rather than additional inputs or parameters\. Route reassignment\. We freeze the trained model and evaluate eight reroutings ofzAz\_\{A\}andzHz\_\{H\}\(Figure[7](https://arxiv.org/html/2607.26599#S5.F7)\)\. The original routing performs best: swapping the channels or removingzAz\_\{A\}from the propensity pathway harms treatment prediction, while incorrect outcome\-side routing degrades CATE estimation\. This supports the intended assignment and heterogeneity roles\. Figure 7\.Route\-reassignment validation\. The original routing performs best overall, while swapping or incorrectly routing the two semantic channels degrades assignment and effect estimation\.Probe analysis\. Lightweight probes on frozen representations provide consistent evidence \(Figure[8](https://arxiv.org/html/2607.26599#S5.F8)\):zAz\_\{A\}better predicts treatment assignment, whereaszHz\_\{H\}better predicts treatment\-dependent responses when treatment–representation interactions are included\. Combining both channels performs best, indicating complementary information\. Figure 8\.Probe\-based role validation\. The assignment\-oriented channel is more informative for treatment prediction, whereas the heterogeneity\-oriented channel is more informative when modeling treatment\-dependent responses\.Representational\-lossiness checks\. Semantic\-corruption results on IHDP \(Appendix[G](https://arxiv.org/html/2607.26599#A7)\) show that CURL performs best: anonymizing feature names and shuffling category labels both reduce its advantage, while random embeddings cause the largest degradation\. This ordering indicates that CURL’s improvement derives in part from semantic structure encoded by the LLM, rather than additional adapter capacity alone; Appendix[H](https://arxiv.org/html/2607.26599#A8)illustrates the channels’ content\. ## 6\.Conclusion We presented CURL, a plug\-in framework for finite\-sample CATE estimation under representational lossiness, where causally sufficient covariates remain difficult to exploit in their raw encoded form\. Rather than using LLMs as unstructured feature generators or direct outcome imputers,CURLuses estimator uncertainty to identify locally unstable units, constructs assignment\- and heterogeneity\-oriented representations from their observed covariates, and integrates them through role\-aware pathways\. Across four benchmarks and ten host learners,CURLimproves performance in most settings, while ablation, refinement, corruption, and role\-validation analyses support uncertainty\-guided allocation and channel separation\.CURLreorganizes observed covariates rather than introducing new unit\-level information; its generated text is a semantic representation and does not address unmeasured confounding or overlap violations\. Future work should examine semantic reliability and uncertainty calibration under domain shift\. Overall,CURLshows how a pretrained semantic prior can support more reliable finite\-sample CATE estimation by making these relations easier to exploit, without changing the observed information set, target estimand, or identifying assumptions\. ## References - Y\. Abbasi\-Yadkori, D\. Pál, and C\. Szepesvári \(2011\)Improved algorithms for linear stochastic bandits\.InAdvances in Neural Information Processing Systems,Vol\.24,pp\. 2312–2320\.Cited by:[§2\.3](https://arxiv.org/html/2607.26599#S2.SS3.p1.1)\. - A\. M\. Alaa and M\. van der Schaar \(2017\)Bayesian inference of individualized treatment effects using multi\-task Gaussian processes\.InAdvances in Neural Information Processing Systems,Cited by:[§2\.3](https://arxiv.org/html/2607.26599#S2.SS3.p1.1)\. - G\. An, S\. Moon, J\. Kim, and H\. O\. Song \(2021\)Uncertainty\-based offline reinforcement learning with diversified Q\-ensemble\.InAdvances in Neural Information Processing Systems,Vol\.34,pp\. 7436–7447\.Cited by:[§2\.3](https://arxiv.org/html/2607.26599#S2.SS3.p1.1)\. - C\. Bai, L\. Wang, Z\. Yang, Z\. Deng, A\. Garg, P\. Liu, and Z\. Wang \(2022\)Pessimistic bootstrapping for uncertainty\-driven offline reinforcement learning\.arXiv preprint arXiv:2202\.11566\.Cited by:[§2\.3](https://arxiv.org/html/2607.26599#S2.SS3.p1.1)\. - T\. Ban, L\. Chen, D\. Lyu, X\. Wang, Q\. Zhu, and H\. Chen \(2025\)LLM\-driven causal discovery via harmonized prior\.IEEE Transactions on Knowledge and Data Engineering\.Cited by:[§2\.2](https://arxiv.org/html/2607.26599#S2.SS2.p1.1)\. - C\. Blundell, J\. Cornebise, K\. Kavukcuoglu, and D\. Wierstra \(2015\)Weight uncertainty in neural networks\.InInternational Conference on Machine Learning,pp\. 1613–1622\.Cited by:[§2\.3](https://arxiv.org/html/2607.26599#S2.SS3.p1.1)\. - J\. M\. Chen, R\. Bhattacharya, and K\. A\. Keith \(2024\)Proximal causal inference with text data\.InAdvances in Neural Information Processing Systems,Vol\.37\.Cited by:[§1](https://arxiv.org/html/2607.26599#S1.p3.1),[§2\.2](https://arxiv.org/html/2607.26599#S2.SS2.p1.1)\. - H\. Chi, H\. Li, W\. Yang, F\. Liu, L\. Lan, X\. Ren, T\. Liu, and B\. Han \(2024\)Unveiling causal reasoning in large language models: reality or mirage?\.InAdvances in Neural Information Processing Systems,Vol\.37\.Cited by:[§2\.2](https://arxiv.org/html/2607.26599#S2.SS2.p1.1)\. - N\. Dhawan, L\. Cotta, K\. Ullrich, R\. G\. Krishnan, and C\. J\. Maddison \(2024\)End\-to\-end causal effect estimation from unstructured natural language data\.InAdvances in Neural Information Processing Systems,Vol\.37\.Cited by:[§1](https://arxiv.org/html/2607.26599#S1.p3.1),[§2\.2](https://arxiv.org/html/2607.26599#S2.SS2.p1.1)\. - H\. Du, Y\. Zheng, B\. Jing, Y\. Zhao, G\. Kou, G\. Liu, T\. Gu, W\. Li, and C\. Yang \(2025\)Causal discovery through synergizing large language model and data\-driven reasoning\.InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining,Cited by:[§1](https://arxiv.org/html/2607.26599#S1.p3.1),[§2\.2](https://arxiv.org/html/2607.26599#S2.SS2.p1.1)\. - Y\. Gal and Z\. Ghahramani \(2016\)Dropout as a bayesian approximation: representing model uncertainty in deep learning\.InProceedings of The 33rd International Conference on Machine Learning,M\. F\. Balcan and K\. Q\. Weinberger \(Eds\.\),Proceedings of Machine Learning Research, Vol\.48,New York, New York, USA,pp\. 1050–1059\.External Links:[Link](https://proceedings.mlr.press/v48/gal16.html)Cited by:[§1](https://arxiv.org/html/2607.26599#S1.p4.1),[§2\.3](https://arxiv.org/html/2607.26599#S2.SS3.p1.1),[§4\.2](https://arxiv.org/html/2607.26599#S4.SS2.p1.7)\. - N\. Hassanpour and R\. Greiner \(2020\)Learning disentangled representations for counterfactual regression\.InInternational Conference on Learning Representations,Cited by:[§A\.3](https://arxiv.org/html/2607.26599#A1.SS3.p1.1)\. - N\. Huynh, J\. Piskorz, J\. Berrevoets, M\. R\. Luyten, and M\. van der Schaar \(2025\)Improving treatment effect estimation with llm\-based data augmentation\.In1st ICML Workshop on Foundation Models for Structured Data,Cited by:[§A\.2](https://arxiv.org/html/2607.26599#A1.SS2.SSS0.Px1.p1.1),[§2\.2](https://arxiv.org/html/2607.26599#S2.SS2.p1.1),[§5\.1](https://arxiv.org/html/2607.26599#S5.SS1.SSS0.Px2.p1.1)\. - G\. W\. Imbens and D\. B\. Rubin \(2015\)Causal inference for statistics, social, and biomedical sciences: an introduction\.Cambridge University Press,New York, NY, USA\.External Links:ISBN 9780521885881Cited by:[§1](https://arxiv.org/html/2607.26599#S1.p1.3),[§3](https://arxiv.org/html/2607.26599#S3.p2.8)\. - A\. Jesson, S\. Mindermann, U\. Shalit, and Y\. Gal \(2020\)Identifying causal\-effect inference failure with uncertainty\-aware models\.Advances in Neural Information Processing Systems33,pp\. 11637–11649\.Cited by:[§1](https://arxiv.org/html/2607.26599#S1.p4.1),[§2\.3](https://arxiv.org/html/2607.26599#S2.SS3.p1.1)\. - Z\. Jin, Y\. Chen, F\. Leeb, L\. Gresele, O\. Kamal, Z\. LYU, K\. Blin, F\. Gonzalez Adauto, M\. Kleiman\-Weiner, M\. Sachan, and B\. Schölkopf \(2023\)CLadder: assessing causal reasoning in language models\.InAdvances in Neural Information Processing Systems,A\. Oh, T\. Naumann, A\. Globerson, K\. Saenko, M\. Hardt, and S\. Levine \(Eds\.\),Vol\.36,pp\. 31038–31065\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2023/file/631bb9434d718ea309af82566347d607-Paper-Conference.pdf)Cited by:[§2\.2](https://arxiv.org/html/2607.26599#S2.SS2.p1.1)\. - A\. Kendall and Y\. Gal \(2017\)What uncertainties do we need in bayesian deep learning for computer vision?\.Advances in neural information processing systems30\.Cited by:[§2\.3](https://arxiv.org/html/2607.26599#S2.SS3.p1.1)\. - I\. Khemakhem, D\. Kingma, R\. Monti, and A\. Hyvarinen \(2020\)Variational autoencoders and nonlinear ica: a unifying framework\.InProceedings of the Twenty Third International Conference on Artificial Intelligence and Statistics,S\. Chiappa and R\. Calandra \(Eds\.\),Proceedings of Machine Learning Research, Vol\.108,pp\. 2207–2217\.External Links:[Link](https://proceedings.mlr.press/v108/khemakhem20a.html)Cited by:[§2\.1](https://arxiv.org/html/2607.26599#S2.SS1.p1.4)\. - E\. Kıcıman, R\. Ness, A\. Sharma, and C\. Tan \(2024\)Causal reasoning and large language models: opening a new frontier for causality\.External Links:2305\.00050,[Link](https://arxiv.org/abs/2305.00050)Cited by:[§1](https://arxiv.org/html/2607.26599#S1.p3.1),[§2\.2](https://arxiv.org/html/2607.26599#S2.SS2.p1.1)\. - S\. R\. Künzel, J\. S\. Sekhon, P\. J\. Bickel, and B\. Yu \(2019\)Metalearners for estimating heterogeneous treatment effects using machine learning\.Proceedings of the National Academy of Sciences116\(10\),pp\. 4156–4165\.Cited by:[§1](https://arxiv.org/html/2607.26599#S1.p1.3),[§1](https://arxiv.org/html/2607.26599#S1.p2.1),[§2\.1](https://arxiv.org/html/2607.26599#S2.SS1.p1.4)\. - B\. Lakshminarayanan, A\. Pritzel, and C\. Blundell \(2017\)Simple and scalable predictive uncertainty estimation using deep ensembles\.Advances in neural information processing systems30\.Cited by:[§2\.3](https://arxiv.org/html/2607.26599#S2.SS3.p1.1)\. - X\. Li, M\. Gong, and L\. Yao \(2024\)Self\-distilled disentangled learning for counterfactual prediction\.InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining,pp\. 1667–1678\.Cited by:[§2\.1](https://arxiv.org/html/2607.26599#S2.SS1.p1.4)\. - X\. Liu, P\. Xu, J\. Wu, J\. Yuan, Y\. Yang, Y\. Zhou, F\. Liu, T\. Guan, H\. Wang, T\. Yu, J\. McAuley, W\. Ai, and F\. Huang \(2025\)Large language models and causal inference in collaboration: a comprehensive survey\.InFindings of the Association for Computational Linguistics: NAACL 2025,L\. Chiruzzo, A\. Ritter, and L\. Wang \(Eds\.\),Albuquerque, New Mexico,pp\. 7683–7699\.External Links:[Link](https://aclanthology.org/2025.findings-naacl.427/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-naacl.427),ISBN 979\-8\-89176\-195\-7Cited by:[§2\.2](https://arxiv.org/html/2607.26599#S2.SS2.p1.1)\. - C\. Louizos, U\. Shalit, J\. Mooij, D\. Sontag, R\. Zemel, and M\. Welling \(2017\)Causal effect inference with deep latent\-variable models\.InProceedings of the 31st International Conference on Neural Information Processing Systems,NIPS’17,Red Hook, NY, USA,pp\. 6449–6459\.External Links:ISBN 9781510860964Cited by:[§1](https://arxiv.org/html/2607.26599#S1.p2.1),[§2\.1](https://arxiv.org/html/2607.26599#S2.SS1.p1.4)\. - D\. J\.C\. MacKay \(1992\)A practical Bayesian framework for backpropagation networks\.Neural Computation4\(3\),pp\. 448–472\.Cited by:[§2\.3](https://arxiv.org/html/2607.26599#S2.SS3.p1.1)\. - H\. Meng, K\. Yang, X\. Peng, and B\. Zheng \(2025\)Deep disentangled representation network for treatment effect estimation\.arXiv preprint arXiv:2507\.06650\.Cited by:[§A\.3](https://arxiv.org/html/2607.26599#A1.SS3.p1.1)\. - X\. Nie and S\. Wager \(2021\)Quasi\-oracle estimation of heterogeneous treatment effects\.Biometrika108\(2\),pp\. 299–319\.External Links:[Document](https://dx.doi.org/10.1093/biomet/asaa076)Cited by:[Table 5](https://arxiv.org/html/2607.26599#A5.T5.14.8.8.5),[§1](https://arxiv.org/html/2607.26599#S1.p1.3),[§2\.1](https://arxiv.org/html/2607.26599#S2.SS1.p1.4),[§3](https://arxiv.org/html/2607.26599#S3.p2.8)\. - S\. Rissanen and P\. Marttinen \(2021\)A critical look at the consistency of causal estimation with deep latent variable models\.InAdvances in Neural Information Processing Systems,Vol\.34\.Cited by:[§2\.1](https://arxiv.org/html/2607.26599#S2.SS1.p1.4)\. - P\. R\. Rosenbaum and D\. B\. Rubin \(1983\)The central role of the propensity score in observational studies for causal effects\.Biometrika70\(1\),pp\. 41–55\.External Links:[Document](https://dx.doi.org/10.1093/biomet/70.1.41)Cited by:[§3](https://arxiv.org/html/2607.26599#S3.p2.8)\. - U\. Shalit, F\. D\. Johansson, and D\. Sontag \(2017\)Estimating individual treatment effect: generalization bounds and algorithms\.InProceedings of the 34th International Conference on Machine Learning,pp\. 3076–3085\.Cited by:[§1](https://arxiv.org/html/2607.26599#S1.p1.3),[§1](https://arxiv.org/html/2607.26599#S1.p2.1),[§2\.1](https://arxiv.org/html/2607.26599#S2.SS1.p1.4)\. - C\. Shi, D\. Blei, and V\. Veitch \(2019\)Adapting neural networks for the estimation of treatment effects\.InAdvances in Neural Information Processing Systems,Vol\.32\.Cited by:[§1](https://arxiv.org/html/2607.26599#S1.p2.1),[§2\.1](https://arxiv.org/html/2607.26599#S2.SS1.p1.4)\. - C\. Tran and E\. Zheleva \(2022\)Improving data\-driven heterogeneous treatment effect estimation under structure uncertainty\.InProceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining,pp\. 1787–1797\.External Links:[Document](https://dx.doi.org/10.1145/3534678.3539444)Cited by:[§2\.1](https://arxiv.org/html/2607.26599#S2.SS1.p1.4)\. - M\. Veljanovski and Z\. Wood\-Doughty \(2024\)DoubleLingo: causal estimation with large language models\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 2: Short Papers\),pp\. 799–807\.Cited by:[§1](https://arxiv.org/html/2607.26599#S1.p3.1),[§2\.2](https://arxiv.org/html/2607.26599#S2.SS2.p1.1)\. - F\. Wang, C\. Chen, W\. Liu, T\. Fan, X\. Liao, Y\. Tan, L\. Qi, and X\. Zheng \(2024\)CE\-RCFR: robust counterfactual regression for consensus\-enabled treatment effect estimation\.InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining,pp\. 3013–3023\.External Links:[Document](https://dx.doi.org/10.1145/3637528.3672054)Cited by:[§2\.1](https://arxiv.org/html/2607.26599#S2.SS1.p1.4)\. - H\. Wang, K\. Kuang, H\. Chi, L\. Yang, M\. Geng, W\. Huang, and W\. Yang \(2023\)Treatment effect estimation with adjustment feature selection\.InProceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining,pp\. 2290–2301\.External Links:[Document](https://dx.doi.org/10.1145/3580305.3599531)Cited by:[§2\.1](https://arxiv.org/html/2607.26599#S2.SS1.p1.4)\. - H\. Wen, T\. Chen, G\. Ye, L\. K\. Chai, S\. Sadiq, and H\. Yin \(2025\)Progressive generalization risk reduction for data\-efficient causal effect estimation\.InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Volume 1,pp\. 1575–1586\.External Links:[Document](https://dx.doi.org/10.1145/3690624.3709305)Cited by:[§2\.3](https://arxiv.org/html/2607.26599#S2.SS3.p1.1)\. - X\. Wen, H\. Zhang, S\. Zheng, W\. Xu, and J\. Bian \(2024\)From supervised to generative: a novel paradigm for tabular deep learning with large language models\.InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining,pp\. 3323–3333\.External Links:[Document](https://dx.doi.org/10.1145/3637528.3671975)Cited by:[§1](https://arxiv.org/html/2607.26599#S1.p3.1)\. - A\. Wu, J\. Yuan, K\. Kuang, B\. Li, R\. Wu, Q\. Zhu, Y\. Zhuang, and F\. Wu \(2023\)Learning decomposed representations for treatment effect estimation\.IEEE Transactions on Knowledge and Data Engineering35\(5\),pp\. 4989–5001\.External Links:[Document](https://dx.doi.org/10.1109/TKDE.2022.3150807)Cited by:[§A\.3](https://arxiv.org/html/2607.26599#A1.SS3.p1.1)\. - P\. A\. Wu and K\. Fukumizu \(2022\)β\\beta\-Intact\-VAE: identifying and estimating causal effects under limited overlap\.InInternational Conference on Learning Representations,Cited by:[§2\.1](https://arxiv.org/html/2607.26599#S2.SS1.p1.4)\. - Y\. Wu, S\. Zhai, N\. Srivastava, J\. Susskind, J\. Zhang, R\. Salakhutdinov, and H\. Goh \(2021\)Uncertainty weighted actor\-critic for offline reinforcement learning\.arXiv preprint arXiv:2105\.08140\.Cited by:[§2\.3](https://arxiv.org/html/2607.26599#S2.SS3.p1.1)\. - P\. Xu, Z\. Wen, H\. Zhao, and Q\. Gu \(2022\)Neural contextual bandits with deep representation and shallow exploration\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=xnYACQquaGV)Cited by:[§2\.3](https://arxiv.org/html/2607.26599#S2.SS3.p1.1)\. - Z\. Xu, D\. Cheng, J\. Li, J\. Liu, L\. Liu, and K\. Yu \(2024\)Causal inference with conditional front\-door adjustment and identifiable variational autoencoder\.InInternational Conference on Learning Representations,B\. Kim, Y\. Yue, S\. Chaudhuri, K\. Fragkiadaki, M\. Khan, and Y\. Sun \(Eds\.\),Vol\.2024,pp\. 18692–18713\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2024/file/5186db2bfac87765a744547e82ac8848-Paper-Conference.pdf)Cited by:[§2\.1](https://arxiv.org/html/2607.26599#S2.SS1.p1.4)\. - W\. Zhang, L\. Liu, and J\. Li \(2021\)Treatment effect estimation with disentangled latent factors\.Proceedings of the AAAI Conference on Artificial Intelligence35\(12\),pp\. 10923–10930\.External Links:[Link](https://ojs.aaai.org/index.php/AAAI/article/view/17304),[Document](https://dx.doi.org/10.1609/aaai.v35i12.17304)Cited by:[§2\.1](https://arxiv.org/html/2607.26599#S2.SS1.p1.4)\. - X\. Zhang, J\. Chen, H\. Wang, H\. Xie, Y\. Liu, J\. C\.S\. Lui, and H\. Li \(2023\)Uncertainty\-aware instance reweighting for off\-policy learning\.InAdvances in Neural Information Processing Systems,Vol\.36\.Cited by:[§2\.3](https://arxiv.org/html/2607.26599#S2.SS3.p1.1)\. - K\. Zhong, F\. Xiao, Y\. Ren, Y\. Liang, W\. Yao, X\. Yang, and L\. Cen \(2022\)DESCN: deep entire space cross networks for individual treatment effect estimation\.InProceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining,pp\. 4612–4620\.External Links:[Document](https://dx.doi.org/10.1145/3534678.3539198)Cited by:[§2\.1](https://arxiv.org/html/2607.26599#S2.SS1.p1.4)\. - D\. Zhou, L\. Li, and Q\. Gu \(2020\)Neural contextual bandits with UCB\-based exploration\.InInternational Conference on Machine Learning,pp\. 11492–11502\.Cited by:[§2\.3](https://arxiv.org/html/2607.26599#S2.SS3.p1.1)\. - Y\. Zhu, J\. Ma, L\. Wu, Q\. Guo, L\. Hong, and J\. Li \(2025\)Causal effect estimation with mixed latent confounders and post\-treatment variables\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=qe1CsfnN1W)Cited by:[§2\.1](https://arxiv.org/html/2607.26599#S2.SS1.p1.4)\. ## Appendix ADataset and Implementation Details This appendix provides additional benchmark, implementation, and reproducibility details omitted from Section[5\.1](https://arxiv.org/html/2607.26599#S5.SS1)\. ### A\.1\.Benchmark Details ##### IHDP IHDP is a semi\-synthetic benchmark based on covariates from the Infant Health and Development Program\. Treatment assignment and potential outcomes follow the standard semi\-synthetic construction, providing ground\-truth individual treatment effect for evaluation\. The covariates describe infant health, maternal background, pregnancy risk factors, and prenatal context\. ##### Adult\. Adult is constructed from the Adult income dataset with a semi\-synthetic treatment and outcome mechanism\. The treatment represents a career\-opportunity exposure related to observed demographic and employment characteristics, while the outcome indicates high\-income status\. The simulator provides ground\-truth individual effects\. ##### Jobs\. Jobs evaluates the effect of participation in a job\-training program\. The observed covariates include demographic characteristics, education, and pre\-treatment earnings histories\. Since individual\-level counterfactual outcomes are unavailable, we evaluate the quality of the estimated treatment\-effect ranking using uplift metrics\. ##### Hillstrom\. Hillstrom is an original randomized email\-marketing experiment\. We compare the Men’s E\-Mail treatment with the no\-email control\. Its covariates describe customer tenure, recency, historical spending, purchase preferences, transaction channel, and geographic segment\. As individual\-level effects are unavailable, evaluation uses uplift metrics\. ### A\.2\.Additional Implementation Details ##### Baselines\. We plugCURLinto ten host learners spanning three families: meta\-learners \(S\-, T\-, and R\-Learner\), balanced\-representation methods \(TARNet, CFRNet, and DragonNet\), and deep latent\-variable methods \(CEVAE, TEDVAE, SDD, and CiVAE\)\. We additionally include two LLM\-only baselines\.LLMprompts the LLM to directly predict CATE from the textualized profile, whileLLM\+MLPuses the frozen LLM as an encoder and trains an MLP prediction head\.GATE\[Huynhet al\.,[2025](https://arxiv.org/html/2607.26599#bib.bib43)\]directly generates counterfactual outcomes from selected covariates\. Additional comparisons with standalone feature\-disentanglement methods are reported in Appendix[A\.3](https://arxiv.org/html/2607.26599#A1.SS3)\. ##### Semantic construction\. For every unit in the newly surfaced query set𝒬\(r\)\\mathcal\{Q\}^\{\(r\)\}, we issue two independent prompts,PA,iP\_\{A,i\}andPH,iP\_\{H,i\}\. They construct assignment\-oriented and heterogeneity\-oriented representations, respectively, from the same observed profile\. The two prompts use separate dialogue states and key\-value caches\. Complete prompt, profile, andinsight\_hinttemplates are given in Appendix[C](https://arxiv.org/html/2607.26599#A3)\. ##### Hyperparameters\. The main experiments useM=30M=30MC\-dropout passes, semantic\-allocation ratioρ=0\.25\\rho=0\.25, and progressive refinement withR=3R=3, unless otherwise stated\. The objective weights areλprop=1\.0\\lambda\_\{\\mathrm\{prop\}\}=1\.0andλortho=0\.1\\lambda\_\{\\mathrm\{ortho\}\}=0\.1\. Qwen2\.5\-7B or Llama3\-8B is selected on the validation set for each benchmark; additional backbones are examined in Section[5\.4](https://arxiv.org/html/2607.26599#S5.SS4)\. ##### Optimization and data splits\. For each benchmark, 80% of the samples are allocated to the training partition, with a validation subset used for host\-model selection, LLM\-backbone selection, hyperparameter selection, and early stopping\. All methods use identical splits under each random seed\. ##### Hardware and reproducibility\. Experiments are run on two GPUs with 48 GB memory each\. Main results are reported over five random seeds\. ### A\.3\.Additional Comparison with Feature\-Disentanglement Baselines We further compare against three standalone feature\-disentanglement baselines: DR\-CFR\[Hassanpour and Greiner,[2020](https://arxiv.org/html/2607.26599#bib.bib49)\], DeR\-CFR\[Wuet al\.,[2023](https://arxiv.org/html/2607.26599#bib.bib50)\], and DDRN\[Menget al\.,[2025](https://arxiv.org/html/2607.26599#bib.bib51)\]\. These methods learn decomposed representations from the observed sample and are trained as independent estimators rather thanCURLhosts\. To keep the comparison compact and consistent across metrics, we report a fixedCURL\-enhanced R\-Learner, denoted byCURL\-R, on Jobs and Adult\. Tables[3](https://arxiv.org/html/2607.26599#A1.T3)and[4](https://arxiv.org/html/2607.26599#A1.T4)show thatCURL\-Rachieves the best result on all reported metrics\. This suggests that pretrained semantic structure provides an inductive bias complementary to feature disentanglement learned solely from the study sample\. Table 3\.Additional comparison with feature\-disentanglement baselines on Jobs\. Higher values are better\.MethodAUUC↑\\uparrowAUQC↑\\uparrowLIFT@30↑\\uparrowLLM336\.0751\.523\.25LLM\+MLP810\.9969\.112\.56DR\-CFR745\.90206\.147\.93DeR\-CFR774\.31209\.347\.68DDRN734\.14228\.017\.63CURL\-R1007\.40240\.128\.14Table 4\.Additional comparison with feature\-disentanglement baselines on Adult\. Lower values are better\.MethodPEHE↓\\downarrowϵATE\\epsilon\_\{\\mathrm\{ATE\}\}↓\\downarrowϵATT\\epsilon\_\{\\mathrm\{ATT\}\}↓\\downarrowLLM57\.3121\.5690\.708LLM\+MLP0\.4290\.1910\.170DR\-CFR0\.0920\.0100\.014DeR\-CFR0\.0890\.0100\.015DDRN0\.0940\.0220\.012CURL\-R0\.0620\.0040\.002 ## Appendix BDetailed Definitions of Evaluation Metrics ##### Uplift metrics \(Jobs and Hillstrom\)\. Individual\-level counterfactual outcomes are unavailable on Jobs and Hillstrom, so pointwise treatment\-effect errors cannot be computed\. We instead evaluate the ranking induced byτ^\(x\)\\widehat\{\\tau\}\(x\)\. Leti\(1\),…,i\(N\)i\_\{\(1\)\},\\ldots,i\_\{\(N\)\}denote the units sorted in descending order ofτ^\(xi\)\\widehat\{\\tau\}\(x\_\{i\}\), and let π\(k\)=\{i\(1\),…,i\(k\)\}\\pi\(k\)=\\\{i\_\{\(1\)\},\\ldots,i\_\{\(k\)\}\\\}be the top\-kkranked units\. Fora∈\{0,1\}a\\in\\\{0,1\\\}, define the cumulative outcome and sample count within the top\-kkprefix as \(17\)Sa\(k\)\\displaystyle S\_\{a\}\(k\)=∑i∈π\(k\)Ti=aYi,\\displaystyle=\\sum\_\{\\begin\{subarray\}\{c\}i\\in\\pi\(k\)\\\\ T\_\{i\}=a\\end\{subarray\}\}Y\_\{i\},Na\(k\)\\displaystyle N\_\{a\}\(k\)=∑i∈π\(k\)𝕀\(Ti=a\)\.\\displaystyle=\\sum\_\{i\\in\\pi\(k\)\}\\mathbb\{I\}\(T\_\{i\}=a\)\.We use the convention \(18\)Y¯a\(k\)=\{Sa\(k\)/Na\(k\),Na\(k\)\>0,0,Na\(k\)=0\.\\overline\{Y\}\_\{a\}\(k\)=\\begin\{cases\}S\_\{a\}\(k\)/N\_\{a\}\(k\),&N\_\{a\}\(k\)\>0,\\\\ 0,&N\_\{a\}\(k\)=0\.\\end\{cases\} The cumulative uplift\-curve ordinate at cutoffkkis \(19\)U\(k\)=k\(Y¯1\(k\)−Y¯0\(k\)\)\.U\(k\)=k\\left\(\\overline\{Y\}\_\{1\}\(k\)\-\\overline\{Y\}\_\{0\}\(k\)\\right\)\.Letxk=k/Nx\_\{k\}=k/N\. We report the sample\-size\-normalized trapezoidal area under this curve: \(20\)AUUC\\displaystyle\\mathrm\{AUUC\}=1NTrapz\(\{\(xk,U\(k\)\)\}k=1N\)\\displaystyle=\\frac\{1\}\{N\}\\operatorname\{Trapz\}\\left\(\\left\\\{\(x\_\{k\},U\(k\)\)\\right\\\}\_\{k=1\}^\{N\}\\right\)\(21\)=12N2∑k=1N−1\[U\(k\)\+U\(k\+1\)\]\.\\displaystyle=\\frac\{1\}\{2N^\{2\}\}\\sum\_\{k=1\}^\{N\-1\}\\left\[U\(k\)\+U\(k\+1\)\\right\]\.The additional factor1/N1/Nreduces the dependence of the reported value on the evaluation\-set size\. For the Qini curve, let \(22\)NT=N1\(N\),NC=N0\(N\),N\_\{T\}=N\_\{1\}\(N\),\\qquad N\_\{C\}=N\_\{0\}\(N\),and assumeNT\>0N\_\{T\}\>0andNC\>0N\_\{C\}\>0\. The Qini ordinate uses the full\-sample treated\-to\-control count ratio: \(23\)Q\(k\)=S1\(k\)−NTNCS0\(k\)\.Q\(k\)=S\_\{1\}\(k\)\-\\frac\{N\_\{T\}\}\{N\_\{C\}\}S\_\{0\}\(k\)\.The reported normalized area under the Qini curve is \(24\)AUQC\\displaystyle\\mathrm\{AUQC\}=1NTrapz\(\{\(xk,Q\(k\)\)\}k=1N\)\\displaystyle=\\frac\{1\}\{N\}\\operatorname\{Trapz\}\\left\(\\left\\\{\(x\_\{k\},Q\(k\)\)\\right\\\}\_\{k=1\}^\{N\}\\right\)\(25\)=12N2∑k=1N−1\[Q\(k\)\+Q\(k\+1\)\]\.\\displaystyle=\\frac\{1\}\{2N^\{2\}\}\\sum\_\{k=1\}^\{N\-1\}\\left\[Q\(k\)\+Q\(k\+1\)\\right\]\.If either treatment arm is absent from the evaluation set, AUQC is defined as zero\. Finally, define the empirical uplift within a ranked prefix as \(26\)L\(k\)=\{S1\(k\)N1\(k\)−S0\(k\)N0\(k\),N1\(k\)\>0andN0\(k\)\>0,0,otherwise\.L\(k\)=\\begin\{cases\}\\dfrac\{S\_\{1\}\(k\)\}\{N\_\{1\}\(k\)\}\-\\dfrac\{S\_\{0\}\(k\)\}\{N\_\{0\}\(k\)\},&N\_\{1\}\(k\)\>0\\ \\text\{and\}\\ N\_\{0\}\(k\)\>0,\\\\\[6\.0pt\] 0,&\\text\{otherwise\}\.\\end\{cases\}Let \(27\)K30=⌊0\.3N⌋K\_\{30\}=\\lfloor 0\.3N\\rfloorand define the full\-sample empirical uplift as \(28\)Lall=S1\(N\)NT−S0\(N\)NC\.L\_\{\\mathrm\{all\}\}=\\frac\{S\_\{1\}\(N\)\}\{N\_\{T\}\}\-\\frac\{S\_\{0\}\(N\)\}\{N\_\{C\}\}\.We report the relative top\-30% lift \(29\)LIFT@30=L\(K30\)Lall\+ε,ε=10−9\.\\mathrm\{LIFT@30\}=\\frac\{L\(K\_\{30\}\)\}\{L\_\{\\mathrm\{all\}\}\+\\varepsilon\},\\qquad\\varepsilon=10^\{\-9\}\.Thus, LIFT@30 is dimensionless and measures the uplift in the highest\-ranked 30% of units relative to the overall empirical uplift\. AUUC, AUQC, and LIFT@30 are higher\-is\-better under the evaluation protocol used in our experiments\. ## Appendix CPrompt Construction and Semantic Extraction This appendix specifies the construction of the dataset\-levelinsight\_hint, the two role\-conditioned promptsPAP\_\{A\}andPHP\_\{H\}, the deterministic profile constructors, and the semantic\-embedding extraction procedure\. ### C\.1\.Insight Hint At initialization, we construct the dataset\-levelinsight\_hintonce from the training set by comparing the initially selected units𝒮\(0\)\\mathcal\{S\}^\{\(0\)\}with the full training population using only observed covariates\. For each numeric field, we compute the relative difference between the two means and retain at most ten fields whose absolute difference exceeds20%20\\%, indicating whether each field is higher or lower in𝒮\(0\)\\mathcal\{S\}^\{\(0\)\}\. This aggregate summary is inserted into the template below, from which the LLM generates a one\-sentence hinthinsh\_\{\\mathrm\{ins\}\}\. The hint is then reused unchanged for that dataset throughout all refinement rounds and contains no unit\-level treatment assignments, outcomes, counterfactuals, or test\-set information\. `Template for insight\_hint` `C\.2\. Role\-Conditioned Prompt Templates For every queried unit ii, the deterministic profile pi=ψ\(xi\)p\_\{i\}=\\psi\(x\_\{i\}\) and fixed hint hinsh\_\{\\mathrm\{ins\}\} are inserted into two independent prompts\. The assignment\-oriented prompt PA,iP\_\{A,i\} focuses on pre\-treatment baseline structure relevant to the shared and assignment\-side components of the estimator\. The heterogeneity\-oriented prompt PH,iP\_\{H,i\} focuses on stable pre\-treatment attributes relevant to treatment\-response variation\. Both prompts explicitly restrict the LLM to semantic relations supported by the observed profile\. Assignment\-oriented prompt PAP\_\{A\}\. PromptA\\mathrm\{Prompt\}\_\{A\}: Pre\-treatment Assignment\-Side Inference Heterogeneity\-oriented prompt PHP\_\{H\} for IHDP, Jobs, and Hillstrom\. PromptH\\mathrm\{Prompt\}\_\{H\}: General Response\-Side Heterogeneity Inference Specialized PHP\_\{H\} for Adult\. Adult uses a domain\-specific variant because its simulated heterogeneity is primarily associated with employment structure, education, and socio\-economic mobility\. PromptH\\mathrm\{Prompt\}\_\{H\} for Adult: Socio\-Economic Modifier Inference Semantic extraction\. For each prompt, the frozen LLM first generates a deterministic textual response\. We then perform a second forward pass over the concatenated prompt and generated response and extract the final\-layer hidden state of the last valid generated\-response token\. This yields \(30\) zA,iLLM=EmbΦ\(PA,i\),zH,iLLM=EmbΦ\(PH,i\),z\_\{A,i\}^\{\\mathrm\{LLM\}\}=\\operatorname\{Emb\}\_\{\\Phi\}\(P\_\{A,i\}\),\\qquad z\_\{H,i\}^\{\\mathrm\{LLM\}\}=\\operatorname\{Emb\}\_\{\\Phi\}\(P\_\{H,i\}\), consistent with Eq\. \(5\)\. The generated text is used only for inspection; the adapter consumes the corresponding hidden\-state embedding\. C\.3\. Dataset\-Specific Profile Templates The profile constructor ψ\\psi deterministically renders each structured covariate vector into natural language\. No treatment assignment, factual outcome, counterfactual outcome, prediction error, or test statistic is included\. A dataset\-level treatment definition is included only to specify the meaning of the intervention\. IHDP profile\. IHDP profile template Here, the pregnancy\-risk statement is formed deterministically from the observed smoking, alcohol\-use, and drug\-use indicators\. Adult profile\. Adult profile template Jobs profile\. Jobs profile template Hillstrom profile\. Hillstrom profile template Appendix D Training Algorithm Algorithm 1 summarizes training\. We use RR to denote the total number of semantic\-assisted progressive optimization rounds, including the first round initialized by the initial semantic allocation\. Accordingly, the rounds are indexed by r=0,…,R−1r=0,\\ldots,R\-1, and the final wrapped estimator is F\(R\)F^\{\(R\)\}\. The initial allocation at r=0r=0 uses the same joint ranking rule as the subsequent rounds, but se,i\(0\)s\_\{e,i\}^\{\(0\)\} is uniformly set to zero\. Hence, the initial ordering reduces to the ordering induced by sτ,i\(0\)s\_\{\\tau,i\}^\{\(0\)\} without introducing a separate allocation rule\. After the first optimization round, every wrapped host contains a trained propensity head\. Therefore, both se,i\(r\)s\_\{e,i\}^\{\(r\)\} and sτ,i\(r\)s\_\{\\tau,i\}^\{\(r\)\} are used in the remaining refinement rounds r=1,…,R−1r=1,\\ldots,R\-1\. The selected set 𝒮\(r\)\\mathcal\{S\}^\{\(r\)\}, cumulative semantic cache 𝒞\(r\)\\mathcal\{C\}^\{\(r\)\}, and newly surfaced query set 𝒬\(r\)\\mathcal\{Q\}^\{\(r\)\} follow Eq\. \(4\)\. The LLM is queried only for 𝒬\(r\)\\mathcal\{Q\}^\{\(r\)\}, so every training unit is queried at most once\. 1 Input: dataset 𝒟\\mathcal\{D\}; host learner; frozen LLM Φ\\Phi; ρ,M,R,λprop,λortho\\rho,M,R,\\lambda\_\{\\mathrm\{prop\}\},\\lambda\_\{\\mathrm\{ortho\}\} Output: trained wrapped estimator F⋆F^\{\\star\} 2 3Fit the unaugmented host using its native objective 4 5Initialize the wrapped estimator F\(0\)F^\{\(0\)\} with zero semantic inputs 6 7𝒞\(−1\)←∅\\mathcal\{C\}^\{\(\-1\)\}\\leftarrow\\varnothing 8 9for r=0,…,R−1r=0,\\ldots,R\-1 do 10 11 if r=0r=0 then /\* Initial semantic allocation \*/ 12 Run MM MC\-dropout passes of F\(0\)F^\{\(0\)\} and compute sτ,i\(0\)s\_\{\\tau,i\}^\{\(0\)\} for all ii 13 14 Set se,i\(0\)←0s\_\{e,i\}^\{\(0\)\}\\leftarrow 0 for all ii 15 16 else /\* Refinement allocation \*/ 17 Run MM MC\-dropout passes of F\(r\)F^\{\(r\)\} 18 19 Compute se,i\(r\)s\_\{e,i\}^\{\(r\)\} and sτ,i\(r\)s\_\{\\tau,i\}^\{\(r\)\} using Eq\. \(2\) 20 21 22 Form 𝒮\(r\)\\mathcal\{S\}^\{\(r\)\} using Eq\. \(3\) 23 24 𝒞\(r\)←𝒞\(r−1\)∪𝒮\(r\)\\mathcal\{C\}^\{\(r\)\}\\leftarrow\\mathcal\{C\}^\{\(r\-1\)\}\\cup\\mathcal\{S\}^\{\(r\)\} 25 26 𝒬\(r\)←𝒮\(r\)∖𝒞\(r−1\)\\mathcal\{Q\}^\{\(r\)\}\\leftarrow\\mathcal\{S\}^\{\(r\)\}\\setminus\\mathcal\{C\}^\{\(r\-1\)\} 27 28 Query the LLM only for units in 𝒬\(r\)\\mathcal\{Q\}^\{\(r\)\} to obtain zA,iLLMz\_\{A,i\}^\{\\mathrm\{LLM\}\} and zH,iLLMz\_\{H,i\}^\{\\mathrm\{LLM\}\} 29 30 Construct zA,i\(r\)z\_\{A,i\}^\{\(r\)\} and zH,i\(r\)z\_\{H,i\}^\{\(r\)\} using Eq\. \(6\) 31 32 Optimize the host, projectors, gates, and propensity head from the current parameters using Eq\. \(12\), obtaining F\(r\+1\)F^\{\(r\+1\)\} 33 34 35F⋆←F\(R\)F^\{\\star\}\\leftarrow F^\{\(R\)\} 36 37return F⋆F^\{\\star\} Algorithm 1 CURL training with progressive refinement Appendix E Adapter Instantiation Across Host Learners For every host, the attached propensity head consumes the assignment\-adjusted representation uAu\_\{A\}, while the host\-specific effect operator consumes uHu\_\{H\}, which already contains uAu\_\{A\} together with the heterogeneity\-oriented feature block\. The host’s native parameterization and main objective are otherwise retained\. Table 5\. Instantiation of CURL across the ten host learners\. The propensity component consumes uAu\_\{A\}, while the effect component consumes uHu\_\{H\}\. The complete objective is ℒmain\+λpropℒprop\+λorthoℒortho\\mathcal\{L\}\_\{\\mathrm\{main\}\}\+\\lambda\_\{\\mathrm\{prop\}\}\\mathcal\{L\}\_\{\\mathrm\{prop\}\}\+\\lambda\_\{\\mathrm\{ortho\}\}\\mathcal\{L\}\_\{\\mathrm\{ortho\}\}\. Pattern Host learner\(s\) Assignment/shared pathway Effect pathway Native main objective Treatment\-conditioned outcome S\-Learner e^←uA\\widehat\{e\}\\leftarrow u\_\{A\} μ^\(uH,t\)\\widehat\{\\mu\}\(u\_\{H\},t\) Factual outcome loss Separate potential\-outcome heads T\-Learner e^←uA\\widehat\{e\}\\leftarrow u\_\{A\} μ^0\(uH\),μ^1\(uH\)\\widehat\{\\mu\}\_\{0\}\(u\_\{H\}\),\\widehat\{\\mu\}\_\{1\}\(u\_\{H\}\) Arm\-specific factual losses Shared representation with outcome heads TARNet, CFRNet, DragonNet e^←uA\\widehat\{e\}\\leftarrow u\_\{A\} μ^0\(uH\),μ^1\(uH\)\\widehat\{\\mu\}\_\{0\}\(u\_\{H\}\),\\widehat\{\\mu\}\_\{1\}\(u\_\{H\}\) Native factual, balancing, or targeted loss Direct CATE R\-Learner e^,μ^←uA\\widehat\{e\},\\widehat\{\\mu\}\\leftarrow u\_\{A\} τ^←uH\\widehat\{\\tau\}\\leftarrow u\_\{H\} Orthogonalized R\-loss \[Nie and Wager, 2021\] Latent\-variable CEVAE, TEDVAE, SDD, CiVAE Propensity and shared baseline components use uAu\_\{A\} Outcome/CATE operator additionally conditions on uHu\_\{H\} Native ELBO or disentanglement objective Figure 9\. Representative CURL instantiations for \(a\) S\-Learner, \(b\) TARNet, \(c\) R\-Learner, and \(d\) TEDVAE\. Across hosts, the propensity pathway consumes uAu\_\{A\}, whereas the host\-specific effect operator consumes uHu\_\{H\}\. The insertion point varies with the host architecture, while the role\-aware interface remains unchanged\. Table 5 summarizes the common interface rather than claiming an identifiable causal decomposition of the host’s internal variables\. The assignment\-oriented representation contributes to the shared state used by the propensity component, while the heterogeneity\-oriented representation is retained as a separate block available only to the effect operator\. This interface accommodates treatment\-conditioned predictors, separate potential\-outcome heads, direct\-CATE learners, and latent\-variable estimators\. Appendix F Additional Analysis of Progressive Refinement We perform an initial semantic allocation followed by five refinement rounds on IHDP and Jobs using S\-Learner, T\-Learner, TARNet, and CFRNet\. For the cache analysis, define \(31\) κ\(r\)=\|𝒞\(r\)\|N\.\\kappa^\{\(r\)\}=\\frac\{\|\\mathcal\{C\}^\{\(r\)\}\|\}\{N\}\. For each refinement round, we also report the mean uncertainty scores \(32\) s¯e\(r\)=1N∑i=1Nse,i\(r\),s¯τ\(r\)=1N∑i=1Nsτ,i\(r\)\.\\overline\{s\}\_\{e\}^\{\(r\)\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}s\_\{e,i\}^\{\(r\)\},\\qquad\\overline\{s\}\_\{\\tau\}^\{\(r\)\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}s\_\{\\tau,i\}^\{\(r\)\}\. Figure 10\. Progressive\-refinement dynamics on IHDP and Jobs\. For each host, the left panel reports the cumulative semantic\-cache ratio from the initial allocation through five refinement rounds\. The right panel reports the mean assignment\-side uncertainty s¯e\(r\)\\overline\{s\}\_\{e\}^\{\(r\)\} and CATE\-side uncertainty s¯τ\(r\)\\overline\{s\}\_\{\\tau\}^\{\(r\)\}\. Error bars and shaded regions indicate variation across runs\. Figure 10 shows that the cache expands rapidly during the early rounds but exhibits progressively smaller marginal growth\. The final cumulative cache ratio ranges from 68\.7%68\.7\\% to 82\.5%82\.5\\% on IHDP and from 51\.2%51\.2\\% to 63\.2%63\.2\\% on Jobs\. Across all eight dataset–host combinations, the cache ratio is unchanged between the fourth and fifth refinement rounds, indicating that few or no previously unqueried units are surfaced at that stage\. Meanwhile, both s¯e\(r\)\\overline\{s\}\_\{e\}^\{\(r\)\} and s¯τ\(r\)\\overline\{s\}\_\{\\tau\}^\{\(r\)\} exhibit an overall downward trend, with only minor intermediate fluctuations\. The simultaneous reduction in predictive uncertainty and saturation of the cache is consistent with the estimator becoming locally more stable while the allocation set approaches a fixed configuration\. This dynamic evidence complements the performance\-based sensitivity analysis of RR in Section 5\.4\. Appendix G Analysis of Semantic Representation We conduct two complementary studies on IHDP to examine whether CURL’s gains arise from pretrained representational structure and whether they persist across training\-set sizes\. Figure 11 summarizes the results\. Figure 11\. Additional validation of representational lossiness on IHDP\. \(a\) Semantic\-corruption results averaged over ten IHDP realizations and five seeds\. Here, Δ=PEHEHost−PEHESetting\\Delta=\\mathrm\{PEHE\}\_\{\\mathrm\{Host\}\}\-\\mathrm\{PEHE\}\_\{\\mathrm\{Setting\}\}, so positive values indicate improvement over the host\. \(b\) PEHE under nested training\-set fractions, averaged over ten realizations and three seeds\. CURL improves both hosts at every fraction\. Semantic corruption\. This experiment distinguishes gains from meaningful pretrained representations from those caused merely by additional vectors, parameters, or textual inputs\. We use the ten IHDP realizations with T\-Learner and CFRNet\. Each realization is evaluated using seeds \{42,52,62,72,82\}\\\{42,52,62,72,82\\\}, giving 10×5=5010\\times 5=50 results per bar\. The five settings are: • Host: the original host learner without CURL; • CURL: the complete method with the original field and category semantics; • Anonymized: field and category names are replaced by fixed identifiers such as Feature F1 and Category C1, while values and category identities remain unchanged; • Shuffled: category names are mapped to other names through a fixed within\-field permutation while retaining the natural\-language input format; • Random: semantic embeddings are replaced by fixed random vectors of the same dimension, while the adapter, gates, and routing remain unchanged\. All settings use identical train/test splits, model seeds, and queried units\. The uncorrupted CURL run first records its query plan, which is then replayed for the other semantic variants\. We report mean PEHE and define \(33\) Δ=PEHEHost−PEHESetting\.\\Delta=\\mathrm\{PEHE\}\_\{\\mathrm\{Host\}\}\-\\mathrm\{PEHE\}\_\{\\mathrm\{Setting\}\}\. Thus, Δ\>0\\Delta\>0 indicates improvement over the host\. CURL achieves the lowest PEHE for both learners\. Anonymized and Shuffled both degrade performance relative to CURL, while Random produces the largest deterioration\. This ordered degradation indicates that CURL’s improvement derives in part from the semantic structure encoded by the LLM representations, rather than merely from additional vectors, adapter parameters, or routing components\. Training\-set fraction\. We next examine whether CURL’s advantage persists as the amount of training data changes\. For each IHDP realization, we construct a fixed treatment\-stratified split containing an 80% training pool and a 20% test set\. From the same training pool, we form deterministic nested subsets at 30%, 50%, 70%, and 100%, such that each smaller subset is contained in the next larger subset\. The displayed comparison uses T\-Learner and TARNet with model seeds \{42,52,62\}\\\{42,52,62\\\}\. Host and CURL use identical subsets, test sets, and model seeds\. The LLM configuration is the same as in the semantic\-corruption experiment\. CURL obtains lower PEHE than its paired host for both learners at every training fraction\. Thus, CURL’s benefit is not restricted to a particular subsample size and remains visible even when all available IHDP training data are used\. Appendix H Case Study: Semantic Interpretation of the Two Channels Figure 12\. Walkthrough of a representative locally unstable Hillstrom unit\. The observed profile, aggregate covariate\-shift hint, assignment\-oriented and heterogeneity\-oriented textual traces, and downstream estimates illustrate how CURL reorganizes information already contained in the structured covariates\. Figure 12 examines a representative Hillstrom customer selected by the uncertainty diagnostic\. The observed profile describes a relatively new customer who last interacted with the brand nine months earlier, has $900\.73 in historical spending concentrated on men’s items, primarily transacts by phone, and resides in a rural area\. These fields are observed in XX, but their semantic relations are not explicit in the original numerical and categorical encoding\. The assignment\-oriented trace emphasizes pre\-treatment baseline structure already supported by the profile, including the relation between customer tenure, inactivity, accumulated value, transaction channel, and geographic context\. This trace should not be interpreted as recovering hidden confounders or as implying non\-random treatment assignment in Hillstrom\. Rather, it provides a semantic organization of the observed baseline profile for the shared and assignment\-side components of the estimator\. The heterogeneity\-oriented trace instead emphasizes relations that may be useful for modeling differential response to the Men’s E\-Mail campaign\. In particular, the combination of men’s\-product purchase history, prolonged inactivity, historical value, and preferred transaction channel provides a more explicit description of potential content match and engagement sensitivity\. These are semantic interpretations of observed fields rather than newly observed individual characteristics\. The textual traces are displayed only for interpretation\. The actual adapter inputs are zA,iLLMz\_\{A,i\}^\{\\mathrm\{LLM\}\} and zH,iLLMz\_\{H,i\}^\{\\mathrm\{LLM\}\}, extracted from the corresponding generated responses\. The case study therefore illustrates how pretrained semantic structure can reshape the effective representation of XX without changing the observed information set, treatment assignment, outcome, estimand, or identifying assumptions\.`
Similar Articles
LLMs on Tabular Data with Limited Semantics: Evidence from Industrial Car Retrofit Prediction
This paper evaluates LLM-based strategies (embedding, prompt, hybrid) against classical tabular models on an industrial car retrofit prediction dataset with hashed categorical features. It finds that tree ensembles outperform LLMs overall, but embeddings and hybrid approaches remain useful, while direct prompting fails without semantic cues.
Uncertainty-Aware Trust Estimation for Multi-LLM Systems via Structured Expert Judgement
This paper introduces an uncertainty-aware trust estimation method for aggregating predictions from multiple LLMs, adapting structured expert judgment with Cooke-style log weighting to penalize overconfident incorrect predictions. Evaluations on MMLU and MMLU-Pro show that this approach achieves superior accuracy-reliability balance under heterogeneous and contaminated expert panels.
Can LLMs Take Retrieved Information with a Grain of Salt?
This paper investigates how large language models adapt to the certainty of retrieved information, identifying systematic limitations in handling uncertainty. It proposes an interaction strategy that reduces obedience errors by 25% without modifying model weights.
GEM: Geometric Entropy Mixing for Optimal LLM Data Curation
GEM reformulates LLM data curation as a variational problem on the hypersphere, using geometric entropy mixing and a minorize-maximize algorithm to discover balanced semantic clusters, achieving state-of-the-art improvements in data mixing strategies by up to 1.2% average downstream accuracy.
Uncertainty Decomposition for Clarification Seeking in LLM Agents
This paper proposes a prompt-based uncertainty decomposition method for LLM agents that separates action confidence from request uncertainty, enabling proactive clarification seeking in underspecified tasks. The method is evaluated on new clarification-augmented benchmarks across five LLM backbones, showing significant improvements.