Observational Policy Ranking for SMB Financial Guidance from Multi-Action Accounting Logs

arXiv cs.LG Papers

Summary

This paper introduces CAR-PL, an observational policy ranking method for recommending business changes to small and medium-sized businesses using multi-action accounting logs, comparing it against several baselines on financial KPI prediction.

arXiv:2608.10050v1 Announce Type: new Abstract: Small and medium-sized businesses need timely financial guidance, yet historical accounting logs record self-selected and often co-occurring business changes rather than randomized recommendations. We formulate this setting as observational policy ranking: from pre-decision financial information, a policy selects one of 34 ledger-derived business-change categories for a target financial KPI. Using 85,078 company-month observations from 7,505 firms, we introduce Covariate-Adjusted Residual Policy Learning (CAR-PL), an action-wise R-learner that operates directly on multi-hot logs and regularizes selection by observational support. We compare CAR-PL with an uplift T-Learner, a conservative contextual value model, a zero-shot LLM, and non-personalized references on company-disjoint held-out firms under a shared model-assisted scoring rule. CAR-PL has the highest Gross Profit point estimate (0.084), the T-Learner has the highest Revenue point estimate (0.085), and the contextual value model has the highest Quick Ratio point estimate (0.062). CAR-PL and the T-Learner are not statistically separated on either growth KPI in matched company-clustered comparisons, while CAR-PL selects 33-34 categories and produces less concentrated selections across the catalog. Outcome-model-only scoring retains the same KPI-level point-estimate leader or top pair, and category rankings remain similar when the all-zero treatment reference is replaced by the most common training co-action pattern. These findings support objective-specific ranking of SMB financial guidance from multi-action accounting logs.
Original Article
View Cached Full Text

Cached at: 08/12/26, 08:27 AM

# Observational Policy Ranking for SMB Financial Guidance from Multi-Action Accounting Logs
Source: [https://arxiv.org/html/2608.10050](https://arxiv.org/html/2608.10050)
Vignesh SubrahmaniamForesight\-AI, IntuitVikas RaturiForesight\-AI, IntuitKamalika DasForesight\-AI, IntuitXiang GaoForesight\-AI, IntuitKratika GuptaForesight\-AI, IntuitRuocheng GuoForesight\-AI, IntuitPadmaja JonnalageddaForesight\-AI, IntuitAnanya PramodForesight\-AI, IntuitSricharan KumarForesight\-AI, Intuit

\(August 2026\)

###### Abstract

Small and medium\-sized businesses need timely financial guidance, yet historical accounting logs record self\-selected and often co\-occurring business changes rather than randomized recommendations\. We formulate this setting as observational policy ranking: from pre\-decision financial information, a policy selects one of 34 ledger\-derived business\-change categories for a target financial KPI\. Using 85,078 company\-month observations from 7,505 firms, we introduce Covariate\-Adjusted Residual Policy Learning \(CAR\-PL\), an action\-wise R\-learner that operates directly on multi\-hot logs and regularizes selection by observational support\. We compare CAR\-PL with an uplift T\-Learner, a conservative contextual value model, a zero\-shot LLM, and non\-personalized references on company\-disjoint held\-out firms under a shared model\-assisted scoring rule\. CAR\-PL has the highest Gross Profit point estimate \(0\.084\), the T\-Learner has the highest Revenue point estimate \(0\.085\), and the contextual value model has the highest Quick Ratio point estimate \(0\.062\)\. CAR\-PL and the T\-Learner are not statistically separated on either growth KPI in matched company\-clustered comparisons, while CAR\-PL selects 33–34 categories and produces less concentrated selections across the catalog\. Outcome\-model\-only scoring retains the same KPI\-level point\-estimate leader or top pair, and category rankings remain similar when the all\-zero treatment reference is replaced by the most common training co\-action pattern\. These findings support objective\-specific ranking of SMB financial guidance from multi\-action accounting logs\.

Keywords:observational policy ranking; financial decision support; multi\-action logs; model\-assisted comparison\.

CCS Concepts:Computing methodologies – Machine learning approaches; Applied computing – Economics\.

## 1Introduction

Small and medium\-sized businesses \(SMBs\) often need advice before financial problems become acute, yet expert advisory capacity is costly and difficult to scale\. Randomized consulting programs show that external guidance can change business practices and outcomes, but such programs are expensive to deliver broadly\[[1](https://arxiv.org/html/2608.10050#bib.bib1)\]\. Accounting systems provide a complementary source of evidence: detailed financial histories together with the business changes that firms subsequently make\. Converting these histories into candidate guidance requires learning from observational, multi\-action records rather than from randomized recommendation labels\.

We formulate the task as*observational policy ranking*\. At a prospective decision boundary, a policy observes information available through montht\+0t\{\+\}0, selects one category from a fixed 34\-category catalog, and targets a subsequent financial KPI\. Historical categories are ledger\-derived business\-change proxies detected during montht\+1t\{\+\}1; outcomes are measured fromt\+2t\{\+\}2throught\+12t\{\+\}12\. The action month is excluded from both state and outcome, producing a clean prospective timeline for every company\-month decision\.

Three properties make the task difficult: \(1\) category occurrence is strongly state dependent, creating systematic selection; \(2\) multiple categories can occur in the same month, complicating focal credit assignment; and \(3\) observational support varies sharply across the 34 categories, making the policy argmax sensitive to rare\-category noise\. The learner must therefore estimate conditional contrasts and select stably across unevenly observed candidates\. Within each KPI, a credible comparison holds the action space, retained test rows, outcome definition, and scoring rule fixed across policy families\.

We compare a conservative contextual value model, an uplift T\-Learner, and CAR\-PL, a covariate\-adjusted action\-wise R\-learner with support shrinkage\. A zero\-shot LLM, a modal constant policy, and a frequency\-weighted random reference provide complementary comparisons\. Within each KPI, every method uses the same catalog and retained company\-disjoint test rows\. A shared model\-assisted focal\-action score combines an outcome\-model contrast with propensity\-weighted residual evidence\.

The point\-estimate pattern is KPI\-specific: CAR\-PL has the highest Gross Profit point estimate, the T\-Learner the highest Revenue point estimate, and the contextual value model the highest Quick Ratio point estimate\. CAR\-PL and the T\-Learner are not statistically separated on either growth KPI in matched company\-clustered comparisons\. On those KPIs, they agree on only 12\.8–14\.2% of company\-month states, while CAR\-PL selects 33–34 categories\. Thus, comparable aggregate scores can arise from markedly different state\-to\-category mappings, making recommendation concentration a substantive dimension of policy comparison\.

Our contributions are threefold\. First, we formulate prospective policy ranking from multi\-action accounting logs, with a fixed pre\-decision, action, and outcome timeline and company\-disjoint evaluation across 7,505 firms\. Second, we introduce CAR\-PL, which combines action\-wise R\-learning with support shrinkage to stabilize one\-of\-34 selection while retaining the original multi\-hot observations\. Third, we benchmark learned, zero\-shot, and non\-personalized policies under a single frozen scorer and report KPI\-dependent point\-estimate leaders, robustness across scoring variants, and sharply different recommendation\-concentration profiles\.

## 2Related Work

Observational policy learning\.Policy\-learning methods use observational data to optimize treatment rules under explicit adjustment and overlap assumptions\. Athey and Wager develop efficient observational policy learning for binary treatment, while Zhou, Athey, and Wager extend doubly robust policy learning to multiple actions\[[2](https://arxiv.org/html/2608.10050#bib.bib2),[3](https://arxiv.org/html/2608.10050#bib.bib3)\]\. Our setting contributes a large industrial application in which the logged representation is multi\-hot and the policy emits one interpretable category\. Work on multiple versions of treatment provides the conceptual basis for defining focal effects when observed treatment categories occur in heterogeneous bundles\[[4](https://arxiv.org/html/2608.10050#bib.bib4)\]\.

Heterogeneous effects and support\-aware selection\.Meta\-learners estimate conditional treated\-control contrasts by reducing heterogeneous\-effect estimation to supervised learning\[[5](https://arxiv.org/html/2608.10050#bib.bib5)\]\. R\-learning residualizes both outcome and treatment against pre\-treatment covariates and estimates the remaining conditional effect signal\[[6](https://arxiv.org/html/2608.10050#bib.bib6)\]\. CAR\-PL builds on this construction with action\-wise heads and a support\-based shrinkage step designed for argmax selection over many categories\. This complements uncertainty\-aware and pessimistic policy\-learning methods that explicitly penalize weakly supported decisions\[[7](https://arxiv.org/html/2608.10050#bib.bib7)\]\.

Offline value learning and financial decisions\.Offline reinforcement learning addresses policy optimization from fixed logs and must control overestimation on unsupported actions\[[8](https://arxiv.org/html/2608.10050#bib.bib8),[9](https://arxiv.org/html/2608.10050#bib.bib9)\]\. CQL implements this principle through a conservative value penalty\[[10](https://arxiv.org/html/2608.10050#bib.bib10)\]; we adapt its one\-step objective as a contextual value baseline\. Machine learning is increasingly used in credit and financial decision systems\[[11](https://arxiv.org/html/2608.10050#bib.bib11),[12](https://arxiv.org/html/2608.10050#bib.bib12)\], and language models are being explored for recommendation and financial advice\[[13](https://arxiv.org/html/2608.10050#bib.bib13),[14](https://arxiv.org/html/2608.10050#bib.bib14)\]\. The zero\-shot policy in our study measures the ranking induced by pretrained business knowledge without fitting to historical outcomes\.

SMB guidance as decision support\.Financial recommendation differs from conventional product recommendation because the output is an operational business change and the outcome unfolds over subsequent accounting periods\. Prior work studies consulting interventions for small firms\[[1](https://arxiv.org/html/2608.10050#bib.bib1)\], algorithmic credit decisions\[[11](https://arxiv.org/html/2608.10050#bib.bib11),[12](https://arxiv.org/html/2608.10050#bib.bib12)\], and financial advice from language models\[[14](https://arxiv.org/html/2608.10050#bib.bib14)\]\. Our study brings these strands together in a repeated company\-month setting with prospective timing and a common interpretable catalog\. It targets the combination left open by these strands: multi\-hot observational actions, a single interpretable policy output, and KPI\-specific evaluation on unseen firms\.

Pre\-decision state and baselinet−10,…,t\+0t\{\-\}10,\\ldots,t\{\+\}0Ledger\-detected focal exposureaction montht\+1t\{\+\}1Post\-action KPI outcomet\+2,…,t\+12t\{\+\}2,\\ldots,t\{\+\}12Financial reports andprior\-action historyNarrative embeddingpolicy stateXXCandidate policyselectsa∈𝒜a\\in\\mathcal\{A\}Shared scoring rulefor selected categoryMean selected\-categorymodel\-assisted score

Figure 1:Prospective timing and shared comparison pipeline\. No action\-month or post\-action information enters a policy state\. Every policy emits one of the same 34 categories, and within each KPI the same scoring rule is applied to the selected category on the same retained test rows\.
## 3Data and Task

### 3\.1Decision Timing, State, and Data Partitions

Let𝒜=\{1,…,34\}\\mathcal\{A\}=\\\{1,\\ldots,34\\\}denote the supported policy catalog and𝒳\\mathcal\{X\}the policy\-state space\. Each observation is a company\-month pairiiwith policy stateXiX\_\{i\}and action\-month exposure vector𝐓i∈\{0,1\}34\\mathbf\{T\}\_\{i\}\\in\\\{0,1\\\}^\{34\}\. The corpus contains 85,078 observations from 7,505 firms, with a median of 12 decision months per firm\. Figure[1](https://arxiv.org/html/2608.10050#S2.F1)shows the fixed timing convention\. The policy state and reward baseline use information throught\+0t\{\+\}0; business changes are detected duringt\+1t\{\+\}1; and monthly KPI values fromt\+2t\{\+\}2throught\+12t\{\+\}12determine the reward\. The action month is excluded\. Observations with an empty prior or outcome window are removed before KPI\-specific denominator filters are applied\.

A frozen, versioned reporting pipeline first derives structured facts and diagnostics from multiple pre\-decision financial reports, including KPI trajectories, product, customer, and vendor concentration, receivables and payables, cash\-flow and debt trends, and detected anomalies\. Multiple LLM calls then synthesize these intermediate summaries into the final natural\-language state document, which also contains month\-by\-month historical KPI tables\. This staged construction accommodates long report histories without passing the complete raw report set to a single model call\. The pipeline was fixed before policy training, receives neither action labels nor post\-decision outcomes, and excludes action\-month transactions\. These are substantial state documents \(median approximately 23,000 characters\), preserving both high\-level interpretation and detailed historical trajectories\. The Qwen3\-Embedding\-0\.6B model encodes each document into 768 dimensions\[[15](https://arxiv.org/html/2608.10050#bib.bib15)\]\. A 44\-dimensional binary history block records prior\-month activity from the full detector vocabulary\. The ten labels below the policy\-output support threshold serve only as lagged state covariates and are never eligible policy outputs\. Concatenation produces the policy stateXi∈𝒳⊆ℝ812X\_\{i\}\\in\\mathcal\{X\}\\subseteq\\mathbb\{R\}^\{812\}, which policy learners use\. Records are de\-identified before modeling, processed in the approved analytics environment, and reported only in aggregate\.

Firm identifiers are assigned by a deterministic pseudorandom hash, so the split is outcome independent and reproducible\. The hash reserves held\-out test firms, corresponding to approximately 10% of retained rows, and divides the remaining development firms into training and validation partitions\. Training rows fit policy parameters and nuisance functions\. Validation rows select model configurations, early stopping, and CAR\-PL’s shrinkage strength\. All partitions, nuisance cross\-fitting folds, and early\-stopping splits are defined at the firm level, so all months of a firm remain in one partition\. The test set is used only for the reported policy scores and confidence intervals\. Table[1](https://arxiv.org/html/2608.10050#S3.T1)reports training and validation jointly as the development population\.

Table 1:KPI\-specific sample sizes after reward construction and filtering\. Row counts are company\-month observations\. Development combines the company\-disjoint training and validation partitions; the final column reports unique firms represented after KPI\-specific filtering\.
### 3\.2Action Catalog

The extraction layer produces 44 raw ledger\-derived labels\. An outcome\-independent support rule retains 34 labels with at least 100 treated and 100 control observations in training; these define the common policy catalog\. Table[2](https://arxiv.org/html/2608.10050#S3.T2)groups them into ten operational domains\. Representative categories include increasing advertising spend, canceling subscriptions, substituting vendors, increasing training spend, securing financing, requiring upfront deposits, switching payment processors, and adopting time tracking\. Directional changes are represented as separate categories where applicable\. The shared vocabulary provides an interpretable action layer for all policy families\.

Table 2:Ten semantic catalog domains with representative categories\.The extraction layer is rule based\. It maps action\-month ledger evidence to categories using account\-code movements, vendor or payment\-processor appearance and disappearance, recurring\-payment patterns, double\-entry signatures, and changes relative to a trailing historical baseline\. Detectors use action\-month ledger activity but not the later KPI windows that define the policy outcome\. Definitions are fixed before policy training and shared across KPIs\. The evidence families cover changes in spending, vendors, financing structure, recurring payments, and operating tools, giving the catalog a consistent accounting interpretation across firms\. Historical labels are ledger\-derived business\-change proxies; for categoryaa,Ti,a=1T\_\{i,a\}=1means that the corresponding detector fired in the action month\. The representation is multi\-hot because several categories may be detected together\.

In a separate precision\-oriented audit, a sample of 1,000 positive company\-month/category detections was reviewed against the underlying ledger evidence by a judge LLM calibrated against human labels\. The resulting audited precision was 95%\. This positive\-only audit estimates detection precision and does not estimate recall\. At inference, the same vocabulary is presented as candidate guidance, and every policy is masked to the 34 supported categories before test\-set scoring\. The support rule is outcome\-independent and shared across KPIs\.

### 3\.3KPI Rewards

We train and evaluate a separate policy for Gross Profit, Revenue, and Quick Ratio\. The reward compares the same 11 calendar months before and after the action month, which aligns seasonality while preserving prospective timing\. Letκi,ℓk\\kappa^\{k\}\_\{i,\\ell\}denote the realized value of KPIkkfor observationiiat relative monthℓ\\ell, and letGk​\(κi,a:bk\)G\_\{k\}\(\\kappa^\{k\}\_\{i,a:b\}\)aggregate a window by summation for Gross Profit and Revenue and by averaging for Quick Ratio\. All three KPIs use the normalized post\-versus\-pre change

rik=Gk​\(κi,2:12k\)−Gk​\(κi,−10:0k\)\|Gk​\(κi,−10:0k\)\|\.r\_\{i\}^\{k\}=\\frac\{G\_\{k\}\(\\kappa^\{k\}\_\{i,2:12\}\)\-G\_\{k\}\(\\kappa^\{k\}\_\{i,\-10:0\}\)\}\{\|G\_\{k\}\(\\kappa^\{k\}\_\{i,\-10:0\}\)\|\}\.\(1\)Rows with undefined or numerically unstable denominators are removed before modeling\. Across all filtering steps, the retained Gross Profit, Revenue, and Quick Ratio samples are 91\.8%, 94\.5%, and 93\.8% of the corpus, respectively \(Table[1](https://arxiv.org/html/2608.10050#S3.T1)\)\. The raw change is capped at\|rik\|≤50\|r\_\{i\}^\{k\}\|\\leq 50and winsorized at the training\-set 5th and 95th percentiles\. We denote the resulting modeling reward byRikR\_\{i\}^\{k\}\. Because each KPI is modeled separately, Sections 4–7 suppresskkand writeRiR\_\{i\}\. Scores are interpreted within each KPI because reward aggregation, filtering, and scale differ\.

### 3\.4Focal\-Action Conditional Effect Target

For categorya∈𝒜a\\in\\mathcal\{A\}, letTi,aT\_\{i,a\}indicate whether the category is present in the action month and letRa​\(t\)R^\{a\}\(t\)denote the potential modeling reward under focal exposureTa=tT\_\{a\}=t\. We define the focal\-action conditional average treatment effect as

τa​\(x\)=𝔼​\[Ra​\(1\)−Ra​\(0\)∣X=x\]\.\\tau\_\{a\}\(x\)=\\mathbb\{E\}\\left\[R^\{a\}\(1\)\-R^\{a\}\(0\)\\mid X=x\\right\]\.\(2\)The target intentionally averages over the co\-action environment observed with and without the focal category, matching the policy task of ranking one category from bundled historical changes\. Under consistency, conditional exchangeability given pre\-decision information, and positivity, the corresponding conditional treated\-control contrast identifiesτa​\(x\)\\tau\_\{a\}\(x\)\[[16](https://arxiv.org/html/2608.10050#bib.bib16)\]\. The target motivates selecting the category with the largest estimated focal contrast for the current pre\-decision state\. A deterministic policyπ:𝒳→𝒜\\pi:\\mathcal\{X\}\\rightarrow\\mathcal\{A\}maps each state to one category, and all candidate policies are compared by the shared test\-set score in Section[5](https://arxiv.org/html/2608.10050#S5)\.

## 4Policies

### 4\.1Contextual Conservative Value Model

For this baseline, the training data are expanded to one row for each pair\(i,a\)\(i,a\)withTi,a=1T\_\{i,a\}=1;A∈𝒜A\\in\\mathcal\{A\}denotes the active category attached to an expanded row\. Company\-months with no active supported category contribute no expanded row\. The contextual value model learnsQθ​\(x,a\)Q\_\{\\theta\}\(x,a\)and selects the largest value over the catalog\. Its objective, with expectations over expanded training rows, combines a CQL\-style log\-sum\-exp penalty with regression to the normalized KPI reward:

minθ\\displaystyle\\min\_\{\\theta\}α​𝔼​\[log​∑a′∈𝒜eQθ​\(X,a′\)−Qθ​\(X,A\)\]\\displaystyle\\alpha\\,\\mathbb\{E\}\\\!\\left\[\\log\\sum\_\{a^\{\\prime\}\\in\\mathcal\{A\}\}e^\{Q\_\{\\theta\}\(X,a^\{\\prime\}\)\}\-Q\_\{\\theta\}\(X,A\)\\right\]\(3\)\+12​𝔼​\[\(Qθ​\(X,A\)−R\)2\]\.\\displaystyle\+\\frac\{1\}\{2\}\\mathbb\{E\}\\\!\\left\[\\bigl\(Q\_\{\\theta\}\(X,A\)\-R\\bigr\)^\{2\}\\right\]\.We use a\[512,512,256\]\[512,512,256\]MLP, conservative weightα=0\.01\\alpha=0\.01, 30,000 updates, batch size 256, and learning rate10−410^\{\-4\}withd3rlpy\[[17](https://arxiv.org/html/2608.10050#bib.bib17)\]\. Each expanded row retains the same state and realized reward, exposing every observed supported category to the conservative objective\. Test\-set scoring returns to the original, unexpanded rows\.

### 4\.2Uplift T\-Learner

For each categoryaa, the T\-Learner fitsμ^a,tTL​\(x\)≈𝔼​\[R∣X=x,Ta=t\]\\hat\{\\mu\}^\{\\mathrm\{TL\}\}\_\{a,t\}\(x\)\\approx\\mathbb\{E\}\[R\\mid X=x,T\_\{a\}=t\]fort∈\{0,1\}t\\in\\\{0,1\\\}and ranks their conditional prediction difference,

δ^a​\(x\)=μ^a,1TL​\(x\)−μ^a,0TL​\(x\),πTL​\(x\)=arg⁡maxa∈𝒜⁡δ^a​\(x\)\.\\hat\{\\delta\}\_\{a\}\(x\)=\\hat\{\\mu\}^\{\\mathrm\{TL\}\}\_\{a,1\}\(x\)\-\\hat\{\\mu\}^\{\\mathrm\{TL\}\}\_\{a,0\}\(x\),\\qquad\\pi\_\{\\mathrm\{TL\}\}\(x\)=\\arg\\max\_\{a\\in\\mathcal\{A\}\}\\hat\{\\delta\}\_\{a\}\(x\)\.\(4\)The regressors use only decision\-time features and the normalized KPI reward\. LightGBM models use 500 boosting rounds, depth 6, shared regularization across actions and KPIs, and validation\-set early stopping\[[18](https://arxiv.org/html/2608.10050#bib.bib18)\]\.

### 4\.3Covariate\-Adjusted Residual Policy Learning

CAR\-PL denotes*Covariate\-Adjusted Residual Policy Learning*, an action\-wise R\-learner with global support shrinkage\. Five\-fold company\-grouped cross\-fitting on the training partition estimates a reward nuisancem^​\(x\)=𝔼​\[R∣X=x\]\\hat\{m\}\(x\)=\\mathbb\{E\}\[R\\mid X=x\]and policy\-learning propensitiese^aPL​\(x\)=Pr⁡\(Ta=1∣X=x\)\\hat\{e\}^\{\\mathrm\{PL\}\}\_\{a\}\(x\)=\\Pr\(T\_\{a\}=1\\mid X=x\)\. We form

R~i=Ri−m^​\(Xi\),T~i,a=Ti,a−e^aPL​\(Xi\),\\widetilde\{R\}\_\{i\}=R\_\{i\}\-\\hat\{m\}\(X\_\{i\}\),\\qquad\\widetilde\{T\}\_\{i,a\}=T\_\{i,a\}\-\\hat\{e\}^\{\\mathrm\{PL\}\}\_\{a\}\(X\_\{i\}\),\(5\)and fit one effect head per category by

δ^a​\(⋅\)=arg⁡ming∈𝒢​∑i\(R~i−T~i,a​g​\(Xi\)\)2\.\\hat\{\\delta\}\_\{a\}\(\\cdot\)=\\arg\\min\_\{g\\in\\mathcal\{G\}\}\\sum\_\{i\}\\left\(\\widetilde\{R\}\_\{i\}\-\\widetilde\{T\}\_\{i,a\}g\(X\_\{i\}\)\\right\)^\{2\}\.\(6\)Here𝒢\\mathcal\{G\}is the class of depth\-4 gradient\-boosted regression heads\. We implement Eq\. \([6](https://arxiv.org/html/2608.10050#S4.E6)\) as weighted regression: the pseudo\-outcome isR~i/T~i,a\\widetilde\{R\}\_\{i\}/\\widetilde\{T\}\_\{i,a\}, and the weight isT~i,a2\\widetilde\{T\}\_\{i,a\}^\{2\}\. Rows with weight below10−410^\{\-4\}are omitted; each head uses 150 rounds\.

Training proceeds in five stages\. We first assign firms to five nuisance folds and fit the reward and propensity models on the complementary firms\. We then generate out\-of\-fold residuals for every training row, fit one weighted effect head per category, and evaluate candidate head settings on the company\-disjoint validation partition\. Finally, the selected support multiplier is applied uniformly across categories before the policy argmax\. The action heads share the nuisance residuals but are fit independently and can be trained in parallel, making the method practical for a catalog with dozens of candidate categories\.

Maximizing over 34 noisy heads can favor low\-frequency categories by chance\. Letna=∑i∈𝒟trainTi,an\_\{a\}=\\sum\_\{i\\in\\mathcal\{D\}\_\{\\mathrm\{train\}\}\}T\_\{i,a\}be categoryaa’s training support\. Validation considersλsup∈\{0,1000,2000,4000\}\\lambda\_\{\\mathrm\{sup\}\}\\in\\\{0,1000,2000,4000\\\}and yieldsλsup=2000\\lambda\_\{\\mathrm\{sup\}\}=2000for all three KPI policies\. CAR\-PL then applies the support multiplier

δ^ashr​\(x\)\\displaystyle\\hat\{\\delta\}\_\{a\}^\{\\mathrm\{shr\}\}\(x\)=δ^a​\(x\)​nana\+λsup,λsup=2000,\\displaystyle=\\hat\{\\delta\}\_\{a\}\(x\)\\frac\{n\_\{a\}\}\{n\_\{a\}\+\\lambda\_\{\\mathrm\{sup\}\}\},\\quad\\lambda\_\{\\mathrm\{sup\}\}=2000,\(7\)πCAR​\(x\)\\displaystyle\\pi\_\{\\mathrm\{CAR\}\}\(x\)=arg⁡maxa⁡δ^ashr​\(x\)\.\\displaystyle=\\arg\\max\_\{a\}\\hat\{\\delta\}\_\{a\}^\{\\mathrm\{shr\}\}\(x\)\.\(8\)The multiplier favors better\-supported categories in the many\-category argmax while retaining state\-specific variation within each head\. CAR\-PL operates on the original, unexpanded multi\-hot rows; within each action head, rows below the weight threshold are omitted\. Because nuisance predictions are out of fold and folds are grouped by company, each effect head is trained on residualized signals from firms not used to fit its nuisance values\.

### 4\.4Reference Policies

The zero\-shot policy usesanthropic\.claude\-opus\-4\-6\-v1\. It receives the full, untruncated pre\-decision narrative \(median approximately 23,000 characters\), the target KPI, and the 34 catalog names\. Its fixed template asks for the single category most likely to improve that KPI and requires the exact category name in JSON\. We generate one completion per state with reasoning effort set tohigh, and every reported output matches an eligible catalog name\.

The modal constant policy recommends the most frequent eligible training category for every state\. It provides a simple non\-personalized comparator\. The random reference samples categoryaaonce per test row with probability proportional to its training frequency, renormalized over the same catalog\. A fixed random seed is used, and the resulting test\-row assignments are held constant across all company\-clustered bootstrap replicates\.

Table 3:Policy and scoring\-model implementation summary\. Model and preprocessing choices are selected with development data and frozen before test\-set scoring\.

## 5Shared Model\-Assisted Policy Comparison

### 5\.1Shared Nuisance Models and Reference Contrast

All policies are compared using the same frozen scorer\. Its propensity model uses the policy\-state representationXiX\_\{i\}, while its outcome model usesWiW\_\{i\}, a structured encoding of KPI histories and accounting covariates from the same pre\-decision records\. Both representations exclude action\-month and post\-action information\. A shared\-trunk MLP estimates scoring propensitiese^asc​\(x\)=Pr⁡\(Ta=1∣X=x\)\\hat\{e\}^\{\\mathrm\{sc\}\}\_\{a\}\(x\)=\\Pr\(T\_\{a\}=1\\mid X=x\)\. The model is fit with company\-grouped folds on the training partition; validation data determine early stopping and the frozen configuration\. The test\-set scoring rule usese~asc​\(x\)=clip⁡\(e^asc​\(x\),0\.05,0\.95\)\\widetilde\{e\}^\{\\mathrm\{sc\}\}\_\{a\}\(x\)=\\operatorname\{clip\}\(\\hat\{e\}^\{\\mathrm\{sc\}\}\_\{a\}\(x\),0\.05,0\.95\)\.

A neural outcome model produces 11 post\-action monthly KPI predictions from pre\-decision covariatesWWand multi\-hot treatment vector𝐭\\mathbf\{t\}\. The implemented fixed KPI\-specific map aggregates those outputs with the pre\-period baseline inWWand applies the reward construction in Eq\. \([1](https://arxiv.org/html/2608.10050#S3.E1)\);μ^​\(w,𝐭\)\\hat\{\\mu\}\(w,\\mathbf\{t\}\)denotes the resulting scalar reward prediction\. The network uses Huber loss, AdamW\[[19](https://arxiv.org/html/2608.10050#bib.bib19)\], and validation\-set early stopping\. For focal categoryaa, we evaluate it at the one\-hot vector𝐮a\\mathbf\{u\}\_\{a\}and at the zero vector:

m^a,1ref​\(w\)=μ^​\(w,𝐮a\),m^a,0ref​\(w\)=μ^​\(w,𝟎\)\.\\hat\{m\}^\{\\mathrm\{ref\}\}\_\{a,1\}\(w\)=\\hat\{\\mu\}\(w,\\mathbf\{u\}\_\{a\}\),\\qquad\\hat\{m\}^\{\\mathrm\{ref\}\}\_\{a,0\}\(w\)=\\hat\{\\mu\}\(w,\\mathbf\{0\}\)\.\(9\)These standardized probes place all categories on one model\-based reference scale and allow a shared scoring model to compare policies that select different categories\.

### 5\.2Bounded Augmented Policy Score

Letclip⁡\(z,ℓ,u\)\\operatorname\{clip\}\(z,\\ell,u\)truncatezzto\[ℓ,u\]\[\\ell,u\]\. For test observationiiand focal categoryaa, define

s^i​\(a\)=\\displaystyle\\hat\{s\}\_\{i\}\(a\)=m^a,1ref​\(Wi\)−m^a,0ref​\(Wi\)\\displaystyle\\;\\hat\{m\}^\{\\mathrm\{ref\}\}\_\{a,1\}\(W\_\{i\}\)\-\\hat\{m\}^\{\\mathrm\{ref\}\}\_\{a,0\}\(W\_\{i\}\)\+clip\(Ti,ae~asc​\(Xi\)\[Ri−m^a,1ref\(Wi\)\]\\displaystyle\+\\operatorname\{clip\}\\Bigg\(\\frac\{T\_\{i,a\}\}\{\\widetilde\{e\}^\{\\mathrm\{sc\}\}\_\{a\}\(X\_\{i\}\)\}\[R\_\{i\}\-\\hat\{m\}^\{\\mathrm\{ref\}\}\_\{a,1\}\(W\_\{i\}\)\]−1−Ti,a1−e~asc​\(Xi\)\[Ri−m^a,0ref\(Wi\)\],−2,2\)\.\\displaystyle\\qquad\-\\frac\{1\-T\_\{i,a\}\}\{1\-\\widetilde\{e\}^\{\\mathrm\{sc\}\}\_\{a\}\(X\_\{i\}\)\}\[R\_\{i\}\-\\hat\{m\}^\{\\mathrm\{ref\}\}\_\{a,0\}\(W\_\{i\}\)\],\-2,2\\Bigg\)\.\(10\)The first term is a shared outcome\-model reference contrast, and the bounded residual term incorporates observed focal\-exposure evidence while limiting the influence of extreme weights and outcomes\. The construction draws on augmented inverse\-propensity scoring used in causal and off\-policy evaluation\[[20](https://arxiv.org/html/2608.10050#bib.bib20),[21](https://arxiv.org/html/2608.10050#bib.bib21),[22](https://arxiv.org/html/2608.10050#bib.bib22)\], with the same clipping, reward transformation, and correction bound applied to every policy\. Because the outcome\-model probes use standardized treatment references rather than each row’s factual co\-action bundle, we interpret Eq\. \([10](https://arxiv.org/html/2608.10050#S5.E10)\) as a common model\-assisted score for matched within\-KPI policy comparisons, rather than as an estimate of deployed policy value\.

The shared scorer is deliberately policy\-agnostic\. For each test state, it computes action\-specific quantities from the same fitted nuisance models and then extracts only the score of the category selected by the candidate policy\. Differences between policies therefore arise from their state\-to\-category mappings rather than from policy\-specific scoring models, coverage filters, or outcome transformations\.

For the current KPI, let𝒟test\\mathcal\{D\}\_\{\\mathrm\{test\}\}be its retained test rows andN=\|𝒟test\|N=\|\\mathcal\{D\}\_\{\\mathrm\{test\}\}\|\. For any realized one\-category mappingπ\\pi, the reported metric is the mean selected\-category score,

S^MA​\(π\)=1N​∑i∈𝒟tests^i​\(π​\(Xi\)\)\.\\widehat\{S\}\_\{\\mathrm\{MA\}\}\(\\pi\)=\\frac\{1\}\{N\}\\sum\_\{i\\in\\mathcal\{D\}\_\{\\mathrm\{test\}\}\}\\hat\{s\}\_\{i\}\\\!\\left\(\\pi\(X\_\{i\}\)\\right\)\.\(11\)

### 5\.3Matched Comparisons and Inference

Every policy is constrained to𝒜\\mathcal\{A\}before inference\. Value\-model heads outside the catalog are masked, the T\-Learner and CAR\-PL take their argmax only over𝒜\\mathcal\{A\}, the LLM prompt contains only catalog names, and the random reference renormalizes its probabilities over the same set\. Consequently, each policy emits one scorable category for every valid test row; no policy\-specific recommendation is dropped after prediction\.

All model selection, early stopping, nuisance fitting, action\-set construction, reward transformation, and scoring constants use the training and validation partitions\. The complete pipeline is frozen before test\-set evaluation\. Because every policy is scored on the same company\-months with the same frozen nuisance models, matched policy differences are the primary comparative quantity\. For two policiesπ1\\pi\_\{1\}andπ2\\pi\_\{2\}, the matched contrast is

Δ^​\(π1,π2\)=1N​∑i∈𝒟test\[s^i​\(π1​\(Xi\)\)−s^i​\(π2​\(Xi\)\)\]\.\\widehat\{\\Delta\}\(\\pi\_\{1\},\\pi\_\{2\}\)=\\frac\{1\}\{N\}\\sum\_\{i\\in\\mathcal\{D\}\_\{\\mathrm\{test\}\}\}\\left\[\\hat\{s\}\_\{i\}\(\\pi\_\{1\}\(X\_\{i\}\)\)\-\\hat\{s\}\_\{i\}\(\\pi\_\{2\}\(X\_\{i\}\)\)\\right\]\.\(12\)We form 95% confidence intervals with 1,000 bootstrap replicates that resample held\-out test firms and retain all of each sampled firm’s months\. Firm\-level resampling preserves within\-firm dependence, including dependence induced by overlapping monthly outcome windows\. For the random reference, the original seeded test\-row assignments remain fixed while firms are resampled\. The same clustered bootstrap is applied to Eq\. \([11](https://arxiv.org/html/2608.10050#S5.E11)\) and Eq\. \([12](https://arxiv.org/html/2608.10050#S5.E12)\)\.

## 6Results

Table[4](https://arxiv.org/html/2608.10050#S6.T4)reports the shared test\-set comparison\. CAR\-PL and the T\-Learner have the two highest point estimates for both growth KPIs, while the contextual value model ranks highest on Quick Ratio\. The modal constant has a higher point estimate than the random reference but remains below the highest\-scoring learned policy on every KPI\.

Table 4:Model\-assisted focal\-action scores on company\-disjoint held\-out test firms\. Cells show mean \[95% company\-clustered bootstrap interval\]\. Bold marks the highest point estimate\. Scores are comparable within, not across, KPIs\.### 6\.1KPI\-Specific Rankings

On Gross Profit, CAR\-PL has the highest point estimate \(0\.084\), followed by the T\-Learner \(0\.074\)\. Their paired difference is 0\.010 with a 95% interval of\[−0\.008,0\.028\]\[\-0\.008,0\.028\], so the two policies are not statistically separated\. CAR\-PL has positive paired margins over the contextual value model \(0\.037,\[0\.016,0\.056\]\[0\.016,0\.056\]\) and random reference \(0\.063,\[0\.041,0\.085\]\[0\.041,0\.085\]\)\. Revenue yields the same top pair in reverse order: the T\-Learner scores 0\.085 and CAR\-PL 0\.079, with a paired difference of−0\.006\-0\.006and interval\[−0\.025,0\.012\]\[\-0\.025,0\.012\]; these policies are again not statistically separated\. CAR\-PL exceeds the contextual value model and random reference on Revenue\. On Quick Ratio, the contextual value model has the highest point estimate \(0\.062\) and positive paired margins over the T\-Learner and random reference\. Thus, the matched comparisons do not separate CAR\-PL and the T\-Learner on either growth KPI; on Quick Ratio, the contextual value model has the highest point estimate and positive reported margins over the T\-Learner and random reference\.

Table 5:Selected matched policy\-score differences\. Intervals use the company\-clustered bootstrap\. Differences use unrounded row\-level scores and may differ slightly from differences between rounded means in Table[4](https://arxiv.org/html/2608.10050#S6.T4)\.
### 6\.2Scoring\-Model Stability and Overlap

Table[6](https://arxiv.org/html/2608.10050#S6.T6)summarizes propensity calibration, overlap, and factual outcome\-model diagnostics\. Across the 34 categories, macro AUPRC is 0\.423/0\.428 \(training/test\) and macro AUROC is 0\.819/0\.819\. Expected calibration error \(ECE\) is 0\.036\. The treated\-side clip rate is 4\.8%, median inverse\-weight effective sample size \(ESS\) is 58% of nominal treated support and 93% of nominal control support, and outcome\-model rank correlations are positive on held\-out test firms for every KPI\.

Table 6:Nuisance\-model and overlap diagnostics\. Propensity AUROC and AUPRC are macro\-averaged across the 34 categories\. Where three KPI\-specific values appear, they are ordered Gross Profit / Revenue / Quick Ratio; other pairs are identified in the row label\. Outcome\-model diagnostics use held\-out test rows\.Table[7](https://arxiv.org/html/2608.10050#S6.T7)gives the fraction of test rows whose residual correction reaches±2\\pm 2\. Saturation is 2\.6–5\.5% on the growth KPIs and 6\.6–12\.9% on Quick Ratio, so the correction bound is inactive on most rows\.

Table 7:Residual\-correction saturation rate at the headline score settings\.Table[8](https://arxiv.org/html/2608.10050#S6.T8)reports the Gross Profit CAR\-PL grid\. Propensity bounds are 0\.01, 0\.05, and 0\.10; correction bounds are 1, 2, 5, and unbounded\. Every score remains positive, ranging from 0\.083 to 0\.118\. On Revenue, CAR\-PL ranges from 0\.077 to 0\.095, and the contextual value model’s Quick Ratio score remains between 0\.06 and 0\.10\. We additionally report an outcome\-model\-only score that omits the residual correction in Eq\. \([10](https://arxiv.org/html/2608.10050#S5.E10)\)\. Separately, we replace the all\-zero treatment reference in Eq\. \([9](https://arxiv.org/html/2608.10050#S5.E9)\) with the modal training co\-action vector, defined as the most common treatment configuration in the training data\. Table[9](https://arxiv.org/html/2608.10050#S6.T9)shows that the outcome\-model\-only analysis retains the same KPI\-level point\-estimate leader or top pair\. Category\-contrast rankings under the two treatment references are strongly correlated \(ρ=0\.884\\rho=0\.884–0\.9710\.971\), and the top\-ranked category is unchanged for every KPI\.

Table 8:Gross Profit CAR\-PL score under propensity and residual\-correction bounds\. The headline setting is\(0\.05,2\)\(0\.05,2\)\.Table 9:Robustness to score component and treatment reference\. Outcome\-model\-only leader\(s\) use only the reference contrast in Eq\. \([10](https://arxiv.org/html/2608.10050#S5.E10)\)\. Zero/modalρ\\rhois the Spearman correlation across 34 category contrasts under the all\-zero and modal co\-action treatment references\.
### 6\.3Policy Behavior and Constant References

Table[10](https://arxiv.org/html/2608.10050#S6.T10)describes test\-set selection concentration\. CAR\-PL uses 33–34 categories with top\-category shares below 19% on all three KPIs\. The T\-Learner is more concentrated on the two growth KPIs, while the zero\-shot LLM assigns 84\.2% of its Revenue recommendations to one broad customer\-growth category\. CAR\-PL and the T\-Learner choose the same category on only 12\.8% of Gross Profit states, 14\.2% of Revenue states, and 10\.7% of Quick Ratio states\. This broader catalog use indicates that CAR\-PL does not obtain its score by collapsing onto a small set of generic categories and produces broader, less concentrated selections across the action catalog\.

Table 10:Test\-set action\-selection concentration\. Entropy is−∑apa​log2⁡pa\-\\sum\_\{a\}p\_\{a\}\\log\_\{2\}p\_\{a\}bits for empirical selection sharespap\_\{a\}\.The modal constant provides a simple non\-personalized comparator\. Its scores are 0\.038 on Gross Profit, 0\.038 on Revenue, and 0\.043 on Quick Ratio, below the highest\-scoring learned policy on every KPI\. Together with the concentration results, this shows that the highest\-scoring learned mappings use a substantially more diverse set of categories than the one\-category modal rule\.

## 7Discussion

### 7\.1KPI\-Dependent Results

Quick Ratio presents a noisier residualized learning setting than the growth KPIs: its outcome\-residual IQR is 0\.779 versus 0\.382 for Revenue, and its CAR\-PL correction saturation is 10\.8% versus 2\.6%\. This pattern may partly explain why the contextual value model has the highest Quick Ratio point estimate\.

On Gross Profit and Revenue, CAR\-PL and the T\-Learner are not statistically separated in matched comparisons while agreeing on only 12\.8% and 14\.2% of states, respectively\. Comparable aggregate scores can therefore arise from substantially different policy mappings, making recommendation coverage and concentration important complements to the mean policy score\. Matched comparisons further show that CAR\-PL exceeds the contextual value model on the growth KPIs, whereas the contextual value model exceeds the T\-Learner and random reference on Quick Ratio\.

### 7\.2CAR\-PL in the Comparison

CAR\-PL’s support\-shrunk action\-wise residual estimates directly address instability from selecting among 34 unevenly supported categories\. On the growth KPIs, CAR\-PL is not statistically separated from the T\-Learner while selecting 33–34 categories and maintaining top\-category shares below 19%\. It therefore combines competitive growth\-KPI performance with broader recommendation coverage across heterogeneous firm states and business domains\.

### 7\.3Learning from Multi\-Action Accounting Logs

The multi\-hot structure is central to the study\. A company may change spending, financing, vendors, and operating tools in the same month, so the learning methods must extract a useful focal signal from repeated bundled observations\. The contextual value model converts each active category into a discrete training row\. The T\-Learner and CAR\-PL retain the original rows and fit one focal category at a time, with CAR\-PL additionally residualizing both outcome and category occurrence against the decision state\. These constructions provide complementary ways to exploit the same logs\.

The treatment\-reference analysis provides a complementary check on learning from bundled logs\. Per\-category outcome\-model contrasts remain strongly rank\-aligned when the all\-zero treatment reference is replaced by the modal training co\-action vector, and the highest\-ranked category is unchanged for every reported KPI\. Together with the outcome\-model\-only comparison, this indicates that the headline KPI\-level point\-estimate leader or top pair is not specific to the residual correction or to one treatment reference\.

### 7\.4Reference Policies and Practical Interpretation

The zero\-shot LLM is competitive on Revenue, but its 84\.2% concentration on one category indicates that generic business priors produce a much narrower policy than outcome\-fitted models\. The comparison distinguishes broad pretrained knowledge from state\-specific policy learning\. The modal constant and frequency\-weighted random references are both below the highest\-scoring learned policy on every KPI, a pattern consistent with outcome fitting adding useful information beyond the historical marginal action distribution\. The single\-category output also fits a human\-review workflow: a policy surfaces one prioritized business\-change category, while an advisor can translate it into a firm\-specific implementation plan\.

### 7\.5Evaluation Design

Four design choices support the comparison\. Every method shares the 34\-category action space and frozen scorer\. Development and testing are company\-disjoint, pairwise claims use matched differences on the same company\-months, and separate policies are learned for each financial objective\. This isolates state\-to\-category mappings from coverage and scorer changes while respecting the different scales and business meanings of Gross Profit, Revenue, and Quick Ratio\. The accounting records and exact detector thresholds are proprietary; the paper therefore reports the partitioning design, prompt configuration, model architectures, hyperparameters, score constants, and aggregate diagnostics to support reimplementation on an equivalent corpus\.

### 7\.6Limitations

The study uses observational accounting data and therefore relies on measured pre\-decision covariates, overlap, and the quality of ledger\-derived category labels\. Co\-occurring business changes are represented through focal categories rather than explicit bundle policies, and the model\-assisted score depends on fitted nuisance models and standardized treatment references\. The company\-disjoint test assesses generalization within the observed environment rather than the effect of deployed recommendations\. Prospective validation is the next step toward operational deployment\.

## 8Conclusion

We presented an observational policy\-ranking framework for SMB financial guidance from multi\-action accounting logs\. Across 85,078 company\-month observations, CAR\-PL and the T\-Learner have the two highest point estimates and are not statistically separated on Gross Profit or Revenue, while the contextual value model has the highest Quick Ratio point estimate\. CAR\-PL combines competitive growth\-KPI performance with broader recommendation coverage, selecting 33–34 categories across held\-out firms\. Outcome\-model\-only scoring retains the same KPI\-level point\-estimate leader or top pair, and category rankings remain similar when the all\-zero treatment reference is replaced by the modal training co\-action vector\. Overall, the findings support objective\-specific comparison of learning strategies and position CAR\-PL as a practical support\-aware method for ranking guidance over large observational action catalogs\.

## References

- Bruhn et al\. \[2018\]Miriam Bruhn, Dean Karlan, and Antoinette Schoar\.The impact of consulting services on small and medium enterprises: Evidence from a randomized trial in mexico\.*Journal of Political Economy*, 126\(2\):635–687, 2018\.doi:10\.1086/696154\.
- Athey and Wager \[2021\]Susan Athey and Stefan Wager\.Policy learning with observational data\.*Econometrica*, 89\(1\):133–161, 2021\.doi:10\.3982/ECTA15732\.
- Zhou et al\. \[2023\]Zhengyuan Zhou, Susan Athey, and Stefan Wager\.Offline multi\-action policy learning: Generalization and optimization\.*Operations Research*, 71\(1\):148–183, 2023\.doi:10\.1287/opre\.2022\.2271\.
- VanderWeele and Hernán \[2013\]Tyler J\. VanderWeele and Miguel A\. Hernán\.Causal inference under multiple versions of treatment\.*Journal of Causal Inference*, 1\(1\):1–20, 2013\.doi:10\.1515/jci\-2012\-0002\.
- Künzel et al\. \[2019\]Sören R\. Künzel, Jasjeet S\. Sekhon, Peter J\. Bickel, and Bin Yu\.Metalearners for estimating heterogeneous treatment effects using machine learning\.*Proceedings of the National Academy of Sciences*, 116\(10\):4156–4165, 2019\.doi:10\.1073/pnas\.1804597116\.
- Nie and Wager \[2021\]Xinkun Nie and Stefan Wager\.Quasi\-oracle estimation of heterogeneous treatment effects\.*Biometrika*, 108\(2\):299–319, 2021\.doi:10\.1093/biomet/asaa076\.
- Jin et al\. \[2025\]Ying Jin, Zhimei Ren, Zhuoran Yang, and Zhaoran Wang\.Policy learning “without” overlap: Pessimism and generalized empirical bernstein’s inequality\.*The Annals of Statistics*, 53\(4\):1483–1512, 2025\.doi:10\.1214/25\-AOS2511\.
- Fujimoto et al\. \[2019\]Scott Fujimoto, David Meger, and Doina Precup\.Off\-policy deep reinforcement learning without exploration\.In*Proceedings of the 36th International Conference on Machine Learning*, volume 97 of*Proceedings of Machine Learning Research*, pages 2052–2062, 2019\.
- Levine et al\. \[2020\]Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu\.Offline reinforcement learning: Tutorial, review, and perspectives on open problems\.*arXiv preprint arXiv:2005\.01643*, 2020\.
- Kumar et al\. \[2020\]Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine\.Conservative Q\-learning for offline reinforcement learning\.In*Advances in Neural Information Processing Systems*, volume 33, pages 1179–1191, 2020\.
- Khandani et al\. \[2010\]Amir E\. Khandani, Adlar J\. Kim, and Andrew W\. Lo\.Consumer credit\-risk models via machine\-learning algorithms\.*Journal of Banking & Finance*, 34\(11\):2767–2787, 2010\.doi:10\.1016/j\.jbankfin\.2010\.06\.001\.
- Fuster et al\. \[2022\]Andreas Fuster, Paul Goldsmith\-Pinkham, Tarun Ramadorai, and Ansgar Walther\.Predictably unequal? the effects of machine learning on credit markets\.*The Journal of Finance*, 77\(1\):5–47, 2022\.doi:10\.1111/jofi\.13090\.
- Dai et al\. \[2023\]Sunhao Dai, Ninglu Shao, Haiyuan Zhao, Weijie Yu, Zihua Si, Chen Xu, Zhongxiang Sun, Xiao Zhang, and Jun Xu\.Uncovering ChatGPT’s capabilities in recommender systems\.*arXiv preprint arXiv:2305\.02182*, 2023\.
- Fieberg et al\. \[2025\]Christian Fieberg, Lars Hornuf, Maximilian Meiler, and David J\. Streich\.Using large language models for financial advice\.CESifo Working Paper 11666, CESifo, 2025\.
- Zhang et al\. \[2025\]Yanzhao Zhang, Mingxin Li, Dingkun Long, Xin Zhang, Huan Lin, Baosong Yang, Pengjun Xie, An Yang, Dayiheng Liu, Junyang Lin, Fei Huang, and Jingren Zhou\.Qwen3 Embedding: Advancing text embedding and reranking through foundation models\.*arXiv preprint arXiv:2506\.05176*, 2025\.
- Rosenbaum and Rubin \[1983\]Paul R\. Rosenbaum and Donald B\. Rubin\.The central role of the propensity score in observational studies for causal effects\.*Biometrika*, 70\(1\):41–55, 1983\.doi:10\.1093/biomet/70\.1\.41\.
- Seno and Imai \[2022\]Takuma Seno and Michita Imai\.d3rlpy: An offline deep reinforcement learning library\.*Journal of Machine Learning Research*, 23\(315\):1–20, 2022\.
- Ke et al\. \[2017\]Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei Ye, and Tie\-Yan Liu\.LightGBM: A highly efficient gradient boosting decision tree\.In*Advances in Neural Information Processing Systems*, volume 30, pages 3146–3154, 2017\.
- Loshchilov and Hutter \[2019\]Ilya Loshchilov and Frank Hutter\.Decoupled weight decay regularization\.In*International Conference on Learning Representations*, 2019\.
- Robins et al\. \[1994\]James M\. Robins, Andrea Rotnitzky, and Lue Ping Zhao\.Estimation of regression coefficients when some regressors are not always observed\.*Journal of the American Statistical Association*, 89\(427\):846–866, 1994\.doi:10\.1080/01621459\.1994\.10476818\.
- Dudík et al\. \[2011\]Miroslav Dudík, John Langford, and Lihong Li\.Doubly robust policy evaluation and learning\.In*Proceedings of the 28th International Conference on Machine Learning*, pages 1097–1104, 2011\.
- Thomas and Brunskill \[2016\]Philip S\. Thomas and Emma Brunskill\.Data\-efficient off\-policy policy evaluation for reinforcement learning\.In*Proceedings of the 33rd International Conference on Machine Learning*, volume 48 of*Proceedings of Machine Learning Research*, pages 2139–2148, 2016\.

Similar Articles

APPO: Agentic Procedural Policy Optimization

Hugging Face Daily Papers

APPO improves multi-turn tool-use in LLM agents by refining branching decisions and credit assignment using fine-grained decision points and procedure-level advantage scaling, outperforming baselines by 4 points on 13 benchmarks.