From Structural Equation Modelling to Double Machine Learning: Robustness Analysis for Survey-Based Research
Summary
This paper develops a staged robustness analysis framework that connects Structural Equation Modelling (SEM), Ordinary Least Squares (OLS), and Double Machine Learning (DML) for survey-based latent-construct research, demonstrated on a FinTech Digital Customer Intimacy survey model. The framework provides a reusable template for researchers to assess stability of findings across different estimation methods.
View Cached Full Text
Cached at: 07/02/26, 05:38 AM
# From Structural Equation Modelling to Double Machine Learning: Robustness Analysis for Survey-Based Research
Source: [https://arxiv.org/html/2607.00512](https://arxiv.org/html/2607.00512)
###### Abstract
Structural equation modelling \(SEM\) is widely used in survey\-based business and information systems research to assess latent constructs and theory\-driven structural relationships\. However, SEM path significance is obtained within a particular model specification and may not show whether findings remain stable under alternative estimation frameworks\. This study develops and demonstrates a staged robustness analysis framework that connects SEM, ordinary least squares \(OLS\) regression, and Double Machine Learning \(DML\)\. SEM is first used to refine the measurement structure and estimate the robustness\-baseline SEM model, in which the full theory\-specified structural path system is retained for downstream robustness analysis before final structural path evaluation\. OLS regression is then applied to SEM\-derived construct scores as a transparent regression benchmark\. Finally, DML\-style residualisation is used to examine whether each tested focal relationship remains stable after flexible machine\-learning\-based adjustment for observed controls\. Learner\-sensitivity checks compare Random Forest, Gradient Boosting, and Support Vector Machine learners, and selected reverse\-direction diagnostics are used to examine directional sensitivity\. The framework is demonstrated using a FinTech Digital Customer Intimacy survey model\. The findings identify which relationships are stable across SEM, OLS, and DML\-style checks, and which require more cautious interpretation\. A reproducible Google Colab workbook and generated result files are publicly available, providing a reusable template that researchers and students can adapt to other survey\-based latent\-construct studies\. The paper contributes a practical robustness workflow and interpretation guide for survey\-based researchers seeking to complement SEM with conventional and machine\-learning\-based robustness checks\.
###### keywords:
Structural Equation Modelling , Double Machine Learning , OLS , Robustness Analysis , Survey Research , Information Systems , FinTech
††journal:Journal to be confirmed\\affiliation
\[UniSQ\]organization=School of Business, Law, Humanities and Pathways, addressline=University of Southern Queensland, city=Springfield, postcode=4300, state=Queensland, country=Australia
## 1Introduction
Survey\-based research frequently relies on latent constructs such as trust, satisfaction, attitude, intention, perceived quality, and customer intimacy\. These constructs are usually measured through multiple survey items rather than observed directly\. Structural equation modelling \(SEM\) is therefore well\-suited to this type of research because it enables researchers to assess the measurement model and estimate theory\-driven structural relationships within a single framework\. However, a significant SEM path coefficient remains a result within a particular model specification\. Researchers may therefore benefit from supplementary robustness analyses that examine whether key relationships remain stable under alternative score representations and estimation assumptions before final model decisions are made\.
This study develops a staged SEM–OLS–DML robustness framework for survey\-based latent\-construct research\. The framework is organised around two analytical layers\. The first layer generates the empirical evidence: the measurement model is refined, the full theory\-specified structural model is re\-estimated as the robustness\-baseline SEM model, construct scores are constructed, and OLS and DML\-style robustness checks are run for all SEM paths before final structural path evaluation; SEM factor scores and mean composite scores are constructed, and OLS and DML\-style robustness checks are run for all tested SEM paths\. The second layer compares and interprets the generated outputs across methods, score representations, nuisance learners, and, where theoretically useful, selected reverse\-direction diagnostic checks\. This two\-layer design separates the production of robustness evidence from the interpretation of convergence, divergence, and sensitivity\.
The paper is motivated by an emerging methodological gap\. Recent studies have introduced Double Machine Learning \(DML\) to information systems research\(Shi et al\.,[2025](https://arxiv.org/html/2607.00512#bib.bib22)\), combined PLS\-SEM with machine\-learning algorithms for predictive and exploratory purposes\(Richter and Tudoran,[2024](https://arxiv.org/html/2607.00512#bib.bib21)\), and applied DML directly in FinTech\-related empirical settings\(Wu et al\.,[2024](https://arxiv.org/html/2607.00512#bib.bib25)\)\. These studies show that machine learning and DML are increasingly relevant to empirical research\. Less attention, however, has been given to using DML as a post\-SEM, path\-by\-path robustness tool for survey\-based latent\-construct models\. This paper addresses that gap by positioning SEM as the primary measurement and theory\-testing framework, OLS as a transparent regression benchmark, and DML as a flexible\-control robustness check for SEM\-implied relationships\.
The study is guided by the following research questions:
- •RQ1:How can SEM, OLS, and DML\-style analysis be integrated into a staged robustness framework for survey\-based latent\-construct research?
- •RQ2:How can a robustness\-baseline SEM model be translated into path\-level OLS and DML\-style specifications before final structural path evaluation?
- •RQ3:How do OLS and DML\-style estimates complement SEM when evaluating the robustness of initially tested structural paths?
- •RQ4:How can researchers interpret convergence, divergence, score sensitivity, learner sensitivity, and selected reverse\-direction diagnostics within the proposed robustness framework?
- •RQ5:In the FinTech DCI demonstration case, which initially tested structural paths remain robust across SEM, OLS, and DML\-style checks, and which paths require further theory, measurement, or directional\-diagnostic discussion?
The paper makes three contributions\. First, it provides a practical workflow for translating SEM\-implied structural paths into OLS and DML\-style robustness checks without treating machine learning as a replacement for SEM\. Second, it clarifies how score\-representation sensitivity and learner sensitivity can be interpreted in survey\-based latent\-construct research\. Third, it demonstrates the workflow using a FinTech Digital Customer Intimacy model in which robustness evidence is generated before final structural paths are evaluated and decisions made, and will provide a Google Colab workbook and generated CSV outputs through a Zenodo archive, allowing researchers and students to adapt the workflow to other survey\-based SEM applications\.
## 2Literature Background and Methodological Positioning
### 2\.1SEM and latent\-construct survey research
Survey\-based studies commonly use multi\-item instruments to measure constructs that cannot be directly observed\. In this setting, SEM is not merely a regression model with named variables\. Its distinctive value lies in modelling latent constructs from multiple observed indicators, assessing measurement quality, and estimating a theory\-specified system of structural paths\(Bollen,[1989](https://arxiv.org/html/2607.00512#bib.bib5); Anderson and Gerbing,[1988](https://arxiv.org/html/2607.00512#bib.bib1); MacKenzie et al\.,[2011](https://arxiv.org/html/2607.00512#bib.bib17)\)\. This measurement role is particularly important in behavioural, business, and information systems research, where construct validity is central to credible theory testing\. The two\-step logic of first assessing the measurement model and then estimating the structural model provides the foundation for using SEM as the primary modelling stage rather than as a preliminary score\-generation tool\(Anderson and Gerbing,[1988](https://arxiv.org/html/2607.00512#bib.bib1)\)\.
Regression\-based robustness checks can still play a useful supplementary role after SEM\. SEM and regression have long been discussed as complementary rather than mutually exclusive approaches in applied information systems research\(Gefen et al\.,[2000](https://arxiv.org/html/2607.00512#bib.bib13)\)\. In the present framework, OLS is used as a transparent benchmark that asks whether SEM\-implied relationships remain visible when latent constructs are represented as respondent\-level construct scores\. This benchmark does not replace SEM; instead, it provides a simpler score\-based comparison for examining whether the structural pattern is visible under conventional linear regression assumptions\.
### 2\.2DML, SEM–ML integration, and adjacent applied research
DML has become influential because it allows a low\-dimensional focal relationship to be estimated while flexible machine\-learning methods approximate nuisance functions for observed controls and co\-predictors\(Chernozhukov et al\.,[2018](https://arxiv.org/html/2607.00512#bib.bib8)\)\. The DML software framework further supports reproducible implementation of this logic in Python\(Bach et al\.,[2022](https://arxiv.org/html/2607.00512#bib.bib2)\)\. Recent information systems work has introduced DML as a flexible semiparametric method for empirical model specification and causal\-inference\-oriented research designs\(Shi et al\.,[2025](https://arxiv.org/html/2607.00512#bib.bib22)\)\. In applied FinTech and sustainability research, DML is often used directly as the main empirical strategy\. For example,Wu et al\. \([2024](https://arxiv.org/html/2607.00512#bib.bib25)\)investigates the relationship between FinTech development and inclusive green growth using city\-level panel data under flexible control adjustment\.
A related but distinct stream combines SEM or PLS\-SEM with machine learning\. For instance,Richter and Tudoran \([2024](https://arxiv.org/html/2607.00512#bib.bib21)\), discuss how PLS\-SEM and selected machine\-learning algorithms can be combined to improve prediction, identify nonlinearities and interactions, and support theoretical insight in business research\. Such work shows the increasing value of SEM–ML integration, but its emphasis is often on predictive, exploratory, or causal\-predictive approaches\. The present study is narrower and more diagnostic\. It asks whether theory\-specified SEM paths remain directionally and statistically stable when re\-examined using OLS and DML\-style robustness checks\.
### 2\.3Why DML is applied after SEM and path by path
The distinction is important because applying DML directly to raw survey items or simple composite scores may underuse one of SEM’s central strengths: the ability to model measurement structure before testing structural relationships\. Therefore, this study first uses SEM to validate the latent\-variable measurement structure and identify the retained theoretical path system\. OLS and DML are then applied on a per\-path basis as supplementary robustness analyses\. Each selected SEM path is re\-examined as a focal relationship, with other theoretically relevant constructs and observed controls treated as adjustment variables\.
This use of DML differs from typical applied DML studies in both purpose and implementation\. In many applications, the main objective is to estimate one focal treatment–outcome relationship while flexibly adjusting for controls\. In contrast, this study uses DML as a robustness diagnostic for a latent\-construct path system\. SEM provides the measurement and theoretical map, OLS provides a transparent regression benchmark, and DML provides a flexible\-control robustness check for individual SEM\-implied paths\.
### 2\.4Causal caution, endogeneity, and directional interpretation
Although DML is often discussed within the causal\-inference literature, the present study uses it as a robustness analysis rather than as definitive causal identification\. The survey data are observational, and the structural paths estimated in SEM, OLS, and DML are therefore interpreted as theoretically grounded associations unless stronger identification assumptions are justified\. This cautious interpretation is consistent with Pearl’s account of causal models and inference\(Pearl,[2009](https://arxiv.org/html/2607.00512#bib.bib18)\)and with Pearl’s hierarchy, which distinguishes associational, interventional, and counterfactual claims\(Bareinboim et al\.,[2022](https://arxiv.org/html/2607.00512#bib.bib3)\)\. Accordingly, the proposed SEM–OLS–DML workflow strengthens robustness assessment, but it does not by itself establish intervention\-level or counterfactual causal claims\.
The same caution applies to exogeneity, endogeneity, and directionality\. In SEM, exogenous and endogenous constructs are defined by their positions in the specified structural model\. This modelling distinction should not be confused with econometric endogeneity, which refers to potential bias arising from omitted variables, simultaneity, reverse causality, measurement error, or related sources\. DML can improve robustness by flexibly adjusting for observed covariates, but it does not eliminate unobserved confounding or prove causal direction in cross\-sectional survey data\. Throughout this study, directional language refers to the theoretically specified and statistically estimated direction of association among constructs, rather than definitive causal direction\. Positive coefficients indicate a direct relationship between the predictor and outcome variables, while negative coefficients indicate an inverse relationship\. For theoretically sensitive paths, reverse\-direction OLS/DML checks may therefore be used as diagnostic evidence of directional robustness or possible reciprocal association, rather than as proof of bidirectional causality\.
## 3Preliminaries and Notation
This section introduces the notation used to connect the SEM, OLS, and DML\-style robustness stages\. The purpose is not to derive a new estimator, but to make explicit how survey indicators, latent construct scores, structural paths, and robustness estimates are represented consistently across the three stages\.
### 3\.1Observed survey data
LetNNandPPdenote the number of respondents and observed survey variables, respectively\. For respondentii, the full observed response vector is denoted by𝐙i\\mathbf\{Z\}\_\{i\}:
𝐙i\\displaystyle\\mathbf\{Z\}\_\{i\}=\[Zi1Zi2⋯ZiP\]′,i=1,…,N\.\\displaystyle=\\begin\{bmatrix\}Z\_\{i1\}&Z\_\{i2\}&\\cdots&Z\_\{iP\}\\end\{bmatrix\}^\{\\prime\},\\quad i=1,\\ldots,N\.\(1\)
Stacking all respondent\-level vectors gives the observed data matrix:
𝐙\\displaystyle\\mathbf\{Z\}=\[𝐙1′𝐙2′⋮𝐙N′\]∈ℝN×P\.\\displaystyle=\\begin\{bmatrix\}\\mathbf\{Z\}\_\{1\}^\{\\prime\}\\\\ \\mathbf\{Z\}\_\{2\}^\{\\prime\}\\\\ \\vdots\\\\ \\mathbf\{Z\}\_\{N\}^\{\\prime\}\\end\{bmatrix\}\\in\\mathbb\{R\}^\{N\\times P\}\.\(2\)
Let𝒞\\mathcal\{C\}denote the set of latent constructs and letℐc\\mathcal\{I\}\_\{c\}denote the set of observed indicators assigned to constructc∈𝒞c\\in\\mathcal\{C\}\. Ifℐc=\{j1,j2,…,jJc\}\\mathcal\{I\}\_\{c\}=\\\{j\_\{1\},j\_\{2\},\\ldots,j\_\{J\_\{c\}\}\\\}, then the observed indicator vector for constructccand respondentiiis:
𝐳c,i\\displaystyle\\mathbf\{z\}\_\{c,i\}=\[Zij1Zij2⋯ZijJc\]′,j1,…,jJc∈ℐc\.\\displaystyle=\\begin\{bmatrix\}Z\_\{ij\_\{1\}\}&Z\_\{ij\_\{2\}\}&\\cdots&Z\_\{ij\_\{J\_\{c\}\}\}\\end\{bmatrix\}^\{\\prime\},\\quad j\_\{1\},\\ldots,j\_\{J\_\{c\}\}\\in\\mathcal\{I\}\_\{c\}\.\(3\)
### 3\.2First\-order and second\-order latent constructs
The SEM measurement model includes both first\-order and second\-order latent constructs\. For a first\-order constructcc, let𝐳c,i\\mathbf\{z\}\_\{c,i\}denote the vector of observed indicators for respondentii, and letωc,i\\omega\_\{c,i\}denote the corresponding latent construct\. The first\-order measurement model can be written as:
𝐳c,i\\displaystyle\\mathbf\{z\}\_\{c,i\}=𝝀cωc,i\+𝐞c,i,\\displaystyle=\\boldsymbol\{\\lambda\}\_\{c\}\\omega\_\{c,i\}\+\\mathbf\{e\}\_\{c,i\},\(4\)
where𝝀c\\boldsymbol\{\\lambda\}\_\{c\}is the vector of factor loadings and𝐞c,i\\mathbf\{e\}\_\{c,i\}is the vector of measurement errors\.
For a second\-order constructhh, the construct is inferred from multiple first\-order latent dimensions\. Let𝝎h,i\(1\)\\boldsymbol\{\\omega\}\_\{h,i\}^\{\(1\)\}denote the vector of first\-order latent dimensions associated with second\-order constructhh, and letωh,i\(2\)\\omega\_\{h,i\}^\{\(2\)\}denote the second\-order latent construct\. The second\-order measurement relationship can be represented as:
𝝎h,i\(1\)\\displaystyle\\boldsymbol\{\\omega\}\_\{h,i\}^\{\(1\)\}=𝝀h\(2\)ωh,i\(2\)\+𝒖h,i,\\displaystyle=\\boldsymbol\{\\lambda\}\_\{h\}^\{\(2\)\}\\omega\_\{h,i\}^\{\(2\)\}\+\\boldsymbol\{u\}\_\{h,i\},\(5\)
where𝝀h\(2\)\\boldsymbol\{\\lambda\}\_\{h\}^\{\(2\)\}contains the second\-order loadings and𝒖h,i\\boldsymbol\{u\}\_\{h,i\}captures first\-order dimension disturbances not explained by the second\-order construct\.
In this study, Perceived Quality \(PQ\) and User Experience \(UX\) are treated as second\-order constructs\. PQ is inferred from system, information, and service quality\. UX is inferred from utilitarian, hedonic, personalised, and security experiences\. Other constructs are treated as first\-order latent constructs measured directly by their observed survey indicators\.
### 3\.3SEM measurement and structural notation
In the SEM stage, the measurement model defines latent constructs, while the structural model specifies their relationships\. Let𝝃i\\boldsymbol\{\\xi\}\_\{i\}denote the vector of structurally exogenous latent constructs for respondentii, and let𝜼i\\boldsymbol\{\\eta\}\_\{i\}denote the vector of structurally endogenous latent constructs\. The combined latent construct vector is:
𝝎i\\displaystyle\\boldsymbol\{\\omega\}\_\{i\}=\[𝝃i′𝜼i′\]′\.\\displaystyle=\\begin\{bmatrix\}\\boldsymbol\{\\xi\}\_\{i\}^\{\\prime\}&\\boldsymbol\{\\eta\}\_\{i\}^\{\\prime\}\\end\{bmatrix\}^\{\\prime\}\.\(6\)
A compact measurement model can be written as:
𝐙i\\displaystyle\\mathbf\{Z\}\_\{i\}=𝚲𝝎i\+𝐞i,\\displaystyle=\\boldsymbol\{\\Lambda\}\\boldsymbol\{\\omega\}\_\{i\}\+\\mathbf\{e\}\_\{i\},\(7\)
where𝚲\\boldsymbol\{\\Lambda\}is the factor\-loading matrix and𝐞i\\mathbf\{e\}\_\{i\}is the vector of measurement errors\.
The structural component of the SEM can be written as:
𝜼i\\displaystyle\\boldsymbol\{\\eta\}\_\{i\}=𝐁𝜼i\+𝚪𝝃i\+𝜻i,\\displaystyle=\\mathbf\{B\}\\boldsymbol\{\\eta\}\_\{i\}\+\\boldsymbol\{\\Gamma\}\\boldsymbol\{\\xi\}\_\{i\}\+\\boldsymbol\{\\zeta\}\_\{i\},\(8\)
where𝐁\\mathbf\{B\}contains relationships among endogenous latent constructs,𝚪\\boldsymbol\{\\Gamma\}contains effects from exogenous to endogenous latent constructs, and𝜻i\\boldsymbol\{\\zeta\}\_\{i\}is the structural disturbance vector\.
The SEM parameter vector is estimated by fitting the model\-implied covariance structure to the observed data:
𝜽^\\displaystyle\\widehat\{\\boldsymbol\{\\theta\}\}=argmin𝜽ℒSEM\(𝜽;𝐙,ℳ,𝒫\),\\displaystyle=\\arg\\min\_\{\\boldsymbol\{\\theta\}\}\\mathcal\{L\}\_\{\\mathrm\{SEM\}\}\\left\(\\boldsymbol\{\\theta\};\\mathbf\{Z\},\\mathcal\{M\},\\mathcal\{P\}\\right\),\(9\)
whereℳ\\mathcal\{M\}denotes the measurement specification and𝒫\\mathcal\{P\}denotes the structural path specification\.
In SEM terminology, structurally exogenous and endogenous constructs are distinguished by their positions in the specified structural model\. Exogenous constructs do not receive incoming structural paths, whereas endogenous constructs are explained by one or more other constructs\. This distinction is important for SEM specification and interpretation\. However, once respondent\-level construct scores are extracted, both exogenous and endogenous constructs can be represented within the same construct\-score matrix for downstream OLS and DML\-style robustness checks\.
### 3\.4Construct\-score representations
After estimating the SEM, respondent\-level latent construct scores are extracted\. Let𝝎^i\\widehat\{\\boldsymbol\{\\omega\}\}\_\{i\}denote the SEM factor\-score vector for respondentii\. A general model\-based factor\-score representation is:
𝝎^i\\displaystyle\\widehat\{\\boldsymbol\{\\omega\}\}\_\{i\}=𝚽ω𝚲′𝚺\(𝜽^\)−1𝐙i,\\displaystyle=\\boldsymbol\{\\Phi\}\_\{\\omega\}\\boldsymbol\{\\Lambda\}^\{\\prime\}\\mathbf\{\\Sigma\}\\left\(\\widehat\{\\boldsymbol\{\\theta\}\}\\right\)^\{\-1\}\\mathbf\{Z\}\_\{i\},\(10\)
where𝚽ω\\boldsymbol\{\\Phi\}\_\{\\omega\}is the latent construct covariance matrix,𝚲\\boldsymbol\{\\Lambda\}is the estimated loading matrix, and𝚺\(𝜽^\)\\mathbf\{\\Sigma\}\(\\widehat\{\\boldsymbol\{\\theta\}\}\)is the model\-implied covariance matrix\.
Stacking all extracted SEM factor scores gives:
𝐅^SEM\\displaystyle\\widehat\{\\mathbf\{F\}\}^\{\\mathrm\{SEM\}\}=\[𝝎^1′𝝎^2′⋮𝝎^N′\]∈ℝN×C,\\displaystyle=\\begin\{bmatrix\}\\widehat\{\\boldsymbol\{\\omega\}\}\_\{1\}^\{\\prime\}\\\\ \\widehat\{\\boldsymbol\{\\omega\}\}\_\{2\}^\{\\prime\}\\\\ \\vdots\\\\ \\widehat\{\\boldsymbol\{\\omega\}\}\_\{N\}^\{\\prime\}\\end\{bmatrix\}\\in\\mathbb\{R\}^\{N\\times C\},\(11\)
whereCCis the number of construct scores used in the downstream robustness checks\.
As a simpler benchmark, mean composite scores are calculated from the retained indicators\. For constructccand respondentii, the composite score is:
Z¯c,iCOMP\\displaystyle\\bar\{Z\}\_\{c,i\}^\{\\mathrm\{COMP\}\}=1Jc∑j∈ℐcZij\.\\displaystyle=\\frac\{1\}\{J\_\{c\}\}\\sum\_\{j\\in\\mathcal\{I\}\_\{c\}\}Z\_\{ij\}\.\(12\)
Stacking the composite scores gives the composite\-score matrix:
𝐅COMP\\displaystyle\\mathbf\{F\}^\{\\mathrm\{COMP\}\}=\[𝐳¯1COMP′𝐳¯2COMP′⋮𝐳¯NCOMP′\]∈ℝN×C\.\\displaystyle=\\begin\{bmatrix\}\\bar\{\\mathbf\{z\}\}\_\{1\}^\{\\mathrm\{COMP\}\\prime\}\\\\ \\bar\{\\mathbf\{z\}\}\_\{2\}^\{\\mathrm\{COMP\}\\prime\}\\\\ \\vdots\\\\ \\bar\{\\mathbf\{z\}\}\_\{N\}^\{\\mathrm\{COMP\}\\prime\}\\end\{bmatrix\}\\in\\mathbb\{R\}^\{N\\times C\}\.\(13\)
The robustness checks therefore use two alternative construct\-score representations:
s\\displaystyle s∈\{SEM,COMP\},\\displaystyle\\in\\left\\\{\\mathrm\{SEM\},\\mathrm\{COMP\}\\right\\\},\(14\)
whereSEM\\mathrm\{SEM\}denotes SEM factor scores andCOMP\\mathrm\{COMP\}denotes mean composite scores\.
SEM factor scores are treated as the primary construct\-score representation because they preserve the fitted SEM measurement model, including estimated factor loadings, latent\-variable structure, and the distinction between first\-order and second\-order constructs\. Mean composite scores are included as a transparent sensitivity benchmark\. Composite scores are widely used in survey\-based research because they are simple, interpretable, and easy to reproduce\. However, they treat retained indicators as equally weighted and do not explicitly account for factor loadings, measurement error, or higher\-order construct structure\. Comparing the two score representations is therefore useful for diagnosing measurement\-representation sensitivity\. The next section uses this notation to specify how tested SEM paths are translated into OLS and DML\-style robustness checks\.
## 4Proposed SEM–OLS–DML Robustness Framework
### 4\.1Methodological rationale
The proposed framework is based on methodological complementarity rather than methodological substitution\. SEM remains the primary modelling stage because it validates the latent measurement structure and estimates theory\-driven structural relationships\. OLS provides a transparent score\-based benchmark by asking whether the SEM\-implied relationships remain visible when respondent\-level construct scores are analysed using conventional regression\. DML\-style residualisation provides an additional robustness layer by asking whether focal path estimates remain stable when the outcome and focal predictor are flexibly adjusted for observed controls and co\-predictors\. The three methods therefore answer different but connected questions: SEM asks whether the latent\-variable model is theoretically and statistically supported; OLS asks whether the same path pattern is visible under a simpler linear score\-based specification; and DML\-style checks ask whether the focal relationships remain stable under flexible control\-function adjustment\.
Operationally, the framework separates the analysis into two layers\. The*result\-generation layer*estimates the SEM model, constructs alternative construct\-score representations, and runs the OLS and DML\-style models across the planned scenarios\. The*comparison\-and\-interpretation layer*then compares the generated outputs across methods, score representations, nuisance learners, and selected reverse\-direction diagnostics\. This separation is useful because the Colab implementation can first export all scenario\-level results as CSV files, after which the paper can analyse convergence, divergence, and sensitivity in an integrated way\.
### 4\.2Workflow overview
The proposed workflow is organised into eight operational stages within these two analytical layers\. Stages 1–7 form the result\-generation layer: they refine the measurement model, establish the robustness\-baseline SEM model, construct factor and composite scores, define path\-level specifications, and run OLS and DML\-style checks across score representations and nuisance learners before final structural path evaluation\. Stage 8 forms the comparison\-and\-interpretation layer: it integrates the evidence by comparing direction, statistical support, score\-representation sensitivity, learner sensitivity, and selected reverse\-direction diagnostics to support final structural model refinement decisions\.
The workflow distinguishes measurement refinement from final structural path evaluation\. The initial conceptual model specifies both the measurement structure and the theory\-driven structural paths among constructs\. The measurement part is first evaluated using item diagnostics, factor loadings, reliability, validity, and model\-fit evidence\. Where justified, indicators or measurement factors are refined so that the constructs used in subsequent analysis are empirically defensible\. However, the theory\-specified structural paths are not removed or modified immediately after SEM estimation solely on the basis of SEM significance\. Instead, after the refined measurement model is obtained, the full theory\-specified structural model is re\-estimated and treated as the robustness\-baseline SEM model\. This model provides the path system that is then translated into OLS and DML\-style specifications\. Final structural model refinement decisions are made only after SEM estimates, OLS benchmarks, DML\-style checks, score sensitivity, learner sensitivity, reverse\-direction diagnostics, and theoretical considerations are examined together\.
Table[1](https://arxiv.org/html/2607.00512#S4.T1)complements the figure by translating the same workflow into operational stages, questions, and outputs that can be followed in empirical applications\.
Table 1:Operational stages of the SEM–OLS–DML robustness workflowStageWorkflow componentMain questionAnalytical output1Initial conceptual model and item diagnosticsAre the proposed constructs, indicators, and theory\-specified structural paths suitable for empirical testing?Initial conceptual model and diagnostic evidence for item/indicator quality2Measurement\-model refinementDo the retained indicators provide an empirically defensible measurement model for the latent constructs?Refined measurement model, retained indicators, reliability, validity, and fit evidence3Robustness\-baseline SEM estimationUsing the refined measurement model, what evidence is obtained when the full theory\-specified structural model is estimated before final structural path evaluation?Robustness\-baseline SEM model, SEM path estimates, and tested structural path system4Construct\-score constructionHow are the refined latent constructs represented for downstream robustness analysis?SEM factor scores and mean composite scores5SEM\-to\-regression path translationHow is each tested SEM path translated into path\-level OLS and DML\-style specifications?Outcome constructYY, focal predictorDD, and control vector𝐗\\mathbf\{X\}for each tested path6OLS robustness benchmarkDo the tested SEM paths remain visible under transparent score\-based linear regression?OLS path\-level robustness estimates across construct\-score representations7DML\-style robustness and sensitivity checksDo the tested SEM paths remain stable after flexible adjustment for SEM co\-predictors and observed controls, and are results sensitive to the learner used?DML\-style estimates across score types and RF, GBM, and SVM learners8Integrated robustness interpretation and final structural path evaluationWhere do SEM, OLS, and DML\-style evidence converge or diverge, and which paths should be retained, reviewed, or treated as directionally sensitive?Integrated robustness evidence, selected reverse\-direction diagnostics, and final structural model refinement decisionsNote\.The workflow separates measurement refinement from structural path evaluation\. Indicators and measurement factors may be refined before the robustness analysis, but theory\-specified structural paths are not modified immediately after SEM estimation\. Final structural model refinement decisions are made only after SEM, OLS, DML\-style results, score sensitivity, learner sensitivity, reverse\-direction diagnostics, and theoretical considerations are reviewed together\.
### 4\.3SEM\-to\-regression translation
A SEM structural path can be translated into an OLS or DML\-style specification by identifying the outcome construct, the focal predictor, and the local co\-predictor/control set\. For a SEM equation of the formY∼D\+XY\\sim D\+X, OLS estimates the score\-based regression ofYYonDDandXX, while DML\-style residualisation partials outXXfrom bothYYandDDbefore estimating the residual relationship\.
A key feature of this translation is that SEM structural equations are translated at the equation level, while the resulting regression coefficients are interpreted at the path level\. One regression equation may therefore contain multiple SEM\-implied paths\. For example, the SEM equationBI←AT\+CABI\\leftarrow AT\+CAcorresponds to one regression model withBIBIas the outcome andATATandCACAas predictors, but it yields two path\-level benchmarks:AT→BIAT\\rightarrow BIandCA→BICA\\rightarrow BI\. Similarly,SA←BI\+UX\+UBSA\\leftarrow BI\+UX\+UBcorresponds to one regression equation but provides three path\-level benchmarks\. This distinction is important because the robustness analysis follows the SEM structural system while allowing each individual path coefficient to be compared across SEM, OLS, and DML\.
Table 2:SEM\-to\-regression translation logic
### 4\.4Path\-level robustness notation
Letk=1,…,Kk=1,\\ldots,Kindex the structural paths from the robustness\-baseline SEM model\. For each pathkkand score representations∈\{SEM,COMP\}s\\in\\\{\\mathrm\{SEM\},\\mathrm\{COMP\}\\\}, letYik\(s\)Y\_\{ik\}^\{\(s\)\}denote the outcome construct score,Dik\(s\)D\_\{ik\}^\{\(s\)\}denote the focal predictor construct score, and𝐗ik\(s\)\\mathbf\{X\}\_\{ik\}^\{\(s\)\}denote the control vector for respondentii\. The control vector may include observed controls and SEM co\-predictors from the same structural equation\. For example, when a SEM equation contains two predictors, the focal predictor is treated asDD, while the other predictor is included in𝐗\\mathbf\{X\}as a co\-predictor control\.
The path\-level notation is:
Yik\(s\)\\displaystyle Y\_\{ik\}^\{\(s\)\}=outcome construct score,\\displaystyle=\\text\{outcome construct score\},\(15\)Dik\(s\)\\displaystyle D\_\{ik\}^\{\(s\)\}=focal predictor construct score,\\displaystyle=\\text\{focal predictor construct score\},𝐗ik\(s\)\\displaystyle\\mathbf\{X\}\_\{ik\}^\{\(s\)\}=observed controls and co\-predictors\.\\displaystyle=\\text\{observed controls and co\-predictors\}\.
### 4\.5OLS robustness benchmark
OLS is included as an intermediate benchmark rather than as a competing method\. Because DML estimates can be less transparent to survey researchers, OLS provides a simple regression\-based reference point between SEM and the more flexible DML analysis\. This benchmark helps distinguish relationships that are sensitive to the move from SEM to construct\-score regression from those that are specifically sensitive to machine\-learning\-based control adjustment\.
Following the SEM\-to\-regression translation described above, OLS is estimated equation by equation using the selected construct\-score representation, and the resulting coefficients are interpreted as path\-level benchmarks\. For each focal SEM pathkk, the corresponding OLS benchmark can be written as:
Yik\(s\)\\displaystyle Y\_\{ik\}^\{\(s\)\}=αk\(s\)\+θOLS,k\(s\)Dik\(s\)\+𝜸k\(s\)′𝐗ik\(s\)\+εik\(s\)\.\\displaystyle=\\alpha\_\{k\}^\{\(s\)\}\+\\theta\_\{\\mathrm\{OLS\},k\}^\{\(s\)\}D\_\{ik\}^\{\(s\)\}\+\\boldsymbol\{\\gamma\}\_\{k\}^\{\(s\)\\prime\}\\mathbf\{X\}\_\{ik\}^\{\(s\)\}\+\\varepsilon\_\{ik\}^\{\(s\)\}\.\(16\)
Here,Yik\(s\)Y\_\{ik\}^\{\(s\)\}is the outcome construct score for respondentiiin focal pathkkunder score representationss,Dik\(s\)D\_\{ik\}^\{\(s\)\}is the focal predictor construct score, and𝐗ik\(s\)\\mathbf\{X\}\_\{ik\}^\{\(s\)\}contains companion predictors or theoretically relevant controls included in the corresponding SEM structural equation\. The coefficientθOLS,k\(s\)\\theta\_\{\\mathrm\{OLS\},k\}^\{\(s\)\}is the OLS benchmark estimate for the focal SEM path, andεik\(s\)\\varepsilon\_\{ik\}^\{\(s\)\}is the regression error term\.
This stage deliberately uses a linear additive score\-based specification\. Its value is therefore transparency: OLS provides a familiar benchmark for examining whether the SEM\-implied path pattern remains visible under conventional regression assumptions before the more flexible DML checks are conducted\.
### 4\.6DML\-style partialling\-out robustness check
The DML\-style robustness check uses a partially linear residualisation logic\. For each tested pathkk, the outcome and focal predictor are modelled as functions of the observed controls and co\-predictors:
Yik\(s\)\\displaystyle Y\_\{ik\}^\{\(s\)\}=θDML,k\(s\)Dik\(s\)\+gk\(s\)\(𝐗ik\(s\)\)\+εik\(s\),\\displaystyle=\\theta\_\{\\mathrm\{DML\},k\}^\{\(s\)\}D\_\{ik\}^\{\(s\)\}\+g\_\{k\}^\{\(s\)\}\\left\(\\mathbf\{X\}\_\{ik\}^\{\(s\)\}\\right\)\+\\varepsilon\_\{ik\}^\{\(s\)\},\(17\)Dik\(s\)\\displaystyle D\_\{ik\}^\{\(s\)\}=mk\(s\)\(𝐗ik\(s\)\)\+νik\(s\)\.\\displaystyle=m\_\{k\}^\{\(s\)\}\\left\(\\mathbf\{X\}\_\{ik\}^\{\(s\)\}\\right\)\+\\nu\_\{ik\}^\{\(s\)\}\.
The nuisance functionsgk\(s\)\(⋅\)g\_\{k\}^\{\(s\)\}\(\\cdot\)andmk\(s\)\(⋅\)m\_\{k\}^\{\(s\)\}\(\\cdot\)are estimated using machine\-learning models\. In contrast with OLS, the DML\-style stage does not require the control functions for the outcome and focal predictor to be linear in𝐗ik\(s\)\\mathbf\{X\}\_\{ik\}^\{\(s\)\}\. The focal path remains partially linear, but the adjustment for observed controls and co\-predictors can be estimated flexibly\. This makes the DML\-style check useful for diagnosing whether the OLS robustness result depends on a simple linear control\-function assumption\. The nuisance functions are represented as:
g^k\(s\)\(𝐗ik\(s\)\)\\displaystyle\\widehat\{g\}\_\{k\}^\{\(s\)\}\\left\(\\mathbf\{X\}\_\{ik\}^\{\(s\)\}\\right\)≈𝔼\[Yik\(s\)∣𝐗ik\(s\)\],\\displaystyle\\approx\\mathbb\{E\}\\left\[Y\_\{ik\}^\{\(s\)\}\\mid\\mathbf\{X\}\_\{ik\}^\{\(s\)\}\\right\],\(18\)m^k\(s\)\(𝐗ik\(s\)\)\\displaystyle\\widehat\{m\}\_\{k\}^\{\(s\)\}\\left\(\\mathbf\{X\}\_\{ik\}^\{\(s\)\}\\right\)≈𝔼\[Dik\(s\)∣𝐗ik\(s\)\]\.\\displaystyle\\approx\\mathbb\{E\}\\left\[D\_\{ik\}^\{\(s\)\}\\mid\\mathbf\{X\}\_\{ik\}^\{\(s\)\}\\right\]\.
The residualised outcome and residualised focal predictor are then:
Y~ik\(s\)\\displaystyle\\widetilde\{Y\}\_\{ik\}^\{\(s\)\}=Yik\(s\)−g^k\(s\)\(𝐗ik\(s\)\),\\displaystyle=Y\_\{ik\}^\{\(s\)\}\-\\widehat\{g\}\_\{k\}^\{\(s\)\}\\left\(\\mathbf\{X\}\_\{ik\}^\{\(s\)\}\\right\),\(19\)D~ik\(s\)\\displaystyle\\widetilde\{D\}\_\{ik\}^\{\(s\)\}=Dik\(s\)−m^k\(s\)\(𝐗ik\(s\)\)\.\\displaystyle=D\_\{ik\}^\{\(s\)\}\-\\widehat\{m\}\_\{k\}^\{\(s\)\}\\left\(\\mathbf\{X\}\_\{ik\}^\{\(s\)\}\\right\)\.
The DML\-style estimate is obtained by regressing the residualised outcome on the residualised focal predictor:
θ^DML,k\(s\)\\displaystyle\\widehat\{\\theta\}\_\{\\mathrm\{DML\},k\}^\{\(s\)\}=\(𝐃~k\(s\)′𝐃~k\(s\)\)−1𝐃~k\(s\)′𝐘~k\(s\)\.\\displaystyle=\\left\(\\widetilde\{\\mathbf\{D\}\}\_\{k\}^\{\(s\)\\prime\}\\widetilde\{\\mathbf\{D\}\}\_\{k\}^\{\(s\)\}\\right\)^\{\-1\}\\widetilde\{\\mathbf\{D\}\}\_\{k\}^\{\(s\)\\prime\}\\widetilde\{\\mathbf\{Y\}\}\_\{k\}^\{\(s\)\}\.\(20\)
In the implementation, cross\-fitting is used to reduce overfitting in the nuisance\-function estimation\. Letℐq\\mathcal\{I\}\_\{q\}denote foldqqand letℐ−q\\mathcal\{I\}\_\{\-q\}denote the training observations outside foldqq\. For learnerℓ\\elland observationsi∈ℐqi\\in\\mathcal\{I\}\_\{q\}, the cross\-fitted residuals are:
Y~ik,ℓ\(s\)\\displaystyle\\widetilde\{Y\}\_\{ik,\\ell\}^\{\(s\)\}=Yik\(s\)−g^k,ℓ,−q\(s\)\(𝐗ik\(s\)\),i∈ℐq,\\displaystyle=Y\_\{ik\}^\{\(s\)\}\-\\widehat\{g\}\_\{k,\\ell,\-q\}^\{\(s\)\}\\left\(\\mathbf\{X\}\_\{ik\}^\{\(s\)\}\\right\),\\quad i\\in\\mathcal\{I\}\_\{q\},\(21\)D~ik,ℓ\(s\)\\displaystyle\\widetilde\{D\}\_\{ik,\\ell\}^\{\(s\)\}=Dik\(s\)−m^k,ℓ,−q\(s\)\(𝐗ik\(s\)\),i∈ℐq\.\\displaystyle=D\_\{ik\}^\{\(s\)\}\-\\widehat\{m\}\_\{k,\\ell,\-q\}^\{\(s\)\}\\left\(\\mathbf\{X\}\_\{ik\}^\{\(s\)\}\\right\),\\quad i\\in\\mathcal\{I\}\_\{q\}\.
Here,ℓ\\ellindexes the machine\-learning learner used to estimate the nuisance functions, and−q\-qindicates that the nuisance model is trained on observations outside foldqq\.
### 4\.7Machine\-learning nuisance learners
To examine whether the DML\-style conclusions depend on one particular nuisance learner, the residualisation step is repeated using multiple learners:
ℒ\\displaystyle\\mathcal\{L\}=\{RF,GBM,SVM\}\.\\displaystyle=\\left\\\{\\mathrm\{RF\},\\mathrm\{GBM\},\\mathrm\{SVM\}\\right\\\}\.\(22\)
For each tested pathkk, score representationss, and learnerℓ∈ℒ\\ell\\in\\mathcal\{L\}, the learner\-specific nuisance functions are:
g^k,ℓ\(s\)\(𝐗ik\(s\)\)\\displaystyle\\widehat\{g\}\_\{k,\\ell\}^\{\(s\)\}\\left\(\\mathbf\{X\}\_\{ik\}^\{\(s\)\}\\right\)≈𝔼\[Yik\(s\)∣𝐗ik\(s\)\],\\displaystyle\\approx\\mathbb\{E\}\\left\[Y\_\{ik\}^\{\(s\)\}\\mid\\mathbf\{X\}\_\{ik\}^\{\(s\)\}\\right\],\(23\)m^k,ℓ\(s\)\(𝐗ik\(s\)\)\\displaystyle\\widehat\{m\}\_\{k,\\ell\}^\{\(s\)\}\\left\(\\mathbf\{X\}\_\{ik\}^\{\(s\)\}\\right\)≈𝔼\[Dik\(s\)∣𝐗ik\(s\)\]\.\\displaystyle\\approx\\mathbb\{E\}\\left\[D\_\{ik\}^\{\(s\)\}\\mid\\mathbf\{X\}\_\{ik\}^\{\(s\)\}\\right\]\.
Random Forest represents the nuisance prediction as an average of regression trees:
f^RF\(𝐗\)\\displaystyle\\widehat\{f\}\_\{\\mathrm\{RF\}\}\\left\(\\mathbf\{X\}\\right\)=1B∑b=1BTb\(𝐗\),\\displaystyle=\\frac\{1\}\{B\}\\sum\_\{b=1\}^\{B\}T\_\{b\}\\left\(\\mathbf\{X\}\\right\),\(24\)
whereTb\(⋅\)T\_\{b\}\(\\cdot\)is the prediction from treebbandBBis the number of trees\.
Gradient Boosting Machine represents the nuisance prediction as an additive sequence of weak learners:
f^GBM\(𝐗\)\\displaystyle\\widehat\{f\}\_\{\\mathrm\{GBM\}\}\\left\(\\mathbf\{X\}\\right\)=f0\(𝐗\)\+∑m=1Mνhm\(𝐗\),\\displaystyle=f\_\{0\}\\left\(\\mathbf\{X\}\\right\)\+\\sum\_\{m=1\}^\{M\}\\nu h\_\{m\}\\left\(\\mathbf\{X\}\\right\),\(25\)
wheref0\(⋅\)f\_\{0\}\(\\cdot\)is the initial prediction function,hm\(⋅\)h\_\{m\}\(\\cdot\)is the weak learner added at iterationmm,MMis the number of boosting iterations, andν\\nuis the learning rate\.
Support Vector Machine regression can be written as:
f^SVM\(𝐗\)\\displaystyle\\widehat\{f\}\_\{\\mathrm\{SVM\}\}\\left\(\\mathbf\{X\}\\right\)=∑r∈𝒮𝒱\(αr−αr∗\)K\(𝐗r,𝐗\)\+b0,\\displaystyle=\\sum\_\{r\\in\\mathcal\{SV\}\}\\left\(\\alpha\_\{r\}\-\\alpha\_\{r\}^\{\*\}\\right\)K\\left\(\\mathbf\{X\}\_\{r\},\\mathbf\{X\}\\right\)\+b\_\{0\},\(26\)
where𝒮𝒱\\mathcal\{SV\}is the set of support vectors,K\(⋅,⋅\)K\(\\cdot,\\cdot\)is the kernel function,αr\\alpha\_\{r\}andαr∗\\alpha\_\{r\}^\{\*\}are support\-vector coefficients, andb0b\_\{0\}is the intercept term\.
These learners are not used to redefine the SEM model, serving solely to estimate nuisance functions within the DML\-style residualisation step\. Learner\-sensitivity analysis examines whether the estimated path direction and statistical support remain stable when the nuisance functions are estimated using RF, GBM, and SVM\.
### 4\.8Directional robustness and reverse\-direction diagnostics
For most tested SEM paths, the robustness checks are conducted in the SEM\-specified direction\. However, selected theoretically sensitive paths may also be examined in the reverse direction as a diagnostic extension\. This is useful when a relationship could plausibly be reciprocal, when the SEM sign is unexpected, or when score\-representation sensitivity suggests that the path may require cautious interpretation\.
For a tested SEM pathD→YD\\rightarrow Y, the same\-direction DML\-style check treatsYYas the outcome andDDas the focal predictor\. The reverse\-direction diagnostic swaps the focal roles:
Dik\(s\)\\displaystyle D\_\{ik\}^\{\(s\)\}=ϕDML,k\(s\)Yik\(s\)\+hk\(s\)\(𝐗ik\(s\)\)\+υik\(s\)\.\\displaystyle=\\phi\_\{\\mathrm\{DML\},k\}^\{\(s\)\}Y\_\{ik\}^\{\(s\)\}\+h\_\{k\}^\{\(s\)\}\\left\(\\mathbf\{X\}\_\{ik\}^\{\(s\)\}\\right\)\+\\upsilon\_\{ik\}^\{\(s\)\}\.\(27\)
The reverse\-direction check is not interpreted as proof of bidirectional causality\. Instead, it is a diagnostic test of directional robustness\. If the SEM\-specified direction is supported but the reverse direction is weak, the original direction has stronger robustness support\. If both directions are supported, the relationship may reflect reciprocal association, simultaneity, or shared antecedents\. If the reverse direction is stronger than the SEM\-specified direction, the original theoretical specification should be interpreted cautiously and may require further longitudinal or experimental validation\.
### 4\.9Robustness interpretation criteria
The robustness checks focus on direction, statistical support, score sensitivity, and learner sensitivity\. Directional stability across SEM, OLS, and DML\-style estimates is defined as:
sign\(β^SEM,k\)\\displaystyle\\mathrm\{sign\}\\left\(\\widehat\{\\beta\}\_\{\\mathrm\{SEM\},k\}\\right\)=sign\(θ^OLS,k\(s\)\)=sign\(θ^DML,k\(s\)\)\.\\displaystyle=\\mathrm\{sign\}\\left\(\\widehat\{\\theta\}\_\{\\mathrm\{OLS\},k\}^\{\(s\)\}\\right\)=\\mathrm\{sign\}\\left\(\\widehat\{\\theta\}\_\{\\mathrm\{DML\},k\}^\{\(s\)\}\\right\)\.\(28\)
Statistical stability requires that the path remains statistically supported across the robustness checks:
pOLS,k\(s\)\\displaystyle p\_\{\\mathrm\{OLS\},k\}^\{\(s\)\}<α,pDML,k\(s\)<α,\\displaystyle<\\alpha,\\quad p\_\{\\mathrm\{DML\},k\}^\{\(s\)\}<\\alpha,\(29\)
whereα\\alphais the chosen significance threshold\.
Score sensitivity is assessed by comparing the SEM factor\-score and mean composite\-score results:
sign\(θ^DML,kSEM\)\\displaystyle\\mathrm\{sign\}\\left\(\\widehat\{\\theta\}\_\{\\mathrm\{DML\},k\}^\{\\mathrm\{SEM\}\}\\right\)=sign\(θ^DML,kCOMP\)\.\\displaystyle=\\mathrm\{sign\}\\left\(\\widehat\{\\theta\}\_\{\\mathrm\{DML\},k\}^\{\\mathrm\{COMP\}\}\\right\)\.\(30\)
If the signs remain stable but statistical support differs across score representations, the path is interpreted as measurement\-representation sensitive\.
Learner sensitivity is assessed by comparing GBM, SVM, and RF:
sign\(θ^DML,k,GBM\(s\)\)\\displaystyle\\mathrm\{sign\}\\left\(\\widehat\{\\theta\}\_\{\\mathrm\{DML\},k,\\mathrm\{GBM\}\}^\{\(s\)\}\\right\)=sign\(θ^DML,k,SVM\(s\)\)=sign\(θ^DML,k,RF\(s\)\)\.\\displaystyle=\\mathrm\{sign\}\\left\(\\widehat\{\\theta\}\_\{\\mathrm\{DML\},k,\\mathrm\{SVM\}\}^\{\(s\)\}\\right\)=\\mathrm\{sign\}\\left\(\\widehat\{\\theta\}\_\{\\mathrm\{DML\},k,\\mathrm\{RF\}\}^\{\(s\)\}\\right\)\.\(31\)
A path is classified as strongly robust when it is directionally and statistically stable across SEM, OLS, and DML\-style checks\. A path is classified as partially robust when the direction is stable but statistical support weakens under one or more robustness specifications\. A path is interpreted cautiously when its sign changes, or when its statistical support depends strongly on the construct\-score representation or nuisance learner\.
Before applying the workflow to the DCI demonstration case, it is necessary to clarify how convergence and divergence across SEM, OLS, and DML\-style checks are interpreted\. The three approaches do not estimate identical models: SEM estimates the latent\-variable path system, OLS translates the tested SEM equations from the robustness\-baseline SEM model into transparent construct\-score regressions, and DML\-style analysis re\-examines each focal path under flexible control\-function adjustment\. Therefore, robustness is assessed through directional consistency, statistical support, score sensitivity, learner sensitivity, and, for selected sensitive paths, reverse\-direction diagnostic evidence\. Table[3](https://arxiv.org/html/2607.00512#S4.T3)summarises the interpretation rules used to classify each path in the empirical demonstration\.
Table 3:Robustness interpretation rules for SEM–OLS–DML comparison
## 5Empirical Demonstration: FinTech Digital Customer Intimacy
The framework is demonstrated using a FinTech Digital Customer Intimacy \(DCI\) survey model,\(Liu et al\.,[2026](https://arxiv.org/html/2607.00512#bib.bib16)\)\. The revised implementation estimates the robustness\-baseline SEM model after measurement refinement and before final structural path evaluation\. This is an important change from a conventional trim\-then\-check workflow\. In the revised workflow, SEM, OLS, DML\-style robustness checks, learner\-sensitivity checks, and reverse\-direction diagnostics are all run on the initially tested SEM path system\. Final path decisions are then left to the end, after all empirical evidence has been considered together\.
The tested SEM model includes 17 structural paths\. Its structural component is:
ATi\\displaystyle AT\_\{i\}=β1CAi\+β2PQi\+ζ1,i,\\displaystyle=\\beta\_\{1\}CA\_\{i\}\+\\beta\_\{2\}PQ\_\{i\}\+\\zeta\_\{1,i\},\(32\)BIi\\displaystyle BI\_\{i\}=β3ATi\+β4CAi\+ζ2,i,\\displaystyle=\\beta\_\{3\}AT\_\{i\}\+\\beta\_\{4\}CA\_\{i\}\+\\zeta\_\{2,i\},UBi\\displaystyle UB\_\{i\}=β5BIi\+ζ3,i,\\displaystyle=\\beta\_\{5\}BI\_\{i\}\+\\zeta\_\{3,i\},UXi\\displaystyle UX\_\{i\}=β6UBi\+β7PQi\+ζ4,i,\\displaystyle=\\beta\_\{6\}UB\_\{i\}\+\\beta\_\{7\}PQ\_\{i\}\+\\zeta\_\{4,i\},SAi\\displaystyle SA\_\{i\}=β8BIi\+β9UXi\+β10UBi\+ζ5,i,\\displaystyle=\\beta\_\{8\}BI\_\{i\}\+\\beta\_\{9\}UX\_\{i\}\+\\beta\_\{10\}UB\_\{i\}\+\\zeta\_\{5,i\},TRi\\displaystyle TR\_\{i\}=β11UBi\+β12SAi\+ζ6,i,\\displaystyle=\\beta\_\{11\}UB\_\{i\}\+\\beta\_\{12\}SA\_\{i\}\+\\zeta\_\{6,i\},ECi\\displaystyle EC\_\{i\}=β13UBi\+β14TRi\+ζ7,i,\\displaystyle=\\beta\_\{13\}UB\_\{i\}\+\\beta\_\{14\}TR\_\{i\}\+\\zeta\_\{7,i\},DCIi\\displaystyle DCI\_\{i\}=β15ECi\+β16TRi\+β17CAi\+ζ8,i\.\\displaystyle=\\beta\_\{15\}EC\_\{i\}\+\\beta\_\{16\}TR\_\{i\}\+\\beta\_\{17\}CA\_\{i\}\+\\zeta\_\{8,i\}\.
Table[4](https://arxiv.org/html/2607.00512#S5.T4)summarises the retained indicators and dimensions used to estimate the tested SEM model\. The DCI construct uses DCI1, DCI23, and DCI5, where DCI23 is the averaged indicator formed from DCI2 and DCI3\.
Table 4:Demonstration case constructs and retained indicators in the tested SEM modelNote\.DCI23 denotes the averaged indicator formed from DCI2 and DCI3\. Alpha is reported for first\-order retained constructs only; second\-order constructs are represented by their first\-order dimensions\.
### 5\.1Robustness\-baseline SEM specification and fit
The robustness\-baseline DCI SEM model provides the common structural path system for all downstream robustness checks\. Table[5](https://arxiv.org/html/2607.00512#S5.T5)reports the fit overview for this model\. The model has acceptable approximate fit for demonstration purposes, with CFI above 0\.90, RMSEA below 0\.08, andχ2/df\\chi^\{2\}/dfbelow 3\. The purpose of the demonstration is not to claim that the tested model is the final substantive model, but to show how empirical evidence can be generated before making final structural model refinement decisions\.
Table 5:Tested SEM model fit overviewModelDoFχ2\\chi^\{2\}χ2\\chi^\{2\}/dfCFITLIRMSEABICTested SEM model4411166\.4212\.6450\.9070\.8950\.063517\.942
Note\.The robustness\-baseline SEM model is the full theory\-driven structural specification estimated after measurement refinement and used to generate SEM, OLS, DML, learner\-sensitivity, and reverse\-diagnostic evidence before final structural model refinement decisions are made\.
Figure 1:The robustness\-baseline SEM structural model used as the basis for the OLS and DML\-style robustness stages\.
### 5\.2SEM structural path evidence
Table[6](https://arxiv.org/html/2607.00512#S5.T6)reports the 17 tested SEM structural paths\. Fourteen paths are statistically significant at the 1% level or better, whileCA→ATCA\\rightarrow AT,TR→ECTR\\rightarrow EC, andCA→DCICA\\rightarrow DCIare not statistically significant in the SEM model\. These non\-significant paths are not removed at this stage\. Instead, they are carried forward into the OLS, DML, learner\-sensitivity, and optional reverse\-diagnostic stages so that the final decision can be based on a fuller evidence base\.
Table[7](https://arxiv.org/html/2607.00512#S5.T7)reports the explained variance for endogenous constructs in the robustness\-baseline SEM model\. Most endogenous constructs show substantial to strong explained variance, indicating that the tested structural specification accounts for a large share of variation in the corresponding construct scores\. Usage Behaviour has comparatively low explained variance, suggesting that behavioural intention alone explains only a limited portion of actual usage behaviour in this model\. The very high explained variance for DCI should be interpreted cautiously, as it may reflect strong shared relational meaning among Digital Customer Intimacy, Emotional Connection, Customer Awareness, and Trust\. TheseR2R^\{2\}values are therefore reported as explanatory diagnostics, not as evidence of causal determination\.
Table 6:Structural path estimates for the robustness\-baseline SEM model using retained indicatorsNote\.The robustness\-baseline SEM model uses the retained/refined measurement indicators after item\-level reliability and measurement checks\. Non\-significant paths are retained at this stage so that path decisions can be informed by SEM, OLS, DML, score\-sensitivity, learner\-sensitivity, and reverse\-diagnostic evidence together\.
### 5\.3Path mapping for OLS and DML\-style robustness checks
Table[8](https://arxiv.org/html/2607.00512#S5.T8)shows how each tested SEM path is translated into the score\-based OLS and DML\-style specifications\. The selection of the focal predictor and control vector follows the structure of the tested SEM equation\. For each SEM\-implied pathD→YD\\rightarrow Y, the outcome construct becomesYY, the predictor corresponding to the path under examination becomes the focal variableDD, and the remaining predictors in the same SEM structural equation are included in𝐗\\mathbf\{X\}as co\-predictor controls\. The observed Fintech Type variable is also included as a control in the score\-based robustness checks\. This design preserves the SEM structural context while allowing each tested path to be examined separately\.
Table 7:Explained variance for endogenous constructs in the robustness\-baseline SEM modelEndogenous constructR2R^\{2\}InterpretationAttitude \(AT\)0\.728Strong explained varianceBehavioural Intention \(BI\)0\.827Strong explained varianceUsage Behaviour \(UB\)0\.065Low explained varianceUser Experience \(UX\)0\.665Substantial explained varianceSatisfaction \(SA\)0\.881Strong explained varianceTrust \(TR\)0\.784Strong explained varianceEmotional Connection \(EC\)0\.606Substantial explained varianceDigital Customer Intimacy \(DCI\)0\.960Very high explained varianceNote\.R2R^\{2\}indicates the proportion of variance in each endogenous construct explained by its predictors in the robustness\-baseline SEM model\. These values describe explanatory association within the specified model and should not be interpreted as causal evidence\.Table 8:Mapping of SEM structural paths to OLS and DML\-style specificationsNote\.YYdenotes the outcome construct score,DDdenotes the focal predictor for the SEM\-implied path being examined, and𝐗\\mathbf\{X\}denotes the control vector used in the OLS and DML\-style robustness checks\. For paths from a multi\-predictor SEM equation,𝐗\\mathbf\{X\}contains the remaining predictors in the same SEM equation plus the observed Fintech Type control\. For single\-predictor equations,𝐗\\mathbf\{X\}contains the observed Fintech Type control\.
## 6Robustness Evaluation
### 6\.1SEM, OLS, and DML comparison using SEM factor scores
Table[9](https://arxiv.org/html/2607.00512#S6.T9)reports the SEM, OLS, and DML learner comparison using SEM factor scores\. The factor\-score results are closely aligned with the SEM measurement model\. Most tested paths are supported across OLS and all three DML learners\. The weakest path isCA→ATCA\\rightarrow AT, which is non\-significant in SEM, OLS, and all three DML learners\. The pathsTR→ECTR\\rightarrow ECandCA→DCICA\\rightarrow DCIare non\-significant in SEM but show some score\-based robustness evidence, suggesting that they should be reviewed rather than automatically removed\.
Table 9:SEM, OLS, and DML learner comparison using SEM factor scoresNote\.OLS and DML estimates use the same path\-level specification as Table[8](https://arxiv.org/html/2607.00512#S5.T8)\. DML estimates are derived from the master DML output using Gradient Boosting Machine \(GBM\), Support Vector Machine regression \(SVM\), and Random Forest \(RF\) nuisance learners\. The final column reports how many of the three DML learners are statistically significant atp<0\.05p<0\.05\.
### 6\.2SEM, OLS, and DML comparison using mean composite scores
Table[10](https://arxiv.org/html/2607.00512#S6.T10)reports the same comparison using mean composite scores\. Most paths remain stable, but several paths become more sensitive\. In particular,UB→SAUB\\rightarrow SAis weak under composite\-score OLS and all three composite\-score DML learners\. TheUB→TRUB\\rightarrow TRpath is supported only by the RF learner under composite scores, whileTR→DCITR\\rightarrow DCIis supported by two of the three composite\-score DML learners\. TheTR→ECTR\\rightarrow ECpath is especially sensitive because SEM estimates it as non\-significant and negative, whereas composite\-score OLS and DML estimate it as positive and significant\. These patterns show why final path decisions should be made only after comparing SEM, OLS, DML, score\-representation, and learner\-sensitivity evidence together\.
Table 10:SEM, OLS, and DML learner comparison using mean composite scoresNote\.OLS and DML estimates use the same path\-level specification as Table[8](https://arxiv.org/html/2607.00512#S5.T8)\. DML estimates are derived from the master DML output using Gradient Boosting Machine \(GBM\), Support Vector Machine regression \(SVM\), and Random Forest \(RF\) nuisance learners\. The final column reports how many of the three DML learners are statistically significant atp<0\.05p<0\.05\.
### 6\.3Reverse\-direction diagnostic results
Table[11](https://arxiv.org/html/2607.00512#S6.T11)summarises the reverse\-direction diagnostics\. The diagnostics include core path pairs that are conceptually important for the DCI model and optional path pairs added because the tested SEM path was non\-significant or because the relationship is useful for model\-decision review\. These checks are interpreted only as directional association diagnostics\. They do not establish causal direction\.
Table 11:Reverse\-direction diagnostic checks for selected path pairsNote\.Entries such as 3/3 indicate that all three DML learners were significant for that direction\. Reverse\-direction diagnostics are reported as association\-level sensitivity checks only\. They do not establish reverse or bidirectional causality\.
The reverse diagnostics add several useful insights\. First,EC↔DCIEC\\leftrightarrow DCIis stable in both directions across both score types, suggesting a strong reciprocal association pattern between emotional connection and digital customer intimacy\. Second,TR↔DCITR\\leftrightarrow DCIis directionally complex: the reverse diagnosticDCI→TRDCI\\rightarrow TRis stable, whileTR→DCITR\\rightarrow DCIis weaker under composite scoring\. Third,UB↔SAUB\\leftrightarrow SAandUB↔TRUB\\leftrightarrow TRare strongly supported under factor scores but weak under composite scores, reinforcing the conclusion that usage\-related relational paths are score\-sensitive\. Fourth, the optional diagnostics forCA↔ATCA\\leftrightarrow AT,CA↔DCICA\\leftrightarrow DCI, andTR↔ECTR\\leftrightarrow ECprovide additional evidence for end\-stage model review rather than automatic path removal\.
### 6\.4Integrated model\-decision evidence
Table[12](https://arxiv.org/html/2607.00512#S6.T12)integrates the SEM, OLS, DML, score\-sensitivity, learner\-sensitivity, and reverse\-diagnostic evidence\. The table should be read as decision support rather than as a mechanical rule for retaining or removing paths\. The strongest candidates for retention are paths that are significant in SEM and stable across OLS, factor\-score DML, composite\-score DML, and learners\. Paths with weak SEM support and weak downstream support, such asCA→ATCA\\rightarrow AT, are candidates for removal, revision, or theoretical reconsideration during final structural path evaluation\. Paths with mixed SEM and score\-based evidence, such asTR→ECTR\\rightarrow ECandCA→DCICA\\rightarrow DCI, require theory and measurement review before a final decision is made\.
Table 12:Integrated evidence summary for model\-decision supportNote\.This table provides model\-decision support, not automatic path\-decision rules\. Final decisions should consider theory, construct validity, SEM fit, parsimony, robustness evidence, and substantive interpretability together\.
Table[13](https://arxiv.org/html/2607.00512#S6.T13)summarises the answers to the main research question, sub\-research questions, and applied demonstration question\.
Table 13:Summary answers to the research questions
## 7Discussion
### 7\.1Methodological implications
The workflow demonstrates a more conservative and transparent way to use SEM–OLS–DML robustness analysis\. Rather than removing structural paths immediately after SEM estimation and then checking only the surviving paths, our implementation tests the full theory\-driven SEM model before final structural path evaluation\. This allows the researcher to review SEM significance, OLS benchmarks, DML\-style estimates, learner sensitivity, score\-representation sensitivity, and reverse\-direction diagnostics together before making final structural decisions\.
This shift strengthens the methodological contribution\. SEM remains central because it validates the measurement model and estimates the theory\-specified latent path system\. OLS and DML\-style checks are not replacements for SEM\. They are supplementary evidence layers that help identify whether a path is strongly supported, score\-sensitive, learner\-sensitive, directionally complex, or empirically weak\. The resulting workflow is therefore closer to a model\-decision support process than a simple robustness appendix\.
### 7\.2How the approach differs from typical DML and SEM–ML use
The proposed approach differs from typical applied DML studies because DML is not used here as the primary estimator for a single observed treatment–outcome relationship\. Instead, SEM first provides the measurement model and tested path system\. DML is then applied path by path to examine whether the SEM\-implied associations remain stable after flexible adjustment for observed controls and co\-predictors\. This adaptation is especially relevant for survey\-based research because direct application of DML to raw items or simple composite scores can underuse SEM’s ability to represent latent constructs and measurement quality\.
The approach also differs from SEM–ML integration work that uses machine learning mainly for prediction, nonlinear discovery, interaction detection, or alternative model exploration\. The purpose here is diagnostic robustness and model\-decision support\. The central question is not whether machine learning can outperform SEM predictively, but whether SEM\-implied associations remain directionally and statistically stable under transparent regression, flexible\-control residualisation, alternative score construction, and alternative nuisance learners\.
### 7\.3Substantive implications for the FinTech DCI case
The demonstration suggests that much of the adoption\-to\-intimacy pathway is stable\. Perceived quality, attitude, customer awareness, behavioural intention, usage behaviour, user experience, satisfaction, trust, emotional connection, and DCI are connected in a broadly coherent sequence\. The strongest downstream relational evidence concerns emotional connection and DCI:EC→DCIEC\\rightarrow DCIremains strongly supported across score types and learners, and the reverse\-direction diagnostic also suggests that emotional connection and digital customer intimacy may be mutually reinforcing\. This is consistent with customer\-experience thinking that treats emotional connection as a deeper relational outcome than satisfaction alone\(Zorfas and Leemon,[2016](https://arxiv.org/html/2607.00512#bib.bib27)\)\.
The usage\-related paths require more nuanced interpretation\. TheUB→SAUB\\rightarrow SApath is negative and significant in SEM and strongly supported under factor\-score OLS and DML, but it becomes weak under composite\-score OLS and DML\. This indicates score\-representation sensitivity rather than a simple conclusion that usage reduces satisfaction\. Prior IS research has treated system usage behaviour as closely related to user satisfaction and even as a potential proxy for satisfaction measurement\(Downing,[1999](https://arxiv.org/html/2607.00512#bib.bib11)\); therefore, the present result should be discussed as a conditional association that may depend on how usage behaviour and satisfaction are operationalised\.
The usage–trust relationship also requires cautious interpretation\. In the tested SEM specification, the model explicitly places usage behaviour as the predictor and trust as the outcome; that is,UB→TRUB\\rightarrow TRtreats usage behaviour as an antecedent of trust\. The positive and significant SEM estimate indicates that higher usage behaviour is associated with higher trust, conditional on the model specification\. This path is strongly supported in the factor\-score robustness analysis, but the composite\-score evidence is weaker, suggesting some sensitivity to how the constructs are represented\. Prior mobile\-application and FinTech studies also support a close relationship between trust and use behaviour, but they often specify the relationship in the opposite direction, with trust acting as an antecedent of behavioural intention or actual use\(Yan et al\.,[2013](https://arxiv.org/html/2607.00512#bib.bib26); Ratnawati et al\.,[2022](https://arxiv.org/html/2607.00512#bib.bib20)\)\. Therefore, theUB↔TRUB\\leftrightarrow TRrelationship is better interpreted as a directionally sensitive association\. Rather than concluding that usage behaviour unidirectionally determines trust, the evidence suggests that usage behaviour and trust may be mutually associated and should be examined further in future longitudinal or process\-oriented research\.
The trust\-to\-intimacy and trust\-to\-emotional\-connection results further illustrate the value of testing before final structural path evaluation\. TheTR→DCITR\\rightarrow DCIpath is negative in SEM and remains negative under most robustness checks, but the reverse diagnosticDCI→TRDCI\\rightarrow TRis also stable\. This suggests directional complexity, possible suppression by emotional connection, or a distinction between functional trust and deeper intimacy\. TheTR→ECTR\\rightarrow ECpath is non\-significant in SEM but becomes significant in score\-based robustness checks, with sign sensitivity between factor and composite scores\. This path should therefore be reviewed theoretically and empirically before a final model decision is made\.
Finally, theCA→ATCA\\rightarrow ATandCA→DCICA\\rightarrow DCIpaths show why the revised workflow is useful\.CA→ATCA\\rightarrow ATis weak in SEM and weak across downstream checks, making it a candidate for removal or theoretical reconsideration if theory is also weak\. By contrast,CA→DCICA\\rightarrow DCIis non\-significant in SEM but shows stronger composite\-score robustness evidence, suggesting that it should be reviewed rather than automatically removed\. These examples demonstrate that path decisions should not be made from SEMpp\-values alone\.
### 7\.4Boundary conditions and limitations
Several limitations should be recognised\. First, DML\-style residualisation adjusts flexibly for observed controls, but it does not remove unobserved confounding by itself\. Second, the DML\-style models used here retain a partially linear focal path: machine learning is used to model the nuisance control functions, not to claim a fully nonlinear causal effect of the focal predictor\. Third, factor\-score results are closely aligned with the SEM measurement model and should be interpreted alongside composite\-score sensitivity checks\. Fourth, very highR2R^\{2\}values in score\-based downstream models should be examined carefully because they may reflect the construction and scaling of SEM factor scores rather than purely substantive explanatory power\. Fifth, reverse\-direction checks should be interpreted only as directional association diagnostics, not as evidence of bidirectional causality\. Longitudinal, experimental, quasi\-experimental, or process\-modelling designs would be needed to make stronger claims about temporal ordering or causal direction\.
## 8Conclusion
This paper proposes and demonstrates a staged SEM–OLS–DML robustness framework for survey\-based latent\-construct research\. The proposed workflow estimates the robustness\-baseline SEM model before final structural path evaluation, then uses OLS, DML\-style residualisation, score\-representation sensitivity, learner sensitivity, and selected reverse\-direction diagnostics to generate model\-decision evidence\. In the FinTech Digital Customer Intimacy demonstration, most tested paths remain stable across methods and score types\. The strongest downstream relational path remainsEC→DCIEC\\rightarrow DCI\. The revised evidence also identifies paths requiring end\-stage review:CA→ATCA\\rightarrow ATshows weak empirical support,TR→ECTR\\rightarrow ECandCA→DCICA\\rightarrow DCIshow mixed SEM and score\-based evidence, andUB→SAUB\\rightarrow SA,UB→TRUB\\rightarrow TR, andTR→DCITR\\rightarrow DCIrequire cautious discussion because of score sensitivity or directional complexity\. The framework provides a practical interpretation guide for researchers seeking to complement SEM with conventional and machine\-learning\-based robustness checks\. The accompanying Colab workbook and Zenodo archive further position the paper as a reusable template for researchers and students who wish to adapt the workflow to other survey\-based SEM applications\.
## Data and Code Availability
The Google Colab workbook, generated result tables, and selected output figures associated with this study are available from Zenodo \([https://doi\.org/10\.5281/zenodo\.21073457](https://doi.org/10.5281/zenodo.21073457)\)\. The archive is intended to serve both as a replication package for the empirical demonstration and as a reusable template for applying the SEM–OLS–DML robustness workflow to other survey\-based latent\-construct datasets\.
## References
- Anderson and Gerbing \(1988\)Anderson, J\. C\., and Gerbing, D\. W\. \(1988\)\. Structural equation modeling in practice: A review and recommended two\-step approach\.Psychological Bulletin, 103\(3\), 411–423\.
- Bach et al\. \(2022\)Bach, P\., Chernozhukov, V\., Kurz, M\. S\., and Spindler, M\. \(2022\)\. DoubleML—An object\-oriented implementation of double machine learning in Python\.Journal of Machine Learning Research, 23\(53\), 1–6\.
- Bareinboim et al\. \(2022\)Bareinboim, E\., Correa, J\. D\., Ibeling, D\., and Icard, T\. \(2022\)\. On Pearl’s hierarchy and the foundations of causal inference\. In H\. Geffner, R\. Dechter, and J\. Y\. Halpern \(Eds\.\),Probabilistic and causal inference: The works of Judea Pearl\(pp\. 507–556\)\. ACM Books\.
- Bhattacherjee \(2001\)Bhattacherjee, A\. \(2001\)\. Understanding information systems continuance: An expectation\-confirmation model\.MIS Quarterly, 25\(3\), 351–370\.
- Bollen \(1989\)Bollen, K\. A\. \(1989\)\.Structural equations with latent variables\. John Wiley & Sons\.
- Breiman \(2001\)Breiman, L\. \(2001\)\. Random forests\.Machine Learning, 45\(1\), 5–32\.
- Brunner \(2023\)Brunner, J\. \(2023\)\.Structural equation models: An open textbook\(Edition 0\.10\)\. Department of Statistical Sciences, University of Toronto\.
- Chernozhukov et al\. \(2018\)Chernozhukov, V\., Chetverikov, D\., Demirer, M\., Duflo, E\., Hansen, C\., Newey, W\., and Robins, J\. \(2018\)\. Double/debiased machine learning for treatment and structural parameters\.The Econometrics Journal, 21\(1\), C1–C68\.
- Cortes and Vapnik \(1995\)Cortes, C\., and Vapnik, V\. \(1995\)\. Support\-vector networks\.Machine Learning, 20\(3\), 273–297\.
- Davis \(1989\)Davis, F\. D\. \(1989\)\. Perceived usefulness, perceived ease of use, and user acceptance of information technology\.MIS Quarterly, 13\(3\), 319–340\.
- Downing \(1999\)Downing, C\. E\. \(1999\)\. System usage behavior as a proxy for user satisfaction: An empirical investigation\.Information & Management, 35\(4\), 203–216\.[https://doi\.org/10\.1016/S0378\-7206\(98\)00090\-1](https://doi.org/10.1016/S0378-7206(98)00090-1)\.
- Friedman \(2001\)Friedman, J\. H\. \(2001\)\. Greedy function approximation: A gradient boosting machine\.The Annals of Statistics, 29\(5\), 1189–1232\.
- Gefen et al\. \(2000\)Gefen, D\., Straub, D\. W\., and Boudreau, M\.\-C\. \(2000\)\. Structural equation modeling and regression: Guidelines for research practice\.Communications of the Association for Information Systems, 4, Article 7\.
- Liu et al\. \(2024a\)Liu, Q\., Chan, K\.\-C\., and Chimhundu, R\. \(2024a\)\. Fintech research: Systematic mapping, classification, and future directions\.Financial Innovation, 10\(1\), Article 24\.
- Liu et al\. \(2024b\)Liu, Q\., Chan, K\.\-C\., and Chimhundu, R\. \(2024b\)\. From customer intimacy to digital customer intimacy\.Journal of Theoretical and Applied Electronic Commerce Research, 19\(4\), 3386–3411\.
- Liu et al\. \(2026\)Liu, Q\., Chan, K\.\-C\., Tiwari, S\., and Chimhundu, R\. \(2026\)\.From adoption to intimacy: Experiential and emotional pathways of digital customer intimacy in FinTech\. Manuscript under review\.
- MacKenzie et al\. \(2011\)MacKenzie, S\. B\., Podsakoff, P\. M\., and Podsakoff, N\. P\. \(2011\)\. Construct measurement and validation procedures in MIS and behavioral research: Integrating new and existing techniques\.MIS Quarterly, 35\(2\), 293–334\.
- Pearl \(2009\)Pearl, J\. \(2009\)\.Causality: Models, reasoning, and inference\(2nd ed\.\)\. Cambridge University Press\.
- Podsakoff et al\. \(2003\)Podsakoff, P\. M\., MacKenzie, S\. B\., Lee, J\.\-Y\., and Podsakoff, N\. P\. \(2003\)\. Common method biases in behavioral research: A critical review of the literature and recommended remedies\.Journal of Applied Psychology, 88\(5\), 879–903\.
- Ratnawati et al\. \(2022\)Ratnawati, S\., Durachman, Y\., and Saputra, A\. \(2022\)\. Analyzing factors influencing intention to use and actual use of mobile fintech applications free interbank money transfer Flip using UTAUT 2 model with trust and perceived security\. In2022 10th International Conference on Cyber and IT Service Management \(CITSM\)\.[https://doi\.org/10\.1109/CITSM56380\.2022\.9935838](https://doi.org/10.1109/CITSM56380.2022.9935838)\.
- Richter and Tudoran \(2024\)Richter, N\. F\., and Tudoran, A\. A\. \(2024\)\. Elevating theoretical insight and predictive accuracy in business research: Combining PLS\-SEM and selected machine learning algorithms\.Journal of Business Research, 173, Article 114453\.
- Shi et al\. \(2025\)Shi, B\., Mao, X\., Yang, M\., and Li, B\. \(2025\)\. What, why, and how: An empiricist’s guide to double/debiased machine learning\.Information Systems Research\. Advance online publication\.
- Treacy and Wiersema \(1993\)Treacy, M\., and Wiersema, F\. \(1993\)\. Customer intimacy and other value disciplines\.Harvard Business Review, 71\(1\), 84–93\.
- Venkatesh et al\. \(2003\)Venkatesh, V\., Morris, M\. G\., Davis, G\. B\., and Davis, F\. D\. \(2003\)\. User acceptance of information technology: Toward a unified view\.MIS Quarterly, 27\(3\), 425–478\.
- Wu et al\. \(2024\)Wu, B\., Ding, Y\., Xie, B\., and Zhang, Y\. \(2024\)\. FinTech and inclusive green growth: A causal inference based on double machine learning\.Sustainability, 16\(22\), Article 9989\.[https://doi\.org/10\.3390/su16229989](https://doi.org/10.3390/su16229989)\.
- Yan et al\. \(2013\)Yan, Z\., Dong, Y\., Niemi, V\., and Yu, G\. \(2013\)\. Exploring trust of mobile applications based on user behaviors: An empirical study\.Journal of Applied Social Psychology, 43\(3\), 638–659\.[https://doi\.org/10\.1111/j\.1559\-1816\.2013\.01044\.x](https://doi.org/10.1111/j.1559-1816.2013.01044.x)\.
- Zorfas and Leemon \(2016\)Zorfas, A\., and Leemon, D\. \(2016, August 29\)\. An emotional connection matters more than customer satisfaction\.Harvard Business Review\.[https://hbr\.org/2016/08/an\-emotional\-connection\-matters\-more\-than\-customer\-satisfaction](https://hbr.org/2016/08/an-emotional-connection-matters-more-than-customer-satisfaction)\.Similar Articles
Beyond the Mean: Three-Axis Fidelity for Aligning LLM-Based Survey Simulators from Small Pilot Data
This paper introduces a three-axis fidelity framework (structural, marginal, individual) to evaluate how well LLMs can simulate survey responses from small pilot data. Using a COVID-19 misinformation survey, it compares prompting, rectification, and fine-tuning approaches, finding that fine-tuning offers balanced fidelity but with variation across subsamples.
Can Large Language Models Revolutionize Survey Research? Experiments with Disaster Preparedness Responses
This paper presents a five-stage framework integrating large language models into survey research, addressing declining response rates, sample bias, and fraudulent completions. Using 2024 Hurricane Milton survey data, the authors propose a theory-informed LLM (A-TLM) that outperforms classical imputation methods in missing-data scenarios and demonstrates manageable hallucination risk through grounded refusal.
Independent study: one LLM misses ~half the code-review defects a multi-model panel catches. Feedback wanted + seeking arXiv endorsement.
An independent researcher's study finds that a single LLM misses about half of code-review defects, while using multiple models from different providers significantly improves coverage, with the biggest gain from adding a second model. The paper seeks feedback and arXiv endorsement.
LPDS: Evaluating LLM Robustness Through Logic-Preserving Difficulty Scaling
Introduces LPDS, a framework to systematically evaluate LLM robustness by scaling difficulty of logic-preserving variations, finding that performance drops up to 5x compared to random sampling and that training on harder variations improves robustness.
Beyond Surface Statistics: Robust Conformal Prediction for LLMs via Internal Representations
This paper proposes a conformal prediction framework for LLMs that leverages internal representations rather than output-level statistics, introducing Layer-Wise Information (LI) scores as nonconformity measures to improve validity-efficiency trade-offs under distribution shift. The method demonstrates stronger robustness to calibration-deployment mismatch compared to text-level baselines across QA benchmarks.