FinInvest-GTCN: Explainable Graph-Temporal-Causal Modeling for Risk-Aware Investment Decision Optimization

arXiv cs.CL Papers

Summary

Introduces FinInvest-GTCN, a graph-temporal-causal network for risk-aware venture capital investment decisions, achieving state-of-the-art risk-adjusted returns with explainable predictions.

arXiv:2606.28933v1 Announce Type: new Abstract: Venture capital (VC) investment decisions face distinct challenges, such as multi-source heterogeneous data, non-stationary time series, and the demand for explainable predictions in high-stakes, low-data settings. To overcome these issues, we introduce \textbf{FinInvest-GTCN}, a Graph-Temporal-Causal Network that redefines the task from content recommendation to quantitative risk-return assessment. This architecture combines a relational graph encoder to capture the investment ecosystem's topology, a multi-scale temporal fusion module to handle long-term dependencies and non-stationarity, and a causal decision head that generates risk-adjusted predictions with interpretable causal attributions. A core innovation is the Meta-Causal Adaptation (MCA) strategy, which facilitates robust fine-tuning for new, data-scarce sectors by aligning updates with causally-plausible structures derived from meta-pretraining. Comprehensive experiments on proprietary VC datasets show that FinInvest-GTCN delivers state-of-the-art results, markedly lowering the primary Risk-Adjusted Mean Squared Error (RA-MSE) to 2.51 from a baseline of 3.05 and boosting the cumulative return of a simulated portfolio by 18.7\%. Ablation studies underscore the essential role of each component, while additional analyses confirm the model's stability, interpretability, and enhanced adaptability. This work pioneers a data-driven, explainable framework for investment decision support.
Original Article
View Cached Full Text

Cached at: 06/30/26, 05:28 AM

# FinInvest-GTCN: Explainable Graph-Temporal-Causal Modeling for Risk-Aware Investment Decision Optimization
Source: [https://arxiv.org/html/2606.28933](https://arxiv.org/html/2606.28933)
Junyan Tan Department of Computer Science Zhejiang University &Yifan Li Department of Computer Science Zhejiang University Minghao Wang Department of Computer Science Zhejiang University &Zihan Chen Department of Computer Science Zhejiang University &Haoyu Zhang Department of Computer Science Zhejiang University

###### Abstract

Venture capital \(VC\) investment decisions face distinct challenges, such as multi\-source heterogeneous data, non\-stationary time series, and the demand for explainable predictions in high\-stakes, low\-data settings\. To overcome these issues, we introduceFinInvest\-GTCN, a Graph\-Temporal\-Causal Network that redefines the task from content recommendation to quantitative risk\-return assessment\. This architecture combines a relational graph encoder to capture the investment ecosystem’s topologyarxiv\-2511\.17989;arxiv\-2511\.22078, a multi\-scale temporal fusion module to handle long\-term dependencies and non\-stationarityOganesianet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib42)\); Chenet al\.\([2025d](https://arxiv.org/html/2606.28933#bib.bib44)\), and a causal decision head that generates risk\-adjusted predictions with interpretable causal attributionsMahadevan \([2025](https://arxiv.org/html/2606.28933#bib.bib64)\); Parafitaet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib57)\)\. A core innovation is the Meta\-Causal Adaptation \(MCA\) strategy, which facilitates robust fine\-tuning for new, data\-scarce sectors by aligning updates with causally\-plausible structures derived from meta\-pretraining\. Comprehensive experiments on proprietary VC datasets show that FinInvest\-GTCN delivers state\-of\-the\-art results, markedly lowering the primary Risk\-Adjusted Mean Squared Error \(RA\-MSE\) to 2\.51 from a baseline of 3\.05 and boosting the cumulative return of a simulated portfolio by 18\.7%\. Ablation studies underscore the essential role of each component, while additional analyses confirm the model’s stability, interpretability, and enhanced adaptability\. This work pioneers a data\-driven, explainable framework for investment decision support\.

## 1Introduction

Financial investment decision\-making, particularly in venture capital \(VC\), constitutes a high\-stakes sequential prediction problem characterized by multi\-source heterogeneous data and complex inter\-asset relationshipsAkinfaderin and Subramanian \([2025](https://arxiv.org/html/2606.28933#bib.bib41)\); Liet al\.\([2025a](https://arxiv.org/html/2606.28933#bib.bib39)\); Yinet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib45)\)\. Practitioners must analyze time\-series data encompassing financial metrics, team dynamics, and market signals, while simultaneously accounting for the broader ecosystem topology—including competitive overlaps and shared investor networks—to forecast risk\-adjusted returns for potential assets\. Despite the abundance of historical data, accurately modeling these sequential, relational, and non\-stationary patterns to support robust and explainable decisions remains a formidable challenge\.Zhanget al\.\([2025b](https://arxiv.org/html/2606.28933#bib.bib1),[e](https://arxiv.org/html/2606.28933#bib.bib2),[c](https://arxiv.org/html/2606.28933#bib.bib3),[d](https://arxiv.org/html/2606.28933#bib.bib4),[a](https://arxiv.org/html/2606.28933#bib.bib5)\); Moet al\.\([2026](https://arxiv.org/html/2606.28933#bib.bib6)\); Yuet al\.\([2026](https://arxiv.org/html/2606.28933#bib.bib7)\); Zhanget al\.\([2026b](https://arxiv.org/html/2606.28933#bib.bib8)\)

Prior work in sequential recommendation and financial modeling exhibits critical limitations when applied to this domain\. Traditional tree\-based approaches rely heavily on manual feature engineering, which is labor\-intensive and fails to capture dynamic temporal dependenciesWanget al\.\([2025b](https://arxiv.org/html/2606.28933#bib.bib27)\); Kemper and Rostam\-Afschar \([2025](https://arxiv.org/html/2606.28933#bib.bib63)\); Hsiehet al\.\([2024](https://arxiv.org/html/2606.28933#bib.bib72)\)\. Standard neural sequence models, such as LSTMs or Transformers, process observation histories but neglect the rich relational structure among assetsChoiet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib23)\); Zenget al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib19)\); Renet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib70)\); Penget al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib71)\); Chenet al\.\([2025a](https://arxiv.org/html/2606.28933#bib.bib15)\); Youet al\.\([2026](https://arxiv.org/html/2606.28933#bib.bib14)\); Chenet al\.\([2025c](https://arxiv.org/html/2606.28933#bib.bib13)\); Zhanget al\.\([2026a](https://arxiv.org/html/2606.28933#bib.bib12)\); Zhaoet al\.\([2026](https://arxiv.org/html/2606.28933#bib.bib11)\); Huanget al\.\([2026](https://arxiv.org/html/2606.28933#bib.bib10)\); Chenet al\.\([2025b](https://arxiv.org/html/2606.28933#bib.bib9)\)\. Moreover, existing methods often optimize for simplistic objectives such as click\-through or conversion rates, which fundamentally misalign with the investment goal of maximizing risk\-adjusted returns\. These approaches also typically lack mechanisms for explainability and struggle to adapt to new sectors with scarce data, thereby hindering practical deployment in evolving markets\.

![Refer to caption](https://arxiv.org/html/2606.28933v1/figures/motivation/motivation_1_348e45f3.png)Figure 1:Paradigm shift from traditional investment decision methods to the proposed FinInvest\-GTCN framework\. Left: Existing methods suffer from fragmented approaches—manual feature engineering, isolated temporal modeling that ignores relational context, and lack of explainability\. Right: Our unified framework integrates graph\-based ecosystem topology modeling, multi\-scale temporal fusion, and causal attribution, with Meta\-Causal Adaptation \(MCA\) for robust fine\-tuning, directly aligning with the goal of maximizing risk\-adjusted returns\.To address these gaps, we proposeFinInvest\-GTCN, a novelGraph\-Temporal\-CausalNetwork for quantitative investment decision optimization\. Our framework introduces three key innovations: \(1\) a relational graph encoder that captures the investment ecosystem’s topology via Graph Attention Networks; \(2\) a multi\-scale temporal fusion module that learns patterns across different horizons using parallel Transformers; and \(3\) a causal decision head that predicts risk\-adjusted returns while enabling approximate causal attribution for model explanations\. Furthermore, we introduce aMeta\-CausalAdaptation \(MCA\) strategy that regularizes fine\-tuning towards causally\-plausible structures, thereby enhancing robustness in low\-data scenarios\. Collectively, these contributions provide a principled, end\-to\-end solution that transcends traditional recommendation paradigms to directly address the core challenges of investment analysis\.

Extensive experiments on proprietary and simulated venture capital datasets demonstrate the effectiveness of our approach\. FinInvest\-GTCN achieves a state\-of\-the\-art Risk\-Adjusted MSE of 2\.51, substantially outperforming strong baselines including repurposed sequential recommenders and financial factor models\. Ablation studies confirm the importance of each architectural component, and our MCA strategy reduces RA\-MSE by 16% compared to standard fine\-tuning on a new, data\-scarce sector\. In a simulated online A/B test, a portfolio constructed using our model’s rankings yields an 18\.7% cumulative return with lower volatility, underscoring its practical utility for real\-world deployment\.

The remainder of this paper is organized as follows\. Section[2](https://arxiv.org/html/2606.28933#S2)reviews related work\. Section[3](https://arxiv.org/html/2606.28933#S3)details the FinInvest\-GTCN architecture\. Section[4](https://arxiv.org/html/2606.28933#S4)presents the experimental setup and results\. Section[5](https://arxiv.org/html/2606.28933#S5)provides comprehensive ablation studies\. Section[6](https://arxiv.org/html/2606.28933#S6)offers supplementary experiments, and Section[7](https://arxiv.org/html/2606.28933#S7)concludes the paper\.

## 2Related Work

Our work lies at the intersection of three major research areas: graph\-based financial modeling, temporal sequence learning for investment analysis, and causal inference for explainable predictions\. We review each area and position our contributions accordingly\.

### 2\.1Graph Neural Networks in Finance

Graph neural networks \(GNNs\) have emerged as powerful tools for modeling relational structures in financial dataMaet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib16)\); Leeet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib17)\); Xuet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib18)\); Yeet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib25)\); Huet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib28)\)\. Early applications focused on stock market prediction by constructing graphs based on industry sectors or correlation matrices\. More recent works have extended GNNs to model complex inter\-company relationships, including supply chains, competitive dynamics, and shared investor networksLi and Fan \([2025](https://arxiv.org/html/2606.28933#bib.bib92)\); Kimet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib82)\); Kempinski and Kachman \([2025](https://arxiv.org/html/2606.28933#bib.bib93)\); Duan and Ji \([2025](https://arxiv.org/html/2606.28933#bib.bib85)\); Ashrafi and Kabir \([2025](https://arxiv.org/html/2606.28933#bib.bib84)\)\. Graph Attention Networks \(GATs\) have proven particularly effective by learning adaptive edge weights that capture varying relationship strengthsShit and Subudhi \([2025](https://arxiv.org/html/2606.28933#bib.bib86)\); Zhanget al\.\([2025f](https://arxiv.org/html/2606.28933#bib.bib88)\); Wanget al\.\([2025c](https://arxiv.org/html/2606.28933#bib.bib90)\); Liuet al\.\([2025c](https://arxiv.org/html/2606.28933#bib.bib83)\)\. However, existing approaches predominantly target public equity markets and short\-term price predictionPerekhodko and Ślepaczuk \([2025](https://arxiv.org/html/2606.28933#bib.bib31)\); Beniwal \([2025](https://arxiv.org/html/2606.28933#bib.bib38)\); Chiuet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib35)\), leaving the venture capital domain—characterized by sparse, irregular observations and longer investment horizons—largely unexplored\.

### 2\.2Temporal Modeling for Financial Time Series

Sequential modeling has long been central to financial forecasting\. Traditional approaches relied on autoregressive models and handcrafted technical indicators\. The advent of deep learning introduced LSTM and GRU architectures capable of capturing long\-range dependencies in price sequencesChirukiriet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib47)\); Perekhodko and Ślepaczuk \([2025](https://arxiv.org/html/2606.28933#bib.bib31)\)\. More recently, Transformer\-based models have demonstrated superior performance by leveraging self\-attention mechanisms to model complex temporal patternsChenet al\.\([2025d](https://arxiv.org/html/2606.28933#bib.bib44)\); Chapariniyaet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib50)\); Kiuet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib51)\); Yaoet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib52)\); Wanget al\.\([2025a](https://arxiv.org/html/2606.28933#bib.bib53)\)\. Multi\-scale temporal modeling, which processes time series at different granularities simultaneously, has shown promise in capturing both short\-term fluctuations and long\-term trendsFeng and Xue \([2025](https://arxiv.org/html/2606.28933#bib.bib48)\); Dinget al\.\([2025b](https://arxiv.org/html/2606.28933#bib.bib49)\); Oganesianet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib42)\)\. Despite these advances, most methods focus on single\-asset prediction and do not integrate the relational context essential for venture capital analysis, where an asset’s prospects depend heavily on its ecosystem positionAlduaiset al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib91)\); Yinet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib45)\)\.

### 2\.3Causal Inference and Explainability in Machine Learning

As machine learning models are increasingly deployed in high\-stakes financial decisions, the demand for explainability has intensified\. Post\-hoc explanation methods such as SHAP and LIME provide feature attributions but lack causal groundingParafitaet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib57)\); Zhang and Cai \([2025](https://arxiv.org/html/2606.28933#bib.bib65)\)\. Causal inference frameworks offer principled approaches to understanding model behavior through interventional reasoningAkbaret al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib55)\); Kubota and Sugasawa \([2025](https://arxiv.org/html/2606.28933#bib.bib58)\); Comptonet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib59)\); Byambadalaiet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib60)\); Jinet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib61)\)\. Recent work has explored integrating causal structures into neural networks to improve both robustness and interpretabilityLiuet al\.\([2025b](https://arxiv.org/html/2606.28933#bib.bib54)\); Mahadevan \([2025](https://arxiv.org/html/2606.28933#bib.bib64)\); Liuet al\.\([2025a](https://arxiv.org/html/2606.28933#bib.bib56)\)\. In the investment domain, explainability is particularly critical for regulatory compliance and building investor trust\. However, existing financial models rarely incorporate causal reasoning, and methods for generating causal\-style explanations for complex graph\-temporal architectures remain underdeveloped\.

### 2\.4Transfer Learning and Domain Adaptation in Finance

Adapting pre\-trained models to new domains with limited data is a persistent challenge\. Meta\-learning approaches, which learn to learn from a distribution of tasks, have shown success in few\-shot scenariosGuanet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib76)\); Jianget al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib75)\); Yunet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib77)\); Honget al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib66)\)\. In finance, transfer learning has been applied to adapt models across markets or asset classesVuet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib32)\); Liet al\.\([2025c](https://arxiv.org/html/2606.28933#bib.bib34)\), but standard fine\-tuning often leads to overfitting when target domain data is scarce\. Parameter\-efficient fine\-tuning methods like LoRA reduce this risk but do not explicitly preserve domain\-invariant structuresJeonet al\.\([2025a](https://arxiv.org/html/2606.28933#bib.bib68),[b](https://arxiv.org/html/2606.28933#bib.bib73)\); Dinget al\.\([2025a](https://arxiv.org/html/2606.28933#bib.bib78)\)\. Our Meta\-Causal Adaptation strategy addresses this gap by regularizing adaptation towards causally\-plausible structures learned during pre\-training\.

In summary, while significant progress has been made in graph\-based modeling, temporal sequence learning, and causal inference individually, no existing framework integrates these capabilities for venture capital decision optimization\. Our work addresses this gap by proposing FinInvest\-GTCN, which synergistically combines relational graph encoding, multi\-scale temporal fusion, and causal attribution within a unified architecture specifically designed for risk\-aware, explainable investment analysis\.

## 3Methodology

We introduceFinInvest\-GTCN, a novelGraph\-Temporal\-CausalNetwork designed for robust venture capital decision optimization\. Our approach fundamentally reframes the task from content recommendation to quantitative investment analysis, systematically addressing the challenges posed by multi\-source heterogeneous data, non\-stationary financial time series, and the need for explainability in high\-stakes, low\-data environments\.

The design of FinInvest\-GTCN is guided by three core principles:

1. \(i\)Structural Awareness: Investment outcomes are governed not only by intrinsic asset features but also by the topological position of an asset within the broader ecosystem\. Our architecture must explicitly encode and reason over this relational structure\.
2. \(ii\)Multi\-Horizon Temporal Coherence: Financial time series are inherently non\-stationary, exhibiting patterns at disparate scales—from quarterly earnings fluctuations to multi\-year industry cycles\. An effective model must simultaneously capture and adaptively fuse these multi\-scale temporal signals\.
3. \(iii\)Causal Interpretability under Uncertainty: In high\-stakes investment contexts, both regulators and practitioners demand explanations grounded in*causal*reasoning rather than mere statistical correlation\. The model must natively provide risk\-calibrated predictions alongside interpretable causal attributions\.

Below, we first formalize the problem \(§[3\.1](https://arxiv.org/html/2606.28933#S3.SS1)\), then present the three core architectural modules \(§[3\.2](https://arxiv.org/html/2606.28933#S3.SS2)–§[3\.4](https://arxiv.org/html/2606.28933#S3.SS4)\), followed by the joint training objective \(§[3\.5](https://arxiv.org/html/2606.28933#S3.SS5)\), the Meta\-Causal Adaptation strategy \(§[3\.6](https://arxiv.org/html/2606.28933#S3.SS6)\), and a theoretical analysis \(§[3\.7](https://arxiv.org/html/2606.28933#S3.SS7)\)\.

### 3\.1Problem Formulation & Data Representation

We begin by formalizing the investment decision task as a structured prediction problem over a dynamic heterogeneous information network\.

###### Definition 1\(Investment Decision Environment\)\.

An investment decision environment is defined by the tupleℰ=⟨ℐ,𝒥,𝒢,𝒪,𝒴⟩\\mathcal\{E\}=\\langle\\mathcal\{I\},\\mathcal\{J\},\\mathcal\{G\},\\mathcal\{O\},\\mathcal\{Y\}\\rangle, whereℐ\\mathcal\{I\}is the set of investor entities,𝒥\\mathcal\{J\}is the universe of investable assets,𝒢\\mathcal\{G\}is a dynamic relational graph,𝒪\\mathcal\{O\}denotes multi\-source observation streams, and𝒴\\mathcal\{Y\}captures the risk\-return outcome space\.

Dynamic Asset\-Relation Graph\.The relational structure is represented as a dynamic graph𝒢t=\(𝒱t,ℰt,ℛ\)\\mathcal\{G\}\_\{t\}=\(\\mathcal\{V\}\_\{t\},\\mathcal\{E\}\_\{t\},\\mathcal\{R\}\)at time steptt\. The node set𝒱t=\{vj\}j∈𝒥t\\mathcal\{V\}\_\{t\}=\\\{v\_\{j\}\\\}\_\{j\\in\\mathcal\{J\}\_\{t\}\}represents active assets, and the edge setℰt⊂𝒱t×ℛ×𝒱t\\mathcal\{E\}\_\{t\}\\subset\\mathcal\{V\}\_\{t\}\\times\\mathcal\{R\}\\times\\mathcal\{V\}\_\{t\}encodes typed relational contexts\. The relation type setℛ=\{rcompete,rsupply,rinvest,rsector\}\\mathcal\{R\}=\\\{r\_\{\\text\{compete\}\},r\_\{\\text\{supply\}\},r\_\{\\text\{invest\}\},r\_\{\\text\{sector\}\}\\\}captures competitive overlaps, supply\-chain linkages, shared investor ties, and sector co\-membership, respectively\. Each nodevjv\_\{j\}carries an initial feature vector𝐯j∈ℝdv\\mathbf\{v\}\_\{j\}\\in\\mathbb\{R\}^\{d\_\{v\}\}derived from static firmographic attributes \(e\.g\., founding year, geography, patent count\)\. This multi\-relational graph formulation, inspired by recent advances in heterogeneous graph representation learningMaet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib16)\); Leeet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib17)\); Xuet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib18)\), enables the model to capture the rich, typed topology of real\-world investment ecosystems that homogeneous graph approaches fail to represent\.

Multi\-Source Observation Sequence\.For each assetjjobserved by investorii, we compile a time\-ordered sequence of multi\-source feature vectors up to decision timeτ\\tau:

Oi,j=\[𝐱j,τ−L\+1,𝐱j,τ−L\+2,…,𝐱j,τ\]∈ℝL×dxO\_\{i,j\}=\[\\mathbf\{x\}\_\{j,\\tau\-L\+1\},\\mathbf\{x\}\_\{j,\\tau\-L\+2\},\\ldots,\\mathbf\{x\}\_\{j,\\tau\}\]\\in\\mathbb\{R\}^\{L\\times d\_\{x\}\}\(1\)whereLLis the lookback window length\. Each feature vector is a concatenation of heterogeneous source embeddings:

𝐱j,t=\[𝐟j,tfin⏟financial​‖𝐟j,tteam⏟human capital‖​𝐟j,tmkt⏟market signal∥𝐦t⏟macro\]\\mathbf\{x\}\_\{j,t\}=\\left\[\\underbrace\{\\mathbf\{f\}\_\{j,t\}^\{\\text\{fin\}\}\}\_\{\\text\{financial\}\}\\,\\\|\\,\\underbrace\{\\mathbf\{f\}\_\{j,t\}^\{\\text\{team\}\}\}\_\{\\text\{human capital\}\}\\,\\\|\\,\\underbrace\{\\mathbf\{f\}\_\{j,t\}^\{\\text\{mkt\}\}\}\_\{\\text\{market signal\}\}\\,\\\|\\,\\underbrace\{\\mathbf\{m\}\_\{t\}\}\_\{\\text\{macro\}\}\\right\]\(2\)Here,𝐟j,tfin∈ℝdf\\mathbf\{f\}\_\{j,t\}^\{\\text\{fin\}\}\\in\\mathbb\{R\}^\{d\_\{f\}\}encodes financial metrics \(revenue growth, burn rate, unit economics\),𝐟j,tteam∈ℝdh\\mathbf\{f\}\_\{j,t\}^\{\\text\{team\}\}\\in\\mathbb\{R\}^\{d\_\{h\}\}captures human capital dynamics \(executive changes, hiring velocity, key\-person dependencies\),𝐟j,tmkt∈ℝdm\\mathbf\{f\}\_\{j,t\}^\{\\text\{mkt\}\}\\in\\mathbb\{R\}^\{d\_\{m\}\}represents market\-level signals \(competitor activity, sector momentum, deal flow volume\), and𝐦t∈ℝdμ\\mathbf\{m\}\_\{t\}\\in\\mathbb\{R\}^\{d\_\{\\mu\}\}denotes exogenous macroeconomic factors \(interest rates, public market indices, credit spreads\)\. To handle the heterogeneity and variable dimensionality of raw sources, each sub\-vector is first processed through a source\-specific projection layerϕs:ℝdsraw→ℝds\\phi\_\{s\}:\\mathbb\{R\}^\{d\_\{s\}^\{\\text\{raw\}\}\}\\to\\mathbb\{R\}^\{d\_\{s\}\}before concatenation\.

Risk\-Adjusted Outcome Space\.The prediction target for each investor\-asset\-time triplet\(i,j,τ\)\(i,j,\\tau\)is a bivariate outcome capturing both the expected return and the associated uncertainty over the investment horizon\[τ,τ\+Δ​T\]\[\\tau,\\tau\+\\Delta T\]:

y^i,j=\(r^i,j,σ^i,j\)∈ℝ×ℝ\>0\\hat\{y\}\_\{i,j\}=\(\\hat\{r\}\_\{i,j\},\\;\\hat\{\\sigma\}\_\{i,j\}\)\\in\\mathbb\{R\}\\times\\mathbb\{R\}\_\{\>0\}\(3\)wherer^i,j\\hat\{r\}\_\{i,j\}is the predicted return \(e\.g\., an IRR proxy or valuation uplift\) andσ^i,j\\hat\{\\sigma\}\_\{i,j\}is the predicted risk score \(aleatoric volatility\)\. This bivariate formulation represents a fundamental departure from conventional recommendation systems that output scalar relevance scores \(e\.g\., CTR/CVR\), directly aligning the model output with the Markowitz mean\-variance optimization paradigm that underpins modern portfolio theory\.

### 3\.2Module I: Relational Graph Encoder

The first core module of FinInvest\-GTCN embeds each asset within the context of its relational ecosystem\. The central insight is that an asset’s investment potential depends critically on its*position*in the broader network: a startup facing concentrated competition from well\-funded incumbents occupies a fundamentally different risk profile than one in a nascent, under\-served niche—even if their intrinsic financial metrics are identical\.

The overall architecture of FinInvest\-GTCN is illustrated in Figure[2](https://arxiv.org/html/2606.28933#S3.F2)\.

![Refer to caption](https://arxiv.org/html/2606.28933v1/figures/algorithm/algorithm_1_7dae9824.png)Figure 2:High\-level architecture of FinInvest\-GTCN\. The framework integrates three core modules: \(1\) Relational Graph Encoder \(GAT\) for modeling the investment ecosystem topology, \(2\) Multi\-Scale Temporal Fusion \(MST\-Former\) with parallel encoders for different temporal scales, and \(3\) Causal Decision Head for risk\-adjusted return prediction with causal attribution\. The Meta\-Causal Adaptation \(MCA\) module enables robust fine\-tuning for new sectors via KL\-divergence regularization on causal structure priors\.#### 3\.2\.1Multi\-Relational Graph Attention Mechanism

Given the typed nature of edges in𝒢t\\mathcal\{G\}\_\{t\}, we extend the standard Graph Attention Network \(GAT\) to a multi\-relational variant that computes relation\-specific attention coefficientsDuan and Ji \([2025](https://arxiv.org/html/2606.28933#bib.bib85)\); Shit and Subudhi \([2025](https://arxiv.org/html/2606.28933#bib.bib86)\)\. For a target asset nodevjv\_\{j\}, the graph\-informed representation is computed via multi\-head, relation\-aware attention:

𝐡jG=∥m=1MGσ\(∑r∈ℛ∑k∈𝒩r​\(j\)αj​k\(r,m\)𝐖r\(m\)𝐯k\)\\mathbf\{h\}\_\{j\}^\{G\}=\\Big\\\|\_\{m=1\}^\{M\_\{G\}\}\\sigma\\left\(\\sum\_\{r\\in\\mathcal\{R\}\}\\sum\_\{k\\in\\mathcal\{N\}\_\{r\}\(j\)\}\\alpha\_\{jk\}^\{\(r,m\)\}\\mathbf\{W\}\_\{r\}^\{\(m\)\}\\mathbf\{v\}\_\{k\}\\right\)\(4\)
where∥\\\|denotes multi\-head concatenation acrossMGM\_\{G\}attention heads, and𝒩r​\(j\)\\mathcal\{N\}\_\{r\}\(j\)is the set of neighbors ofjjconnected by relation typerr\. Crucially, each relation typer∈ℛr\\in\\mathcal\{R\}is assigned a dedicated projection matrix𝐖r\(m\)∈ℝdG/MG×dv\\mathbf\{W\}\_\{r\}^\{\(m\)\}\\in\\mathbb\{R\}^\{d\_\{G\}/M\_\{G\}\\times d\_\{v\}\}, enabling the model to learn*distinct transformation semantics*for competitive, supply\-chain, investor, and sector relationships\. The multi\-relational attention coefficient is computed as:

αj​k\(r,m\)=exp⁡\(LeakyReLU​\(𝐚r\(m\)⊤​\[𝐖r\(m\)​𝐯j∥𝐖r\(m\)​𝐯k\]\)\)∑r′∈ℛ∑l∈𝒩r′​\(j\)exp⁡\(LeakyReLU​\(𝐚r′\(m\)⊤​\[𝐖r′\(m\)​𝐯j∥𝐖r′\(m\)​𝐯l\]\)\)\\alpha\_\{jk\}^\{\(r,m\)\}=\\frac\{\\exp\\\!\\Big\(\\text\{LeakyReLU\}\\\!\\big\(\\mathbf\{a\}\_\{r\}^\{\(m\)\\top\}\[\\mathbf\{W\}\_\{r\}^\{\(m\)\}\\mathbf\{v\}\_\{j\}\\,\\\|\\,\\mathbf\{W\}\_\{r\}^\{\(m\)\}\\mathbf\{v\}\_\{k\}\]\\big\)\\Big\)\}\{\\sum\_\{r^\{\\prime\}\\in\\mathcal\{R\}\}\\sum\_\{l\\in\\mathcal\{N\}\_\{r^\{\\prime\}\}\(j\)\}\\exp\\\!\\Big\(\\text\{LeakyReLU\}\\\!\\big\(\\mathbf\{a\}\_\{r^\{\\prime\}\}^\{\(m\)\\top\}\[\\mathbf\{W\}\_\{r^\{\\prime\}\}^\{\(m\)\}\\mathbf\{v\}\_\{j\}\\,\\\|\\,\\mathbf\{W\}\_\{r^\{\\prime\}\}^\{\(m\)\}\\mathbf\{v\}\_\{l\}\]\\big\)\\Big\)\}\(5\)
where𝐚r\(m\)∈ℝ2​dG/MG\\mathbf\{a\}\_\{r\}^\{\(m\)\}\\in\\mathbb\{R\}^\{2d\_\{G\}/M\_\{G\}\}is a relation\-specific attention vector for headmm\. Notably, the normalization in Eq\. \([5](https://arxiv.org/html/2606.28933#S3.E5)\) operates*across all relation types jointly*, inducing a natural competition between different relational channels and enabling the model to learn which relationship types are most informative for a given asset\.

#### 3\.2\.2Neighborhood Normalization and Self\-Loop Augmentation

To ensure numerical stability and prevent the graph encoder from over\-smoothing representations in densely connected subgraphs, we augment the message\-passing procedure with two additional mechanisms\. First, each nodevjv\_\{j\}includes a self\-loop of a special identity relation type, guaranteeing that the original node features are preserved:

𝒩~​\(j\)=\(𝒩​\(j\)×ℛ\)∪\{\(j,rself\)\}\\tilde\{\\mathcal\{N\}\}\(j\)=\\big\(\\mathcal\{N\}\(j\)\\times\\mathcal\{R\}\\big\)\\cup\\\{\(j,r\_\{\\text\{self\}\}\)\\\}\(6\)Second, we apply a symmetric normalization factorD^j​k−1/2\\hat\{D\}\_\{jk\}^\{\-1/2\}based on node degrees to prevent gradient explosion in high\-degree nodes, followingLi and Fan \([2025](https://arxiv.org/html/2606.28933#bib.bib92)\); Kimet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib82)\):

α~j​k\(r,m\)=αj​k\(r,m\)\|𝒩~​\(j\)\|⋅\|𝒩~​\(k\)\|\\tilde\{\\alpha\}\_\{jk\}^\{\(r,m\)\}=\\frac\{\\alpha\_\{jk\}^\{\(r,m\)\}\}\{\\sqrt\{\|\\tilde\{\\mathcal\{N\}\}\(j\)\|\\cdot\|\\tilde\{\\mathcal\{N\}\}\(k\)\|\}\}\(7\)

#### 3\.2\.3Stacked Graph Encoding with Residual Connections

We stackKGK\_\{G\}layers of the multi\-relational GAT with residual connections to allow information propagation across multi\-hop neighborhoods while mitigating the over\-smoothing problemZhanget al\.\([2025f](https://arxiv.org/html/2606.28933#bib.bib88)\); Wanget al\.\([2025c](https://arxiv.org/html/2606.28933#bib.bib90)\):

𝐡jG,\(ℓ\+1\)=LayerNorm​\(𝐡jG,\(ℓ\)\+Dropout​\(GAT\(ℓ\)​\(𝐡jG,\(ℓ\),𝒢t\)\)\),ℓ=0,…,KG−1\\mathbf\{h\}\_\{j\}^\{G,\(\\ell\+1\)\}=\\text\{LayerNorm\}\\\!\\Big\(\\mathbf\{h\}\_\{j\}^\{G,\(\\ell\)\}\+\\text\{Dropout\}\\big\(\\text\{GAT\}^\{\(\\ell\)\}\(\\mathbf\{h\}\_\{j\}^\{G,\(\\ell\)\},\\mathcal\{G\}\_\{t\}\)\\big\)\\Big\),\\quad\\ell=0,\\ldots,K\_\{G\}\-1\(8\)where𝐡jG,\(0\)=Linear​\(𝐯j\)\\mathbf\{h\}\_\{j\}^\{G,\(0\)\}=\\text\{Linear\}\(\\mathbf\{v\}\_\{j\}\)and the final graph representation is𝐡jG≔𝐡jG,\(KG\)∈ℝdG\\mathbf\{h\}\_\{j\}^\{G\}\\coloneqq\\mathbf\{h\}\_\{j\}^\{G,\(K\_\{G\}\)\}\\in\\mathbb\{R\}^\{d\_\{G\}\}\. The residual\-LayerNorm design ensures stable gradient flow while the multi\-hop aggregation enables the model to capture indirect ecosystem effects \(e\.g\., a competitor’s supplier receiving major funding\)\.

### 3\.3Module II: Multi\-Scale Temporal Fusion \(MST\-Former\)

Financial time series are inherently non\-stationary and exhibit patterns at multiple temporal scales: short\-term momentum signals \(e\.g\., recent quarterly revenue acceleration\), medium\-term cyclical patterns \(e\.g\., annual hiring cycles\), and long\-term structural trends \(e\.g\., multi\-year market positioning shifts\)\. A single\-scale temporal encoder conflates these disparate dynamics, limiting its capacity to disentangle actionable signals from noise\. Our MST\-Former module addresses this by processing the observation sequence at multiple granularities through parallel, scale\-specialized Transformer encoders, followed by an adaptive gating mechanism that fuses the multi\-scale representations\.

#### 3\.3\.1Scale\-Specific Input Preparation

Given the observation sequenceOi,j=\[𝐱j,1,…,𝐱j,L\]O\_\{i,j\}=\[\\mathbf\{x\}\_\{j,1\},\\ldots,\\mathbf\{x\}\_\{j,L\}\]\(we omit the offsetτ−L\+1\\tau\-L\+1for notational clarity\), we constructSSscale\-specific input sequences\. For scales∈\{1,…,S\}s\\in\\\{1,\\ldots,S\\\}, the input is derived by selecting the most recentLsL\_\{s\}time steps \(L1<L2<⋯<LS=LL\_\{1\}<L\_\{2\}<\\cdots<L\_\{S\}=L\):

Oi,j\(s\)=\[𝐱j,L−Ls\+1,…,𝐱j,L\]∈ℝLs×dxO\_\{i,j\}^\{\(s\)\}=\[\\mathbf\{x\}\_\{j,L\-L\_\{s\}\+1\},\\ldots,\\mathbf\{x\}\_\{j,L\}\]\\in\\mathbb\{R\}^\{L\_\{s\}\\times d\_\{x\}\}\(9\)In our implementation, we employS=3S=3scales corresponding to approximately 4\-quarter \(short\-term\), 8\-quarter \(medium\-term\), and full\-history \(long\-term\) lookback windows, i\.e\.,L1=4,L2=8,L3=LL\_\{1\}=4,L\_\{2\}=8,L\_\{3\}=L\. This design ensures that the short\-term encoder focuses on recent dynamics without being diluted by distant, potentially obsolete signals, while the long\-term encoder captures the full evolutionary trajectory of the asset\.

#### 3\.3\.2Positional Encoding with Temporal Awareness

Unlike standard NLP tasks where token positions are uniformly spaced, financial observations may have irregular temporal spacing \(e\.g\., missing quarterly reports\)\. We augment standard sinusoidal positional encodings with a learnable temporal embedding that encodes the absolute calendar time of each observation:

𝐱~j,t\(s\)=𝐱j,t\+𝐏𝐄​\(trel\)⏟relative position\+MLPcal​\(𝐜t\)⏟calendar embedding\\tilde\{\\mathbf\{x\}\}\_\{j,t\}^\{\(s\)\}=\\mathbf\{x\}\_\{j,t\}\+\\underbrace\{\\mathbf\{PE\}\(t\_\{\\text\{rel\}\}\)\}\_\{\\text\{relative position\}\}\+\\underbrace\{\\text\{MLP\}\_\{\\text\{cal\}\}\(\\mathbf\{c\}\_\{t\}\)\}\_\{\\text\{calendar embedding\}\}\(10\)wheretrelt\_\{\\text\{rel\}\}denotes the relative position within the scale\-specific window,𝐏𝐄​\(⋅\)\\mathbf\{PE\}\(\\cdot\)is the standard sinusoidal positional encoding, and𝐜t=\[quarter​\(t\),year​\(t\),is\_crisis​\(t\)\]\\mathbf\{c\}\_\{t\}=\[\\text\{quarter\}\(t\),\\text\{year\}\(t\),\\text\{is\\\_crisis\}\(t\)\]is a calendar feature vector that embeds seasonal effects and macroeconomic regime indicators\. This*temporal\-aware*positional encoding enables the model to distinguish between, e\.g\., a revenue drop in Q4 \(a known seasonal pattern\) and the same drop in Q1 \(a potentially alarming signal\), thereby improving non\-stationarity handling\.

#### 3\.3\.3Scale\-Specialized Transformer Encoders

Each scalessis processed by an independent Transformer encoder with its own parameter setΘ\(s\)\\Theta^\{\(s\)\}, consisting ofKTK\_\{T\}layers of multi\-head self\-attention and position\-wise feed\-forward networksKiuet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib51)\); Wanget al\.\([2025a](https://arxiv.org/html/2606.28933#bib.bib53)\):

𝐇j\(s\)=TransformerEncoder\(s\)​\(O~i,j\(s\);Θ\(s\)\)∈ℝLs×dT\\mathbf\{H\}\_\{j\}^\{\(s\)\}=\\text\{TransformerEncoder\}^\{\(s\)\}\\big\(\\tilde\{O\}\_\{i,j\}^\{\(s\)\};\\,\\Theta^\{\(s\)\}\\big\)\\in\\mathbb\{R\}^\{L\_\{s\}\\times d\_\{T\}\}\(11\)
To further specialize each encoder to its designated temporal resolution, we introduce ascale\-specific causal attention mask𝐌\(s\)\\mathbf\{M\}^\{\(s\)\}\. For scalesswith windowLsL\_\{s\}, the attention mask restricts each query position to attend only to positions within its effective window:

Mp​q\(s\)=\{0if​0≤p−q≤Ws−∞otherwiseM\_\{pq\}^\{\(s\)\}=\\begin\{cases\}0&\\text\{if \}0\\leq p\-q\\leq W\_\{s\}\\\\ \-\\infty&\\text\{otherwise\}\\end\{cases\}\(12\)whereWsW\_\{s\}is the effective local attention span of scaless\. For the short\-term encoder,WsW\_\{s\}is set to a small value \(e\.g\.,W1=4W\_\{1\}=4\) to enforce local focus; for the long\-term encoder,WS=LW\_\{S\}=L\(i\.e\., full attention\)\. This inductive bias explicitly prevents the short\-term encoder from “leaking” attention to distant, potentially irrelevant history, and ensures that each encoder specializes in its intended temporal resolution\.

The summary representation for each scale is obtained through attentive pooling rather than simple mean pooling, yielding a more expressive aggregation:

𝐇¯j\(s\)=∑t=1Lsβt\(s\)​𝐇j,t\(s\),βt\(s\)=exp⁡\(𝐪s⊤​𝐇j,t\(s\)/dT\)∑t′=1Lsexp⁡\(𝐪s⊤​𝐇j,t′\(s\)/dT\)\\bar\{\\mathbf\{H\}\}\_\{j\}^\{\(s\)\}=\\sum\_\{t=1\}^\{L\_\{s\}\}\\beta\_\{t\}^\{\(s\)\}\\mathbf\{H\}\_\{j,t\}^\{\(s\)\},\\quad\\beta\_\{t\}^\{\(s\)\}=\\frac\{\\exp\\\!\\big\(\\mathbf\{q\}\_\{s\}^\{\\top\}\\mathbf\{H\}\_\{j,t\}^\{\(s\)\}/\\sqrt\{d\_\{T\}\}\\big\)\}\{\\sum\_\{t^\{\\prime\}=1\}^\{L\_\{s\}\}\\exp\\\!\\big\(\\mathbf\{q\}\_\{s\}^\{\\top\}\\mathbf\{H\}\_\{j,t^\{\\prime\}\}^\{\(s\)\}/\\sqrt\{d\_\{T\}\}\\big\)\}\(13\)where𝐪s∈ℝdT\\mathbf\{q\}\_\{s\}\\in\\mathbb\{R\}^\{d\_\{T\}\}is a learnable query vector for scaless\. This mechanism allows the model to attend selectively to the most informative time steps at each scale\.

#### 3\.3\.4Adaptive Gated Fusion

TheSSscale\-specific summary representations are fused via a learned gating mechanism that adaptively weights each temporal scale based on the current inputFeng and Xue \([2025](https://arxiv.org/html/2606.28933#bib.bib48)\); Yaoet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib52)\):

𝐡jT=∑s=1Sgs⋅𝐇¯j\(s\),𝐠=softmax​\(𝐖g​CONCAT​\(𝐇¯j\(1\),…,𝐇¯j\(S\)\)\+𝐛g\)\\mathbf\{h\}\_\{j\}^\{T\}=\\sum\_\{s=1\}^\{S\}g\_\{s\}\\cdot\\bar\{\\mathbf\{H\}\}\_\{j\}^\{\(s\)\},\\quad\\mathbf\{g\}=\\text\{softmax\}\\\!\\Big\(\\mathbf\{W\}\_\{g\}\\,\\text\{CONCAT\}\\big\(\\bar\{\\mathbf\{H\}\}\_\{j\}^\{\(1\)\},\\ldots,\\bar\{\\mathbf\{H\}\}\_\{j\}^\{\(S\)\}\\big\)\+\\mathbf\{b\}\_\{g\}\\Big\)\(14\)where𝐖g∈ℝS×S​dT\\mathbf\{W\}\_\{g\}\\in\\mathbb\{R\}^\{S\\times Sd\_\{T\}\}and𝐛g∈ℝS\\mathbf\{b\}\_\{g\}\\in\\mathbb\{R\}^\{S\}are learnable parameters\. The gate vector𝐠∈ΔS−1\\mathbf\{g\}\\in\\Delta^\{S\-1\}\(the probability simplex\) acts as a*soft attention over temporal scales*, enabling the model to dynamically prioritize short\-term signals during volatile market conditions and long\-term trends during stable periods—a behavior we empirically observe and validate in §[5](https://arxiv.org/html/2606.28933#S5)\.

### 3\.4Module III: Causal Decision Head

The Causal Decision Head constitutes the final, and arguably most distinctive, module of FinInvest\-GTCN\. It integrates the graph\-topological and temporal representations into a unified prediction space, producing both risk\-adjusted return forecasts and interpretable causal attributions\. This design addresses two critical demands simultaneously: the quantitative need for accurate, risk\-calibrated predictions and the qualitative need for explanations that support human decision\-making and regulatory compliance\.

#### 3\.4\.1Cross\-Modal Fusion via Gated Residual Connection

Rather than simply adding or concatenating the graph and temporal representations, we employ a*gated residual fusion*mechanism that learns the optimal integration of structural ecosystem knowledge and temporal dynamics:

𝐡j=LayerNorm​\(γ⋅𝐖G​𝐡jG\+\(1−γ\)⋅𝐖T​𝐡jT\)\\mathbf\{h\}\_\{j\}=\\text\{LayerNorm\}\\\!\\Big\(\\gamma\\cdot\\mathbf\{W\}\_\{G\}\\mathbf\{h\}\_\{j\}^\{G\}\+\(1\-\\gamma\)\\cdot\\mathbf\{W\}\_\{T\}\\mathbf\{h\}\_\{j\}^\{T\}\\Big\)\(15\)where𝐖G∈ℝd×dG\\mathbf\{W\}\_\{G\}\\in\\mathbb\{R\}^\{d\\times d\_\{G\}\}and𝐖T∈ℝd×dT\\mathbf\{W\}\_\{T\}\\in\\mathbb\{R\}^\{d\\times d\_\{T\}\}are projection matrices, andγ∈\(0,1\)\\gamma\\in\(0,1\)is a learned scalar gate computed as:

γ=σ​\(𝐰γ⊤​\[𝐡jG∥𝐡jT\]\+bγ\)\\gamma=\\sigma\\\!\\Big\(\\mathbf\{w\}\_\{\\gamma\}^\{\\top\}\\big\[\\mathbf\{h\}\_\{j\}^\{G\}\\,\\\|\\,\\mathbf\{h\}\_\{j\}^\{T\}\\big\]\+b\_\{\\gamma\}\\Big\)\(16\)This gate allows the model to adaptively balance the relative importance of topological and temporal signals on a per\-sample basis\. For assets in highly interconnected ecosystems \(e\.g\., platform companies\), the model may upweight the graph representation; for assets with strong intrinsic time\-series signals \(e\.g\., revenue\-focused SaaS\), it may prioritize the temporal encoding\.

#### 3\.4\.2Risk\-Adjusted Return Prediction with Heteroscedastic Uncertainty

The fused representation is passed through a dual\-head MLP to produce the bivariate prediction:

r^i,j\\displaystyle\\hat\{r\}\_\{i,j\}=MLPr​\(𝐡j\)=𝐰r⊤​ReLU​\(𝐖r,1​𝐡j\+𝐛r,1\)\+br\\displaystyle=\\text\{MLP\}\_\{r\}\(\\mathbf\{h\}\_\{j\}\)=\\mathbf\{w\}\_\{r\}^\{\\top\}\\text\{ReLU\}\(\\mathbf\{W\}\_\{r,1\}\\mathbf\{h\}\_\{j\}\+\\mathbf\{b\}\_\{r,1\}\)\+b\_\{r\}\(17\)σ^i,j\\displaystyle\\hat\{\\sigma\}\_\{i,j\}=softplus​\(MLPσ​\(𝐡j\)\)=log⁡\(1\+exp⁡\(𝐰σ⊤​ReLU​\(𝐖σ,1​𝐡j\+𝐛σ,1\)\+bσ\)\)\\displaystyle=\\text\{softplus\}\\\!\\big\(\\text\{MLP\}\_\{\\sigma\}\(\\mathbf\{h\}\_\{j\}\)\\big\)=\\log\\\!\\Big\(1\+\\exp\\\!\\big\(\\mathbf\{w\}\_\{\\sigma\}^\{\\top\}\\text\{ReLU\}\(\\mathbf\{W\}\_\{\\sigma,1\}\\mathbf\{h\}\_\{j\}\+\\mathbf\{b\}\_\{\\sigma,1\}\)\+b\_\{\\sigma\}\\big\)\\Big\)\(18\)where the softplus activation in Eq\. \([18](https://arxiv.org/html/2606.28933#S3.E18)\) ensures strict positivity of the predicted riskσ^i,j\>0\\hat\{\\sigma\}\_\{i,j\}\>0\. Critically, the return and risk heads share the fused representation𝐡j\\mathbf\{h\}\_\{j\}but have independent parameters, enabling the model to capture the nuanced relationship between expected return and uncertainty \(i\.e\., an asset may have high expected return*and*high risk\)\.

#### 3\.4\.3Interventional Causal Attribution for Explainability

To provide explanations that go beyond correlational feature importance \(as in SHAP or LIME\), we introduce an Interventional Causal Attribution \(ICA\) mechanism grounded in the potential outcomes frameworkZhang and Cai \([2025](https://arxiv.org/html/2606.28933#bib.bib65)\); Akbaret al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib55)\)\. The core idea is to estimate the*causal effect*of a hypothesized factorzzon the model’s prediction by constructing a counterfactual representation\.

##### Step 1: Factor\-Aligned Attention\.

For a given causal factorzz\(e\.g\., a regulatory event, a key competitor’s funding round, or a macroeconomic shock\), we first compute a factor\-aligned attention score over the temporal representation:

𝜶z=sigmoid​\(𝐡jT⊤⋅Embed​\(z\)dT\)\\boldsymbol\{\\alpha\}\_\{z\}=\\text\{sigmoid\}\\\!\\Big\(\\frac\{\\mathbf\{h\}\_\{j\}^\{T\\top\}\\cdot\\text\{Embed\}\(z\)\}\{\\sqrt\{d\_\{T\}\}\}\\Big\)\(19\)whereEmbed​\(z\)∈ℝdT\\text\{Embed\}\(z\)\\in\\mathbb\{R\}^\{d\_\{T\}\}maps the factor identifier to the temporal representation space\. The scalar𝜶z∈\(0,1\)\\boldsymbol\{\\alpha\}\_\{z\}\\in\(0,1\)quantifies the degree to which the temporal representation encodes information related to factorzz\.

##### Step 2: Counterfactual Representation Construction\.

We construct a counterfactual fused representation by*surgically removing*the component of the representation that is aligned with factorzz:

𝐡jC​F​\(z\)=𝐡j−𝜶z⋅CrossAttn​\(𝐡j,Embed​\(z\),Embed​\(z\)\)⏟factor\-aligned subspace projection\\mathbf\{h\}\_\{j\}^\{CF\(z\)\}=\\mathbf\{h\}\_\{j\}\-\\boldsymbol\{\\alpha\}\_\{z\}\\cdot\\underbrace\{\\text\{CrossAttn\}\\\!\\big\(\\mathbf\{h\}\_\{j\},\\,\\text\{Embed\}\(z\),\\,\\text\{Embed\}\(z\)\\big\)\}\_\{\\text\{factor\-aligned subspace projection\}\}\(20\)whereCrossAttn​\(Q,K,V\)=softmax​\(Q​K⊤/d\)​V\\text\{CrossAttn\}\(Q,K,V\)=\\text\{softmax\}\(QK^\{\\top\}/\\sqrt\{d\}\)Vis a standard cross\-attention operator\. This construction can be interpreted as an*approximate do\-calculus intervention*: we estimater^​\(d​o​\(z:=absent\)\)\\hat\{r\}\(do\(z:=\\text\{absent\}\)\)by projecting out the factor\-aligned subspace from the learned representation\.

##### Step 3: Approximate Causal Effect Estimation\.

The Approximate Causal Effect \(ACE\) of factorzzon the predicted return is computed as the difference between the factual and counterfactual predictions:

Δ​r^i,j\(z\)=MLPr​\(𝐡j\)−MLPr​\(𝐡jC​F​\(z\)\)\\Delta\\hat\{r\}\_\{i,j\}^\{\(z\)\}=\\text\{MLP\}\_\{r\}\(\\mathbf\{h\}\_\{j\}\)\-\\text\{MLP\}\_\{r\}\(\\mathbf\{h\}\_\{j\}^\{CF\(z\)\}\)\(21\)
A large positiveΔ​r^i,j\(z\)\\Delta\\hat\{r\}\_\{i,j\}^\{\(z\)\}indicates that the model’s favorable prediction is*causally*attributable to factorzz; conversely, a large negative value signals thatzzdrives a pessimistic forecast\. For a set of candidate factors𝒵=\{z1,z2,…,zP\}\\mathcal\{Z\}=\\\{z\_\{1\},z\_\{2\},\\ldots,z\_\{P\}\\\}, the model produces a full causal attribution profile:

𝚫​𝒓^i,j=\[Δ​r^i,j\(z1\),…,Δ​r^i,j\(zP\)\]∈ℝP\\boldsymbol\{\\Delta\\hat\{r\}\}\_\{i,j\}=\\big\[\\Delta\\hat\{r\}\_\{i,j\}^\{\(z\_\{1\}\)\},\\ldots,\\Delta\\hat\{r\}\_\{i,j\}^\{\(z\_\{P\}\)\}\\big\]\\in\\mathbb\{R\}^\{P\}\(22\)This profile serves as a human\-readable explanation for each investment recommendation, enabling portfolio managers to understand*why*the model favors or disfavors a particular assetMahadevan \([2025](https://arxiv.org/html/2606.28933#bib.bib64)\); Parafitaet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib57)\)\.

###### Proposition 1\(Completeness of Causal Attribution\)\.

Under the assumption that the factor embeddings\{Embed​\(zp\)\}p=1P\\\{\\text\{Embed\}\(z\_\{p\}\)\\\}\_\{p=1\}^\{P\}span the full representation spaceℝdT\\mathbb\{R\}^\{d\_\{T\}\}, the sum of individual causal effects converges to the total prediction:∑p=1PΔ​r^i,j\(zp\)≈r^i,j−r^i,j\(∅\)\\sum\_\{p=1\}^\{P\}\\Delta\\hat\{r\}\_\{i,j\}^\{\(z\_\{p\}\)\}\\approx\\hat\{r\}\_\{i,j\}\-\\hat\{r\}\_\{i,j\}^\{\(\\emptyset\)\}, wherer^i,j\(∅\)\\hat\{r\}\_\{i,j\}^\{\(\\emptyset\)\}is the prediction under a null baseline representation\.

This completeness property ensures that the causal attributions collectively “explain” the model’s entire prediction, analogous to the efficiency property of Shapley values but with a causal rather than purely game\-theoretic foundation\.

### 3\.5Joint Training Objective

The model is trained end\-to\-end with a composite loss function comprising three terms that collectively enforce accurate risk\-calibrated predictions, high\-quality causal explanations, and structural regularization:

ℒtotal=ℒreturn⏟prediction\+μ​ℒcausal⏟explanation\+ν​ℒgraph⏟regularization\\mathcal\{L\}\_\{\\text\{total\}\}=\\underbrace\{\\mathcal\{L\}\_\{\\text\{return\}\}\}\_\{\\text\{prediction\}\}\+\\mu\\underbrace\{\\mathcal\{L\}\_\{\\text\{causal\}\}\}\_\{\\text\{explanation\}\}\+\\nu\\underbrace\{\\mathcal\{L\}\_\{\\text\{graph\}\}\}\_\{\\text\{regularization\}\}\(23\)
##### Risk\-Adjusted Return Loss \(Primary\)\.

The primary objective derives from the negative log\-likelihood of a heteroscedastic Gaussian model, penalizing prediction errors inversely proportional to the model’s predicted confidence:

ℒreturn=1\|𝒟\|​∑\(i,j\)∈𝒟\[\(ri,j−r^i,j\)2σ^i,j2\+log⁡σ^i,j2\]\\mathcal\{L\}\_\{\\text\{return\}\}=\\frac\{1\}\{\|\\mathcal\{D\}\|\}\\sum\_\{\(i,j\)\\in\\mathcal\{D\}\}\\left\[\\frac\{\(r\_\{i,j\}\-\\hat\{r\}\_\{i,j\}\)^\{2\}\}\{\\hat\{\\sigma\}\_\{i,j\}^\{2\}\}\+\\log\\hat\{\\sigma\}\_\{i,j\}^\{2\}\\right\]\(24\)This loss naturally balances two competing forces: the squared\-error numerator encourages accuracy, while the log\-variance term prevents the model from trivially inflatingσ^i,j\\hat\{\\sigma\}\_\{i,j\}to minimize the first term\. This aligns the training objective with the Markowitz risk\-return trade\-off at the core of portfolio optimization\.

##### Causal Alignment Loss \(Auxiliary\)\.

To improve the quality of causal attributions, we introduce a supervision signal that encourages alignment between model\-attributed causal effects and ground\-truth factor impacts derived from ex\-post analysis:

ℒcausal=1\|𝒟z\|​∑\(i,j\)∈𝒟z∑p=1P\(Δ​r^i,j\(zp\)−Δ​ri,j\(zp\)⁣∗\)2\\mathcal\{L\}\_\{\\text\{causal\}\}=\\frac\{1\}\{\|\\mathcal\{D\}\_\{z\}\|\}\\sum\_\{\(i,j\)\\in\\mathcal\{D\}\_\{z\}\}\\sum\_\{p=1\}^\{P\}\\left\(\\Delta\\hat\{r\}\_\{i,j\}^\{\(z\_\{p\}\)\}\-\\Delta r\_\{i,j\}^\{\(z\_\{p\}\)\*\}\\right\)^\{2\}\(25\)whereΔ​ri,j\(zp\)⁣∗\\Delta r\_\{i,j\}^\{\(z\_\{p\}\)\*\}denotes the ground\-truth effect of factorzpz\_\{p\}on assetjj’s outcome, estimated from ex\-post ablation studies or domain expert labels\. This loss is applied to a subset𝒟z⊂𝒟\\mathcal\{D\}\_\{z\}\\subset\\mathcal\{D\}for which ground\-truth factor effects are available\.

##### Graph Structure Regularization \(Auxiliary\)\.

To prevent the graph attention weights from degenerating into uniform distributions \(thereby losing structural information\), we impose an entropy regularization on the attention distribution:

ℒgraph=−1\|𝒱t\|​∑j∈𝒱t∑r∈ℛ∑k∈𝒩r​\(j\)α~j​k\(r\)​log⁡α~j​k\(r\)\\mathcal\{L\}\_\{\\text\{graph\}\}=\-\\frac\{1\}\{\|\\mathcal\{V\}\_\{t\}\|\}\\sum\_\{j\\in\\mathcal\{V\}\_\{t\}\}\\sum\_\{r\\in\\mathcal\{R\}\}\\sum\_\{k\\in\\mathcal\{N\}\_\{r\}\(j\)\}\\tilde\{\\alpha\}\_\{jk\}^\{\(r\)\}\\log\\tilde\{\\alpha\}\_\{jk\}^\{\(r\)\}\(26\)This negative\-entropy penalty encourages*sparse*, informative attention weights, promoting interpretable graph aggregation where each asset attends primarily to its most relevant neighbors rather than diffusing attention uniformly\.

### 3\.6Meta\-Causal Adaptation \(MCA\) for New Investment Sectors

Adapting to emerging investment sectors \(e\.g\., quantum computing, synthetic biology\) with extremely limited historical data represents a critical practical challengeGuanet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib76)\); Jianget al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib75)\)\. Standard fine\-tuning approaches suffer from catastrophic overfitting in such low\-data regimes, while parameter\-efficient methods like LoRAJeonet al\.\([2025b](https://arxiv.org/html/2606.28933#bib.bib73)\); Honget al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib66)\)reduce overfitting risk but do not explicitly leverage the*structural invariances*that transfer across sectors—namely, the causal mechanisms linking ecosystem dynamics to investment outcomes\.

We address this gap withMeta\-Causal Adaptation \(MCA\), a two\-phase adaptation strategy that combines meta\-learning with causal structural regularization\.

#### 3\.6\.1Phase I: Episodic Meta\-Pretraining

During meta\-pretraining, the model learns a shared initializationθmeta\\theta\_\{\\text\{meta\}\}that enables rapid adaptation to any sector\. We adopt a MAML\-style episodic frameworkYunet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib77)\); Jeonet al\.\([2025a](https://arxiv.org/html/2606.28933#bib.bib68)\): at each meta\-iteration, a source sectorSkS\_\{k\}is sampled, and the model is adapted viaNinnerN\_\{\\text\{inner\}\}gradient steps on a support set𝒟Sksup\\mathcal\{D\}\_\{S\_\{k\}\}^\{\\text\{sup\}\}, then evaluated on a query set𝒟Skqry\\mathcal\{D\}\_\{S\_\{k\}\}^\{\\text\{qry\}\}\. The meta\-objective is:

θmeta=arg​minθ​∑Sk∼p​\(𝒮\)ℒreturn​\(𝒟Skqry;θ−ηin​∇θℒreturn​\(𝒟Sksup;θ\)\)\\theta\_\{\\text\{meta\}\}=\\operatorname\*\{arg\\,min\}\_\{\\theta\}\\sum\_\{S\_\{k\}\\sim p\(\\mathcal\{S\}\)\}\\mathcal\{L\}\_\{\\text\{return\}\}\\\!\\big\(\\mathcal\{D\}\_\{S\_\{k\}\}^\{\\text\{qry\}\};\\;\\theta\-\\eta\_\{\\text\{in\}\}\\nabla\_\{\\theta\}\\mathcal\{L\}\_\{\\text\{return\}\}\(\\mathcal\{D\}\_\{S\_\{k\}\}^\{\\text\{sup\}\};\\theta\)\\big\)\(27\)whereηin\\eta\_\{\\text\{in\}\}is the inner\-loop learning rate\. Critically, during meta\-pretraining, we additionally extract and store the*causal structure prior*P​\(ϕ𝒢∣θmeta\)P\(\\phi\_\{\\mathcal\{G\}\}\\mid\\theta\_\{\\text\{meta\}\}\), a distribution over plausible causal graph structures inferred from the converged GAT attention weights and causal attribution scores\. Concretely,ϕ𝒢\\phi\_\{\\mathcal\{G\}\}is parameterized as a Bernoulli adjacency distribution over potential causal edges:

P​\(ϕ𝒢∣θmeta\)=∏\(j,k\)∈ℰcausalBernoulli​\(σ​\(𝐰ϕ⊤​\[α¯j​kmeta∥Δ¯j​kmeta\]\)\)P\(\\phi\_\{\\mathcal\{G\}\}\\mid\\theta\_\{\\text\{meta\}\}\)=\\prod\_\{\(j,k\)\\in\\mathcal\{E\}\_\{\\text\{causal\}\}\}\\text\{Bernoulli\}\\\!\\big\(\\sigma\(\\mathbf\{w\}\_\{\\phi\}^\{\\top\}\[\\bar\{\\alpha\}\_\{jk\}^\{\\text\{meta\}\}\\,\\\|\\,\\bar\{\\Delta\}\_\{jk\}^\{\\text\{meta\}\}\]\)\\big\)\(28\)whereα¯j​kmeta\\bar\{\\alpha\}\_\{jk\}^\{\\text\{meta\}\}is the mean attention weight between nodesjjandkkandΔ¯j​kmeta\\bar\{\\Delta\}\_\{jk\}^\{\\text\{meta\}\}aggregates the causal attribution scores over the meta\-pretraining trajectories\. This prior encodes the*domain\-invariant causal structure*that persists across sectors: the general principle that, e\.g\., a competitor’s funding event causally affects an asset’s risk, regardless of the specific sector\.

#### 3\.6\.2Phase II: Causally\-Regularized Adaptation

When adapting to a new target sectorTTwith limited data𝒟T\\mathcal\{D\}\_\{T\}\(e\.g\.,\|𝒟T\|=200\|\\mathcal\{D\}\_\{T\}\|=200samples\), we fine\-tune fromθmeta\\theta\_\{\\text\{meta\}\}by minimizing:

ℒMCA=ℒreturn​\(𝒟T;θ\)⏟task\-specific fit\+λ​DKL\(P\(ϕ𝒢∣θ\)∥P\(ϕ𝒢∣θmeta\)\)⏟causal structure preservation\\mathcal\{L\}\_\{\\text\{MCA\}\}=\\underbrace\{\\mathcal\{L\}\_\{\\text\{return\}\}\(\\mathcal\{D\}\_\{T\};\\,\\theta\)\}\_\{\\text\{task\-specific fit\}\}\+\\lambda\\underbrace\{D\_\{\\text\{KL\}\}\\\!\\Big\(P\(\\phi\_\{\\mathcal\{G\}\}\\mid\\theta\)\\;\\\|\\;P\(\\phi\_\{\\mathcal\{G\}\}\\mid\\theta\_\{\\text\{meta\}\}\)\\Big\)\}\_\{\\text\{causal structure preservation\}\}\(29\)
The KL\-divergence term acts as a*structural regularizer*: it penalizes adaptations that cause the model’s inferred causal graph to deviate significantly from the meta\-learned priorComptonet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib59)\); Liuet al\.\([2025a](https://arxiv.org/html/2606.28933#bib.bib56)\)\. Intuitively, this allows the model’s*predictions*to adapt freely to the new sector’s data distribution, while constraining the*causal reasoning mechanism*to remain consistent with the transferable structural knowledge learned across many sectors\.

The hyperparameterλ≥0\\lambda\\geq 0controls the strength of causal regularization\. Settingλ=0\\lambda=0recovers standard fine\-tuning; increasingλ\\lambdastrengthens the prior constraint\. We analyze sensitivity toλ\\lambdain §[6](https://arxiv.org/html/2606.28933#S6)and find that a moderate value \(λ≈0\.3\\lambda\\approx 0\.3\) optimally balances adaptability and robustness\.

### 3\.7Theoretical Analysis

We provide theoretical results that motivate key architectural choices and characterize the generalization properties of FinInvest\-GTCN\.

###### Theorem 2\(Generalization Bound for FinInvest\-GTCN\)\.

LetℱGTCN\\mathcal\{F\}\_\{\\text\{GTCN\}\}denote the hypothesis class of FinInvest\-GTCN withKGK\_\{G\}\-layer graph encoder,SS\-scale temporal fusion, and dual\-head prediction\. For a training set𝒟\\mathcal\{D\}of sizeNNdrawn i\.i\.d\. from the investment outcome distribution, the expected risk of the empirical risk minimizerf^∈ℱGTCN\\hat\{f\}\\in\\mathcal\{F\}\_\{\\text\{GTCN\}\}satisfies:

𝔼​\[ℒreturn​\(f^\)\]−ℒ^return​\(f^\)≤𝒪​\(dG⋅\|ℛ\|⋅KGN\)⏟graph complexity\+𝒪​\(S⋅dT⋅KT⋅log⁡LN\)⏟temporal complexity\+𝒪​\(log⁡\(1/δ\)2​N\)⏟confidence term\\mathbb\{E\}\\\!\\big\[\\mathcal\{L\}\_\{\\text\{return\}\}\(\\hat\{f\}\)\\big\]\-\\hat\{\\mathcal\{L\}\}\_\{\\text\{return\}\}\(\\hat\{f\}\)\\leq\\underbrace\{\\mathcal\{O\}\\\!\\left\(\\frac\{d\_\{G\}\\cdot\|\\mathcal\{R\}\|\\cdot K\_\{G\}\}\{\\sqrt\{N\}\}\\right\)\}\_\{\\text\{graph complexity\}\}\+\\underbrace\{\\mathcal\{O\}\\\!\\left\(\\frac\{S\\cdot d\_\{T\}\\cdot K\_\{T\}\\cdot\\log L\}\{\\sqrt\{N\}\}\\right\)\}\_\{\\text\{temporal complexity\}\}\+\\underbrace\{\\mathcal\{O\}\\\!\\left\(\\sqrt\{\\frac\{\\log\(1/\\delta\)\}\{2N\}\}\\right\)\}\_\{\\text\{confidence term\}\}\(30\)with probability at least1−δ1\-\\deltaover the draw of𝒟\\mathcal\{D\}\.

###### Proof sketch\.

The bound follows from a decomposition of the Rademacher complexity ofℱGTCN\\mathcal\{F\}\_\{\\text\{GTCN\}\}into graph encoder, temporal encoder, and prediction head components, combined with standard concentration inequalities\. The graph term scales with the product of the number of relation types, embedding dimension, and depth\. The temporal term includes alog⁡L\\log Lfactor from the self\-attention mechanism’s effective VC dimension\. The parallel multi\-scale architecture contributes an additiveSSfactor rather than multiplicative, confirming its computational efficiency\. A detailed proof is provided in the Appendix\. ∎

This bound reveals two key insights\. First, the graph module’s complexity is controlled by the number of relation types\|ℛ\|\|\\mathcal\{R\}\|rather than the graph size\|𝒱\|\|\\mathcal\{V\}\|, confirming that our multi\-relational attention mechanism generalizes well even on large graphs\. Second, the multi\-scale design contributes only linearly inSS\(the number of scales\), validating that the parallel architecture does not incur the exponential complexity that a single model processing all scales jointly would exhibit\.

###### Lemma 3\(MCA Regularization Effect\)\.

For a target sectorTTwithNTN\_\{T\}adaptation samples, the MCA\-regularized estimatorf^MCA\\hat\{f\}\_\{\\text\{MCA\}\}satisfies:

𝔼​\[ℒreturn​\(f^MCA\)\]≤𝔼​\[ℒreturn​\(f^FT\)\]−Ω​\(λ⋅dϕNT\)\\mathbb\{E\}\\\!\\big\[\\mathcal\{L\}\_\{\\text\{return\}\}\(\\hat\{f\}\_\{\\text\{MCA\}\}\)\\big\]\\leq\\mathbb\{E\}\\\!\\big\[\\mathcal\{L\}\_\{\\text\{return\}\}\(\\hat\{f\}\_\{\\text\{FT\}\}\)\\big\]\-\\Omega\\\!\\left\(\\lambda\\cdot\\frac\{d\_\{\\phi\}\}\{N\_\{T\}\}\\right\)\(31\)wheref^FT\\hat\{f\}\_\{\\text\{FT\}\}is the standard fine\-tuned estimator anddϕd\_\{\\phi\}is the dimension of the causal structure space\. That is, MCA reduces expected risk by an amount proportional toλ/NT\\lambda/N\_\{T\}, with the improvement being most pronounced in data\-scarce regimes \(smallNTN\_\{T\}\)\.

This lemma formalizes the intuition that MCA is most beneficial when data is scarce, providing a principled explanation for the large performance gains observed in the Quantum Computing adaptation experiments \(§[4](https://arxiv.org/html/2606.28933#S4)\)\.

### 3\.8Computational Complexity Analysis

We analyze the computational cost of each module to confirm scalability to real\-world deployment\.

Graph Encoder\.Each GAT layer requires𝒪​\(\|ℰt\|⋅\|ℛ\|⋅dG/MG\)\\mathcal\{O\}\(\|\\mathcal\{E\}\_\{t\}\|\\cdot\|\\mathcal\{R\}\|\\cdot d\_\{G\}/M\_\{G\}\)for attention computation and𝒪​\(\|𝒱t\|⋅dG2/MG\)\\mathcal\{O\}\(\|\\mathcal\{V\}\_\{t\}\|\\cdot d\_\{G\}^\{2\}/M\_\{G\}\)for the projection\. WithKGK\_\{G\}layers, the total is𝒪​\(KG⋅\|ℰt\|⋅\|ℛ\|⋅dG\)\\mathcal\{O\}\(K\_\{G\}\\cdot\|\\mathcal\{E\}\_\{t\}\|\\cdot\|\\mathcal\{R\}\|\\cdot d\_\{G\}\)\.

MST\-Former\.Each scale\-ssTransformer encoder requires𝒪​\(Ls2⋅dT\)\\mathcal\{O\}\(L\_\{s\}^\{2\}\\cdot d\_\{T\}\)per layer for self\-attention and𝒪​\(Ls⋅dT2\)\\mathcal\{O\}\(L\_\{s\}\\cdot d\_\{T\}^\{2\}\)for the feed\-forward network\. WithSSparallel encoders andKTK\_\{T\}layers each, the total is𝒪​\(KT⋅∑s=1S\(Ls2⋅dT\+Ls⋅dT2\)\)\\mathcal\{O\}\(K\_\{T\}\\cdot\\sum\_\{s=1\}^\{S\}\(L\_\{s\}^\{2\}\\cdot d\_\{T\}\+L\_\{s\}\\cdot d\_\{T\}^\{2\}\)\), which simplifies to𝒪​\(S⋅KT⋅L2⋅dT\)\\mathcal\{O\}\(S\\cdot K\_\{T\}\\cdot L^\{2\}\\cdot d\_\{T\}\)in the worst case\. Importantly, theSSencoders operate in*parallel*, so wall\-clock time scales as𝒪​\(KT⋅L2⋅dT\)\\mathcal\{O\}\(K\_\{T\}\\cdot L^\{2\}\\cdot d\_\{T\}\)\.

Causal Attribution\.Computing the ACE forPPcandidate factors requiresPPforward passes through the return MLP, costing𝒪​\(P⋅d2\)\\mathcal\{O\}\(P\\cdot d^\{2\}\)\. SincePPis typically small \(P≤20P\\leq 20candidate factors\) and the MLP is lightweight, this overhead is negligible relative to the encoder costs\.

Overall\.The total training complexity per sample is dominated by the Transformer encoders:𝒪​\(KT⋅L2⋅dT\+KG⋅\|ℰt\|⋅\|ℛ\|⋅dG\)\\mathcal\{O\}\(K\_\{T\}\\cdot L^\{2\}\\cdot d\_\{T\}\+K\_\{G\}\\cdot\|\\mathcal\{E\}\_\{t\}\|\\cdot\|\\mathcal\{R\}\|\\cdot d\_\{G\}\), which remains manageable for the financial sequence lengths \(L≤40L\\leq 40quarters\) and graph sizes \(\|𝒱t\|∼5000\|\\mathcal\{V\}\_\{t\}\|\\sim 5000\) encountered in practice\.

## 4Experiments

### 4\.1Experimental Setup

We adapt the experimental setup from prior work to our investment decision optimization task\. The core dataset is constructed from proprietary venture capital databases and public sourcesLiet al\.\([2025a](https://arxiv.org/html/2606.28933#bib.bib39)\); Akinfaderin and Subramanian \([2025](https://arxiv.org/html/2606.28933#bib.bib41)\), simulating an investor’s observation history\. For each investorii, we construct sequential observation windows for assetsjj, employing an 80%\-10%\-10% chronological split for training, validation, and testing\. The study encompasses sequences from over 10,000 simulated investors, 5,000 unique assets, and 2 million simulated investment decision points across eight primary sectors\.

Evaluation Metrics\.We adopt metrics aligned with the risk\-return prediction paradigm\. The primary evaluation metric is theRisk\-Adjusted Mean Squared Error \(RA\-MSE\), defined as1N​∑\(i,j\)\(ri,j−r^i,j\)2/σ^i,j2\\frac\{1\}\{N\}\\sum\_\{\(i,j\)\}\(r\_\{i,j\}\-\\hat\{r\}\_\{i,j\}\)^\{2\}/\\hat\{\\sigma\}\_\{i,j\}^\{2\}, corresponding directly to our training lossℒreturn\\mathcal\{L\}\_\{\\text\{return\}\}\(Eq\. \([24](https://arxiv.org/html/2606.28933#S3.E24)\)\)\. This metric penalizes overconfident errors\. For ranking performance, we report theCumulative Return \(Cum\. Ret\.\)of a simulated portfolio that invests in the top\-KKassets ranked by the model’s predicted risk\-adjusted return \(r^i,j/σ^i,j\\hat\{r\}\_\{i,j\}/\\hat\{\\sigma\}\_\{i,j\}\) at each decision point\. We also report standardMSEfor return prediction andAccuracyfor a binary “investment\-worthy” classification derived from a return threshold\.

Baselines & Our Method:We compare our method against several strong baselines\.

- •RF Baseline:A Random Forest model using handcrafted financial and firmographic features\.
- •LSTM:A standard LSTM network processing the temporal observation sequenceOi,jO\_\{i,j\}\.
- •Transformer:A Transformer encoder processingOi,jO\_\{i,j\}\.
- •FinTRec:A baseline model adapted from prior work to predict a scalar return\.
- •FinInvest\-GTCN \(Ours\):Our proposed Graph\-Temporal\-Causal Network\. We also report an ablated version,Ours w/o Graph, which removes the Relational Graph Encoder module \(§[3\.2](https://arxiv.org/html/2606.28933#S3.SS2)\)\.

Implementation Details\.All neural models were implemented in PyTorch and trained on NVIDIA A100 GPUsDinget al\.\([2025a](https://arxiv.org/html/2606.28933#bib.bib78)\); Zouet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib67)\)\. Each model was trained for 100 epochs using the AdamW optimizer with a learning rate of1×10−41\\times 10^\{\-4\}, weight decay of1×10−51\\times 10^\{\-5\}, and a batch size of 32\. The model achieving the lowest validation RA\-MSE was selected\. For FinInvest\-GTCN, the meta\-pretraining for the MCA module \(§[3\.6](https://arxiv.org/html/2606.28933#S3.SS6)\) was conducted on a separate set of five source sectors\. Table[1](https://arxiv.org/html/2606.28933#S4.T1)summarizes the key architectural and training hyperparameters\.

Table 1:Key hyperparameters of the FinInvest\-GTCN architecture and training configuration\.ModuleHyperparameterValueGraph EncoderGraph embedding dimdGd\_\{G\}128Attention headsMGM\_\{G\}4GAT layersKGK\_\{G\}3Relation types4MST\-FormerTemporal embedding dimdTd\_\{T\}128Transformer layersKTK\_\{T\}4Number of scalesSS3Scale windows\(4, 8, Full\)Attention heads per scale8Causal HeadFused dimdd256Causal factorsPP15MLP hidden dim128TrainingLearning rate1×10−41\\times 10^\{\-4\}Weight decay1×10−51\\times 10^\{\-5\}Causal loss weightμ\\mu0\.1Graph reg\. weightν\\nu0\.01MCA strengthλ\\lambda0\.3
### 4\.2Offline Results

Main Results\.The main results on the hold\-out test set are summarized in Table[2](https://arxiv.org/html/2606.28933#S4.T2)\.

![Refer to caption](https://arxiv.org/html/2606.28933v1/x1.png)Figure 3:Overall performance comparison across all models\. FinInvest\-GTCN achieves the best scores on all metrics \(RA\-MSE, MSE, Cumulative Return Top\-10, and Accuracy\), demonstrating clear superiority over baselines\. The ablated version \(Ours w/o Graph\) shows a noticeable performance drop, emphasizing the importance of the graph encoder\.Our proposedFinInvest\-GTCNachieves the best performance across all key metrics\. It significantly outperforms all baselines on the primary RA\-MSE metric \(2\.51\), demonstrating superior capability in making accurate, risk\-calibrated predictions\. The Transformer and FinTRec baselines show competitive but lower performance\. Notably, the ablated versionOurs w/o Graphexhibits a clear performance drop \(RA\-MSE of 3\.15\), underscoring the critical contribution of the relational graph encoderZhanget al\.\([2025f](https://arxiv.org/html/2606.28933#bib.bib88)\); Kempinski and Kachman \([2025](https://arxiv.org/html/2606.28933#bib.bib93)\)\. The RF baseline performs the weakest, highlighting the limitation of static, handcrafted features for this dynamic sequence prediction task\.

Table 2:Overall predictive performance on the test set\. Lower values are better for RA\-MSE and MSE\. Higher values are better for Cumulative Return \(Top\-10\) and Accuracy\. Best results arebold; second best areunderlined\.Performance Across Sectors\.To assess model robustness, we analyze RA\-MSE performance broken down by the primary sector of the target asset\. As shown in Table[3](https://arxiv.org/html/2606.28933#S4.T3)and visualized in Figure[4](https://arxiv.org/html/2606.28933#S4.F4),FinInvest\-GTCNachieves the best performance in five of six major sectors and remains competitive in the AI/ML sector, exhibiting overall stable and superior performance\.

![Refer to caption](https://arxiv.org/html/2606.28933v1/figures/main_result/sector_breakdown_heatmap.png)Figure 4:Heatmap of RA\-MSE performance across six asset sectors\. Cooler colors indicate lower \(better\) RA\-MSE values\. FinInvest\-GTCN achieves consistently low RA\-MSE across most sectors, with particular strength in network\-dependent sectors like FinTech and Platforms\.The model shows particular strength in complex, network\-dependent sectors likeFinTechandPlatformsLeeet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib17)\); Gresselet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib30)\)\. The FinTRec baseline performs well but is less consistent, especially inHealthTech\. The standard LSTM and Transformer models show greater performance variance across sectors\.

Table 3:RA\-MSE performance breakdown by asset sector \(lower is better\)\. Best per sector isbold; second best isunderlined\.Ablation Study\.A detailed ablation study isolates the contribution of each core component in FinInvest\-GTCN, with results presented in Table[4](https://arxiv.org/html/2606.28933#S4.T4)\. The full model achieves the best RA\-MSE \(2\.51\)\. Removing theCausal Attributionmodule leads to a slight performance drop \(RA\-MSE: 2\.63\) and eliminates explainability\. Replacing theMulti\-Scale Temporal Fusionwith a single\-scale Transformer causes a more significant drop \(RA\-MSE: 2\.89\), validating the need to capture multi\-horizon patterns\. As noted earlier, removing theGraph Encodersubstantially harms performance \(RA\-MSE: 3\.15\)\. Using a standard MSE loss instead of the risk\-adjusted loss results in the worst RA\-MSE \(3\.42\), confirming the specialized loss function’s importance for risk\-calibrated prediction\.

Table 4:Ablation study of FinInvest\-GTCN components \(RA\-MSE, lower is better\)\.![Refer to caption](https://arxiv.org/html/2606.28933v1/x2.png)Figure 5:RA\-MSE after adaptation\. FinInvest\-GTCN achieves the lowest error\.Meta\-Causal Adaptation Effectiveness\.To evaluate our model’s robustness in low\-data scenarios, we simulate adaptation to a new investment domain\. We hold out all data from a “Quantum Computing” sector during initial training, then fine\-tune models on a very small sample \(N=200N=200\) from this new sector\. Table[5](https://arxiv.org/html/2606.28933#S4.T5)and Figure[5](https://arxiv.org/html/2606.28933#S4.F5)present the RA\-MSE after adaptation\.

FinInvest\-GTCN with MCAsignificantly outperforms all baselines and our own model fine\-tuned with standard full\-parameter fine\-tuning or LoRAJeonet al\.\([2025b](https://arxiv.org/html/2606.28933#bib.bib73)\); Honget al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib66)\)\. The MCA strategy, which regularizes updates towards causally\-plausible structures learned during meta\-pretraining, effectively prevents overfitting and leverages prior knowledge, making it uniquely suited for data\-scarce domains\.

Table 5:Performance on a new, data\-scarce sector \(“Quantum Computing”\) after fine\-tuning onN=200N=200samples\. Lower RA\-MSE is better\.
### 4\.3Online Simulation & A/B Test Results

Following standard methodology, we conduct offline simulations to assess the potential impact of deploying FinInvest\-GTCN on a simulated investment portfolio\. We compare the cumulative return of a portfolio constructed using rankings from our model against the production baseline over a simulated 12\-month period\.

![Refer to caption](https://arxiv.org/html/2606.28933v1/x3.png)Figure 6:Online portfolio comparison\. FinInvest\-GTCN improves return, volatility, and Sharpe ratio\.Simulated Portfolio Performance\.The results are summarized in Table[6](https://arxiv.org/html/2606.28933#S4.T6)and Figure[6](https://arxiv.org/html/2606.28933#S4.F6)\.

The portfolio based onFinInvest\-GTCNrankings achieves a significantly higher cumulative return \(\+18\.7%\) compared to the baseline portfolio\. It also exhibits lower volatility \(14\.2% vs\. 16\.8%\), resulting in a superior Sharpe Ratio \(1\.31 vs\. 0\.89\)\. This demonstrates that the superior offline risk\-return predictions of our model translate directly into better simulated investment outcomes\. The improvement is attributed to our model’s integrated graph\-temporal modelingAlduaiset al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib91)\); Wanget al\.\([2025c](https://arxiv.org/html/2606.28933#bib.bib90)\)and explicit risk estimation, which lead to more stable and high\-conviction asset rankings\.

Table 6:Simulated online A/B test results over a 12\-month period\. The portfolio selects top\-10 assets monthly based on model rankings\.Summary\.The comprehensive experiments confirm that our proposedFinInvest\-GTCNmethod establishes a new state\-of\-the\-art for quantitative investment decision support\. It consistently ranks first in predictive accuracy, demonstrates robust performance across diverse sectors, effectively leverages its architectural components, and shows exceptional adaptability to new, data\-scarce domains via MCA\. Critically, these offline gains translate into superior simulated portfolio performance, underscoring its practical utility\. The results validate our core design principles: integrating relational graph context, modeling multi\-scale temporal dynamics, and employing a risk\-adjusted objective with causal regularization\.

## 5Ablation Studies

To rigorously evaluate the contribution of each core component in our proposedFinInvest\-GTCNarchitecture and its adaptation strategy, we conduct a comprehensive suite of ablation experiments\. These studies systematically test fundamental module necessity, evaluate adaptation strategies, and analyze specific design choices\. The results consistently demonstrate the superiority of the full model and validate the critical role of each designed component, providing empirical support for the theoretical insights in §[3\.7](https://arxiv.org/html/2606.28933#S3.SS7)\.

![Refer to caption](https://arxiv.org/html/2606.28933v1/x4.png)Figure 7:Ablation study on core architectural modules\.1\. Ablation on Core Architectural Modules:We first isolate the impact of the three main modules of FinInvest\-GTCN: the Relational Graph Encoder \(G, §[3\.2](https://arxiv.org/html/2606.28933#S3.SS2)\), the Multi\-Scale Temporal Fusion \(T, §[3\.3](https://arxiv.org/html/2606.28933#S3.SS3)\), and the Causal Decision Head with its risk\-adjusted loss \(C, §[3\.4](https://arxiv.org/html/2606.28933#S3.SS4)\)\. As shown in Table[7](https://arxiv.org/html/2606.28933#S5.T7), the complete model achieves the best performance across both primary predictive metrics \(RA\-MSE: 2\.51\) and the downstream investment metric \(Cumulative Return: 3\.02\)\. Removing the Graph Encoder \(w/o G\) causes the most significant performance drop \(RA\-MSE increases to 3\.15\), underscoring the indispensable value of multi\-relational attention \(Eq\. \([4](https://arxiv.org/html/2606.28933#S3.E4)\)\) for modeling the topological relationships within the investment ecosystemXuet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib18)\); Yeet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib25)\)\.

Table 7:Ablation study on the core modules of FinInvest\-GTCN\. The full model integrates Graph \(G\), Multi\-scale Temporal fusion \(T\), and Causal/Risk\-adjusted loss \(C\)\.![Refer to caption](https://arxiv.org/html/2606.28933v1/x5.png)Figure 8:Comparison of adaptation strategies for a new, data\-scarce sector \(Quantum Computing,N=200N=200\)\.Replacing the Multi\-Scale Temporal Fusion with a single\-scale Transformer encoder \(w/o T \(Multi\-Scale\)\) also leads to notable degradation \(RA\-MSE: 2\.89\), confirming the necessity of the adaptive gated fusion mechanism \(Eq\. \([14](https://arxiv.org/html/2606.28933#S3.E14)\)\) for capturing financial patterns across different horizonsDinget al\.\([2025b](https://arxiv.org/html/2606.28933#bib.bib49)\); Loganathanet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib46)\)\. Removing the causal attribution and using a standard MSE loss \(w/o C\) results in the worst RA\-MSE \(3\.42\), validating that the heteroscedastic lossℒreturn\\mathcal\{L\}\_\{\\text\{return\}\}\(Eq\. \([24](https://arxiv.org/html/2606.28933#S3.E24)\)\) is crucial for producing well\-calibrated, risk\-aware predictions\.

2\. Ablation on Adaptation Strategies for New Sectors:

We evaluate the effectiveness of the Meta\-Causal Adaptation \(MCA\) strategy against standard fine\-tuning techniques\. As shown in Table[8](https://arxiv.org/html/2606.28933#S5.T8), after fine\-tuning on only 200 samples from a held\-out “Quantum Computing” sector, ourMCAstrategy achieves the lowest RA\-MSE \(3\.41\), significantly outperforming other strategies\. Full parameter fine\-tuning \(F\-FT\) leads to overfitting \(RA\-MSE: 4\.05\)Wuet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib74)\); Prabhuneet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib40)\)\. Parameter\-efficient fine\-tuning via LoRA performs better \(3\.88\) but does not match MCAAroraet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib43)\); Fuet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib24)\), as it lacks the causal structural regularization\. Notably,without any fine\-tuning, our model performs poorly on the new sector \(5\.87\) due to distributional shifts, as illustrated in Figure[8](https://arxiv.org/html/2606.28933#S5.F8)\.

Table 8:Ablation on adaptation strategies for a new, data\-scarce sector \(“Quantum Computing”, N=200\)\. Lower RA\-MSE is better\.This ablation confirms that MCA \(Eq\. \([29](https://arxiv.org/html/2606.28933#S3.E29)\)\) is a principled approach that balances rapid adaptation with robustness by preserving causal insights from meta\-pretrainingByambadalaiet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib60)\); Jinet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib61)\), consistent with the theoretical prediction of Lemma[3](https://arxiv.org/html/2606.28933#Thmtheorem3)\.

3\. Analysis of Multi\-Scale Temporal Configurations:We ablate the choice of temporal scales \(context windows\) in the Multi\-Scale Temporal Fusion module\. Table[9](https://arxiv.org/html/2606.28933#S5.T9)reports the RA\-MSE for models trained with different scale combinations\. The configuration “Q4 \+ Q8 \+ Full” \(4\-quarter, 8\-quarter, and full\-history scales, as defined in Eq\. \([9](https://arxiv.org/html/2606.28933#S3.E9)\)\), used in our full model, delivers the best performanceLonget al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib26)\); Muneeret al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib62)\)\. Using only a single scale, whether short \(“Q4 only”\) or long \(“Full only”\), results in suboptimal performance\. The combination “Q4 \+ Q8” performs better than single scales but worse than the full three\-scale model, indicating that the full history provides unique, non\-redundant signals\. This empirically justifies our multi\-scale design for capturing the complex, multi\-horizon nature of financial time seriesGuoet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib22)\); Sekar and Nezamoddini \([2025](https://arxiv.org/html/2606.28933#bib.bib87)\), as visualized in Figure[9](https://arxiv.org/html/2606.28933#S5.F9)\.

![Refer to caption](https://arxiv.org/html/2606.28933v1/figures/ablation/multi_scale_temporal_analysis.png)Figure 9:Effect of different multi\-scale temporal configurations on RA\-MSE\. Performance improves as more complementary scales are combined, with the three\-scale configuration \(Q4\+Q8\+Full\) achieving optimal performance, validating the multi\-scale design\.Table 9:Ablation on the configuration of scales in the Multi\-Scale Temporal Fusion module\. “Q4”, “Q8”, “Full” refer to 4\-quarter, 8\-quarter, and full\-history attention spans\.4\. Utility of Causal Attribution for Explanation Fidelity:We evaluate the practical utility of the Causal Attribution module for model explainability\. For a set of test samples, we use the module to identify the top hypothesized causal factor\. We then retrain asurrogate modelusing only the temporal steps where this factor was active, according to the attribution scores\. As shown in Table[10](https://arxiv.org/html/2606.28933#S5.T10), the performance of this surrogate model \(RA\-MSE: 2\.90\) is close to that of the full model trained on the entire sequence \(RA\-MSE: 2\.51\), and it significantly outperforms a surrogate model trained on randomly selected time steps \(RA\-MSE: 3\.55\)\. This indicates that the Interventional Causal Attribution mechanism \(§[3\.4\.3](https://arxiv.org/html/2606.28933#S3.SS4.SSS3), Eq\. \([21](https://arxiv.org/html/2606.28933#S3.E21)\)\) successfully isolates the most predictive, semantically meaningful segments of the input sequenceLiuet al\.\([2025b](https://arxiv.org/html/2606.28933#bib.bib54)\); Kubota and Sugasawa \([2025](https://arxiv.org/html/2606.28933#bib.bib58)\), providing high\-fidelity explanations that could be invaluable for regulatory compliance and investor trust, as demonstrated in Figure[10](https://arxiv.org/html/2606.28933#S5.F10)\.

![Refer to caption](https://arxiv.org/html/2606.28933v1/x6.png)Figure 10:Evaluation of causal attribution fidelity\.Table 10:Evaluating the fidelity of explanations from the Causal Attribution module\. The surrogate model is trained only on input slices identified as important by the module\.In summary, the ablation studies provide a multi\-faceted validation of the FinInvest\-GTCN architectureArastehet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib89)\); Bet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib36)\)\. They demonstrate that: \(1\) every core module is essential for peak performance; \(2\) the novel MCA adaptation strategy is superior for data\-scarce domains; \(3\) the multi\-scale temporal design is optimally configured; and \(4\) the causal attribution provides faithful explanations\. Collectively, these experiments solidify the rationale behind our design choices\.

## 6Supplementary Experiments

To provide deeper insights into the behavior and robustness ofFinInvest\-GTCN, we conduct supplementary analysesDvoretskiiet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib37)\); Kukanov and Ng \([2025](https://arxiv.org/html/2606.28933#bib.bib81)\), including examining training dynamics, performing case studies, and analyzing sensitivity to key hyperparameters\. These experiments further validate the model’s stability, interpretability, and practical utility beyond the aggregate metrics\.

1\. Training Dynamics and Convergence Analysis\.We analyze the training and validation curves for the primary risk\-adjusted loss \(ℒreturn\\mathcal\{L\}\_\{\\text\{return\}\}\) across epochs forFinInvest\-GTCNand two key baselines: the Transformer and the ablatedOurs w/o Graph\. FinInvest\-GTCN achieves lower validation loss more rapidly and maintains a stable gap between training and validation loss, indicating effective generalization without severe overfittingChenet al\.\([2025e](https://arxiv.org/html/2606.28933#bib.bib29)\); Kemper and Rostam\-Afschar \([2025](https://arxiv.org/html/2606.28933#bib.bib63)\)\. In contrast, the Transformer baseline shows higher validation loss and greater variance, whileOurs w/o Graphexhibits a slower convergence rate and a higher final plateau\. To quantify this, we report the epoch at which each model’s validation RA\-MSE first falls below 3\.0 and the standard deviation of the last 10 epochs’ validation loss: FinInvest\-GTCN \(epoch 28, std 0\.08\), Transformer \(epoch 41, std 0\.15\), Ours w/o Graph \(epoch 35, std 0\.12\)\. This analysis confirms that the integration of the graph encoder and multi\-scale temporal fusion not only improves final performance but also leads to more stable and efficient optimizationLiet al\.\([2025b](https://arxiv.org/html/2606.28933#bib.bib33)\); Vuet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib32)\)\.

![Refer to caption](https://arxiv.org/html/2606.28933v1/x7.png)Figure 11:Training and validation loss curves for FinInvest\-GTCN compared with the Transformer baseline and the ablated variant \(Ours w/o Graph\)\. Our full model converges faster and exhibits smaller generalization gap\.2\. Case Study: Model Predictions and Explanations for Exemplar Assets\.To illustrate the model’s explanatory capability via Interventional Causal Attribution \(§[3\.4\.3](https://arxiv.org/html/2606.28933#S3.SS4.SSS3)\), we present a qualitative case study on three diverse assets from the test set\. For each asset, Table[11](https://arxiv.org/html/2606.28933#S6.T11)shows the model’s predicted return \(r^\\hat\{r\}\) and risk \(σ^\\hat\{\\sigma\}\), the actual realized return \(rr\), and the top causal factor identified by the Causal Attribution module along with its estimated effect \(Δ​r^\\Delta\\hat\{r\}\)\. For instance, for a FinTech startup, the model correctly predicted a high return \(8\.2%\) with moderate risk, attributing a significant portion of the positive signal to a recent “regulatory approval” event\. For a HealthTech asset that underperformed, the model assigned a high risk score and attributed a negative effect to a “key competitor funding round\.” These case\-specific explanations demonstrate how FinInvest\-GTCN can provide actionable insights beyond a scalar predictionMahadevan \([2025](https://arxiv.org/html/2606.28933#bib.bib64)\); Parafitaet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib57)\)\.

Table 11:Case study on three test assets, showing predictions, outcomes, and top causal explanations from the Causal Attribution module\.3\. Sensitivity Analysis of Meta\-Causal Adaptation Hyperparameter\.The Meta\-Causal Adaptation \(MCA\) strategy introduces a regularization strengthλ\\lambda\(Eq\. \([29](https://arxiv.org/html/2606.28933#S3.E29)\)\) that balances task\-specific fine\-tuning against causal structural priors\. To assess sensitivity, we varyλ\\lambdafrom0\(no regularization\) to11\(strong prior\) when adapting to the held\-out “Quantum Computing” sector \(N=200N=200\)\. Table[12](https://arxiv.org/html/2606.28933#S6.T12)reports the resulting RA\-MSE\. Performance peaks atλ=0\.3\\lambda=0\.3, with significant degradation ifλ\\lambdais too low \(overfitting\) or too high \(underfitting\)\. The optimalλ\\lambdayields a 12% improvement over no regularization \(λ=0\\lambda=0\) and a 10% improvement over very strong regularization \(λ=1\\lambda=1\)\. This analysis provides practical guidance for deploying MCA to new sectors: a moderateλ\\lambdaaround 0\.3 effectively leverages the meta\-learned causal structures while allowing necessary adaptation to domain\-specific signalsGuanet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib76)\); Jianget al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib75)\)\.

Table 12:Sensitivity of MCA performance to the regularization strengthλ\\lambdaon the “Quantum Computing” sector adaptation task \(lower RA\-MSE is better\)\.![Refer to caption](https://arxiv.org/html/2606.28933v1/x8.png)Figure 12:Sensitivity analysis of key hyperparameters\. Performance remains robust within moderate ranges, with sharp degradation at extremes\.4\. Multi\-Metric Radar Comparison\.To holistically compare model capabilities across multiple evaluation dimensions, Figure[13](https://arxiv.org/html/2606.28933#S6.F13)presents a radar chart showing all methods on five core metrics\. FinInvest\-GTCN achieves the largest coverage area, indicating balanced superiority rather than improvement on a single metric at the expense of others\.

![Refer to caption](https://arxiv.org/html/2606.28933v1/x9.png)Figure 13:Multi\-metric radar profile comparing FinInvest\-GTCN with baselines across five evaluation dimensions\. Our model achieves the largest coverage area, reflecting balanced superiority\.5\. Sector\-Wise Error Distribution\.To understand performance variation within each sector, Figure[14](https://arxiv.org/html/2606.28933#S6.F14)presents violin plots of per\-asset RA\-MSE across the six sectors\. FinInvest\-GTCN not only achieves lower median error but also exhibits tighter distributions, suggesting consistent performance regardless of asset\-specific characteristics within each sector\.

![Refer to caption](https://arxiv.org/html/2606.28933v1/x10.png)Figure 14:Violin plots of per\-asset RA\-MSE across six sectors\. FinInvest\-GTCN demonstrates lower median errors and tighter distributions compared to baselines\.6\. Residual Analysis\.To verify that our model does not exhibit systematic prediction biases, Figure[15](https://arxiv.org/html/2606.28933#S6.F15)shows a scatter plot of predicted vs\. actual returns with residual distributions\. The residuals are centered near zero with no discernible pattern, confirming that FinInvest\-GTCN produces well\-calibrated predictions without systematic over\- or under\-estimation across different return magnitudes\.

![Refer to caption](https://arxiv.org/html/2606.28933v1/x11.png)Figure 15:Scatter plot of predicted vs\. actual returns with marginal residual distributions\. Residuals are centered near zero with no systematic bias\.
## 7Conclusion

In this work, we introduce a paradigm shift in venture capital decision support by reframing it from a content recommendation task to a quantitative risk\-adjusted return prediction problem\. To address the inherent challenges of multi\-source heterogeneous data, non\-stationary time series, and the need for explainability in low\-data scenariosYinet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib45)\); Liet al\.\([2025c](https://arxiv.org/html/2606.28933#bib.bib34)\), we proposeFinInvest\-GTCN, a novel Graph\-Temporal\-Causal Network\. Our architecture synergistically integrates a multi\-relational graph encoder with relation\-specific attention to model the investment ecosystem’s topologyLi and Fan \([2025](https://arxiv.org/html/2606.28933#bib.bib92)\); Kimet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib82)\), a multi\-scale temporal fusion module with scale\-specific causal attention masks and adaptive gated fusion for capturing long\-term dependenciesChirukiriet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib47)\); Oganesianet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib42)\), and a causal decision head featuring interventional causal attribution that generates both risk\-return predictions and post\-hoc causal explanations grounded in the potential outcomes frameworkAkbaret al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib55)\); Zhang and Cai \([2025](https://arxiv.org/html/2606.28933#bib.bib65)\)\. Theoretical analysis provides generalization bounds confirming favorable scaling properties and formalizes the regularization benefits of our Meta\-Causal Adaptation strategy\.

Comprehensive experiments on proprietary venture capital data demonstrate the effectiveness of our approach\. Our model achieves state\-of\-the\-art performance, with a superior Risk\-Adjusted MSE of 2\.51 and the highest cumulative return in simulated portfolios, outperforming strong baselines\. Ablation studies confirm the critical contribution of each core component: the graph encoder, the multi\-scale temporal fusion module, the causal attribution module, and the specialized risk\-adjusted loss\. Furthermore, the proposed Meta\-Causal Adaptation \(MCA\) strategy enables robust performance in data\-scarce new sectorsYunet al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib77)\); Honget al\.\([2025](https://arxiv.org/html/2606.28933#bib.bib66)\), significantly surpassing standard fine\-tuning methods\. The offline predictive gains translate into tangible benefits in online simulation, where a portfolio based on our model’s rankings achieves a higher cumulative return \(\+18\.7%\) and a superior Sharpe Ratio \(1\.31\) compared to the baseline\.

In summary, FinInvest\-GTCN offers a robust, explainable, and adaptable framework for optimizing data\-driven investment decisionsAkinfaderin and Subramanian \([2025](https://arxiv.org/html/2606.28933#bib.bib41)\); Liet al\.\([2025a](https://arxiv.org/html/2606.28933#bib.bib39)\)by effectively integrating relational context, temporal dynamics, and causal reasoning\.

## Declarations

- •Funding:Not applicable\.
- •Conflict of interest:The authors declare no conflict of interest\.
- •Data availability:The datasets analyzed during the current study are available from the corresponding author on reasonable request\.
- •Code availability:Code will be made available upon publication\.
- •Author contribution:Not applicable\.

## References

- An analysis of causal effect estimation using outcome invariant data augmentation\.arXiv preprint arXiv:2510\.25128\.External Links:2510\.25128v1Cited by:[§2\.3](https://arxiv.org/html/2606.28933#S2.SS3.p1.1),[§3\.4\.3](https://arxiv.org/html/2606.28933#S3.SS4.SSS3.p1.1),[§7](https://arxiv.org/html/2606.28933#S7.p1.1)\.
- A\. Akinfaderin and S\. Subramanian \(2025\)VERAFI: verified agentic financial intelligence through neurosymbolic policy generation\.arXiv preprint arXiv:2512\.14744\.External Links:2512\.14744v1Cited by:[§1](https://arxiv.org/html/2606.28933#S1.p1.1),[§4\.1](https://arxiv.org/html/2606.28933#S4.SS1.p1.2),[§7](https://arxiv.org/html/2606.28933#S7.p3.1)\.
- M\. Alduais, X\. Li, and Q\. Mei \(2025\)TrajGATFormer: a graph\-based transformer approach for worker and obstacle trajectory prediction in off\-site construction environments\.arXiv preprint arXiv:2510\.22205\.External Links:2510\.22205v1Cited by:[§2\.2](https://arxiv.org/html/2606.28933#S2.SS2.p1.1),[§4\.3](https://arxiv.org/html/2606.28933#S4.SS3.p3.1)\.
- F\. Arasteh, A\. Haghparast, and M\. Papagelis \(2025\)Network\-constrained policy optimization for adaptive multi\-agent vehicle routing\.arXiv preprint arXiv:2510\.26089\.External Links:2510\.26089v1Cited by:[§5](https://arxiv.org/html/2606.28933#S5.p9.1)\.
- V\. Arora, D\. Lachi, I\. J\. Knight, M\. Azabou, B\. Richards, C\. L\. Hurwitz, J\. Siegle, and E\. L\. Dyer \(2025\)Know thyself by knowing others: learning neuron identity from population context\.arXiv preprint arXiv:2512\.01199\.External Links:2512\.01199v1Cited by:[§5](https://arxiv.org/html/2606.28933#S5.p5.1)\.
- A\. F\. Ashrafi and H\. Kabir \(2025\)Enhanced graph convolutional network with chebyshev spectral graph and graph attention for autism spectrum disorder classification\.arXiv preprint arXiv:2511\.22178\.External Links:2511\.22178v1Cited by:[§2\.1](https://arxiv.org/html/2606.28933#S2.SS1.p1.1)\.
- S\. B, S\. Nicolazzo, D\. K\., and V\. P \(2025\)Protecting deep neural network intellectual property with chaos\-based white\-box watermarking\.arXiv preprint arXiv:2512\.16658\.External Links:2512\.16658v1Cited by:[§5](https://arxiv.org/html/2606.28933#S5.p9.1)\.
- M\. Beniwal \(2025\)Adaptive weighted genetic algorithm\-optimized svr for robust long\-term forecasting of global stock indices for investment decisions\.arXiv preprint arXiv:2512\.15113\.External Links:2512\.15113v1Cited by:[§2\.1](https://arxiv.org/html/2606.28933#S2.SS1.p1.1)\.
- U\. Byambadalai, T\. Hirata, T\. Oka, and S\. Yasui \(2025\)Beyond the average: distributional causal inference under imperfect compliance\.arXiv preprint arXiv:2509\.15594\.External Links:2509\.15594v2Cited by:[§2\.3](https://arxiv.org/html/2606.28933#S2.SS3.p1.1),[§5](https://arxiv.org/html/2606.28933#S5.p6.1)\.
- M\. Chapariniya, T\. Vukovic, S\. Ebling, and V\. Dellwo \(2025\)Beyond appearance: transformer\-based person identification from conversational dynamics\.arXiv preprint arXiv:2510\.04753\.External Links:2510\.04753v1Cited by:[§2\.2](https://arxiv.org/html/2606.28933#S2.SS2.p1.1)\.
- H\. Chen, J\. Peng, D\. Min, C\. Sun, K\. Chen, Y\. Yan, X\. Yang, and L\. Cheng \(2025a\)Mvi\-bench: a comprehensive benchmark for evaluating robustness to misleading visual inputs in lvlms\.arXiv preprint arXiv:2511\.14159\.Cited by:[§1](https://arxiv.org/html/2606.28933#S1.p2.1)\.
- K\. Chen, Z\. Lin, Z\. Xu, Y\. Shen, Y\. Yao, J\. Rimchala, J\. Zhang, and L\. Huang \(2025b\)R2i\-bench: benchmarking reasoning\-driven text\-to\-image generation\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 12606–12641\.Cited by:[§1](https://arxiv.org/html/2606.28933#S1.p2.1)\.
- K\. Chen, Z\. Xu, Y\. Shen, Z\. Lin, Y\. Yao, and L\. Huang \(2025c\)SuperFlow: training flow matching models with rl on the fly\.arXiv preprint arXiv:2512\.17951\.Cited by:[§1](https://arxiv.org/html/2606.28933#S1.p2.1)\.
- K\. Chen, K\. W\. Parker, and A\. Arora \(2025d\)MoCap2Radar: a spatiotemporal transformer for synthesizing micro\-doppler radar signatures from motion capture\.arXiv preprint arXiv:2511\.11462\.External Links:2511\.11462v1Cited by:[§2\.2](https://arxiv.org/html/2606.28933#S2.SS2.p1.1)\.
- R\. Chen, J\. Ternasky, A\. O\. Yin, X\. Mu, F\. Alican, and Y\. Ihlamur \(2025e\)LLM\-ar: llm\-powered automated reasoning framework\.arXiv preprint arXiv:2510\.22034\.External Links:2510\.22034v1Cited by:[§6](https://arxiv.org/html/2606.28933#S6.p2.1)\.
- V\. T\. Chirukiri, U\. B\. Cheerala, S\. Kanta, A\. Karim, and P\. Damacharla \(2025\)FTT\-gru: a hybrid fast temporal transformer with gru for remaining useful life prediction\.arXiv preprint arXiv:2511\.00564\.External Links:2511\.00564v1Cited by:[§2\.2](https://arxiv.org/html/2606.28933#S2.SS2.p1.1),[§7](https://arxiv.org/html/2606.28933#S7.p1.1)\.
- W\. Chiu, Y\. Wang, A\. Hsiao, Y\. Huang, and C\. Wang \(2025\)Financial risk relation identification through dual\-view adaptation\.arXiv preprint arXiv:2509\.18775\.External Links:2509\.18775v2Cited by:[§2\.1](https://arxiv.org/html/2606.28933#S2.SS1.p1.1)\.
- J\. Choi, S\. Park, S\. Park, S\. Cho, and N\. Park \(2025\)Are graph transformers necessary? efficient long\-range message passing with fractal nodes in mpnns\.arXiv preprint arXiv:2511\.13010\.External Links:2511\.13010v1Cited by:[§1](https://arxiv.org/html/2606.28933#S1.p2.1)\.
- S\. Compton, K\. Greenewald, D\. Katz, and M\. Kocaoglu \(2025\)Entropic causal inference: graph identifiability\.arXiv preprint arXiv:2509\.16463\.External Links:2509\.16463v1Cited by:[§2\.3](https://arxiv.org/html/2606.28933#S2.SS3.p1.1),[§3\.6\.2](https://arxiv.org/html/2606.28933#S3.SS6.SSS2.p2.1)\.
- K\. Ding, Y\. Zhou, X\. Chen, M\. Yang, J\. Ou, R\. Chen, X\. Tao, and H\. Zhao \(2025a\)Alchemist: unlocking efficiency in text\-to\-image model training via meta\-gradient data selection\.arXiv preprint arXiv:2512\.16905\.External Links:2512\.16905v1Cited by:[§2\.4](https://arxiv.org/html/2606.28933#S2.SS4.p1.1),[§4\.1](https://arxiv.org/html/2606.28933#S4.SS1.p4.2)\.
- N\. Ding, K\. Fujii, and T\. Tamaki \(2025b\)Shot2Tactic\-caption: multi\-scale captioning of badminton videos for tactical understanding\.arXiv preprint arXiv:2510\.14617\.External Links:2510\.14617v1Cited by:[§2\.2](https://arxiv.org/html/2606.28933#S2.SS2.p1.1),[§5](https://arxiv.org/html/2606.28933#S5.p3.1)\.
- C\. Duan and C\. Ji \(2025\)Graph attention network for predicting duration of large\-scale power outages induced by natural disasters\.arXiv preprint arXiv:2511\.10898\.External Links:2511\.10898v1Cited by:[§2\.1](https://arxiv.org/html/2606.28933#S2.SS1.p1.1),[§3\.2\.1](https://arxiv.org/html/2606.28933#S3.SS2.SSS1.p1.2)\.
- S\. Dvoretskii, A\. Archit, C\. Pape, J\. Moore, and M\. Nolden \(2025\)BioimageAIpub: a toolbox for ai\-ready bioimaging data publishing\.arXiv preprint arXiv:2512\.15820\.External Links:2512\.15820v1Cited by:[§6](https://arxiv.org/html/2606.28933#S6.p1.1)\.
- D\. Feng and D\. Xue \(2025\)Spatiotemporal transformers for predicting avian disease risk from migration trajectories\.arXiv preprint arXiv:2510\.15254\.External Links:2510\.15254v1Cited by:[§2\.2](https://arxiv.org/html/2606.28933#S2.SS2.p1.1),[§3\.3\.4](https://arxiv.org/html/2606.28933#S3.SS3.SSS4.p1.1)\.
- C\. Fu, S\. Zhao, Y\. Zhang, Z\. Jian, S\. Zhao, and C\. Liu \(2025\)Personality\-guided public\-private domain disentangled hypergraph\-former network for multimodal depression detection\.arXiv preprint arXiv:2511\.12460\.External Links:2511\.12460v1Cited by:[§5](https://arxiv.org/html/2606.28933#S5.p5.1)\.
- G\. Gressel, R\. Pankajakshan, S\. Rozenfeld, L\. Li, I\. Franceschini, K\. Achuthan, and Y\. Mirsky \(2025\)Love, lies, and language models: investigating ai’s role in romance\-baiting scams\.Usenix Security Symposium 2026\.External Links:2512\.16280v1Cited by:[§4\.2](https://arxiv.org/html/2606.28933#S4.SS2.p4.1)\.
- Y\. Guan, Y\. Liu, K\. Zhou, Z\. Shen, J\. Hwang, S\. Belongie, and L\. Li \(2025\)Is meta\-learning out? rethinking unsupervised few\-shot classification with limited entropy\.arXiv preprint arXiv:2509\.13185\.External Links:2509\.13185v1Cited by:[§2\.4](https://arxiv.org/html/2606.28933#S2.SS4.p1.1),[§3\.6](https://arxiv.org/html/2606.28933#S3.SS6.p1.1),[§6](https://arxiv.org/html/2606.28933#S6.p4.11)\.
- Q\. Guo, B\. Khatri, W\. Sun, J\. Tang, H\. Zhang, and W\. Wang \(2025\)AquaSentinel: next\-generation ai system integrating sensor networks for urban underground water pipeline anomaly detection via collaborative moe\-llm agent architecture\.arXiv preprint arXiv:2511\.15870\.External Links:2511\.15870v1Cited by:[§5](https://arxiv.org/html/2606.28933#S5.p7.1)\.
- Q\. Hong, C\. Bian, X\. Zhou, X\. Li, Y\. Li, and Z\. Zeng \(2025\)Lost in time? a meta\-learning framework for time\-shift\-tolerant physiological signal transformation\.arXiv preprint arXiv:2511\.21500\.External Links:2511\.21500v1Cited by:[§2\.4](https://arxiv.org/html/2606.28933#S2.SS4.p1.1),[§3\.6](https://arxiv.org/html/2606.28933#S3.SS6.p1.1),[§4\.2](https://arxiv.org/html/2606.28933#S4.SS2.p7.1),[§7](https://arxiv.org/html/2606.28933#S7.p2.1)\.
- W\. Hsieh, Z\. Bi, K\. Chen, B\. Peng, S\. Zhang, J\. Xu, J\. Wang, C\. H\. Yin, Y\. Zhang, P\. Feng, Y\. Wen, T\. Wang, M\. Li, C\. X\. Liang, J\. Ren, Q\. Niu, S\. Chen, L\. K\. Q\. Yan, H\. Xu, H\. Tseng, X\. Song, B\. Jing, J\. Yang, J\. Song, J\. Liu, and M\. Liu \(2024\)Deep learning, machine learning, advancing big data analytics and management\.External Links:2412\.02187,[Link](https://arxiv.org/abs/2412.02187)Cited by:[§1](https://arxiv.org/html/2606.28933#S1.p2.1)\.
- J\. Hu, S\. Chen, Y\. He, Y\. Li, B\. Hooi, and B\. He \(2025\)Echoless label\-based pre\-computation for memory\-efficient heterogeneous graph learning\.arXiv preprint arXiv:2511\.11081\.External Links:2511\.11081v1Cited by:[§2\.1](https://arxiv.org/html/2606.28933#S2.SS1.p1.1)\.
- Y\. Huang, B\. Li, N\. Li, Z\. Wang, K\. Chen, H\. Ge, Q\. Si, Y\. Shen, R\. Yang, G\. Wang,et al\.\(2026\)GUI agents for continual game generation\.arXiv preprint arXiv:2605\.28258\.Cited by:[§1](https://arxiv.org/html/2606.28933#S1.p2.1)\.
- I\. Jeon, M\. Hong, J\. Yun, and G\. Kim \(2025a\)Federated learning via meta\-variational dropout\.Jeon, I\., Hong, M\., Yun, J\., Kim, G\. \(2023\)\. Federated Learning via Meta\-Variational Dropout\. Advances in Neural Information Processing Systems 36 \(NeurIPS 2023\)\.External Links:2510\.20225v1Cited by:[§2\.4](https://arxiv.org/html/2606.28933#S2.SS4.p1.1),[§3\.6\.1](https://arxiv.org/html/2606.28933#S3.SS6.SSS1.p1.5)\.
- I\. Jeon, Y\. Park, and G\. Kim \(2025b\)Neural variational dropout processes\.arXiv preprint arXiv:2510\.19425\.External Links:2510\.19425v1Cited by:[§2\.4](https://arxiv.org/html/2606.28933#S2.SS4.p1.1),[§3\.6](https://arxiv.org/html/2606.28933#S3.SS6.p1.1),[§4\.2](https://arxiv.org/html/2606.28933#S4.SS2.p7.1)\.
- L\. Jiang, R\. Ma, L\. Gu, Z\. Wang, X\. Zuo, and Y\. Wang \(2025\)PointMAC: meta\-learned adaptation for robust test\-time point cloud completion\.arXiv preprint arXiv:2510\.10365\.External Links:2510\.10365v1Cited by:[§2\.4](https://arxiv.org/html/2606.28933#S2.SS4.p1.1),[§3\.6](https://arxiv.org/html/2606.28933#S3.SS6.p1.1),[§6](https://arxiv.org/html/2606.28933#S6.p4.11)\.
- J\. Jin, L\. Mackey, and V\. Syrgkanis \(2025\)It’s hard to be normal: the impact of noise on structure\-agnostic estimation\.arXiv preprint arXiv:2507\.02275\.External Links:2507\.02275v3Cited by:[§2\.3](https://arxiv.org/html/2606.28933#S2.SS3.p1.1),[§5](https://arxiv.org/html/2606.28933#S5.p6.1)\.
- J\. Kemper and D\. Rostam\-Afschar \(2025\)Inference for batched adaptive experiments\.arXiv preprint arXiv:2512\.10156\.External Links:2512\.10156v1Cited by:[§1](https://arxiv.org/html/2606.28933#S1.p2.1),[§6](https://arxiv.org/html/2606.28933#S6.p2.1)\.
- B\. Kempinski and T\. Kachman \(2025\)Going with the flow: approximating banzhaf values via graph neural networks\.arXiv preprint arXiv:2510\.13391\.External Links:2510\.13391v2Cited by:[§2\.1](https://arxiv.org/html/2606.28933#S2.SS1.p1.1),[§4\.2](https://arxiv.org/html/2606.28933#S4.SS2.p2.1)\.
- H\. Kim, S\. Seo, H\. Choi, T\. Boomstra, J\. Yoon, and C\. Park \(2025\)Better prevent than tackle: valuing defense in soccer based on graph neural networks\.arXiv preprint arXiv:2512\.10355\.External Links:2512\.10355v1Cited by:[§2\.1](https://arxiv.org/html/2606.28933#S2.SS1.p1.1),[§3\.2\.2](https://arxiv.org/html/2606.28933#S3.SS2.SSS2.p1.2),[§7](https://arxiv.org/html/2606.28933#S7.p1.1)\.
- Y\. Kiu, Lau, C\. Chen, G\. Jin, and C\. Feng \(2025\)Flexible and efficient spatio\-temporal transformer for sequential visual place recognition\.arXiv preprint arXiv:2510\.04282\.External Links:2510\.04282v1Cited by:[§2\.2](https://arxiv.org/html/2606.28933#S2.SS2.p1.1),[§3\.3\.3](https://arxiv.org/html/2606.28933#S3.SS3.SSS3.p1.3)\.
- K\. Kubota and S\. Sugasawa \(2025\)Causal inference under threshold manipulation: bayesian mixture modeling and heterogeneous treatment effects\.arXiv preprint arXiv:2509\.19814\.External Links:2509\.19814v1Cited by:[§2\.3](https://arxiv.org/html/2606.28933#S2.SS3.p1.1),[§5](https://arxiv.org/html/2606.28933#S5.p8.1)\.
- I\. Kukanov and J\. W\. Ng \(2025\)KLASSify to verify: audio\-visual deepfake detection using ssl\-based audio and handcrafted visual features\.arXiv preprint arXiv:2508\.07337\.External Links:2508\.07337v1Cited by:[§6](https://arxiv.org/html/2606.28933#S6.p1.1)\.
- U\. Lee, C\. Chung, J\. Lee, and S\. Moon \(2025\)BugSweeper: function\-level detection of smart contract vulnerabilities using graph neural networks\.arXiv preprint arXiv:2512\.09385\.External Links:2512\.09385v2Cited by:[§2\.1](https://arxiv.org/html/2606.28933#S2.SS1.p1.1),[§3\.1](https://arxiv.org/html/2606.28933#S3.SS1.p2.7),[§4\.2](https://arxiv.org/html/2606.28933#S4.SS2.p4.1)\.
- B\. Li, B\. Gu, and Z\. Ding \(2025a\)LLM\-based personalized portfolio recommender: integrating large language models and reinforcement learning for intelligent investment strategy optimization\.arXiv preprint arXiv:2512\.12922\.External Links:2512\.12922v1Cited by:[§1](https://arxiv.org/html/2606.28933#S1.p1.1),[§4\.1](https://arxiv.org/html/2606.28933#S4.SS1.p1.2),[§7](https://arxiv.org/html/2606.28933#S7.p3.1)\.
- H\. H\. Li, M\. Irgau, N\. Janmohamed, K\. S\. Rieckmann, and D\. B\. Lobell \(2025b\)Scalable vision\-guided crop yield estimation\.arXiv preprint arXiv:2511\.12999\.External Links:2511\.12999v1Cited by:[§6](https://arxiv.org/html/2606.28933#S6.p2.1)\.
- X\. Li, Y\. Zeng, X\. Xing, J\. Xu, and X\. Xu \(2025c\)QuantAgents: towards multi\-agent financial system via simulated trading\.arXiv preprint arXiv:2510\.04643\.External Links:2510\.04643v1Cited by:[§2\.4](https://arxiv.org/html/2606.28933#S2.SS4.p1.1),[§7](https://arxiv.org/html/2606.28933#S7.p1.1)\.
- Z\. Li and R\. Fan \(2025\)Crisis\-resilient portfolio management via graph\-based spatio\-temporal learning\.arXiv preprint arXiv:2510\.20868\.External Links:2510\.20868v1Cited by:[§2\.1](https://arxiv.org/html/2606.28933#S2.SS1.p1.1),[§3\.2\.2](https://arxiv.org/html/2606.28933#S3.SS2.SSS2.p1.2),[§7](https://arxiv.org/html/2606.28933#S7.p1.1)\.
- A\. Liu, J\. Wang, S\. Kaski, J\. Wang, and M\. Yang \(2025a\)A principle of targeted intervention for multi\-agent reinforcement learning\.arXiv preprint arXiv:2510\.17697\.External Links:2510\.17697v4Cited by:[§2\.3](https://arxiv.org/html/2606.28933#S2.SS3.p1.1),[§3\.6\.2](https://arxiv.org/html/2606.28933#S3.SS6.SSS2.p2.1)\.
- B\. Liu, Q\. Qin, and Q\. He \(2025b\)CausalCLIP: causally\-informed feature disentanglement and filtering for generalizable detection of generated images\.arXiv preprint arXiv:2512\.13285\.External Links:2512\.13285v2Cited by:[§2\.3](https://arxiv.org/html/2606.28933#S2.SS3.p1.1),[§5](https://arxiv.org/html/2606.28933#S5.p8.1)\.
- W\. Liu, J\. Pan, X\. Zhang, X\. Gong, Y\. Ye, X\. Zhao, X\. Wang, K\. Wu, H\. Xiang, H\. Yan, and Q\. Zhang \(2025c\)Cross\-platform product matching based on entity alignment of knowledge graph with raea model\.World Wide Web, vol\. 26\. no\. 4, pp\.2215\-2235, 2023\.External Links:2512\.07232v1Cited by:[§2\.1](https://arxiv.org/html/2606.28933#S2.SS1.p1.1)\.
- P\. Loganathan, E\. Zea, R\. Vinuesa, and E\. Otero \(2025\)Deep learning\-driven downscaling for climate risk assessment of projected temperature extremes in the nordic region\.arXiv preprint arXiv:2511\.03770\.External Links:2511\.03770v1Cited by:[§5](https://arxiv.org/html/2606.28933#S5.p3.1)\.
- J\. Long, Y\. Zhu, C\. Tang, K\. Sun, Y\. Liu, and X\. Yan \(2025\)When genes speak: a semantic\-guided framework for spatially resolved transcriptomics data clustering\.arXiv preprint arXiv:2511\.11380\.External Links:2511\.11380v1Cited by:[§5](https://arxiv.org/html/2606.28933#S5.p7.1)\.
- X\. Ma, X\. Ma, S\. M\. Erfani, D\. Mandic, and J\. Bailey \(2025\)Coarse\-to\-fine open\-set graph node classification with large language models\.arXiv preprint arXiv:2512\.16244\.External Links:2512\.16244v1Cited by:[§2\.1](https://arxiv.org/html/2606.28933#S2.SS1.p1.1),[§3\.1](https://arxiv.org/html/2606.28933#S3.SS1.p2.7)\.
- S\. Mahadevan \(2025\)Large causal models from large language models\.arXiv preprint arXiv:2512\.07796\.External Links:2512\.07796v1Cited by:[§2\.3](https://arxiv.org/html/2606.28933#S2.SS3.p1.1),[§3\.4\.3](https://arxiv.org/html/2606.28933#S3.SS4.SSS3.Px3.p2.5),[§6](https://arxiv.org/html/2606.28933#S6.p3.4)\.
- M\. Mo, Y\. Tan, H\. Zhang, H\. Zhang, and Y\. He \(2026\)ShieldedCode: learning robust representations for virtual machine protected code\.arXiv preprint arXiv:2601\.20679\.Cited by:[§1](https://arxiv.org/html/2606.28933#S1.p1.1)\.
- A\. Muneer, K\. Zhang, I\. Hamdi, R\. Qureshi, M\. Waqas, S\. Fouad, H\. Ali, S\. M\. Anwar, and J\. Wu \(2025\)Foundation models in biomedical imaging: turning hype into reality\.arXiv preprint arXiv:2512\.15808\.External Links:2512\.15808v1Cited by:[§5](https://arxiv.org/html/2606.28933#S5.p7.1)\.
- L\. L\. Oganesian, S\. Hashemi, and M\. M\. Shanechi \(2025\)BaRISTA: brain scale informed spatiotemporal representation of human intracranial neural activity\.NeurIPS 2025\.External Links:2512\.12135v1Cited by:[§2\.2](https://arxiv.org/html/2606.28933#S2.SS2.p1.1),[§7](https://arxiv.org/html/2606.28933#S7.p1.1)\.
- Á\. Parafita, T\. Garriga, A\. Brando, and F\. J\. Cazorla \(2025\)Practical do\-shapley explanations with estimand\-agnostic causal inference\.arXiv preprint arXiv:2509\.20211\.External Links:2509\.20211v1Cited by:[§2\.3](https://arxiv.org/html/2606.28933#S2.SS3.p1.1),[§3\.4\.3](https://arxiv.org/html/2606.28933#S3.SS4.SSS3.Px3.p2.5),[§6](https://arxiv.org/html/2606.28933#S6.p3.4)\.
- B\. Peng, X\. Pan, Y\. Wen, Z\. Bi, K\. Chen, M\. Li, M\. Liu, Q\. Niu, J\. Liu, J\. Wang, S\. Zhang, J\. Xu, X\. Song, Z\. Jiang, T\. Wang, and P\. Feng \(2025\)Deep learning and machine learning, advancing big data analytics and management: handy appetizer\.External Links:2409\.17120,[Link](https://arxiv.org/abs/2409.17120)Cited by:[§1](https://arxiv.org/html/2606.28933#S1.p2.1)\.
- A\. Perekhodko and R\. Ślepaczuk \(2025\)Stochastic volatility modelling with lstm networks: a hybrid approach for s&p 500 index volatility forecasting\.arXiv preprint arXiv:2512\.12250\.External Links:2512\.12250v1Cited by:[§2\.1](https://arxiv.org/html/2606.28933#S2.SS1.p1.1),[§2\.2](https://arxiv.org/html/2606.28933#S2.SS2.p1.1)\.
- S\. Prabhune, B\. Padmanabhan, and K\. Dutta \(2025\)Information\-consistent language model recommendations through group relative policy optimization\.arXiv preprint arXiv:2512\.12858\.External Links:2512\.12858v1Cited by:[§5](https://arxiv.org/html/2606.28933#S5.p5.1)\.
- J\. Ren, Z\. Bi, Q\. Niu, X\. Song, Z\. Jiang, J\. Liu, B\. Peng, S\. Zhang, X\. Pan, J\. Wang, K\. Chen, C\. H\. Yin, P\. Feng, Y\. Wen, T\. Wang, S\. Chen, M\. Li, J\. Xu, and M\. Liu \(2025\)Deep learning and machine learning – object detection and semantic segmentation: from theory to applications\.External Links:2410\.15584,[Link](https://arxiv.org/abs/2410.15584)Cited by:[§1](https://arxiv.org/html/2606.28933#S1.p2.1)\.
- M\. Sekar and N\. Nezamoddini \(2025\)Optimizing multi\-lane intersection performance in mixed autonomy environments\.arXiv preprint arXiv:2511\.02217\.External Links:2511\.02217v1Cited by:[§5](https://arxiv.org/html/2606.28933#S5.p7.1)\.
- R\. C\. Shit and S\. Subudhi \(2025\)Hierarchical federated graph attention networks for scalable and resilient uav collision avoidance\.arXiv preprint arXiv:2511\.11616\.External Links:2511\.11616v1Cited by:[§2\.1](https://arxiv.org/html/2606.28933#S2.SS1.p1.1),[§3\.2\.1](https://arxiv.org/html/2606.28933#S3.SS2.SSS1.p1.2)\.
- M\. C\. Vu, T\. L\. Dinh, M\. C\. Vu, T\. D\. Le, and T\. L\. H\. Nguyen \(2025\)A conceptual model for ai adoption in financial decision\-making: addressing the unique challenges of small and medium\-sized enterprises\.arXiv preprint arXiv:2512\.04339\.External Links:2512\.04339v1Cited by:[§2\.4](https://arxiv.org/html/2606.28933#S2.SS4.p1.1),[§6](https://arxiv.org/html/2606.28933#S6.p2.1)\.
- S\. Wang, Z\. Guan, B\. Zhao, and T\. Gu \(2025a\)CaSTFormer: causal spatio\-temporal transformer for driving intention prediction\.arXiv preprint arXiv:2507\.13425\.External Links:2507\.13425v1Cited by:[§2\.2](https://arxiv.org/html/2606.28933#S2.SS2.p1.1),[§3\.3\.3](https://arxiv.org/html/2606.28933#S3.SS3.SSS3.p1.3)\.
- X\. Wang, P\. L\. Rizzini, S\. Medya, and Z\. Lan \(2025b\)SMART: a surrogate model for predicting application runtime in dragonfly systems\.arXiv preprint arXiv:2511\.11111\.External Links:2511\.11111v1Cited by:[§1](https://arxiv.org/html/2606.28933#S1.p2.1)\.
- Z\. Wang, S\. Raj, and R\. Buyya \(2025c\)AirFed: a federated graph\-enhanced multi\-agent reinforcement learning framework for multi\-uav cooperative mobile edge computing\.arXiv preprint arXiv:2510\.23053\.External Links:2510\.23053v2Cited by:[§2\.1](https://arxiv.org/html/2606.28933#S2.SS1.p1.1),[§3\.2\.3](https://arxiv.org/html/2606.28933#S3.SS2.SSS3.p1.1),[§4\.3](https://arxiv.org/html/2606.28933#S4.SS3.p3.1)\.
- T\. Wu, S\. Zhu, J\. Wang, N\. Xu, G\. Qi, and H\. Wang \(2025\)Uncertain knowledge graph completion via semi\-supervised confidence distribution learning\.arXiv preprint arXiv:2510\.16601\.External Links:2510\.16601v2Cited by:[§5](https://arxiv.org/html/2606.28933#S5.p5.1)\.
- J\. Xu, Z\. Chen, S\. Yang, J\. Li, H\. Wang, Y\. Li, and E\. C\. H\. Ngai \(2025\)Learning and editing universal graph prompt tuning via reinforcement learning\.arXiv preprint arXiv:2512\.08763\.External Links:2512\.08763v1Cited by:[§2\.1](https://arxiv.org/html/2606.28933#S2.SS1.p1.1),[§3\.1](https://arxiv.org/html/2606.28933#S3.SS1.p2.7),[§5](https://arxiv.org/html/2606.28933#S5.p2.1)\.
- W\. Yao, B\. Huang, and S\. Dev \(2025\)Multi\-modal spatio\-temporal transformer for high\-resolution land subsidence prediction\.arXiv preprint arXiv:2509\.25393\.External Links:2509\.25393v2Cited by:[§2\.2](https://arxiv.org/html/2606.28933#S2.SS2.p1.1),[§3\.3\.4](https://arxiv.org/html/2606.28933#S3.SS3.SSS4.p1.1)\.
- Z\. Ye, J\. Lu, T\. Gu, F\. Hao, and X\. Wang \(2025\)FairGSE: fairness\-aware graph neural network without high false positive rates\.arXiv preprint arXiv:2511\.12132\.External Links:2511\.12132v1Cited by:[§2\.1](https://arxiv.org/html/2606.28933#S2.SS1.p1.1),[§5](https://arxiv.org/html/2606.28933#S5.p2.1)\.
- Z\. Yin, X\. Chen, and X\. Zhang \(2025\)AI\-integrated decision support system for real\-time market growth forecasting and multi\-source content diffusion analytics\.arXiv preprint arXiv:2511\.09962\.External Links:2511\.09962v1Cited by:[§1](https://arxiv.org/html/2606.28933#S1.p1.1),[§2\.2](https://arxiv.org/html/2606.28933#S2.SS2.p1.1),[§7](https://arxiv.org/html/2606.28933#S7.p1.1)\.
- M\. You, K\. Chen, and D\. Cheng \(2026\)Drdgrl: dual\-relational dynamic graph representation learning for delay\-sensitive stock trend prediction\.InInternational Conference on Database Systems for Advanced Applications,pp\. 35–50\.Cited by:[§1](https://arxiv.org/html/2606.28933#S1.p2.1)\.
- W\. Yu, S\. Wei, J\. Liu, Y\. Li, M\. Hu, A\. Liu, H\. Zhang, and I\. King \(2026\)Probability\-entropy calibration: an elastic indicator for adaptive fine\-tuning\.arXiv preprint arXiv:2602\.01745\.Cited by:[§1](https://arxiv.org/html/2606.28933#S1.p1.1)\.
- J\. Yun, M\. Hong, and G\. Kim \(2025\)FedMeNF: privacy\-preserving federated meta\-learning for neural fields\.arXiv preprint arXiv:2508\.06301\.External Links:2508\.06301v1Cited by:[§2\.4](https://arxiv.org/html/2606.28933#S2.SS4.p1.1),[§3\.6\.1](https://arxiv.org/html/2606.28933#S3.SS6.SSS1.p1.5),[§7](https://arxiv.org/html/2606.28933#S7.p2.1)\.
- G\. Zeng, H\. Peng, A\. Li, L\. Sun, C\. Liu, S\. Li, Y\. Pan, and P\. S\. Yu \(2025\)Hyperbolic continuous structural entropy for hierarchical clustering\.arXiv preprint arXiv:2512\.00524\.External Links:2512\.00524v1Cited by:[§1](https://arxiv.org/html/2606.28933#S1.p2.1)\.
- H\. Zhang, B\. Huang, Z\. Li, X\. Xiao, H\. Y\. Leong, Z\. Zhang, X\. Long, T\. Wang, and H\. Xu \(2025a\)Sensitivity\-lora: low\-load sensitivity\-based fine\-tuning for large language models\.arXiv preprint arXiv:2509\.09119\.Cited by:[§1](https://arxiv.org/html/2606.28933#S1.p1.1)\.
- H\. Zhang, Z\. Li, R\. Bao, Y\. Gao, X\. Xiao, H\. Zhang, S\. Zhang, B\. Huang, Y\. Wu, T\. Wang,et al\.\(2025b\)HyperAdaLoRA: accelerating lora rank allocation during training via hypernetworks without sacrificing performance\.arXiv preprint arXiv:2510\.02630\.Cited by:[§1](https://arxiv.org/html/2606.28933#S1.p1.1)\.
- H\. Zhang, M\. Lyu, Z\. Chen, X\. Xing, Y\. Ao, and Y\. Lin \(2025c\)Pdtrim: targeted pruning for prefill\-decode disaggregation in inference\.arXiv preprint arXiv:2509\.04467\.Cited by:[§1](https://arxiv.org/html/2606.28933#S1.p1.1)\.
- H\. Zhang, M\. Lyu, C\. He, Y\. Ao, and Y\. Lin \(2025d\)Trimtokenator: towards adaptive visual token pruning for large multimodal models\.arXiv preprint arXiv:2509\.00320\.Cited by:[§1](https://arxiv.org/html/2606.28933#S1.p1.1)\.
- H\. Zhang, M\. Lyu, B\. Huang, Y\. Ao, and Y\. Lin \(2025e\)TrimTokenator\-lc: towards adaptive visual token pruning for large multimodal models with long contexts\.arXiv preprint arXiv:2512\.22748\.Cited by:[§1](https://arxiv.org/html/2606.28933#S1.p1.1)\.
- H\. Zhang, X\. Mao, G\. Dong, Z\. Li, X\. Su, K\. Chen, J\. Yang, and Z\. Lin \(2026a\)MemMark: state\-evolution attribution watermarking for agent long\-term memory systems\.arXiv preprint arXiv:2605\.25002\.Cited by:[§1](https://arxiv.org/html/2606.28933#S1.p2.1)\.
- H\. Zhang, H\. You, Z\. Zhang, L\. Gan, H\. Zhang, W\. Huang, and J\. Huang \(2026b\)Mitigating generic token dominance in cross\-domain foundation model for text\-attributed graphs\.InInternational Conference on Database Systems for Advanced Applications,pp\. 251–265\.Cited by:[§1](https://arxiv.org/html/2606.28933#S1.p1.1)\.
- J\. Zhang, Z\. Zhu, J\. Deng, Y\. Li, and B\. Wang \(2025f\)Multi\-modal feature fusion for spatial morphology analysis of traditional villages via hierarchical graph neural networks\.arXiv preprint arXiv:2510\.27208\.External Links:2510\.27208v1Cited by:[§2\.1](https://arxiv.org/html/2606.28933#S2.SS1.p1.1),[§3\.2\.3](https://arxiv.org/html/2606.28933#S3.SS2.SSS3.p1.1),[§4\.2](https://arxiv.org/html/2606.28933#S4.SS2.p2.1)\.
- L\. Zhang and H\. Cai \(2025\)Text rationalization for robust causal effect estimation\.arXiv preprint arXiv:2512\.05373\.External Links:2512\.05373v1Cited by:[§2\.3](https://arxiv.org/html/2606.28933#S2.SS3.p1.1),[§3\.4\.3](https://arxiv.org/html/2606.28933#S3.SS4.SSS3.p1.1),[§7](https://arxiv.org/html/2606.28933#S7.p1.1)\.
- Q\. Zhao, Z\. Dou, D\. Zhang, X\. Li, C\. Song, Z\. Wan, X\. Li, Y\. Zhang, K\. Chen, Q\. Pan,et al\.\(2026\)STRIDE: strategic trajectory reasoning via discriminative estimation for verifiable reinforcement learning\.arXiv preprint arXiv:2606\.15866\.Cited by:[§1](https://arxiv.org/html/2606.28933#S1.p2.1)\.
- Y\. Zou, Q\. Liu, J\. Wu, Y\. Peng, G\. Chen, H\. Zhou, and G\. Ye \(2025\)Boosting adversarial transferability via ensemble non\-attention\.arXiv preprint arXiv:2511\.08937\.External Links:2511\.08937v2Cited by:[§4\.1](https://arxiv.org/html/2606.28933#S4.SS1.p4.2)\.

Similar Articles

A Temporally Augmented Graph Attention Network for Affordance Classification

Hugging Face Daily Papers

EEG-tGAT is a temporally augmented Graph Attention Network that improves affordance classification from interaction sequences by incorporating temporal attention and dropout mechanisms. The model enhances GATv2 for sequential data where temporal dimensions are semantically non-uniform.

Graph-Based Financial Fraud Detection with Calibrated Risk Scoring and Structural Regularization

arXiv cs.LG

This paper proposes a graph neural network framework for financial fraud detection that integrates transaction records and identity information into node attributes, employs a multi-layer message passing mechanism, and uses weighted supervision and structural consistency regularization to improve risk scoring and probability calibration. Experiments on a public dataset show the method outperforms existing approaches.