Deep Reinforcement Learning for Reliability Based Bi-Objective Portfolio Optimization

arXiv cs.LG Papers

Summary

This paper proposes a deep reinforcement learning framework (MORP-DRL) for multi-objective reliability-based portfolio optimization, jointly optimizing expected return and downside risk using CVaR and EVaR under practical constraints, and demonstrates performance on global equity indices across different market regimes.

arXiv:2607.06610v1 Announce Type: new Abstract: Portfolio optimization under uncertainty is inherently a multi-objective decision problem involving complex interactions among return, risk, market dynamics, and practical investment constraints. Existing reliability based portfolio optimization approaches primarily rely on static optimization frameworks and often fail to capture sequential decision making, tail risk, and market frictions such as transaction costs. To address these limitations, we propose a deep reinforcement learning framework for multi-objective reliability based portfolio optimization (MORP-DRL). The proposed framework jointly optimizes expected return and downside risk using three complementary risk measures: variance, Conditional Value-at-Risk (CVaR), and Entropic Value-at-Risk (EVaR). To model uncertainty and heavy-tailed market behavior, asset returns are represented using GARCH(1,1), Extreme Value Theory, and a t-copula dependence structure, while realistic scenarios are generated through quasi-Monte Carlo simulation. A Proximal Policy Optimization (PPO) based strategy is developed under practical constraints including transaction costs and portfolio bounds, and is benchmarked against NSGA-II. Experiments on ten global equity indices across pre-COVID, COVID, and post-COVID market regimes demonstrate that MORP-DRL achieves competitive risk-return performance, reduced downside risk during periods of market stress, and scalability to high-dimensional portfolio settings.
Original Article
View Cached Full Text

Cached at: 07/09/26, 07:41 AM

# Deep Reinforcement Learning for Reliability Based Bi-Objective Portfolio Optimization
Source: [https://arxiv.org/html/2607.06610](https://arxiv.org/html/2607.06610)
Sounaq Das Indian Institute of Management Amritsar Amritsar, India dassounaq@gmail\.com &Tanmay Sen SQC & OR Unit Indian Statistical Institute Kolkata Kolkata, India tanmay\.sen@isical\.ac\.in &Raghu Nandan Sengupta Department of Management Sciences Indian Institute of Technology Kanpur Kanpur – 208 016, India raghus@iitk\.ac\.in &Aditya Gupta McKinsey and Company Gurgaon, India aditya\_gupta\-guva@mckinsey\.com

###### Abstract

Portfolio optimization under uncertainty is inherently a multi\-objective decision problem involving complex interactions among return, risk, market dynamics, and practical investment constraints\. Existing reliability based portfolio optimization approaches primarily rely on static optimization frameworks and often fail to capture sequential decision making, tail risk, and market frictions such as transaction costs\. To address these limitations, we propose a deep reinforcement learning framework for multi\-objective reliability based portfolio optimization \(MORP\-DRL\)\. The proposed framework jointly optimizes expected return and downside risk using three complementary risk measures: variance, Conditional Value\-at\-Risk \(CVaR\), and Entropic Value\-at\-Risk \(EVaR\)\. To model uncertainty and heavy\-tailed market behavior, asset returns are represented using GARCH\(1,1\), Extreme Value Theory, and att\-copula dependence structure, while realistic scenarios are generated through quasi\-Monte Carlo simulation\. A Proximal Policy Optimization \(PPO\) based strategy is developed under practical constraints including transaction costs and portfolio bounds, and is benchmarked against NSGA\-II\. Experiments on ten global equity indices across pre\-COVID, COVID, and post\-COVID market regimes demonstrate that MORP\-DRL achieves competitive risk\-return performance, reduced downside risk during periods of market stress, and scalability to high\-dimensional portfolio settings\.

Keywords:Bi\-objective, Portfolio Optimization, Deep Reinforcement Learning, Extreme Value Theory \(EVT\), Entropic Value\-at\-Risk \(EVaR\), NSGA\-II , QMC, t\-copula

## 1Introduction

Portfolio optimization is one of the most fundamental problems in quantitative finance and investment management\. The primary objective is to allocate capital among different financial assets in a way that balances expected return and investment risk while satisfying practical investment constraints\. The seminal work of MarkowitzMarkowitz \([1952a](https://arxiv.org/html/2607.06610#bib.bib23)\); Fabozziet al\.\([2008](https://arxiv.org/html/2607.06610#bib.bib24)\)introduced the modern portfolio theory \(MPT\), which formulated portfolio selection as a mean\-variance optimization problem\. This framework established the foundation of modern portfolio management by emphasizing diversification and the trade\-off between return and risk\. Subsequent developments, including the capital asset Pricing model \(CAPM\)Fama and French \([2004](https://arxiv.org/html/2607.06610#bib.bib10)\), Black\-Litterman model, and factor based investment strategies, further extended classical portfolio theory by incorporating systematic market risk and investor preferences\. Despite their importance, these traditional models rely on several simplifying assumptions, such as normally distributed returns, linear dependence structures, constant covariance matrices, and frictionless markets\. However, financial markets behave very differently\. Asset returns frequently exhibit sudden fluctuations, changing dependence patterns, periods of high volatility, and extreme market movements, particularly during financial crises and uncertain economic conditions\. As a result, traditional variance\-based approaches may not adequately capture downside and extreme market risks\. To address these limitations, alternative downside risk measures have been extensively investigated in the literature\. Value\-at\-Risk \(VaR\)Linsmeier and Pearson \([2000](https://arxiv.org/html/2607.06610#bib.bib9)\)emerged as one of the earliest and most widely adopted downside risk measures in financial risk management\. Later,Rockafellar and Uryasev \([2002](https://arxiv.org/html/2607.06610#bib.bib30)\)introduced Conditional Value\-at\-Risk \(CVaR\), which measures the expected loss beyond the VaR threshold and satisfies the properties of a coherent risk measure\. More recently, Entropic Value\-at\-Risk \(EVaR\)Ramoset al\.\([2023](https://arxiv.org/html/2607.06610#bib.bib26)\); Righi and Borenstein \([2018](https://arxiv.org/html/2607.06610#bib.bib27)\)and related tail\-risk measures have gained attention due to their stronger theoretical properties and ability to capture extreme losses under non\-Gaussian market conditions\.

At the same time, practical portfolio optimization problems have become increasingly challenging due to the incorporation of realistic investment constraints such as transaction costs, cardinality restrictions, minimum and maximum holding limits, and budget constraints\. The inclusion of such constraints often leads to highly nonlinear and non\-convex optimization problems that are difficult to solve using conventional convex optimization techniques\. To overcome these challenges, several studies have explored metaheuristic and evolutionary optimization methods, including Genetic Algorithms \(GA\), Particle Swarm Optimization \(PSO\), Ant Colony Optimization \(ACO\), and Non\-dominated Sorting Genetic Algorithm\-II \(NSGA\-II\)Lokeet al\.\([2023](https://arxiv.org/html/2607.06610#bib.bib32)\); Saloet al\.\([2024](https://arxiv.org/html/2607.06610#bib.bib33)\); Ertenlice and Kalayci \([2018](https://arxiv.org/html/2607.06610#bib.bib35)\)\. These approaches have demonstrated strong capabilities in handling multi\-objective optimization problems and complex feasible regions\.

In recent years, the rapid growth of financial data and increasing market complexity have encouraged the use of Machine Learning \(ML\) and Deep Learning \(DL\)Goodfellowet al\.\([2016](https://arxiv.org/html/2607.06610#bib.bib3)\)techniques in portfolio optimizationFischer and Krauss \([2018](https://arxiv.org/html/2607.06610#bib.bib1)\); Beheraet al\.\([2023](https://arxiv.org/html/2607.06610#bib.bib2)\)\. Early ML based approaches primarily focused on predicting asset returns using methods such as Support Vector Machines, Random Forests, and Artificial Neural Networks\. Although these approaches improved predictive capability, portfolio decisions were typically generated through separate static optimization procedures and lacked adaptability to rapidly changing market conditions\. Furthermore, these methods often failed to capture the sequential and dynamic nature of investment decisions\. To address these limitations, Deep Reinforcement Learning \(DRL\)Arulkumaranet al\.\([2017](https://arxiv.org/html/2607.06610#bib.bib11)\); Choudharyet al\.\([2025](https://arxiv.org/html/2607.06610#bib.bib4)\)has emerged as a promising paradigm for dynamic portfolio management by learning allocation strategies directly through interactions with market environments\. Unlike supervised learning methods, DRL can optimize long term investment objectives while continuously adapting to evolving market conditions\. Recent studies have shown the effectiveness of actor\-critic architectures, Deep Q\-Networks, and policy gradient methods in financial applicationsGuet al\.\([2020](https://arxiv.org/html/2607.06610#bib.bib29)\)\. Despite recent progress, several important limitations remain insufficiently addressed in the existing literature\. First, many DRL based portfolio optimization frameworks either neglect transaction costs or model them in a simplified manner, resulting in unrealistic rebalancing strategies\. Second, although tail\-risk measures such as CVaR and EVaR have been investigated independentlyChoudharyet al\.\([2026](https://arxiv.org/html/2607.06610#bib.bib8)\), their integration within a reliability based portfolio optimization framework remains limited\. Third, most existing DRL approaches do not explicitly incorporate probabilistic reliability constraints to ensure robust portfolio feasibility under uncertain market conditions and nonlinear dependence structures\. Finally, comparative studies between classical evolutionary optimization approaches and DRL based optimization across different market regimes remain relatively scarce\.

Motivated by these research gaps, this study proposes a reliability based Bi\-objective portfolio optimization framework integrated with Deep Reinforcement Learning\. The proposed framework incorporates transaction costs, probabilistic reliability constraints, and multiple downside risk measures within a unified optimization setting\. In particular, three complementary portfolio optimization models are developed using variance, Conditional Value\-at\-Risk \(CVaR\), and Entropic Value\-at\-Risk \(EVaR\) as alternative risk measures\. To capture nonlinear dependence and tail co\-movement among financial assets, Quasi\-Monte Carlo \(QMC\) simulation is combined with att\-copula based dependence structure for reliability estimation under uncertainty\.

Furthermore, a PPO\-based reinforcement learning framework is developed to dynamically optimize portfolio allocation policies under realistic market constraints\. The proposed approach is evaluated using a diversified portfolio consisting of major global equity indices across three distinct market regimes, namely the pre\-COVID, COVID, and post\-COVID periods\. The performance of the proposed DRL framework is compared with equal weight benchmark portfolios and the classical NSGA\-II multi\-objective optimization algorithm\.

The major contributions of this work are summarized below:

1. 1\.We develop a reliability based bi\-objective portfolio optimization framework for practical portfolio rebalancing under transaction cost constraints using three complementary risk measures: variance, Conditional Value\-at\-Risk \(CVaR\), and Entropic Value\-at\-Risk \(EVaR\)\.
2. 2\.Reliability constraints are evaluated using Quasi\-Monte Carlo \(QMC\) simulation together with a t\-copula dependence model, enabling the framework to capture uncertainty, nonlinear dependence, and tail dependence among global financial assets\.
3. 3\.We develop a Proximal Policy Optimization \(PPO\) based deep reinforcement learning framework that learns adaptive portfolio allocation policies under joint risk and reliability constraints, and compare its performance with the classical NSGA\-II multi\-objective optimization algorithm
4. 4\.Extensive experiments are conducted across pre\-COVID, COVID, and post\-COVID market regimes\. The proposed approaches are benchmarked against equal\-weight portfolios and NSGA\-II under variance, CVaR, and EVaR risk measures\.
5. 5\.To evaluate the scalability of the proposed framework, additional experiments are performed on the FTSE100 universe, showing that both NSGA\-II and PPO remain effective in high\-dimensional portfolio optimization while exhibiting distinct portfolio allocation characteristics\.

The remainder of this paper is organized as follows\. Section 2 presents a comprehensive literature survey on recent advances and research trends in portfolio optimization\. Section 3 formulates the portfolio optimization problem and introduces the background concepts and risk measures considered in this study\. Section 4 describes the proposed bi\-objective portfolio optimization models\. Section 5 presents the proposed methodology, including the PPO\-based deep reinforcement learning framework\. Section 6 details the experimental setup, data preprocessing procedures, and implementation settings\. Section 7 reports the experimental results and discussion, including a comparative analysis across different market regimes\. Finally, Section 8 concludes the paper and highlights potential directions for future research\.

## 2Literature Survey

Portfolio optimization has been one of the central problems in financial decision making since the seminal work ofMarkowitz \([1952b](https://arxiv.org/html/2607.06610#bib.bib7)\), who introduced the mean–variance framework to balance expected return and risk\. Over the years, several extensions of the classical portfolio optimization problem have been proposed, including multi\-objective formulations, stochastic programming approaches, and portfolio models incorporating transaction costs and practical investment constraintsKolmet al\.\([2014](https://arxiv.org/html/2607.06610#bib.bib12)\); Meghwani and Thakur \([2018](https://arxiv.org/html/2607.06610#bib.bib13)\)\. Although these traditional optimization frameworks are mathematically rigorous, they are typically based on static assumptions and often struggle to capture the highly dynamic and uncertain nature of real financial markets\. To overcome these limitations, researchers increasingly explored machine learning \(ML\) and deep learning \(DL\) methods for financial forecasting and portfolio management\. Early studies primarily focused on predicting stock returns, volatility, or market trends using neural networks and deep architecturesFischer and Krauss \([2018](https://arxiv.org/html/2607.06610#bib.bib1)\); Beheraet al\.\([2023](https://arxiv.org/html/2607.06610#bib.bib2)\)\. However, in many of these approaches, prediction and portfolio allocation are treated as separate tasks\. Specifically, the models first estimate future returns or volatility and subsequently perform portfolio optimization, which may lead to suboptimal decision\-making under rapidly changing market conditions\.

Recently, Deep Reinforcement Learning \(DRL\) has emerged as a promising framework for portfolio optimization due to its ability to model sequential decision\-making under uncertainty\. Unlike traditional optimization methods, DRL enables an agent to continuously interact with the market environment and learn adaptive portfolio allocation policies over time\.Chauet al\.\([2025](https://arxiv.org/html/2607.06610#bib.bib47)\)provides a theoretical foundation for reinforcement learning in portfolio optimization by reformulating continuous\-time portfolio selection as an entropy\-regularized sequential decision\-making problem\. The proposed framework demonstrates the feasibility of integrating reinforcement learning with stochastic optimal control under realistic portfolio constraints, thereby motivating subsequent DRL\-based portfolio optimization research\.Chakraborty \([2019](https://arxiv.org/html/2607.06610#bib.bib14)\)investigated the use of DRL algorithms for generating profitable trading strategies within a Markov Decision Process \(MDP\) framework\. Li et al\.Liet al\.\([2019](https://arxiv.org/html/2607.06610#bib.bib15)\)proposed a deep reinforcement learning ensemble framework combining PPO, A2C, and DDPG for stock trading\. Their method achieved improved risk adjusted returns and outperformed traditional portfolio allocation strategies\. Similarly,Gaoet al\.\([2020](https://arxiv.org/html/2607.06610#bib.bib16)\)proposed a DQN\-based portfolio management framework with discretized portfolio weights and dueling network architectures to improve trading performance\. Hierarchical DRL architectures incorporating transaction costs were further explored inGaoet al\.\([2021](https://arxiv.org/html/2607.06610#bib.bib17)\)\. More recent studies have focused on improving robustness, scalability, and risk sensitivity in portfolio optimization\.Jianget al\.\([2024](https://arxiv.org/html/2607.06610#bib.bib18)\)proposed a model free DRL framework for dynamic portfolio allocation in high\-dimensional financial markets by integrating transaction costs and investor risk aversion into a mean\-variance reward function\.Abolmakaremet al\.\([2023](https://arxiv.org/html/2607.06610#bib.bib19)\)developed predictive multi\-period multi\-objective portfolio optimization models using deep learning techniques to forecast future market behavior\.Ndikum and Ndikum \([2024](https://arxiv.org/html/2607.06610#bib.bib20)\)introduced an industry grade DRL framework incorporating sim to real methodologies and realistic trading constraints for robust portfolio optimization across multiple asset classes\. Furthermore,Yanet al\.\([2024](https://arxiv.org/html/2607.06610#bib.bib21)\)proposed a deep portfolio optimization framework that explicitly incorporates transaction costs and risk\-aware reward functions into reinforcement learning\-based portfolio management\.

Several recent studies have investigated advanced reinforcement learning formulations tailored to financial markets\. The first work in this regard was done byJang and Seong \([2023](https://arxiv.org/html/2607.06610#bib.bib46)\)\. their work lies in integrating Modern Portfolio Theory \(MPT\) with deep reinforcement learning through a multimodal tensor\-based framework\. The paper combines technical indicators and asset correlation information using Tucker tensor decomposition within a DDPG\-based portfolio optimization model, enabling dynamic portfolio allocation that captures both temporal market patterns and cross\-asset dependencies\. The key contribution ofYuet al\.\([2019](https://arxiv.org/html/2607.06610#bib.bib45)\)lies in integrating model\-based deep reinforcement learning with portfolio optimization, enabling the agent to learn market dynamics through an internal environment model rather than relying solely on historical interactions\. This improves sample efficiency and allows more stable long\-term portfolio allocation decisions under changing market conditions\. For example,Jianget al\.\([2024](https://arxiv.org/html/2607.06610#bib.bib18)\)developed a model free deep reinforcement learning framework for dynamic portfolio allocation in high\-dimensional financial markets, whileYanet al\.\([2024](https://arxiv.org/html/2607.06610#bib.bib21)\)incorporated transaction costs and risk\-aware reward functions into deep portfolio optimization\. These studies highlight the importance of designing portfolio learning frameworks that can effectively capture market dynamics, asset dependencies, and realistic trading constraints\.

Despite the significant progress in DRL\-based portfolio optimization, most existing studies primarily focus on return maximization and variance\-based risk control, while limited attention has been given to reliability\-based optimization and tail\-risk\-aware decision\-making\. In parallel, reliability\-based portfolio optimization frameworks have been studied in the operations research literature\. In particular,Senguptaet al\.\([2024](https://arxiv.org/html/2607.06610#bib.bib22)\)developed a bi\-objective reliability\-based portfolio optimization framework under uncertainty\. However, their approach relies on static optimization techniques and does not consider sequential portfolio decision\-making through reinforcement learning\. Moreover, existing DRL based portfolio optimization methods generally employ standard risk measures such as variance or Sharpe ratio and rarely incorporate coherent tail risk measures such as Conditional Value\-at\-Risk \(CVaR\) and Entropic Value\-at\-Risk \(EVaR\) within a reliability constrained framework\. In addition, practical aspects such as transaction costs are often included only as penalty terms in the reward function, without integrating them into probabilistic reliability constraints\.

Motivated by these gaps, this paper proposes a novel framework termed MORP\-DRL for multi\-objective reliability based portfolio optimization using deep reinforcement learning\. The proposed framework integrates reliability constraints, transaction costs, and multiple risk measures including variance, CVaR, and EVaR within a unified DRL based sequential decision making framework\. Unlike existing static reliability based optimization methods and conventional DRL portfolio models, the proposed approach simultaneously addresses tail risk control, probabilistic reliability guarantees, and dynamic portfolio rebalancing under realistic market conditions\.

## 3Problem Formulation

### 3\.1Risk Measures

Risk management plays an important role in portfolio optimization under uncertain and volatile market environments\. In this study, we consider three commonly used downside risk measures: Value\-at\-Risk \(VaR\), Conditional Value\-at\-Risk \(CVaR\), and Entropic Value\-at\-Risk \(EVaR\), which capture different aspects of tail risk\.

For a portfolio returnRpR\_\{p\}and confidence levelα\\alpha,VaRestimates the loss threshold exceeded with probability1−α1\-\\alpha:VaRα=inf\{l:P​\(Rp≤−l\)≥1−α\}\\text\{VaR\}\_\{\\alpha\}=\\inf\\\{l:P\(R\_\{p\}\\leq\-l\)\\geq 1\-\\alpha\\\}\.

CVaRextends VaR by measuring the expected loss beyond the VaR threshold:CVaRα=E​\[−Rp∣−Rp≥VaRα\]\\text\{CVaR\}\_\{\\alpha\}=E\[\-R\_\{p\}\\mid\-R\_\{p\}\\geq\\text\{VaR\}\_\{\\alpha\}\]\. Unlike VaR, CVaR satisfies coherence properties and captures tail losses more effectively\.

EVaRprovides an exponential upper bound on tail risk:EVaRα​\(X\)=infz\>0\{1z​ln⁡\(E​\[ez​X\]α\)\}\\text\{EVaR\}\_\{\\alpha\}\(X\)=\\inf\_\{z\>0\}\\left\\\{\\frac\{1\}\{z\}\\ln\\left\(\\frac\{E\[e^\{zX\}\]\}\{\\alpha\}\\right\)\\right\\\}\. Owing to its strong theoretical properties and sensitivity to extreme downside risk, EVaR serves as an effective alternative risk measure for reliability based and robust portfolio optimization frameworks\. Both CVaR and EVaR are coherent risk measures satisfying the following properties, whereρα​\(⋅\)\\rho\_\{\\alpha\}\(\\cdot\)denotes a generic coherent risk measure at confidence levelα\\alpha:

Monotonicity:X≤Y⇒ρα​\(X\)≤ρα​\(Y\),\\displaystyle X\\leq Y\\;\\Rightarrow\\;\\rho\_\{\\alpha\}\(X\)\\leq\\rho\_\{\\alpha\}\(Y\),Translation Invariance:ρα​\(X\+c\)=ρα​\(X\)\+c,\\displaystyle\\rho\_\{\\alpha\}\(X\+c\)=\\rho\_\{\\alpha\}\(X\)\+c,Positive Homogeneity:ρα​\(λ​X\)=λ​ρα​\(X\),λ≥0,\\displaystyle\\rho\_\{\\alpha\}\(\\lambda X\)=\\lambda\\,\\rho\_\{\\alpha\}\(X\),\\qquad\\lambda\\geq 0,Subadditivity:ρα​\(X\+Y\)≤ρα​\(X\)\+ρα​\(Y\)\.\\displaystyle\\rho\_\{\\alpha\}\(X\+Y\)\\leq\\rho\_\{\\alpha\}\(X\)\+\\rho\_\{\\alpha\}\(Y\)\.These risk measures are subsequently incorporated into the proposed portfolio optimization framework for evaluating downside risk\.

### 3\.2Reliability Based Portfolio Optimization

Reliability Based Design Optimization \(RBDO\)Huet al\.\([2024](https://arxiv.org/html/2607.06610#bib.bib6)\); Senguptaet al\.\([2023](https://arxiv.org/html/2607.06610#bib.bib5)\)incorporates uncertainty directly into optimization through probabilistic constraints\. Rather than enforcing deterministic feasibility, RBDO seeks solutions that satisfy constraints with a prescribed reliability level under uncertain conditions\. In portfolio optimization, uncertainty naturally arises from fluctuating asset returns and market volatility\. Therefore, portfolio decisions should remain feasible with high probability rather than under a single deterministic realization\. Following the RBDO framework, the portfolio optimization problem can be formulated as

max𝐱\\displaystyle\\max\_\{\\mathbf\{x\}\}Rp=∑i=1nE​\(ri\)​xi\\displaystyle R\_\{p\}=\\sum\_\{i=1\}^\{n\}E\(r\_\{i\}\)x\_\{i\}\(3\)s\.t\.P​\(gj​\(𝐱,r\)≥0\)≥βj,j=1,…,J,\\displaystyle P\\\!\\left\(g\_\{j\}\(\\mathbf\{x\},r\)\\geq 0\\right\)\\geq\\beta\_\{j\},\\quad j=1,\\dots,J,∑i=1nxi=1,\\displaystyle\\sum\_\{i=1\}^\{n\}x\_\{i\}=1,xiL≤xi≤xiU,i=1,…,n\.\\displaystyle x\_\{i\}^\{L\}\\leq x\_\{i\}\\leq x\_\{i\}^\{U\},\\quad i=1,\\dots,n\.
wherexix\_\{i\}denotes the portfolio weight of theit​hi^\{th\}asset,rir\_\{i\}represents the return of theit​hi^\{th\}asset, andgj​\(𝐱,r\)g\_\{j\}\(\\mathbf\{x\},r\)denotes thejt​hj^\{th\}portfolio constraint under uncertain market conditions\. The parameterβj∈\(0,1\)\\beta\_\{j\}\\in\(0,1\)represents the prescribed reliability level, where larger values imply lower probabilities of constraint violation\. To incorporate realistic portfolio rebalancing effects, proportional transaction costs based on changes in portfolio allocations are also considered within the return formulationJanaet al\.\([2009](https://arxiv.org/html/2607.06610#bib.bib41)\); Chen \([2015](https://arxiv.org/html/2607.06610#bib.bib25)\)\.

## 4Proposed Bi\-Objective Portfolio Optimization Models

Most existing reliability based multi\-objective portfolio optimization \(MORBPO\) modelsSenguptaet al\.\([2024](https://arxiv.org/html/2607.06610#bib.bib22)\)ignore transaction costs during portfolio rebalancing, leading to unrealistic trading behavior and overestimated returns\. To address this limitation, proportional transaction costs are explicitly incorporated into the proposed framework\. To the best of our knowledge, transaction cost aware reliability based portfolio optimization under variance , CVaR , and EVaR based risk formulations has not been explored in existing literature\.

The notations used throughout the proposed models are summarized below:

- •NN: Total number of assets considered in the portfolio\.
- •wiw\_\{i\}: Portfolio weight allocated to theit​hi^\{th\}asset, wherei=1,2,…,Ni=1,2,\\dots,Nand0≤wi≤10\\leq w\_\{i\}\\leq 1\.
- •wi,minw\_\{i,\\min\}andwi,maxw\_\{i,\\max\}: Minimum and maximum allowable investment weights for theit​hi^\{th\}asset, respectively\.
- •wi0w\_\{i\}^\{0\}: Initial portfolio weight of theit​hi^\{th\}asset before rebalancing\.
- •r¯i\\bar\{r\}\_\{i\}: Expected return of theit​hi^\{th\}asset\.
- •ri,tr\_\{i,t\}: Return of theit​hi^\{th\}asset at time periodtt\.
- •σ^i,j\\hat\{\\sigma\}\_\{i,j\}: Estimated covariance between the returns of assetsiiandjj\.
- •α\\alpha: Confidence level associated with the VaR/CVaR/EVaR calculations\.
- •TT: Number of scenarios or historical observations used for risk estimation\.
- •γ\\gamma: Auxiliary variable associated with the VaR threshold in the CVaR and EVaR formulations\.
- •β1\\beta\_\{1\}: Reliability level associated with the portfolio return constraint\.
- •β2\\beta\_\{2\}: Reliability level associated with the portfolio risk constraint\.
- •rp∗r\_\{p\}^\{\*\}: Minimum target portfolio return specified by the investor\.
- •σp2⁣∗\\sigma\_\{p\}^\{2\*\}: Maximum acceptable portfolio variance in Model A\.
- •C​V​a​R∗CVaR^\{\*\}: Maximum acceptable CVaR threshold in Model B\.
- •E​V​a​R∗EVaR^\{\*\}: Maximum acceptable EVaR threshold in Model C\.
- •kik\_\{i\}: Proportional transaction cost coefficient associated with buying or selling theit​hi^\{th\}asset\.
- •\(x\)\+=max⁡\(x,0\)\(x\)^\{\+\}=\\max\(x,0\): Positive\-part operator used in the CVaR formulation\.

To capture different aspects of portfolio risk under uncertainty, we develop three reliability based bi\-objective portfolio optimization models\. Although all three formulations aim to maximize expected portfolio return while minimizing portfolio risk, they differ in the choice of risk measure employed\. Specifically, Model A uses portfolio variance as the risk metric, Model B incorporates Conditional Value\-at\-Risk \(CVaR\) to capture downside tail risk, and Model C employs Entropic Value\-at\-Risk \(EVaR\), which provides a tighter and more conservative characterization of extreme losses\.

##### Model A: Variance\-Based Formulation

Model A extends the bi\-objective reliability\-based portfolio optimization framework proposed inSenguptaet al\.\([2024](https://arxiv.org/html/2607.06610#bib.bib22)\)by incorporating proportional transaction costs into the return reliability constraint\. Reliability constraints ensure that the net portfolio return, after accounting for proportional transaction costs, exceeds a target returnrp∗r\_\{p\}^\{\*\}with confidence levelβ1\\beta\_\{1\}, while the portfolio variance remains below a prescribed thresholdσp2⁣∗\\sigma\_\{p\}^\{2\*\}with confidence levelβ2\\beta\_\{2\}\. It is further subject to budget and bound constraints\.

max𝐰\\displaystyle\\max\_\{\\mathbf\{w\}\}∑i=1Nwi​r¯i,min𝐰​∑i=1N∑j=1Nwi​wj​σ^i,j\\displaystyle\\sum\_\{i=1\}^\{N\}w\_\{i\}\\bar\{r\}\_\{i\},\\qquad\\min\_\{\\mathbf\{w\}\}\\sum\_\{i=1\}^\{N\}\\sum\_\{j=1\}^\{N\}w\_\{i\}w\_\{j\}\\hat\{\\sigma\}\_\{i,j\}\(6\)s\.t\.Pr⁡\[∑i=1Nwi​ri,t−∑i=1Nki​\|wi−wi0\|≥rp∗\]≥β1,\\displaystyle\\Pr\\\!\\left\[\\sum\_\{i=1\}^\{N\}w\_\{i\}r\_\{i,t\}\-\\sum\_\{i=1\}^\{N\}k\_\{i\}\|w\_\{i\}\-w\_\{i\}^\{0\}\|\\geq r\_\{p\}^\{\*\}\\right\]\\geq\\beta\_\{1\},Pr⁡\[∑i=1N∑j=1Nwi​wj​σ^i,j≤σp2⁣∗\]≥β2,\\displaystyle\\Pr\\\!\\left\[\\sum\_\{i=1\}^\{N\}\\sum\_\{j=1\}^\{N\}w\_\{i\}w\_\{j\}\\hat\{\\sigma\}\_\{i,j\}\\leq\\sigma\_\{p\}^\{2\*\}\\right\]\\geq\\beta\_\{2\},∑i=1Nwi=1,\\displaystyle\\sum\_\{i=1\}^\{N\}w\_\{i\}=1,0≤wi,min≤wi≤wi,max≤1,i=1,…,N\.\\displaystyle 0\\leq w\_\{i,\\min\}\\leq w\_\{i\}\\leq w\_\{i,\\max\}\\leq 1,\\quad i=1,\\dots,N\.

##### Model B: CVaR\-Based Formulation:

Unlike variance, which treats gains and losses symmetrically, CVaR focuses on downside risk by measuring expected losses beyond the VaR threshold\. This is also in extension to the framework used inSenguptaet al\.\([2024](https://arxiv.org/html/2607.06610#bib.bib22)\)Therefore, Model B maximizes expected return while minimizing CVaR under reliability and transaction cost constraints\.

max𝐰\\displaystyle\\max\_\{\\mathbf\{w\}\}∑i=1Nwi​r¯i,min𝐰⁡\{1α​T​∑t=1T\(γ−∑i=1Nwi​ri,t\)\+\+γ\}\\displaystyle\\sum\_\{i=1\}^\{N\}w\_\{i\}\\bar\{r\}\_\{i\},\\qquad\\min\_\{\\mathbf\{w\}\}\\left\\\{\\frac\{1\}\{\\alpha T\}\\sum\_\{t=1\}^\{T\}\\left\(\\gamma\-\\sum\_\{i=1\}^\{N\}w\_\{i\}r\_\{i,t\}\\right\)^\{\+\}\+\\gamma\\right\\\}\(7\)s\.t\.Pr⁡\[∑i=1Nwi​ri,t−∑i=1Nki​\|wi−wi0\|≥rp∗\]≥β1,\\displaystyle\\Pr\\\!\\left\[\\sum\_\{i=1\}^\{N\}w\_\{i\}r\_\{i,t\}\-\\sum\_\{i=1\}^\{N\}k\_\{i\}\|w\_\{i\}\-w\_\{i\}^\{0\}\|\\geq r\_\{p\}^\{\*\}\\right\]\\geq\\beta\_\{1\},Pr⁡\[1α​T​∑t=1T\(γ−∑i=1Nwi​ri,t\)\+\+γ≤C​V​a​R∗\]≥β2,\\displaystyle\\Pr\\\!\\left\[\\frac\{1\}\{\\alpha T\}\\sum\_\{t=1\}^\{T\}\\left\(\\gamma\-\\sum\_\{i=1\}^\{N\}w\_\{i\}r\_\{i,t\}\\right\)^\{\+\}\+\\gamma\\leq CVaR^\{\*\}\\right\]\\geq\\beta\_\{2\},∑i=1Nwi=1,\\displaystyle\\sum\_\{i=1\}^\{N\}w\_\{i\}=1,0≤wi,min≤wi≤wi,max≤1,i=1,…,N\.\\displaystyle 0\\leq w\_\{i,\\min\}\\leq w\_\{i\}\\leq w\_\{i,\\max\}\\leq 1,\\quad i=1,\\dots,N\.

##### Model C: EVaR\-Based Formulation:

Unlike CVaR, which measures average tail losses, EVaR provides a tighter and more conservative characterization of extreme risk\. Therefore, Model C maximizes expected return while minimizing EVaR under reliability and transaction cost constraints and is an extension of the work done inSenguptaet al\.\([2024](https://arxiv.org/html/2607.06610#bib.bib22)\)

max𝐰\\displaystyle\\max\_\{\\mathbf\{w\}\}∑i=1Nwi​r¯i,min𝐰⁡γ​ln⁡\(1T​∑t=1Texp⁡\(−γ−1​∑i=1Nwi​ri,t\)\)\\displaystyle\\sum\_\{i=1\}^\{N\}w\_\{i\}\\bar\{r\}\_\{i\},\\qquad\\min\_\{\\mathbf\{w\}\}\\gamma\\ln\\left\(\\frac\{1\}\{T\}\\sum\_\{t=1\}^\{T\}\\exp\\left\(\-\\gamma^\{\-1\}\\sum\_\{i=1\}^\{N\}w\_\{i\}r\_\{i,t\}\\right\)\\right\)\(8\)s\.t\.Pr⁡\[∑i=1Nwi​ri,t−∑i=1Nki​\|wi−wi0\|≥rp∗\]≥β1,\\displaystyle\\Pr\\\!\\left\[\\sum\_\{i=1\}^\{N\}w\_\{i\}r\_\{i,t\}\-\\sum\_\{i=1\}^\{N\}k\_\{i\}\|w\_\{i\}\-w\_\{i\}^\{0\}\|\\geq r\_\{p\}^\{\*\}\\right\]\\geq\\beta\_\{1\},Pr⁡\[γ​ln⁡\(1T​∑t=1Texp⁡\(−γ−1​∑i=1Nwi​ri,t\)\)≤E​V​a​Rp∗\]≥β2,\\displaystyle\\Pr\\\!\\left\[\\gamma\\ln\\left\(\\frac\{1\}\{T\}\\sum\_\{t=1\}^\{T\}\\exp\\left\(\-\\gamma^\{\-1\}\\sum\_\{i=1\}^\{N\}w\_\{i\}r\_\{i,t\}\\right\)\\right\)\\leq EVaR\_\{p\}^\{\*\}\\right\]\\geq\\beta\_\{2\},∑i=1Nwi=1,\\displaystyle\\sum\_\{i=1\}^\{N\}w\_\{i\}=1,0≤wi,min≤wi≤wi,max≤1,i=1,…,N\.\\displaystyle 0\\leq w\_\{i,\\min\}\\leq w\_\{i\}\\leq w\_\{i,\\max\}\\leq 1,\\quad i=1,\\dots,N\.

## 5Proposed Methodology

We propose a reliability aware multi\-objective deep reinforcement learning framework \(MORP\-DRL\) for sequential portfolio optimization under uncertainty\. The framework integrates transaction costs, probabilistic reliability constraints, and multiple risk measures within a Proximal Policy Optimization \(PPO\)Schulmanet al\.\([2017](https://arxiv.org/html/2607.06610#bib.bib42)\)based Actor\-Critic architecture\. Unlike static portfolio optimization models, the proposed framework enables dynamic portfolio rebalancing by continuously adapting allocations according to evolving market conditions\.

### 5\.1MDP Based Portfolio Optimization Framework

The portfolio optimization problem is formulated as a Markov Decision Process \(MDP\) represented by\(𝒮,𝒜,𝒫,R,γ\),\(\\mathcal\{S\},\\mathcal\{A\},\\mathcal\{P\},R,\\gamma\),where𝒮\\mathcal\{S\},𝒜\\mathcal\{A\},𝒫\\mathcal\{P\},RR, andγ\\gammadenote the state space, action space, transition dynamics, reward function, and discount factor, respectively\.

##### State Space \(𝒮\\mathcal\{S\}\)\.

The state at timettis defined asst=\[𝐩t,𝐈t,𝐰t−1,ct\]s\_\{t\}=\[\\mathbf\{p\}\_\{t\},\\mathbf\{I\}\_\{t\},\\mathbf\{w\}\_\{t\-1\},c\_\{t\}\], where𝐩t\\mathbf\{p\}\_\{t\}denotes normalized asset prices,𝐈t\\mathbf\{I\}\_\{t\}represents market indicators and statistical features,𝐰t−1\\mathbf\{w\}\_\{t\-1\}corresponds to previous portfolio weights, andctc\_\{t\}denotes available cash\. Including𝐰t−1\\mathbf\{w\}\_\{t\-1\}enables the agent to account for portfolio rebalancing and transaction costs\.

##### Action Space \(𝒜\\mathcal\{A\}\)\.

The action corresponds to portfolio allocation weights

at=𝐰t=\[w1,t,w2,t,…,wN,t\],a\_\{t\}=\\mathbf\{w\}\_\{t\}=\[w\_\{1,t\},w\_\{2,t\},\\dots,w\_\{N,t\}\],subject to

∑i=1Nwi,t=1,wi,min≤wi,t≤wi,max\.\\sum\_\{i=1\}^\{N\}w\_\{i,t\}=1,\\qquad w\_\{i,\\min\}\\leq w\_\{i,t\}\\leq w\_\{i,\\max\}\.Since portfolio weights are continuous, PPO is adopted to solve the optimization problem\.

##### Transition Dynamics \(𝒫\\mathcal\{P\}\)\.

The transition dynamics describe the evolution fromsts\_\{t\}tost\+1s\_\{t\+1\}after taking actionata\_\{t\}\. Since financial market dynamics are unknown, transitions are learned implicitly from historical data and simulated scenarios\. Asset returns are modeled using GARCH\(1,1\), extreme losses are characterized through EVT, and cross\-asset dependence is captured using att\-copula\. QMC simulation is then employed to generate training scenarios\.

##### Reward Function \(RR\)\.

The reward jointly captures return improvement, risk\-adjusted performance, and reliability satisfaction\(1\):

rt=Δ​St\+Δ​Rt\+ℛtrel,\\displaystyle r\_\{t\}=\\Delta S\_\{t\}\+\\Delta R\_\{t\}\+\\mathcal\{R\}\_\{t\}^\{\\mathrm\{rel\}\},\(1\)where

Δ​St=\(St−Sbase\)×100,St=Rtℛt,\\Delta S\_\{t\}=\(S\_\{t\}\-S^\{\\mathrm\{base\}\}\)\\times 100,\\qquad S\_\{t\}=\\frac\{R\_\{t\}\}\{\\mathcal\{R\}\_\{t\}\},andℛt\\mathcal\{R\}\_\{t\}denotes the selected portfolio risk measure\. The return improvement is defined as

Δ​Rt=\(Rt−Rtbase\)×100,\\Delta R\_\{t\}=\(R\_\{t\}\-R\_\{t\}^\{\\mathrm\{base\}\}\)\\times 100,with portfolio net return

Rt=𝐰t⊤​𝝁t−T​Ct,R\_\{t\}=\\mathbf\{w\}\_\{t\}^\{\\top\}\\boldsymbol\{\\mu\}\_\{t\}\-TC\_\{t\},whereT​CtTC\_\{t\}represents transaction costs\. The reliability component is given by

ℛtrel=Ψ​\(ρtret,β1\)\+Ψ​\(ρtrisk,β2\),\\mathcal\{R\}\_\{t\}^\{\\mathrm\{rel\}\}=\\Psi\(\\rho\_\{t\}^\{\\mathrm\{ret\}\},\\beta\_\{1\}\)\+\\Psi\(\\rho\_\{t\}^\{\\mathrm\{risk\}\},\\beta\_\{2\}\),whereρtret\\rho\_\{t\}^\{\\mathrm\{ret\}\}andρtrisk\\rho\_\{t\}^\{\\mathrm\{risk\}\}denote the estimated return and risk reliability probabilities, respectively\. The reliability function is defined as

Ψ​\(ρ,β\)=\{30​\(ρ−β\),ρ≥β,−100​\(β−ρ\),ρ<β\.\\Psi\(\\rho,\\beta\)=\\begin\{cases\}30\(\\rho\-\\beta\),&\\rho\\geq\\beta,\\\\ \-100\(\\beta\-\\rho\),&\\rho<\\beta\.\\end\{cases\}

##### Reliability Constraints:

To account for uncertainty, reliability measures are estimated using simulated market scenarios and incorporated into the reward design\. Let\{rt\(m\)\}m=1M\\\{r\_\{t\}^\{\(m\)\}\\\}\_\{m=1\}^\{M\}denoteMMscenarios generated using the proposed GARCH, EVT,tt\-copula, and QMC framework\. The empirical return reliability is estimated as

ρtret=P^return=1M​∑m=1M𝟏​\(∑i=1Nwi​ri,t\(m\)−∑i=1Nki​\|wi−wi0\|≥rp∗\),\\rho\_\{t\}^\{\\mathrm\{ret\}\}=\\hat\{P\}\_\{\\text\{return\}\}=\\frac\{1\}\{M\}\\sum\_\{m=1\}^\{M\}\\mathbf\{1\}\\left\(\\sum\_\{i=1\}^\{N\}w\_\{i\}r\_\{i,t\}^\{\(m\)\}\-\\sum\_\{i=1\}^\{N\}k\_\{i\}\|w\_\{i\}\-w\_\{i\}^\{0\}\|\\geq r\_\{p\}^\{\*\}\\right\),
while the empirical risk reliability is computed as

ρtrisk=P^risk=1M​∑m=1M𝟏​\(ℛ\(m\)​\(w\)≤ℛ∗\)\.\\rho\_\{t\}^\{\\mathrm\{risk\}\}=\\hat\{P\}\_\{\\text\{risk\}\}=\\frac\{1\}\{M\}\\sum\_\{m=1\}^\{M\}\\mathbf\{1\}\\left\(\\mathcal\{R\}^\{\(m\)\}\(w\)\\leq\\mathcal\{R\}^\{\*\}\\right\)\.
Here,ℛ\(m\)​\(w\)\\mathcal\{R\}^\{\(m\)\}\(w\)denotes the scenario\-wise portfolio risk measure \(variance, CVaR, or EVaR\), andℛ∗\\mathcal\{R\}^\{\*\}is the corresponding risk threshold\. These estimated reliabilities are compared against predefined confidence levels:

ρtret≥β1,ρtrisk≥β2\.\\rho\_\{t\}^\{\\mathrm\{ret\}\}\\geq\\beta\_\{1\},\\qquad\\rho\_\{t\}^\{\\mathrm\{risk\}\}\\geq\\beta\_\{2\}\.
The resulting reliability estimates are incorporated into the reward componentℛtrel\\mathcal\{R\}\_\{t\}^\{\\mathrm\{rel\}\}throughΨ​\(⋅\)\\Psi\(\\cdot\), thereby encouraging reliability feasible and risk aware portfolio decisions\.

![Refer to caption](https://arxiv.org/html/2607.06610v1/diagram1.png)Figure 1:Workflow of the Proposed PPO\-Based Portfolio Optimization Model

### 5\.2Proximal Policy Optimization \(PPO\) Agent

The proposed MORP–DRL framework employs a PPO\-based Actor–Critic architecture consisting of a policy network \(Actor\) and a value network \(Critic\)\. The Actor generates portfolio allocation decisions, while the Critic estimates the expected cumulative reward of the current market state\. The reward signal incorporates portfolio return improvement, risk\-adjusted performance, and reliability satisfaction, enabling the agent to learn reliability\-aware portfolio strategies\. Since portfolio allocation involves continuous decision variables, PPO is adopted due to its effectiveness in continuous action spaces and stable policy updatesSchulmanet al\.\([2017](https://arxiv.org/html/2607.06610#bib.bib42)\)\. Policy optimization is performed using the clipped surrogate objective\(2\):

LCLIP​\(θ\)=𝔼t​\[min⁡\(rt​\(θ\)​A^t,clip​\(rt​\(θ\),1−ϵ,1\+ϵ\)​A^t\)\]−c1​LV​F\+c2​HL^\{\\mathrm\{CLIP\}\}\(\\theta\)=\\mathbb\{E\}\_\{t\}\\left\[\\min\\left\(r\_\{t\}\(\\theta\)\\hat\{A\}\_\{t\},\\;\\mathrm\{clip\}\\\!\\left\(r\_\{t\}\(\\theta\),\\,1\-\\epsilon,\\,1\+\\epsilon\\right\)\\hat\{A\}\_\{t\}\\right\)\\right\]\-c\_\{1\}L^\{VF\}\+c\_\{2\}H\(2\)wherert​\(θ\)=πθ​\(at\|st\)πθold​\(at\|st\)r\_\{t\}\(\\theta\)=\\frac\{\\pi\_\{\\theta\}\(a\_\{t\}\|s\_\{t\}\)\}\{\\pi\_\{\\theta\_\{\\text\{old\}\}\}\(a\_\{t\}\|s\_\{t\}\)\}denotes the policy probability ratio between the updated and previous policies, andA^t\\hat\{A\}\_\{t\}represents the estimated advantage function\. The clipping operation limits excessively large policy updates and improves training stability\. Here,LV​F=𝔼t​\[\(Vϕ​\(st\)−Gt\)2\]L^\{VF\}=\\mathbb\{E\}\_\{t\}\[\(V\_\{\\phi\}\(s\_\{t\}\)\-G\_\{t\}\)^\{2\}\]denotes the value function \(critic\) loss, which minimizes the prediction error between the estimated state value and the discounted return, whileHHdenotes the policy entropy that encourages exploration and prevents premature convergence\. The coefficientsc1c\_\{1\}andc2c\_\{2\}control the relative importance of the value loss and entropy regularization, respectively\. Algorithm[1](https://arxiv.org/html/2607.06610#alg1)outlines the complete training procedure of the proposed MORP–DRL framework\. The overall workflow of the proposed PPO\-based portfolio optimization model is illustrated in Figure[1](https://arxiv.org/html/2607.06610#S5.F1)\.

Algorithm 1PPO\-Based Reliability\-Aware Multi\-objective Portfolio Optimization1:Initialize policy network

πθ​\(a\|s\)\\pi\_\{\\theta\}\(a\|s\)and value network

Vϕ​\(s\)V\_\{\\phi\}\(s\)
2:Initialize PPO hyperparameters

\(η,ϵ,γ\)\(\\eta,\\epsilon,\\gamma\)
3:Set reliability thresholds

β1,β2\\beta\_\{1\},\\beta\_\{2\}
4:Define portfolio bounds

wi,min,wi,maxw\_\{i,\\min\},w\_\{i,\\max\}
5:foreach training episodedo

6:Generate market scenarios using GARCH\(1,1\), EVT,

tt\-copula, and QMC

7:Observe initial state

s0=\[𝐩0,𝐈0,𝐰−1,c0\]s\_\{0\}=\[\\mathbf\{p\}\_\{0\},\\mathbf\{I\}\_\{0\},\\mathbf\{w\}\_\{\-1\},c\_\{0\}\]
8:Initialize trajectory buffer

𝒟\\mathcal\{D\}
9:for

t=0,1,…,T−1t=0,1,\\dots,T\-1do

10:Sample action:

at∼Dirichlet​\(πθ​\(st\)\)a\_\{t\}\\sim\\text\{Dirichlet\}\(\\pi\_\{\\theta\}\(s\_\{t\}\)\)
11:Project portfolio weights onto feasible simplex:

wi,min≤wi,t≤wi,max,∑i=1Nwi,t=1w\_\{i,\\min\}\\leq w\_\{i,t\}\\leq w\_\{i,\\max\},\\qquad\\sum\_\{i=1\}^\{N\}w\_\{i,t\}=1
12:Execute action and observe next state

st\+1s\_\{t\+1\}
13:Compute net return

Rt=𝐰t⊤​𝝁t−T​CtR\_\{t\}=\\mathbf\{w\}\_\{t\}^\{\\top\}\\boldsymbol\{\\mu\}\_\{t\}\-TC\_\{t\}
14:Compute risk measure

ℛt∈\{Variance,CVaR,EVaR\}\\mathcal\{R\}\_\{t\}\\in\\\{\\text\{Variance,CVaR,EVaR\}\\\}
15:Estimate reliability probabilities:

ρtr​e​t,ρtr​i​s​k\\rho\_\{t\}^\{ret\},\\;\\rho\_\{t\}^\{risk\}
16:Compute reward

rtr\_\{t\}using Eq\. \([1](https://arxiv.org/html/2607.06610#S5.E1)\)

17:Store

\(st,at,rt,st\+1\)\(s\_\{t\},a\_\{t\},r\_\{t\},s\_\{t\+1\}\)in

𝒟\\mathcal\{D\}
18:Set

st←st\+1s\_\{t\}\\leftarrow s\_\{t\+1\}
19:endfor

20:Compute discounted cumulative returns:

Gt=∑k=0T−tγk​rt\+kG\_\{t\}=\\sum\_\{k=0\}^\{T\-t\}\\gamma^\{k\}r\_\{t\+k\}
21:Estimate advantages

A^t=Gt−Vϕ​\(st\)\\hat\{A\}\_\{t\}=G\_\{t\}\-V\_\{\\phi\}\(s\_\{t\}\)
22:foreach PPO update epochdo

23:Compute policy ratio

rt​\(θ\)=πθ​\(at\|st\)πθo​l​d​\(at\|st\)r\_\{t\}\(\\theta\)=\\frac\{\\pi\_\{\\theta\}\(a\_\{t\}\|s\_\{t\}\)\}\{\\pi\_\{\\theta\_\{old\}\}\(a\_\{t\}\|s\_\{t\}\)\}
24:Update using clipped objective Eq\. \([2](https://arxiv.org/html/2607.06610#S5.E2)\)

25:Update policy and value networks

26:endfor

27:endfor

28:Return trained policy network

πθ\\pi\_\{\\theta\}

## 6Experiments

### 6\.1Data Collection and Pre\-processing

The empirical analysis was conducted using a diversified set of major global equity market indices representing different geographic regions and economic environments\. The dataset includes the NIFTY 50 \(India\), S&P 500 \(United States\), FTSE 100 \(United Kingdom\), DAX \(Germany\), Nikkei 225 \(Japan\), SSE Composite \(China\), Hang Seng Index \(Hong Kong\), KOSPI \(South Korea\), CAC 40 \(France\), and ASX 200 \(Australia\)\. The inclusion of these indices enables the proposed framework to capture heterogeneous market dynamics and evaluate portfolio behavior under varying economic conditions\. In addition to the index\-level experiments, a separate large\-scale portfolio optimization experiment was conducted using the constituent stocks of the FTSE 100 index in order to assess the scalability and robustness of the proposed methodology in higher\-dimensional portfolio settings\. Historical daily market data were collected using theyfinanceAPI\. The analysis period spans from January 2018 to December 2023 and is divided into three distinct market regimes\. The first period, referred to as the*Pre\-COVID*regime, covers January 2018 to December 2019 and represents relatively stable market conditions prior to the pandemic\. The second period corresponds to the*COVID*regime \(January 2020 to December 2021\), characterized by higher volatility and significant market disruptions caused by the global pandemic\. The final period, denoted as the*Post\-COVID*regime, spans January 2022 to December 2023 and captures the subsequent market recovery and stabilization phase\.Daily closing prices were transformed into logarithmic returns to ensure time\-additive return aggregation\. The log\-return of assetiiat timettis computed asri,t=ln⁡\(Pi,tPi,t−1\),r\_\{i,t\}=\\ln\\left\(\\frac\{P\_\{i,t\}\}\{P\_\{i,t\-1\}\}\\right\),wherePi,tP\_\{i,t\}denotes the closing price of assetiiat timett\. Financial return series typically exhibit volatility clustering and conditional heteroskedasticity\. To model these characteristics, a GARCH\(1,1\) model was fitted to the return series of each asset\. The GARCH framework enables time\-varying estimation of conditional volatility and provides a more realistic representation of financial market dynamics\. To capture extreme market movements and tail\-risk behavior, Extreme Value Theory \(EVT\) was employed\. In particular, excess losses beyond the 95th percentile threshold were modeled using the Generalized Pareto Distribution \(GPD\)\. This approach facilitates more reliable estimation of downside risk measures such as CVaR and EVaR, which are incorporated within the proposed portfolio optimization framework\. Expected asset returns were estimated using bootstrap resampling in order to account for sampling uncertainty\. For each asset, 1000 bootstrap samples were generated from the historical return distribution, and the mean of the resampled returns was used as the expected return estimate\. The resulting values were annualized for consistency with the portfolio evaluation metrics\.Portfolio risk was estimated using the empirical covariance matrix of daily returns\. To obtain annualized risk estimates, the covariance matrix was scaled by a factor of 252\.

### 6\.2Experimental Setup

The proposed framework is evaluated under practical portfolio constraints, where the asset weights satisfy

0\.03≤wi≤0\.35,∑i=1Nwi=1,0\.03\\leq w\_\{i\}\\leq 0\.35,\\qquad\\sum\_\{i=1\}^\{N\}w\_\{i\}=1,with reliability thresholds fixed atβ1=β2=0\.65\\beta\_\{1\}=\\beta\_\{2\}=0\.65\. Portfolio rebalancing incorporates proportional transaction costs of 2 basis points at each trading step based on changes in portfolio weights between consecutive periods\. The PPO agent adopts an actor–critic architecture and is trained using the AdamW optimizer with a learning rate of3×10−43\\times 10^\{\-4\}\. The discount factor and clipping parameter are set toγ=0\.99\\gamma=0\.99andϵ=0\.2\\epsilon=0\.2, respectively, and the agent is trained for 1000 episodes using the PPO clipped surrogate objective \([2](https://arxiv.org/html/2607.06610#S5.E2)\)\.

### 6\.3Benchmarks and Performance Evaluation

The proposed PPO framework is compared with two benchmark strategies: \(i\) an equal\-weighted \(EW\) portfolio, wherewi=1/Nw\_\{i\}=1/Nfor all assets, and \(ii\) an NSGA\-II optimized portfolio solved under the same return, risk, reliability, transaction cost, budget, and bound constraints\. Performance is evaluated across the Pre\-COVID, COVID, and Post\-COVID market regimes using return, volatility, Sharpe ratio, and the corresponding risk measure \(Variance, CVaR, or EVaR\)\. In addition, Pareto frontiers are constructed for all optimization models to analyze the trade\-off between expected portfolio return and the selected risk measure\.

#### 6\.3\.1Performance Metrics

The performance of the proposed MORP\-DRL framework and benchmark strategies is evaluated using several standard portfolio performance measures\. These metrics assess not only profitability, but also the associated risk, downside exposure, and computational efficiency of the portfolio optimization framework under different market regimes\.

##### Scalability to FTSE100 Constituents:

To ensure scalability in high\-dimensional portfolio settings, we modify the original quasi\-Monte Carlo \(QMC\) t\-copula scenario generation procedure\. The baseline approach constructs a single Sobol sequence over the full joint space of assets and time, resulting in a dimensionality ofN×TN\\times T, which quickly exceeds the practical limits of Sobol sequences \(≈\\approx21,201 dimensions\) for large portfolios\. To address this, we adopt a time\-decomposed QMC strategy, wherein Sobol sequences are generated independently at each time step with dimensionalityNN, and subsequently transformed via a Student\-t copula using Cholesky\-based correlation embedding\. This approach preserves cross\-sectional dependence across assets while significantly reducing computational complexity and avoiding high\-dimensional degeneration of low\-discrepancy sequences\. Although this introduces an approximation by relaxing inter\-temporal dependence, it enables efficient and stable scenario generation for large\-scale portfolios such as FTSE 100, making it well\-suited for reinforcement learning\-based optimization frameworks\.

## 7Results and Discussion

Tables[1](https://arxiv.org/html/2607.06610#S7.T1)\-[3](https://arxiv.org/html/2607.06610#S7.T3)together with the corresponding Pareto frontiers provide a detailed comparison between the baseline portfolio, NSGA\-II optimization, and PPO\-based reinforcement learning optimization under different market regimes and risk measures\.

##### Variance\-Based Portfolio Optimization \(Model A\)\.

Table[1](https://arxiv.org/html/2607.06610#S7.T1)together with Figures[2](https://arxiv.org/html/2607.06610#S7.F2)–[3](https://arxiv.org/html/2607.06610#S7.F3)summarizes the performance of the variance\-based portfolio optimization framework across the three market regimes\. In all periods, both NSGA\-II and PPO produce portfolios that outperform the equal\-weight benchmark, demonstrating the effectiveness of optimization based on the mean\-variance criterion\. During the Pre\-COVID period, NSGA\-II achieves the highest annualized return \(19\.63%\) together with the highest Sharpe ratio \(2\.231\), whereas PPO attains slightly lower volatility \(8\.61%\) and portfolio variance \(0\.0074\)\. As shown in Figures[2](https://arxiv.org/html/2607.06610#S7.F2)and[3](https://arxiv.org/html/2607.06610#S7.F3), the Pareto frontiers are concentrated within a low\-variance region, reflecting the relatively stable market conditions\. During the COVID period, the Pareto frontiers shift towards portfolios with substantially higher expected returns accompanied by increased variance, illustrating the elevated market uncertainty\. NSGA\-II marginally outperforms PPO by achieving the highest return \(29\.58%\), the highest Sharpe ratio \(1\.894\), and the lowest portfolio variance \(0\.0244\), while PPO provides comparable performance with only a slight increase in risk\. In the Post\-COVID period, the feasible return–variance region becomes narrower than during the pandemic, indicating a smaller range of attainable risk\-return trade\-offs\. PPO achieves the highest annualized return \(13\.08%\), whereas NSGA\-II maintains lower volatility, lower portfolio variance, and the highest Sharpe ratio \(1\.250\), indicating superior risk adjusted performance\. In all three market regimes, the optimized portfolios clearly dominate the equal\-weight benchmark in terms of return\-risk characteristics\. NSGA\-II consistently provides stronger risk\-adjusted performance under the variance objective, whereas PPO generates competitive portfolios with only marginal differences in return and risk\.

##### CVaR\-Based Portfolio Optimization \(Model B\):

Table[2](https://arxiv.org/html/2607.06610#S7.T2)together with Figures[4](https://arxiv.org/html/2607.06610#S7.F4)–[5](https://arxiv.org/html/2607.06610#S7.F5)summarizes the performance of the CVaR\-based portfolio optimization model, where Conditional Value\-at\-Risk \(CVaR\) is used to explicitly control downside tail risk\. Across all three market regimes, both NSGA\-II and PPO substantially outperform the equal\-weight benchmark in terms of return and risk\-adjusted performance, demonstrating the effectiveness of incorporating downside\-risk considerations into portfolio optimization\. During the Pre\-COVID period, NSGA\-II achieves the highest Sharpe ratio \(2\.249\) while maintaining the lowest CVaR \(0\.2281\), indicating the most favorable balance between return and downside risk\. PPO attains a comparable annualized return \(19\.57%\) but with slightly higher volatility and CVaR\. As illustrated in Figures[4\(a\)](https://arxiv.org/html/2607.06610#S7.F4.sf1)and[5\(a\)](https://arxiv.org/html/2607.06610#S7.F5.sf1), the Pareto frontiers are concentrated within a relatively narrow return\-CVaR region, reflecting the comparatively stable market environment\. During the COVID period, the Pareto frontiers shift towards portfolios with substantially higher expected returns accompanied by increased downside risk, illustrating the challenging market conditions\. PPO identifies the highest\-return portfolio \(33\.28%\), but this improvement is accompanied by the highest CVaR \(0\.4601\)\. In contrast, NSGA\-II achieves the highest Sharpe ratio \(1\.914\) with a lower CVaR \(0\.4462\), providing a more balanced trade\-off between return and downside risk\. Moreover, the PPO Pareto frontier spans a wider range of return\-CVaR combinations, indicating greater flexibility in generating portfolios corresponding to different levels of downside\-risk tolerance\. In the Post\-COVID period, the Pareto frontiers become more concentrated than during the pandemic, suggesting a narrower range of feasible return\-risk trade\-offs\. NSGA\-II marginally outperforms PPO in annualized return \(13\.88% versus 13\.64%\), Sharpe ratio \(1\.277 versus 1\.241\), and CVaR \(0\.2912 versus 0\.2935\), although the differences between the two approaches are relatively small\. In all three market regimes, both optimization methods produce portfolios that dominate the equal\-weight benchmark by achieving superior return\-risk characteristics\. The results indicate that NSGA\-II consistently produces portfolios with lower downside risk and better risk\-adjusted performance under the CVaR objective, whereas PPO exhibits greater flexibility by identifying higher return portfolios during periods of elevated market uncertainty, albeit at the expense of increased tail risk\.

##### EVaR\-Based Portfolio Optimization \(Model C\):

Table[3](https://arxiv.org/html/2607.06610#S7.T3)together with Figures[6](https://arxiv.org/html/2607.06610#S7.F6)–[7](https://arxiv.org/html/2607.06610#S7.F7)presents the performance of the EVaR\-based portfolio optimization model, where Entropic Value\-at\-Risk \(EVaR\) provides a conservative measure of downside risk by emphasizing extreme loss events\. Across all market regimes, both NSGA\-II and PPO outperform the equal\-weight benchmark, although they exhibit different risk\-return characteristics\. During the Pre\-COVID period, NSGA\-II identifies the highest\-return portfolio \(22\.36%\), whereas PPO achieves the highest Sharpe ratio \(2\.241\) by substantially reducing portfolio volatility\. Figures[6\(a\)](https://arxiv.org/html/2607.06610#S7.F6.sf1)and[7\(a\)](https://arxiv.org/html/2607.06610#S7.F7.sf1)show that both methods generate Pareto\-efficient portfolios that clearly dominate the equal\-weight benchmark\. During the COVID period, the Pareto frontiers shift toward portfolios with higher expected returns, reflecting the increased downside risk associated with market uncertainty\. NSGA\-II achieves the highest annualized return \(33\.55%\) together with the highest Sharpe ratio \(1\.906\), while PPO produces portfolios with lower volatility, indicating a more conservative risk profile despite a modest reduction in return\. The broader range of feasible portfolios obtained by PPO suggests greater flexibility in exploring different return\-risk trade\-offs under stressed market conditions\. In the Post\-COVID period, the performance differences between the two methods become smaller\. NSGA\-II continues to achieve the highest annualized return \(14\.42%\), whereas PPO maintains lower portfolio volatility and attains the highest Sharpe ratio \(1\.251\), indicating superior risk\-adjusted performance\. In all three market regimes, both optimization approaches substantially outperform the equal\-weight benchmark\. The EVaR based results reveal a consistent trade\-off between return maximization and conservative risk management\. NSGA\-II prioritizes higher expected returns, whereas PPO consistently favors lower portfolio volatility and improved risk\-adjusted performance\.

##### Comparative Behaviour of NSGA\-II and PPO:

Across the three risk measures and market regimes, NSGA\-II and PPO exhibit distinct optimization characteristics\. NSGA\-II generally achieves stronger risk adjusted performance, consistently requiring substantially lower computational time while producing Pareto\-efficient portfolios with competitive or superior returns across most experimental settings\. PPO, although computationally more expensive due to repeated policy updates and environment interactions, produces portfolios with performance comparable to NSGA\-II and, in several cases, attains either the highest annualized return \(e\.g\., the CVaR\-based COVID scenario and the variance\-based Post\-COVID scenario\) or the highest Sharpe ratio \(under the EVaR formulation during the Pre\-COVID and Post\-COVID periods\)\. During periods of elevated market uncertainty, particularly under the EVaR framework, PPO tends to favour portfolios with lower volatility, whereas NSGA\-II more frequently identifies portfolios with higher expected returns\. Across all experiments, both optimization methods consistently outperform the equal\-weight benchmark, demonstrating the effectiveness of multi\-objective optimization for portfolio selection\. NSGA\-II offers a computationally efficient optimization strategy with strong and consistent risk\-adjusted performance, while PPO provides a competitive learning\-based alternative capable of adapting portfolio decisions through interaction with the market environment\.

The differences in optimization behaviour are further illustrated by the portfolio allocation results reported in Appendix A \(Tables A1–A3\), which present the optimal asset weights under the variance, CVaR, and EVaR formulations across the three market regimes\.

##### FTSE100 Scalability Results:

The FTSE100 experiments demonstrate that the proposed framework remains effective in high\-dimensional portfolio optimization settings\. As summarized in Table[4](https://arxiv.org/html/2607.06610#S7.T4), both PPO and NSGA\-II consistently outperform the equal\-weight benchmark across all three market regimes and risk measures, confirming the scalability of the proposed optimization framework to a larger asset universe\. Under the variance formulation, NSGA\-II achieves the highest Sharpe ratios in all market regimes, exceeding 2\.0 throughout the study, while PPO also provides substantial improvements over the baseline\. Similar trends are observed under the CVaR and EVaR formulations, where both optimization methods consistently achieve higher Sharpe ratios than the equal\-weight portfolio, with NSGA\-II maintaining the strongest risk\-adjusted performance across all periods\. Although PPO does not surpass NSGA\-II in terms of Sharpe ratio, it consistently delivers competitive performance while employing a learning\-based optimization strategy\. These results demonstrate that both approaches scale effectively to large portfolio optimization problems, with NSGA\-II providing superior risk\-adjusted performance and computational efficiency, whereas PPO offers a flexible reinforcement learning framework capable of adapting portfolio allocation policies as market conditions evolve\.

Additional portfolio concentration statistics are reported in Appendix A \(Tables A4, A6, and A7\)\. These results show that NSGA\-II generally allocates a larger proportion of portfolio weight to a smaller subset of assets, whereas PPO distributes weights across a larger number of assets, leading to more diversified portfolio allocations\. This behaviour is consistent across the variance, CVaR, and EVaR formulations and provides further evidence of the distinct optimization characteristics of the two approaches in high\-dimensional portfolio settings\.

Model A

Table 1:Portfolio Performance Comparison Across Market Regimes — Variance Risk Measure![Refer to caption](https://arxiv.org/html/2607.06610v1/pre_covid_nsga_variance.png)\(a\)Pre COVID
![Refer to caption](https://arxiv.org/html/2607.06610v1/covid_nsga_variance.png)\(b\)COVID
![Refer to caption](https://arxiv.org/html/2607.06610v1/post_covid_nsga_variance.png)\(c\)Post COVID

Figure 2:NSGA\-II based Pareto frontiers under the Variance risk measure for \(a\) Pre COVID, \(b\) COVID, and \(c\) Post COVID periods\.![Refer to caption](https://arxiv.org/html/2607.06610v1/pre_covid_ppo_variance.png)\(a\)Pre COVID
![Refer to caption](https://arxiv.org/html/2607.06610v1/covid_pareto_ppo_variance.png)\(b\)COVID
![Refer to caption](https://arxiv.org/html/2607.06610v1/post_covid_pareto_ppo_variance.png)\(c\)Post COVID

Figure 3:PPO based Pareto frontiers under the Variance risk measure for \(a\) Pre COVID, \(b\) COVID, and \(c\) Post COVID periods\.Model B

Table 2:Portfolio Performance Comparison Across Market Regimes — CVaR Risk Measure![Refer to caption](https://arxiv.org/html/2607.06610v1/pre_covid_nsga_CVaR.png)\(a\)Pre COVID
![Refer to caption](https://arxiv.org/html/2607.06610v1/covid_nsga_CVaR.png)\(b\)COVID
![Refer to caption](https://arxiv.org/html/2607.06610v1/post_covid_nsga_CVaR.png)\(c\)Post COVID

Figure 4:NSGA\-II based Pareto frontiers under the CVaR risk measure for \(a\) Pre COVID, \(b\) COVID, and \(c\) Post COVID periods\.![Refer to caption](https://arxiv.org/html/2607.06610v1/pre_covid_ppo_cvar.png)\(a\)Pre COVID
![Refer to caption](https://arxiv.org/html/2607.06610v1/covid_ppo_cvar.png)\(b\)COVID
![Refer to caption](https://arxiv.org/html/2607.06610v1/post_covid_ppo_cvar.png)\(c\)Post COVID

Figure 5:Pareto frontiers obtained using PPO under the CVaR risk measure for \(a\) Pre COVID, \(b\) COVID, and \(c\) Post COVID periods\.Model C

Table 3:Portfolio Performance Comparison Across Market Regimes — EVaR Risk Measure![Refer to caption](https://arxiv.org/html/2607.06610v1/pre_covid_nsga_evar.png)\(a\)Pre COVID
![Refer to caption](https://arxiv.org/html/2607.06610v1/covid_nsga_evar.png)\(b\)COVID
![Refer to caption](https://arxiv.org/html/2607.06610v1/post_covid_nsga_evar.png)\(c\)Post COVID

Figure 6:Pareto frontiers obtained using NSGA\-II under the EVaR risk measure for \(a\) Pre COVID, \(b\) COVID, and \(c\) Post COVID periods\.![Refer to caption](https://arxiv.org/html/2607.06610v1/pre_covid_ppo_evar.png)\(a\)Pre COVID
![Refer to caption](https://arxiv.org/html/2607.06610v1/covid_ppo_evar.png)\(b\)COVID
![Refer to caption](https://arxiv.org/html/2607.06610v1/post_covid_ppo_evar.png)\(c\)Post COVID

Figure 7:Pareto frontiers obtained using PPO under the EVaR risk measure for \(a\) Pre COVID, \(b\) COVID, and \(c\) Post COVID periods\.Table 4:FTSE100 Performance Comparison Across Risk Measures

## 8Conclusion and Future Work

This paper proposes a bi\-objective reliability\-based portfolio optimization framework in which portfolio allocation is learned as a sequential decision making problem using deep reinforcement learning\. The framework jointly optimizes expected return and downside risk while accounting for market uncertainty, transaction costs, and reliability constraints under realistic financial conditions\. The empirical study, conducted on ten major global equity indices across the pre\-COVID, COVID, and post\-COVID periods, demonstrates that the PPO\-based policy consistently produces competitive risk\-return portfolios and adapts effectively to changing market regimes\. The results show that the framework provides improved control of downside and extreme\-tail risk under the CVaR and EVaR formulations, leading to more stable portfolio performance during periods of market stress\. Pareto frontier analysis further indicates that the learned policies achieve an effective balance between return maximization and risk minimization while satisfying the prescribed reliability constraints\. Comparison with NSGA\-II shows a clear trade\-off between the two optimization approaches\. NSGA\-II is computationally efficient and often attains higher returns through concentrated portfolio allocations, whereas PPO learns adaptive and more diversified investment strategies that respond effectively to evolving market conditions\. The framework also scales effectively to high\-dimensional portfolio optimization, as demonstrated on the FTSE100 constituent stocks using a modified time decomposed quasi\-Monte Carlo Student\-ttcopula for efficient scenario generation\. Future research may extend the framework to dynamic multi\-period portfolio optimization, incorporate additional market frictions such as liquidity and price impact, and investigate hybrid reinforcement learning and evolutionary optimization methods for large\-scale portfolio selection\. Further improvements may include integrating macroeconomic indicators, news sentiment, and alternative data sources to enable more informed portfolio decisions under rapidly changing market conditions\. The use of multi\-agent reinforcement learning and distributional reinforcement learning may also provide improved modeling of market interactions and uncertainty\. Finally, extending the framework to multi\-asset portfolios comprising equities, fixed income securities, commodities, and cryptocurrencies would further enhance its practical applicability for real\-world investment management\.

## References

- \[1\]\(2023\)Predictive multi\-period multi\-objective portfolio optimization based on higher order moments: deep learning approach\.Computers & industrial engineering183,pp\. 109450\.Cited by:[§2](https://arxiv.org/html/2607.06610#S2.p2.1)\.
- \[2\]K\. Arulkumaran, M\. P\. Deisenroth, M\. Brundage, and A\. A\. Bharath\(2017\)Deep reinforcement learning: a brief survey\.IEEE signal processing magazine34\(6\),pp\. 26–38\.Cited by:[§1](https://arxiv.org/html/2607.06610#S1.p3.1)\.
- \[3\]J\. Behera, A\. K\. Pasayat, H\. Behera, and P\. Kumar\(2023\)Prediction based mean\-value\-at\-risk portfolio optimization using machine learning regression algorithms for multi\-national stock markets\.Engineering Applications of Artificial Intelligence120,pp\. 105843\.Cited by:[§1](https://arxiv.org/html/2607.06610#S1.p3.1),[§2](https://arxiv.org/html/2607.06610#S2.p1.1)\.
- \[4\]S\. Chakraborty\(2019\)Capturing financial markets to apply deep reinforcement learning\.arXiv preprint arXiv:1907\.04373\.Cited by:[§2](https://arxiv.org/html/2607.06610#S2.p2.1)\.
- \[5\]H\. Chau, D\. Nguyen, and T\. Nguyen\(2025\)Continuous\-time optimal investment with portfolio constraints: a reinforcement learning approach\.European Journal of Operational Research\.Cited by:[§2](https://arxiv.org/html/2607.06610#S2.p2.1)\.
- \[6\]W\. Chen\(2015\)Artificial bee colony algorithm for constrained possibilistic portfolio optimization problem\.Physica A: Statistical Mechanics and its Applications429,pp\. 125–139\.Cited by:[§3\.2](https://arxiv.org/html/2607.06610#S3.SS2.p3.7)\.
- \[7\]H\. Choudhary, A\. Orra, K\. Sahoo, and M\. Thakur\(2025\)Risk\-adjusted deep reinforcement learning for portfolio optimization: a multi\-reward approach\.International Journal of Computational Intelligence Systems18\(1\),pp\. 126\.Cited by:[§1](https://arxiv.org/html/2607.06610#S1.p3.1)\.
- \[8\]H\. Choudhary, A\. Orra, M\. Thakur, X\. Gao, and P\. K\. Sahu\(2026\)A cvar\-constrained safe reinforcement learning framework with action repair for practical portfolio optimization\.IEEE Transactions on Artificial Intelligence\.Cited by:[§1](https://arxiv.org/html/2607.06610#S1.p3.1)\.
- \[9\]O\. Ertenlice and C\. B\. Kalayci\(2018\)A survey of swarm intelligence for portfolio optimization: algorithms and applications\.Swarm and evolutionary computation39,pp\. 36–52\.Cited by:[§1](https://arxiv.org/html/2607.06610#S1.p2.1)\.
- \[10\]F\. J\. Fabozzi, H\. M\. Markowitz, and F\. Gupta\(2008\)Portfolio selection\.Handbook of finance2,pp\. 3–13\.Cited by:[§1](https://arxiv.org/html/2607.06610#S1.p1.1)\.
- \[11\]E\. F\. Fama and K\. R\. French\(2004\)The capital asset pricing model: theory and evidence\.Journal of economic perspectives18\(3\),pp\. 25–46\.Cited by:[§1](https://arxiv.org/html/2607.06610#S1.p1.1)\.
- \[12\]T\. Fischer and C\. Krauss\(2018\)Deep learning with long short\-term memory networks for financial market predictions\.European journal of operational research270\(2\),pp\. 654–669\.Cited by:[§1](https://arxiv.org/html/2607.06610#S1.p3.1),[§2](https://arxiv.org/html/2607.06610#S2.p1.1)\.
- \[13\]Y\. Gao, Z\. Gao, Y\. Hu, S\. Song, Z\. Jiang, and J\. Su\(2021\)A framework of hierarchical deep q\-network for portfolio management\.\.InICAART \(2\),pp\. 132–140\.Cited by:[§2](https://arxiv.org/html/2607.06610#S2.p2.1)\.
- \[14\]Z\. Gao, Y\. Gao, Y\. Hu, Z\. Jiang, and J\. Su\(2020\)Application of deep q\-network in portfolio management\.In2020 5th IEEE International Conference on Big Data Analytics \(ICBDA\),pp\. 268–275\.Cited by:[§2](https://arxiv.org/html/2607.06610#S2.p2.1)\.
- \[15\]I\. Goodfellow, Y\. Bengio, A\. Courville, and Y\. Bengio\(2016\)Deep learning\.Vol\.1,MIT press Cambridge\.Cited by:[§1](https://arxiv.org/html/2607.06610#S1.p3.1)\.
- \[16\]S\. Gu, B\. Kelly, and D\. Xiu\(2020\)Empirical asset pricing via machine learning\.The Review of Financial Studies33\(5\),pp\. 2223–2273\.Cited by:[§1](https://arxiv.org/html/2607.06610#S1.p3.1)\.
- \[17\]W\. Hu, S\. Cheng, J\. Yan, J\. Cheng, X\. Peng, H\. Cho, and I\. Lee\(2024\)Reliability\-based design optimization: a state\-of\-the\-art review of its methodologies, applications, and challenges\.Structural and Multidisciplinary Optimization67\(9\),pp\. 168\.Cited by:[§3\.2](https://arxiv.org/html/2607.06610#S3.SS2.p1.1)\.
- \[18\]P\. Jana, T\. Roy, and S\. Mazumder\(2009\)Multi\-objective possibilistic model for portfolio selection with transaction cost\.Journal of computational and applied mathematics228\(1\),pp\. 188–196\.Cited by:[§3\.2](https://arxiv.org/html/2607.06610#S3.SS2.p3.7)\.
- \[19\]J\. Jang and N\. Seong\(2023\)Deep reinforcement learning for stock portfolio optimization by connecting with modern portfolio theory\.Expert Systems with Applications218,pp\. 119556\.Cited by:[§2](https://arxiv.org/html/2607.06610#S2.p3.1)\.
- \[20\]Y\. Jiang, J\. Olmo, and M\. Atwi\(2024\)Deep reinforcement learning for portfolio selection\.Global Finance Journal62,pp\. 101016\.Cited by:[§2](https://arxiv.org/html/2607.06610#S2.p2.1),[§2](https://arxiv.org/html/2607.06610#S2.p3.1)\.
- \[21\]P\. N\. Kolm, R\. Tütüncü, and F\. J\. Fabozzi\(2014\)60 years of portfolio optimization: practical challenges and current trends\.European Journal of Operational Research234\(2\),pp\. 356–371\.Cited by:[§2](https://arxiv.org/html/2607.06610#S2.p1.1)\.
- \[22\]Y\. Li, P\. Ni, and V\. Chang\(2019\)Application of deep reinforcement learning in stock trading strategies and stock forecasting\.Computing\.Cited by:[§2](https://arxiv.org/html/2607.06610#S2.p2.1)\.
- \[23\]T\. J\. Linsmeier and N\. D\. Pearson\(2000\)Value at risk\.Financial analysts journal56\(2\),pp\. 47–67\.Cited by:[§1](https://arxiv.org/html/2607.06610#S1.p1.1)\.
- \[24\]Z\. Loke, S\. L\. Goh, G\. Kendall, S\. Abdullah, and N\. Sabar\(2023\-01\)Portfolio optimisation problem: a taxonomic review of solution methodologies\.IEEE AccessPP,pp\. 1–1\.External Links:[Document](https://dx.doi.org/10.1109/ACCESS.2023.3263198)Cited by:[§1](https://arxiv.org/html/2607.06610#S1.p2.1)\.
- \[25\]H\. Markowitz\(1952\)Portfolio selection\.The Journal of Finance7\(1\),pp\. 77–91\.Cited by:[§1](https://arxiv.org/html/2607.06610#S1.p1.1)\.
- \[26\]H\. Markowitz\(1952\)Portfolio selection\.The Journal of Finance7\(1\),pp\. 77–91\.Cited by:[§2](https://arxiv.org/html/2607.06610#S2.p1.1)\.
- \[27\]S\. S\. Meghwani and M\. Thakur\(2018\)Multi\-objective heuristic algorithms for practical portfolio optimization and rebalancing with transaction cost\.Applied Soft Computing67,pp\. 865–894\.Cited by:[§2](https://arxiv.org/html/2607.06610#S2.p1.1)\.
- \[28\]P\. Ndikum and S\. Ndikum\(2024\)Advancing investment frontiers: industry\-grade deep reinforcement learning for portfolio optimization\.arXiv preprint arXiv:2403\.07916\.Cited by:[§2](https://arxiv.org/html/2607.06610#S2.p2.1)\.
- \[29\]H\. P\. Ramos, M\. B\. Righi, P\. C\. Guedes, and F\. M\. Müller\(2023\)A comparison of risk measures for portfolio optimization with cardinality constraints\.Expert Systems with Applications228,pp\. 120412\.Cited by:[§1](https://arxiv.org/html/2607.06610#S1.p1.1)\.
- \[30\]M\. B\. Righi and D\. Borenstein\(2018\)A simulation comparison of risk measures for portfolio optimization\.Finance Research Letters24,pp\. 105–112\.Cited by:[§1](https://arxiv.org/html/2607.06610#S1.p1.1)\.
- \[31\]R\. T\. Rockafellar and S\. Uryasev\(2002\)Conditional value\-at\-risk for general loss distributions\.Journal of banking & finance26\(7\),pp\. 1443–1471\.Cited by:[§1](https://arxiv.org/html/2607.06610#S1.p1.1)\.
- \[32\]A\. Salo, M\. Doumpos, J\. Liesiö, and C\. Zopounidis\(2024\)Fifty years of portfolio optimization\.European Journal of Operational Research318\(1\),pp\. 1–18\.Cited by:[§1](https://arxiv.org/html/2607.06610#S1.p2.1)\.
- \[33\]J\. Schulman, F\. Wolski, P\. Dhariwal, A\. Radford, and O\. Klimov\(2017\)Proximal policy optimization algorithms\.ArXivabs/1707\.06347\.External Links:[Link](https://api.semanticscholar.org/CorpusID:28695052)Cited by:[§5\.2](https://arxiv.org/html/2607.06610#S5.SS2.p1.7),[§5](https://arxiv.org/html/2607.06610#S5.p1.1)\.
- \[34\]R\. N\. Sengupta, A\. Gupta, S\. Mukherjee, and G\. Weiss\(2024\)Bi\-objective reliability based optimization: an application to investment analysis\.Annals of Operations Research333\(1\),pp\. 47–78\.Cited by:[§2](https://arxiv.org/html/2607.06610#S2.p4.1),[§4](https://arxiv.org/html/2607.06610#S4.SS0.SSS0.Px1.p1.4),[§4](https://arxiv.org/html/2607.06610#S4.SS0.SSS0.Px2.p1.1),[§4](https://arxiv.org/html/2607.06610#S4.SS0.SSS0.Px3.p1.1),[§4](https://arxiv.org/html/2607.06610#S4.p1.1)\.
- \[35\]R\. N\. Sengupta, R\. Seth, and P\. Winker\(2023\)Reliability in portfolio optimization using uncertain estimates\.Sankhya B85\(Suppl 1\),pp\. 199–233\.Cited by:[§3\.2](https://arxiv.org/html/2607.06610#S3.SS2.p1.1)\.
- \[36\]R\. Yan, J\. Jin, and K\. Han\(2024\)Reinforcement learning for deep portfolio optimization\.Electronic Research Archive32\(9\),pp\. 5176\.Cited by:[§2](https://arxiv.org/html/2607.06610#S2.p2.1),[§2](https://arxiv.org/html/2607.06610#S2.p3.1)\.
- \[37\]P\. Yu, J\. S\. Lee, I\. Kulyatin, Z\. Shi, and S\. Dasgupta\(2019\)Model\-based deep reinforcement learning for dynamic portfolio optimization\.arXiv preprint arXiv:1901\.08740\.Cited by:[§2](https://arxiv.org/html/2607.06610#S2.p3.1)\.

## Appendix AAppendix

Table A1:Portfolio Weights Comparison Across Market Regimes – Variance OptimizationTable A2:Portfolio Weights Comparison Across Market Regimes – CVaR OptimizationTable A3:Portfolio Weights Comparison Across Market Regimes – EVaR OptimizationTable A4:Portfolio Concentration and Weight Distribution\(FTSE 100 Variance Risk Measure\)Table A5:Top 10 Holdings \(PPO Optimized Portfolio\)Table A6:Portfolio Concentration and Weight Distribution \(FTSE100 CVaR Risk Measure\)Table A7:Portfolio Concentration and Diversification Analysis\(FTSE100 EVaR Risk Measure\)

Similar Articles

Robust Peak-cost Constrained Reinforcement Learning

arXiv cs.LG

This paper studies robust peak-cost constrained reinforcement learning, addressing limitations of standard CMDPs by controlling the maximum cost along a trajectory and considering dynamics uncertainty. The authors show zero duality gap may not hold and propose a surrogate optimization framework with robust value estimation.

Cross-Epoch Adaptive Rollout Optimization for RL Post-Training

arXiv cs.LG

This paper presents CERO, a cross-epoch adaptive rollout optimization method for RL post-training of LLMs, which allocates a fixed rollout budget across prompts and epochs using Bayesian posterior variance to maximize sample efficiency, achieving theoretical regret bounds and outperforming GRPO on mathematical reasoning tasks.