Q400 : Reinforcement Learning-baxsed Portfolio Optimization under Dynamic Market Conditions
Thesis > Central Library of Shahrood University > Computer Engineering > MSc > 2026
Authors:
Abstarct: Portfolio optimization stands as a fundamental challenge in investment management, aiming to optimally allocate assets to maximize returns while controlling risk. In recent years, machine learning-baxsed methods, particularly reinforcement learning (RL), have emerged as innovative approaches for solving decision-making problems in dynamic environments. This study presents a reinforcement learning-baxsed frxamework for portfolio optimization within the Iranian capital market. We employ the Proximal Policy Optimization (PPO) algorithm within an Actor–Critic architecture to learn the optimal asset allocation policy. The model’s state space incorporates price variables of selected stocks, as well as latent market factors extracted via a Dynamic Factor Model (DFM), providing the learning agent with a comprehensive understanding of overall market conditions. Furthermore, to enhance the training process and improve learning stability, we utilize a transfer learning approach. This technique is a crucial step toward model generalizability, particularly in emerging markets characterized by limitations in high-quality historical data and significant structural volatility. Indeed, by leveraging prior knowledge from similar environments, this architecture significantly accelerates the agent’s convergence speed and prevents overfitting to market noise. The dataset comprises information on 25 stock symbols from 2019 to 2024, encompassing approximately 30,000 observations and 1,199 trading days. The proposed model’s performance is evaluated using metrics such as cumulative return, Sharpe ratio, and maximum drawdown, and compared against benchmark methods including the classical Markowitz model, the equal-weight strategy, and a vanilla PPO algorithm. Empirical results demonstrate that the proposed model achieves more competitive and stable performance in portfolio management, showing superior results in risk management and volatility reduction compared to benchmark methods. Under volatile market conditions, the intelligent agent successfully prevents significant capital losses while maintaining profitability through smart weight rebalancing. The findings of this research suggest that the integration of reinforcement learning, dynamic market factors, and transfer learning provides an effective frxamework for intelligent decision-making in asset allocation, serving as a foundation for developing advanced algorithmic trading systems in Iran.
Keywords:
#Keywords: Portfolio Optimization #Reinforcement Learning #PPO Algorithm #Dynamic Factor Model (DFM) #Transfer Learning #Iranian Capital Market. Keeping place: Central Library of Shahrood University
Visitor:
Visitor: