勷勤数学•专家报告-李迅

勷勤数学•专家报告


题      目:Dynamic mean--variance portfolio selection with no-shorting constraints and unknown investment opportunity sets


报  告  人: 李迅 教授  (邀请人:骆其伦)

                                         香港理工大学


时      间: 7月29日  16:00-17:00

          

地     点:数科院东楼401


报告人简介:

       Xun Li received BSc from Department of Mathematics at Shanghai University of Science and Technology in 1992, and obtained MSc in Department of Mathematics at Shanghai University in 1995. He completed his PhD in Department of Systems Engineering and Engineering Management at the Chinese University of Hong Kong in 2000, and he stayed with the same department as a postdoctoral research fellow until 2001. From 2001 to 2003, he was a postdoctoral fellow in the Mathematical and Computational Finance Laboratory at University of Calgary. From 2003 to 2007, he was a visiting fellow in Department of Mathematics at the National University of Singapore. He joined Department of Applied Mathematics at the Hong Kong Polytechnic University as Assistant Professor in 2007, Associate Professor in 2013, and is currently Professor. His main research areas are stochastic control and applied probability with financial applications, and he has published in journals such as SIAM Journal on Control and Optimization, Annals of Applied Probability, Journal of Differential Equations, IEEE Transactions on Automatic Control, Automatica, Mathematical Finance, Finance and Stochastics, and Quantitative Finance.



摘      要:

       This work is the first to treat constrained continuous-time MV portfolio selection in the realm of reinforcement learning. We study continuous-time mean--variance portfolio selection with no-shorting constraints and unknown investment opportunity sets from a reinforcement learning (RL) perspective. The problem is a constrained stochastic linear--quadratic control problem for which the entropy-regularized exploratory formulation of Wang et al. (2020) leads to difficulty in theoretical analysis, because enforcing the constraint on the support of randomized policies nullifies the tractable Gaussian exploration. To tackle this challenge, we introduce an auxiliary exploratory problem without entropy in which exploratory policies are still Gaussian whose samples may violate the no-shorting requirement but their means satisfy it. We then prove that, for a suitable choice of exploration variance, the mean of the optimal Gaussian policy of the auxiliary problem coincides with the optimal policy of the original problem. Motivated by this theoretical result, we develop a model-free RL algorithm that learns the optimal policy of the auxiliary (and hence the original) problem directly from trajectory data without estimating the investment opportunity set. A numerical example demonstrates the performance of the proposed algorithm.

 (Joint with Yutian Wang and Xun Yu Zhou.)


       


          欢迎老师、同学们参加、交流!