[1] D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot et al., “Mastering the game of go with deep neural networks and tree search,” nature, vol. 529, no. 7587, pp. 484–489, 2016.
[2] O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev et al., “Grandmaster level in StarCraft II using multiagent reinforcement learning,” nature, vol. 575, no. 7782, pp. 350–354, 2019.
[3] S. Levine, C. Finn, T. Darrell, and P. Abbeel, “End-to-end training of deep visuomotor policies,” Journal of Machine Learning Research, vol. 17, no. 39, pp. 1–40, 2016.
[4] S. Gu, E. Holly, T. P. Lillicrap, and S. Levine, “Deep reinforcement learning for robotic manipulation,” arXiv preprint arXiv:1610.00633, vol. 1, no. 1, 2016.
[5] F.-M. Luo, T. Xu, H. Lai, X.-H. Chen, W. Zhang, and Y. Yu, “A survey on model-based reinforcement learning,” Science China Information Sciences, vol. 67, no. 2, p. 121101, 2024.
[6] V. Konda and J. Tsitsiklis, “Actor-critic algorithms,” Advances in neural information processing systems, vol. 12, 1999.
[7] T. Degris, P. M. Pilarski, and R. S. Sutton, “Model-free reinforcement learning with continuous action in practice,” in 2012 American control conference (ACC). IEEE, 2012, pp. 2177–2182.
[8] T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra, “Continuous control with deep reinforcement learning,” arXiv preprint arXiv:1509.02971, 2015.
[9] S. Fujimoto, H. Hoof, and D. Meger, “Addressing function approximation error in actor-critic methods,” in International conference on machine learning. PMLR, 2018, pp. 1587–1596.
[10] T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft Actor-Critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,” pp. 1861–1870, 2018.
[11] M. G. Bellemare, W. Dabney, and R. Munos, “A distributional perspective on reinforcement learning,” in International conference on machine learning. PMLR, 2017, pp. 449–458.
[12] M. G. Bellemare, W. Dabney, and M. Rowland, Distributional Reinforcement Learning. MIT Press, 2023, http://www.distributional-rl.org.
[13] X. Ma, L. Xia, Z. Zhou, J. Yang, and Q. Zhao, “Dsac: Distributional soft actor-critic for risk-sensitive reinforcement learning,” arXiv preprint arXiv:2004.14547, 2020.
[14] J. Duan, Y. Guan, S. E. Li, Y. Ren, Q. Sun, and B. Cheng, “Distributional soft actor-critic: Off-policy reinforcement learning for addressing value estimation errors,” IEEE Transactions on Neural Networks and Learning Systems, vol. 33, no. 11, pp. 6584–6598, 2022.
[15] M. Rowland, M. Bellemare, W. Dabney, R. Munos, and Y. W. Teh, “An analysis of categorical distributional reinforcement learning,” in International Conference on Artificial Intelligence and Statistics. PMLR, 2018, pp. 29–37.
[16] W. Dabney, M. Rowland, M. Bellemare, and R. Munos, “Distributional reinforcement learning with quantile regression,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 32, no. 1, 2018.
[17] N. Mohammadpour, M. Fozi, M. M. Ebadzadeh, A. Azimi, and A. K. Iglie, “Proximal policy optimization with adaptive generalized advantage estimate,” in Proceedings of the First International Conference on Machine Learning and Knowledge Discovery (MLKD 2024), 2024, pp. 453–458. [Online]. Available: https://mlkd.aut.ac.ir/proceedings/2024/paper/4B.7.pdf
[18] N. Mohammadpour, M. Fozi, M. M. Ebadzadeh, A. Azimi, and A. Kamalie Iglie, “Proximal policy optimization with adaptive generalized advantage estimate: critic-aware refinements,” Journal of Mathematical Modeling, 2025.
[19] J. Liu, X. Gu, and S. Liu, “Policy optimization reinforcement learning with entropy regularization,” arXiv preprint arXiv:1912.01557, 2019.
[20] T. Haarnoja, H. Tang, P. Abbeel, and S. Levine, “Reinforcement learning with deep energy-based policies,” in International conference on machine learning. PMLR, 2017, pp. 1352–1361.
[21] B. D. Ziebart, A. L. Maas, J. A. Bagnell, A. K. Dey et al., “Maximum entropy inverse reinforcement learning.” in AAAI, vol. 8. Chicago, IL, USA, 2008, pp. 1433–1438.
[22] B. D. Ziebart, J. A. Bagnell, and A. K. Dey, “Modeling interaction via the principle of maximum causal entropy,” 2010.
[23] R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine learning, vol. 8, no. 3, pp. 229–256, 1992.
[24] K. Lee, S. Kim, S. Lim, S. Choi, and S. Oh, “Tsallis reinforcement learning: A unified framework for maximum entropy reinforcement learning,” arXiv preprint arXiv:1902.00137, 2019.
[25] J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz, “Trust region policy optimization,” in International conference on machine learning. PMLR, 2015, pp. 1889–1897.
[26] J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal policy optimization algorithms,” arXiv preprint arXiv:1707.06347, 2017.
[27] J. Duan, W. Wang, L. Xiao, J. Gao, and S. E. Li, “DSAC-T: Distributional soft actor-critic with three refinements,” arXiv preprint arXiv:2310.05858, 2023.