TY - GEN
T1 - Soft Value Iteration for Bellman Equations via Maximum Entropy Reinforcement Learning
AU - Devadas, Raghavendra M.
AU - Preethi, null
AU - Sapna, R.
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - This work evaluates the effectiveness of entropy-regularized Reinforcement Learning (RL) by contrasting Soft Value Iteration with conventional Bellman-based approaches. Based on the Maximum Entropy principle, our new Soft RL method showcases a clear performance gain in state value estimations. Experimentation within the FrozenLake-v1 environment identifies that Soft RL consistently returns higher and more consistent state values between around 2.5 and 3.5, as opposed to the 0.0 and 1.0 range of Traditional RL. Visualizations of state evaluation using heatmap representations further demonstrate the difference, with Soft RL displaying uniformly higher state assessments, reflecting robustness as well as stability. Difference analysis exemplifies consistent gaps in values among states with peaks of up to 2.9 in key regions, reflecting the higher expressiveness of the entropy-regularized approach. The findings confirm that Soft Value Iteration, based on the maximization of entropy, provides an improved, more robust framework for solving the Bellman equations for reinforcement learning problems.
AB - This work evaluates the effectiveness of entropy-regularized Reinforcement Learning (RL) by contrasting Soft Value Iteration with conventional Bellman-based approaches. Based on the Maximum Entropy principle, our new Soft RL method showcases a clear performance gain in state value estimations. Experimentation within the FrozenLake-v1 environment identifies that Soft RL consistently returns higher and more consistent state values between around 2.5 and 3.5, as opposed to the 0.0 and 1.0 range of Traditional RL. Visualizations of state evaluation using heatmap representations further demonstrate the difference, with Soft RL displaying uniformly higher state assessments, reflecting robustness as well as stability. Difference analysis exemplifies consistent gaps in values among states with peaks of up to 2.9 in key regions, reflecting the higher expressiveness of the entropy-regularized approach. The findings confirm that Soft Value Iteration, based on the maximization of entropy, provides an improved, more robust framework for solving the Bellman equations for reinforcement learning problems.
UR - https://www.scopus.com/pages/publications/105014351638
UR - https://www.scopus.com/pages/publications/105014351638#tab=citedBy
U2 - 10.1109/I2CACIS65476.2025.11100438
DO - 10.1109/I2CACIS65476.2025.11100438
M3 - Conference contribution
AN - SCOPUS:105014351638
T3 - 2025 IEEE International Conference on Automatic Control and Intelligent Systems, I2CACIS 2025 - Proceedings
SP - 361
EP - 365
BT - 2025 IEEE International Conference on Automatic Control and Intelligent Systems, I2CACIS 2025 - Proceedings
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2025 IEEE International Conference on Automatic Control and Intelligent Systems, I2CACIS 2025
Y2 - 27 June 2025 through 28 June 2025
ER -