Skip to main navigation Skip to search Skip to main content

Soft Value Iteration for Bellman Equations via Maximum Entropy Reinforcement Learning

  • Raghavendra M. Devadas*
  • , Preethi
  • , R. Sapna
  • *Corresponding author for this work

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract

This work evaluates the effectiveness of entropy-regularized Reinforcement Learning (RL) by contrasting Soft Value Iteration with conventional Bellman-based approaches. Based on the Maximum Entropy principle, our new Soft RL method showcases a clear performance gain in state value estimations. Experimentation within the FrozenLake-v1 environment identifies that Soft RL consistently returns higher and more consistent state values between around 2.5 and 3.5, as opposed to the 0.0 and 1.0 range of Traditional RL. Visualizations of state evaluation using heatmap representations further demonstrate the difference, with Soft RL displaying uniformly higher state assessments, reflecting robustness as well as stability. Difference analysis exemplifies consistent gaps in values among states with peaks of up to 2.9 in key regions, reflecting the higher expressiveness of the entropy-regularized approach. The findings confirm that Soft Value Iteration, based on the maximization of entropy, provides an improved, more robust framework for solving the Bellman equations for reinforcement learning problems.

Original languageEnglish
Title of host publication2025 IEEE International Conference on Automatic Control and Intelligent Systems, I2CACIS 2025 - Proceedings
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages361-365
Number of pages5
ISBN (Electronic)9798331542948
DOIs
Publication statusPublished - 2025
Event2025 IEEE International Conference on Automatic Control and Intelligent Systems, I2CACIS 2025 - Kuala Lumpur, Malaysia
Duration: 27-06-202528-06-2025

Publication series

Name2025 IEEE International Conference on Automatic Control and Intelligent Systems, I2CACIS 2025 - Proceedings

Conference

Conference2025 IEEE International Conference on Automatic Control and Intelligent Systems, I2CACIS 2025
Country/TerritoryMalaysia
CityKuala Lumpur
Period27-06-2528-06-25

All Science Journal Classification (ASJC) codes

  • Artificial Intelligence
  • Computer Science Applications
  • Information Systems
  • Control and Optimization

Fingerprint

Dive into the research topics of 'Soft Value Iteration for Bellman Equations via Maximum Entropy Reinforcement Learning'. Together they form a unique fingerprint.

Cite this