Abstract
Traditional scalar-based reinforcement learning (RL) methods like Advantage Actor-Critic (A2C) estimate only the expected return, which overlooks the full return distribution. This limitation often leads to suboptimal and unstable policies, especially in high-speed, continuous action spaces. To address these limitations, this paper proposes a Distributional A2C (DA2C) algorithm inspired by the principles of Distributional Soft Actor-Critic (DSAC). Instead of relying on scalar value estimation, proposed DA2C models return distributions as Gaussian, allowing the algorithm to learn the mean and variance of returns. This representation allows the agent to make more accurate value estimates, reducing the risk of over/under estimation. Consequently, policy updates become more stable, precise, and resilient to environmental variability. The proposed method is evaluated using a real-time simulated fruit-cutting task, which is a high-speed control problem that requires quick and precise decision-making. Simulation results show that the proposed DA2C outperforms the traditional A2C baseline across key metrics. Notably, it achieves higher slicing precision, faster response, and enhanced policy robustness, showcasing the benefits of incorporating distributional value estimation in complex, time-sensitive RL applications.
| Original language | English |
|---|---|
| Title of host publication | Proceedings - 17th International Conference on Information Technology and Electrical Engineering, ICITEE 2025 |
| Publisher | Institute of Electrical and Electronics Engineers Inc. |
| ISBN (Electronic) | 9798331599263 |
| DOIs | |
| Publication status | Published - 2025 |
| Event | 17th International Conference on Information Technology and Electrical Engineering, ICITEE 2025 - Bangkok, Thailand Duration: 20-10-2025 → 21-10-2025 |
Conference
| Conference | 17th International Conference on Information Technology and Electrical Engineering, ICITEE 2025 |
|---|---|
| Country/Territory | Thailand |
| City | Bangkok |
| Period | 20-10-25 → 21-10-25 |
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 7 Affordable and Clean Energy
All Science Journal Classification (ASJC) codes
- Computer Science Applications
- Information Systems
- Energy Engineering and Power Technology
- Renewable Energy, Sustainability and the Environment
- Electrical and Electronic Engineering
- Media Technology
Fingerprint
Dive into the research topics of 'Improving Dynamic Task Performance with Distributional Reinforcement Learning: A RealTime Fruit Slicing Control Case Study'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver