Potential-based reward shaping using state–space segmentation for efficiency in reinforcement learning

Bal, Melis; AYDIN, HÜSEYİN; İYİGÜN, CEM; POLAT, FARUK

doi:10.1016/j.future.2024.03.057

Potential-based reward shaping using state–space segmentation for efficiency in reinforcement learning

Bal M. İ., AYDIN H., İYİGÜN C., POLAT F.

Future Generation Computer Systems, cilt.157, ss.469-484, 2024 (SCI-Expanded, Scopus)

Yayın Türü: Makale / Tam Makale
Cilt numarası: 157
Basım Tarihi: 2024
Doi Numarası: 10.1016/j.future.2024.03.057
Dergi Adı: Future Generation Computer Systems
Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Applied Science & Technology Source, Business Source Elite, Business Source Premier, Compendex, Computer & Applied Sciences, INSPEC, zbMATH
Sayfa Sayıları: ss.469-484
Anahtar Kelimeler: Potential-based reward shaping, Reinforcement learning, Reward shaping, Sparse rewards, State–space segmentation
Orta Doğu Teknik Üniversitesi Adresli: Evet

Özet

Reinforcement Learning (RL) algorithms encounter slow learning in environments with sparse explicit reward structures due to the limited feedback available on the agent's behavior. This problem is exacerbated particularly in complex tasks with large state and action spaces. To address this inefficiency, in this paper, we propose a novel approach based on potential-based reward-shaping using state–space segmentation to decompose the task and to provide more frequent feedback to the agent. Our approach involves extracting state–space segments by formulating the problem as a minimum cut problem on a transition graph, constructed using the agent's experiences during interactions with the environment via the Extended Segmented Q-Cut algorithm. Subsequently, these segments are leveraged in the agent's learning process through potential-based reward shaping. Our experimentation on benchmark problem domains with sparse rewards demonstrated that our proposed method effectively accelerates the agent's learning without compromising computation time while upholding the policy invariance principle.