Research / Peer reviewed
Abstract
Reinforcement learning algorithms require many samples when solving complex hierarchical tasks with sparse and delayed rewards. For such complex tasks, the recently proposed RUDDER uses reward redistribution to leverage steps in the Q-function that are associated with accomplishing sub-tasks. However, often only few episodes with high rewards are available as demonstrations since current exploration strategies cannot discover them in reasonable time. In this work, we introduce Align-RUDDER, which utilizes a profile model for reward redistribution that is obtained from multiple sequence alignment of demonstrations. Consequently, Align-RUDDER employs reward redistribution effectively and, thereby, drastically improves learning on few demonstrations. Align-RUDDER outperforms competitors on complex artificial tasks with delayed rewards and few demonstrations. On the Minecraft ObtainDiamond task, Align-RUDDER is able to mine a diamond, though not frequently. Code is available at github.com/ml-jku/align-rudder.
Cite this paper
@inproceedings{patil2022align,
author = {Patil, Vihang and Hofmarcher, Markus and Dinu, Marius-Constantin and Dorfer, Matthias and Blies, Patrick and Brandstetter, Johannes and Arjona-Medina, José and Hochreiter, Sepp},
title = {{Align-RUDDER}: Learning From Few Demonstrations by Reward Redistribution},
booktitle = {ICML},
year = {2022},
url = {https://www.dinu.at/research/align-rudder-learning-from-few-demonstrations-by-reward-redistribution},
}Vihang Patil, Markus Hofmarcher, Marius-Constantin Dinu, Matthias Dorfer, Patrick Blies, Johannes Brandstetter, José Arjona-Medina, Sepp Hochreiter. (2022). Align-RUDDER: Learning From Few Demonstrations by Reward Redistribution. ICML. https://www.dinu.at/research/align-rudder-learning-from-few-demonstrations-by-reward-redistribution
TY - CPAPER AU - Patil, Vihang AU - Hofmarcher, Markus AU - Dinu, Marius-Constantin AU - Dorfer, Matthias AU - Blies, Patrick AU - Brandstetter, Johannes AU - Arjona-Medina, José AU - Hochreiter, Sepp TI - Align-RUDDER: Learning From Few Demonstrations by Reward Redistribution T2 - ICML PY - 2022 UR - https://www.dinu.at/research/align-rudder-learning-from-few-demonstrations-by-reward-redistribution ER -
More peer reviewed
- CoLLAs · 2024 SymbolicAI: A framework for logic-based approaches combining generative models and solvers
- Ph.D. Thesis, Johannes Kepler University · 2024 Parameter Choice and Neuro-Symbolic Approaches for Deep Domain-Invariant Learning
- ICLR · 2023 Addressing Parameter Choice Issues in Unsupervised Domain Adaptation by Aggregation
- CoLLAs · 2022 A Dataset Perspective on Offline Reinforcement Learning