Research / Peer reviewed
InfODist: Online distillation with Informative rewards improves generalization in Curriculum Learning
Rahul Siripurapu, Vihang Patil, Kajetan Schweighofer, Marius-Constantin Dinu, Thomas Schmied, Luis Ferro, Markus Holzleitner, Hamid Eghbal-Zadeh, Michael Kopp, Sepp Hochreiter
- Venue
- Deep RL Workshop, NeurIPS
- Year
- 2022
- Citations
- 4
- Length
- 12 pages
Abstract
Curriculum learning (CL) is an essential part of human learning, just as reinforcement learning (RL) is. However, CL agents that are trained using RL with neural networks produce limited generalization to later tasks in the curriculum. We show that online distillation using learned informative rewards tackles this problem. Here, we consider a reward to be informative if it is positive when the agent makes progress towards the goal and negative otherwise. Thus, an informative reward allows an agent to learn immediately to avoid states which are irrelevant to the task. And, the value and policy networks do not utilize their limited capacity to fit targets for these irrelevant states. Consequently, this improves generalization to later tasks. Our contributions: First, we propose InfODist, an online distillation method that makes use of informative rewards to significantly improve generalization in CL. Second, we show that training with informative rewards ameliorates the capacity loss phenomenon that was previously attributed to non-stationarities during the training process. Third, we show that learning from task-irrelevant states explains the capacity loss and subsequent impaired generalization. In conclusion, our work is a crucial step toward scaling curriculum learning to complex real world tasks.
Cite this paper
@inproceedings{siripurapu2022infodist,
author = {Siripurapu, Rahul and Patil, Vihang and Schweighofer, Kajetan and Dinu, Marius-Constantin and Schmied, Thomas and Ferro, Luis and Holzleitner, Markus and Eghbal-Zadeh, Hamid and Kopp, Michael and Hochreiter, Sepp},
title = {{InfODist}: Online distillation with Informative rewards improves generalization in Curriculum Learning},
booktitle = {Deep RL Workshop, NeurIPS},
year = {2022},
url = {https://www.dinu.at/research/infodist-online-distillation-with-informative-rewards-improves-generaliz},
}Rahul Siripurapu, Vihang Patil, Kajetan Schweighofer, Marius-Constantin Dinu, Thomas Schmied, Luis Ferro, Markus Holzleitner, Hamid Eghbal-Zadeh, Michael Kopp, Sepp Hochreiter. (2022). InfODist: Online distillation with Informative rewards improves generalization in Curriculum Learning. Deep RL Workshop, NeurIPS. https://www.dinu.at/research/infodist-online-distillation-with-informative-rewards-improves-generaliz
TY - CPAPER AU - Siripurapu, Rahul AU - Patil, Vihang AU - Schweighofer, Kajetan AU - Dinu, Marius-Constantin AU - Schmied, Thomas AU - Ferro, Luis AU - Holzleitner, Markus AU - Eghbal-Zadeh, Hamid AU - Kopp, Michael AU - Hochreiter, Sepp TI - InfODist: Online distillation with Informative rewards improves generalization in Curriculum Learning T2 - Deep RL Workshop, NeurIPS PY - 2022 UR - https://www.dinu.at/research/infodist-online-distillation-with-informative-rewards-improves-generaliz ER -
More peer reviewed
- CoLLAs · 2024 SymbolicAI: A framework for logic-based approaches combining generative models and solvers
- Ph.D. Thesis, Johannes Kepler University · 2024 Parameter Choice and Neuro-Symbolic Approaches for Deep Domain-Invariant Learning
- ICLR · 2023 Addressing Parameter Choice Issues in Unsupervised Domain Adaptation by Aggregation
- CoLLAs · 2022 A Dataset Perspective on Offline Reinforcement Learning