Probabilistic inference for determining options in reinforcement learning.
Saved in:
| Title: | Probabilistic inference for determining options in reinforcement learning. |
|---|---|
| Authors: | Daniel, Christian Christian.Daniel@de.bosch.com, van Hoof, Herke1, Peters, Jan, Neumann, Gerhard1 |
| Source: | Machine Learning. Sep2016, Vol. 104 Issue 2-3, p337-357. 21p. |
| Subjects: | Probabilistic inference, Reinforcement learning, Machine learning, Partially observable Markov decision processes, Algorithms |
| Abstract: | Tasks that require many sequential decisions or complex solutions are hard to solve using conventional reinforcement learning algorithms. Based on the semi Markov decision process setting (SMDP) and the option framework, we propose a model which aims to alleviate these concerns. Instead of learning a single monolithic policy, the agent learns a set of simpler sub-policies as well as the initiation and termination probabilities for each of those sub-policies. While existing option learning algorithms frequently require manual specification of components such as the sub-policies, we present an algorithm which infers all relevant components of the option framework from data. Furthermore, the proposed approach is based on parametric option representations and works well in combination with current policy search methods, which are particularly well suited for continuous real-world tasks. We present results on SMDPs with discrete as well as continuous state-action spaces. The results show that the presented algorithm can combine simple sub-policies to solve complex tasks and can improve learning performance on simpler tasks. [ABSTRACT FROM AUTHOR] |
| Copyright of Machine Learning is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
| FullText | Links: – Type: pdflink Text: Availability: 0 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 117381140 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Probabilistic inference for determining options in reinforcement learning. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Daniel%2C+Christian%22">Daniel, Christian</searchLink><i> Christian.Daniel@de.bosch.com</i><br /><searchLink fieldCode="AR" term="%22van+Hoof%2C+Herke%22">van Hoof, Herke</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22Peters%2C+Jan%22">Peters, Jan</searchLink><br /><searchLink fieldCode="AR" term="%22Neumann%2C+Gerhard%22">Neumann, Gerhard</searchLink><relatesTo>1</relatesTo> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Machine+Learning%22">Machine Learning</searchLink>. Sep2016, Vol. 104 Issue 2-3, p337-357. 21p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Probabilistic+inference%22">Probabilistic inference</searchLink><br /><searchLink fieldCode="DE" term="%22Reinforcement+learning%22">Reinforcement learning</searchLink><br /><searchLink fieldCode="DE" term="%22Machine+learning%22">Machine learning</searchLink><br /><searchLink fieldCode="DE" term="%22Partially+observable+Markov+decision+processes%22">Partially observable Markov decision processes</searchLink><br /><searchLink fieldCode="DE" term="%22Algorithms%22">Algorithms</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: Tasks that require many sequential decisions or complex solutions are hard to solve using conventional reinforcement learning algorithms. Based on the semi Markov decision process setting (SMDP) and the option framework, we propose a model which aims to alleviate these concerns. Instead of learning a single monolithic policy, the agent learns a set of simpler sub-policies as well as the initiation and termination probabilities for each of those sub-policies. While existing option learning algorithms frequently require manual specification of components such as the sub-policies, we present an algorithm which infers all relevant components of the option framework from data. Furthermore, the proposed approach is based on parametric option representations and works well in combination with current policy search methods, which are particularly well suited for continuous real-world tasks. We present results on SMDPs with discrete as well as continuous state-action spaces. The results show that the presented algorithm can combine simple sub-policies to solve complex tasks and can improve learning performance on simpler tasks. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of Machine Learning is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=117381140 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1007/s10994-016-5580-x Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 21 StartPage: 337 Subjects: – SubjectFull: Probabilistic inference Type: general – SubjectFull: Reinforcement learning Type: general – SubjectFull: Machine learning Type: general – SubjectFull: Partially observable Markov decision processes Type: general – SubjectFull: Algorithms Type: general Titles: – TitleFull: Probabilistic inference for determining options in reinforcement learning. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Daniel, Christian – PersonEntity: Name: NameFull: van Hoof, Herke – PersonEntity: Name: NameFull: Peters, Jan – PersonEntity: Name: NameFull: Neumann, Gerhard IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 09 Text: Sep2016 Type: published Y: 2016 Identifiers: – Type: issn-print Value: 08856125 Numbering: – Type: volume Value: 104 – Type: issue Value: 2-3 Titles: – TitleFull: Machine Learning Type: main |
| ResultId | 1 |