Efficient Actor-Critic Algorithm with Hierarchical Model Learning and Planning.
Saved in:
| Title: | Efficient Actor-Critic Algorithm with Hierarchical Model Learning and Planning. |
|---|---|
| Authors: | Zhong, Shan1,2, Liu, Quan1,3,4, Fu, QiMing5 |
| Source: | Computational Intelligence & Neuroscience. 9/26/2016, p1-15. 15p. |
| Subjects: | Mathematical models of learning, Mathematical models, Planning, Stochastic convergence, Reinforcement learning, Hierarchical Bayes model |
| Abstract: | To improve the convergence rate and the sample efficiency, two efficient learning methods AC-HMLP and RAC-HMLP (AC-HMLP with l2-regularization) are proposed by combining actor-critic algorithm with hierarchical model learning and planning. The hierarchical models consisting of the local and the global models, which are learned at the same time during learning of the value function and the policy, are approximated by local linear regression (LLR) and linear function approximation (LFA), respectively. Both the local model and the global model are applied to generate samples for planning; the former is used only if the state-prediction error does not surpass the threshold at each time step, while the latter is utilized at the end of each episode. The purpose of taking both models is to improve the sample efficiency and accelerate the convergence rate of the whole algorithm through fully utilizing the local and global information. Experimentally, AC-HMLP and RAC-HMLP are compared with three representative algorithms on two Reinforcement Learning (RL) benchmark problems. The results demonstrate that they perform best in terms of convergence rate and sample efficiency. [ABSTRACT FROM AUTHOR] |
| Copyright of Computational Intelligence & Neuroscience is the property of Wiley-Blackwell and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
| FullText | Links: – Type: pdflink Text: Availability: 0 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 118500455 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Efficient Actor-Critic Algorithm with Hierarchical Model Learning and Planning. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Zhong%2C+Shan%22">Zhong, Shan</searchLink><relatesTo>1,2</relatesTo><br /><searchLink fieldCode="AR" term="%22Liu%2C+Quan%22">Liu, Quan</searchLink><relatesTo>1,3,4</relatesTo><br /><searchLink fieldCode="AR" term="%22Fu%2C+QiMing%22">Fu, QiMing</searchLink><relatesTo>5</relatesTo> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Computational+Intelligence+%26+Neuroscience%22">Computational Intelligence & Neuroscience</searchLink>. 9/26/2016, p1-15. 15p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Mathematical+models+of+learning%22">Mathematical models of learning</searchLink><br /><searchLink fieldCode="DE" term="%22Mathematical+models%22">Mathematical models</searchLink><br /><searchLink fieldCode="DE" term="%22Planning%22">Planning</searchLink><br /><searchLink fieldCode="DE" term="%22Stochastic+convergence%22">Stochastic convergence</searchLink><br /><searchLink fieldCode="DE" term="%22Reinforcement+learning%22">Reinforcement learning</searchLink><br /><searchLink fieldCode="DE" term="%22Hierarchical+Bayes+model%22">Hierarchical Bayes model</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: To improve the convergence rate and the sample efficiency, two efficient learning methods AC-HMLP and RAC-HMLP (AC-HMLP with l2-regularization) are proposed by combining actor-critic algorithm with hierarchical model learning and planning. The hierarchical models consisting of the local and the global models, which are learned at the same time during learning of the value function and the policy, are approximated by local linear regression (LLR) and linear function approximation (LFA), respectively. Both the local model and the global model are applied to generate samples for planning; the former is used only if the state-prediction error does not surpass the threshold at each time step, while the latter is utilized at the end of each episode. The purpose of taking both models is to improve the sample efficiency and accelerate the convergence rate of the whole algorithm through fully utilizing the local and global information. Experimentally, AC-HMLP and RAC-HMLP are compared with three representative algorithms on two Reinforcement Learning (RL) benchmark problems. The results demonstrate that they perform best in terms of convergence rate and sample efficiency. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of Computational Intelligence & Neuroscience is the property of Wiley-Blackwell and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=118500455 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1155/2016/4824072 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 15 StartPage: 1 Subjects: – SubjectFull: Mathematical models of learning Type: general – SubjectFull: Mathematical models Type: general – SubjectFull: Planning Type: general – SubjectFull: Stochastic convergence Type: general – SubjectFull: Reinforcement learning Type: general – SubjectFull: Hierarchical Bayes model Type: general Titles: – TitleFull: Efficient Actor-Critic Algorithm with Hierarchical Model Learning and Planning. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Zhong, Shan – PersonEntity: Name: NameFull: Liu, Quan – PersonEntity: Name: NameFull: Fu, QiMing IsPartOfRelationships: – BibEntity: Dates: – D: 26 M: 09 Text: 9/26/2016 Type: published Y: 2016 Identifiers: – Type: issn-print Value: 16875265 Titles: – TitleFull: Computational Intelligence & Neuroscience Type: main |
| ResultId | 1 |