Efficient Actor-Critic Algorithm with Hierarchical Model Learning and Planning.

Saved in:
Bibliographic Details
Title: Efficient Actor-Critic Algorithm with Hierarchical Model Learning and Planning.
Authors: Zhong, Shan1,2, Liu, Quan1,3,4, Fu, QiMing5
Source: Computational Intelligence & Neuroscience. 9/26/2016, p1-15. 15p.
Subjects: Mathematical models of learning, Mathematical models, Planning, Stochastic convergence, Reinforcement learning, Hierarchical Bayes model
Abstract: To improve the convergence rate and the sample efficiency, two efficient learning methods AC-HMLP and RAC-HMLP (AC-HMLP with l2-regularization) are proposed by combining actor-critic algorithm with hierarchical model learning and planning. The hierarchical models consisting of the local and the global models, which are learned at the same time during learning of the value function and the policy, are approximated by local linear regression (LLR) and linear function approximation (LFA), respectively. Both the local model and the global model are applied to generate samples for planning; the former is used only if the state-prediction error does not surpass the threshold at each time step, while the latter is utilized at the end of each episode. The purpose of taking both models is to improve the sample efficiency and accelerate the convergence rate of the whole algorithm through fully utilizing the local and global information. Experimentally, AC-HMLP and RAC-HMLP are compared with three representative algorithms on two Reinforcement Learning (RL) benchmark problems. The results demonstrate that they perform best in terms of convergence rate and sample efficiency. [ABSTRACT FROM AUTHOR]
Copyright of Computational Intelligence & Neuroscience is the property of Wiley-Blackwell and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Links:
  – Type: pdflink
Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 118500455
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Efficient Actor-Critic Algorithm with Hierarchical Model Learning and Planning.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Zhong%2C+Shan%22">Zhong, Shan</searchLink><relatesTo>1,2</relatesTo><br /><searchLink fieldCode="AR" term="%22Liu%2C+Quan%22">Liu, Quan</searchLink><relatesTo>1,3,4</relatesTo><br /><searchLink fieldCode="AR" term="%22Fu%2C+QiMing%22">Fu, QiMing</searchLink><relatesTo>5</relatesTo>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Computational+Intelligence+%26+Neuroscience%22">Computational Intelligence & Neuroscience</searchLink>. 9/26/2016, p1-15. 15p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Mathematical+models+of+learning%22">Mathematical models of learning</searchLink><br /><searchLink fieldCode="DE" term="%22Mathematical+models%22">Mathematical models</searchLink><br /><searchLink fieldCode="DE" term="%22Planning%22">Planning</searchLink><br /><searchLink fieldCode="DE" term="%22Stochastic+convergence%22">Stochastic convergence</searchLink><br /><searchLink fieldCode="DE" term="%22Reinforcement+learning%22">Reinforcement learning</searchLink><br /><searchLink fieldCode="DE" term="%22Hierarchical+Bayes+model%22">Hierarchical Bayes model</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: To improve the convergence rate and the sample efficiency, two efficient learning methods AC-HMLP and RAC-HMLP (AC-HMLP with l2-regularization) are proposed by combining actor-critic algorithm with hierarchical model learning and planning. The hierarchical models consisting of the local and the global models, which are learned at the same time during learning of the value function and the policy, are approximated by local linear regression (LLR) and linear function approximation (LFA), respectively. Both the local model and the global model are applied to generate samples for planning; the former is used only if the state-prediction error does not surpass the threshold at each time step, while the latter is utilized at the end of each episode. The purpose of taking both models is to improve the sample efficiency and accelerate the convergence rate of the whole algorithm through fully utilizing the local and global information. Experimentally, AC-HMLP and RAC-HMLP are compared with three representative algorithms on two Reinforcement Learning (RL) benchmark problems. The results demonstrate that they perform best in terms of convergence rate and sample efficiency. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Computational Intelligence & Neuroscience is the property of Wiley-Blackwell and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=118500455
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1155/2016/4824072
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 15
        StartPage: 1
    Subjects:
      – SubjectFull: Mathematical models of learning
        Type: general
      – SubjectFull: Mathematical models
        Type: general
      – SubjectFull: Planning
        Type: general
      – SubjectFull: Stochastic convergence
        Type: general
      – SubjectFull: Reinforcement learning
        Type: general
      – SubjectFull: Hierarchical Bayes model
        Type: general
    Titles:
      – TitleFull: Efficient Actor-Critic Algorithm with Hierarchical Model Learning and Planning.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Zhong, Shan
      – PersonEntity:
          Name:
            NameFull: Liu, Quan
      – PersonEntity:
          Name:
            NameFull: Fu, QiMing
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 26
              M: 09
              Text: 9/26/2016
              Type: published
              Y: 2016
          Identifiers:
            – Type: issn-print
              Value: 16875265
          Titles:
            – TitleFull: Computational Intelligence & Neuroscience
              Type: main
ResultId 1