Adaptive Smoothing for Path Integral Control.
Saved in:
| Title: | Adaptive Smoothing for Path Integral Control. |
|---|---|
| Authors: | Thalmeier, Dominik1 d.thalmeier@science.ru.nl, Kappen, Hilbert J.1 b.kappen@science.ru.nl, Totaro, Simone2 simone.totaro@gmail.com, Gómez, Vicenç2 vicen.gomez@upf.edu |
| Source: | Journal of Machine Learning Research. 2020, Vol. 21 Issue 189-216, p1-37. 37p. |
| Subjects: | Path integrals, Reinforcement learning, Algorithms, Cost functions, Dynamical systems |
| Abstract: | In Path Integral control problems a representation of an optimally controlled dynamical system can be formally computed and serve as a guidepost to learn a parametrized policy. The Path Integral Cross-Entropy (PICE) method tries to exploit this, but is hampered by poor sample efficiency. We propose a model-free algorithm called ASPIC (Adaptive Smoothing of Path Integral Control) that applies an inf-convolution to the cost function to speedup convergence of policy optimization. We identify PICE as the infinite smoothing limit of such technique and show that the sample efficiency problems that PICE suffers disappear for finite levels of smoothing. For zero smoothing, ASPIC becomes a greedy optimization of the cost, which is the standard approach in current reinforcement learning. ASPIC adapts the smoothness parameter to keep the variance of the gradient estimator at a predefined level, independently of the number of samples. We show analytically and empirically that intermediate levels of smoothing are optimal, which renders the new method superior to both PICE and direct cost optimization. [ABSTRACT FROM AUTHOR] |
| Copyright of Journal of Machine Learning Research is the property of Microtome Publishing and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
| FullText | Text: Availability: 0 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 146744432 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Adaptive Smoothing for Path Integral Control. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Thalmeier%2C+Dominik%22">Thalmeier, Dominik</searchLink><relatesTo>1</relatesTo><i> d.thalmeier@science.ru.nl</i><br /><searchLink fieldCode="AR" term="%22Kappen%2C+Hilbert+J%2E%22">Kappen, Hilbert J.</searchLink><relatesTo>1</relatesTo><i> b.kappen@science.ru.nl</i><br /><searchLink fieldCode="AR" term="%22Totaro%2C+Simone%22">Totaro, Simone</searchLink><relatesTo>2</relatesTo><i> simone.totaro@gmail.com</i><br /><searchLink fieldCode="AR" term="%22Gómez%2C+Vicenç%22">Gómez, Vicenç</searchLink><relatesTo>2</relatesTo><i> vicen.gomez@upf.edu</i> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Journal+of+Machine+Learning+Research%22">Journal of Machine Learning Research</searchLink>. 2020, Vol. 21 Issue 189-216, p1-37. 37p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Path+integrals%22">Path integrals</searchLink><br /><searchLink fieldCode="DE" term="%22Reinforcement+learning%22">Reinforcement learning</searchLink><br /><searchLink fieldCode="DE" term="%22Algorithms%22">Algorithms</searchLink><br /><searchLink fieldCode="DE" term="%22Cost+functions%22">Cost functions</searchLink><br /><searchLink fieldCode="DE" term="%22Dynamical+systems%22">Dynamical systems</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: In Path Integral control problems a representation of an optimally controlled dynamical system can be formally computed and serve as a guidepost to learn a parametrized policy. The Path Integral Cross-Entropy (PICE) method tries to exploit this, but is hampered by poor sample efficiency. We propose a model-free algorithm called ASPIC (Adaptive Smoothing of Path Integral Control) that applies an inf-convolution to the cost function to speedup convergence of policy optimization. We identify PICE as the infinite smoothing limit of such technique and show that the sample efficiency problems that PICE suffers disappear for finite levels of smoothing. For zero smoothing, ASPIC becomes a greedy optimization of the cost, which is the standard approach in current reinforcement learning. ASPIC adapts the smoothness parameter to keep the variance of the gradient estimator at a predefined level, independently of the number of samples. We show analytically and empirically that intermediate levels of smoothing are optimal, which renders the new method superior to both PICE and direct cost optimization. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of Journal of Machine Learning Research is the property of Microtome Publishing and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=146744432 |
| RecordInfo | BibRecord: BibEntity: Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 37 StartPage: 1 Subjects: – SubjectFull: Path integrals Type: general – SubjectFull: Reinforcement learning Type: general – SubjectFull: Algorithms Type: general – SubjectFull: Cost functions Type: general – SubjectFull: Dynamical systems Type: general Titles: – TitleFull: Adaptive Smoothing for Path Integral Control. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Thalmeier, Dominik – PersonEntity: Name: NameFull: Kappen, Hilbert J. – PersonEntity: Name: NameFull: Totaro, Simone – PersonEntity: Name: NameFull: Gómez, Vicenç IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 08 Text: 2020 Type: published Y: 2020 Identifiers: – Type: issn-print Value: 15324435 Numbering: – Type: volume Value: 21 – Type: issue Value: 189-216 Titles: – TitleFull: Journal of Machine Learning Research Type: main |
| ResultId | 1 |