Safe Reinforcement Learning for Optimal Tracking of Continuous‐Time Nonlinear Systems.
Saved in:
| Title: | Safe Reinforcement Learning for Optimal Tracking of Continuous‐Time Nonlinear Systems. |
|---|---|
| Authors: | Kanso, Soha1 (AUTHOR), Shekhar Jha, Mayank1 (AUTHOR) mayank-shekhar.jha@univ-lorraine.fr, Theilliol, Didier1 (AUTHOR) |
| Source: | International Journal of Robust & Nonlinear Control. Jul2026, Vol. 36 Issue 10, p5378-5391. 14p. |
| Subjects: | Continuous time systems, Reinforcement learning, Quadratic programming, Feedback control systems, Stability of linear systems, Safety regulations, Artificial neural networks |
| Abstract: | This work develops a novel off‐policy safe reinforcement learning (RL) approach for optimal tracking of continuous‐time nonlinear systems, affine in control input. The main contribution consists of the synthesis of an optimal tracker under safety guarantees. A novel formulation is developed enabling optimal tracking of references while satisfying state‐based safety constraints. The tracking error and the state dynamics are considered to form an augmented system, facilitating this dual objective with the primary goal being to guarantee the safety without compromising the system performance. To this end, the safety is achieved during the exploration phase, by dynamically adjusting control inputs that are solutions of quadratic programming (QP) problem that incorporates zeroing control barrier function (ZCBF) conditions. Additionally, the safety during exploitation (operational phase) of the learned policy is strengthened by integrating a reciprocal control barrier function (RCBF) into the cost function, leading to an effective trade‐off between safety and system performance. Neural networks are employed to approximate the optimal control law, and novel mathematically rigorous proofs are developed to guarantee the safety, the stability, and the convergence towards optimality. Finally, the effectiveness of the approach is assessed using a simulation example. [ABSTRACT FROM AUTHOR] |
| Copyright of International Journal of Robust & Nonlinear Control is the property of Wiley-Blackwell and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
| FullText | Text: Availability: 0 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 194235571 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Safe Reinforcement Learning for Optimal Tracking of Continuous‐Time Nonlinear Systems. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Kanso%2C+Soha%22">Kanso, Soha</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Shekhar+Jha%2C+Mayank%22">Shekhar Jha, Mayank</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> mayank-shekhar.jha@univ-lorraine.fr</i><br /><searchLink fieldCode="AR" term="%22Theilliol%2C+Didier%22">Theilliol, Didier</searchLink><relatesTo>1</relatesTo> (AUTHOR) – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22International+Journal+of+Robust+%26+Nonlinear+Control%22">International Journal of Robust & Nonlinear Control</searchLink>. Jul2026, Vol. 36 Issue 10, p5378-5391. 14p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Continuous+time+systems%22">Continuous time systems</searchLink><br /><searchLink fieldCode="DE" term="%22Reinforcement+learning%22">Reinforcement learning</searchLink><br /><searchLink fieldCode="DE" term="%22Quadratic+programming%22">Quadratic programming</searchLink><br /><searchLink fieldCode="DE" term="%22Feedback+control+systems%22">Feedback control systems</searchLink><br /><searchLink fieldCode="DE" term="%22Stability+of+linear+systems%22">Stability of linear systems</searchLink><br /><searchLink fieldCode="DE" term="%22Safety+regulations%22">Safety regulations</searchLink><br /><searchLink fieldCode="DE" term="%22Artificial+neural+networks%22">Artificial neural networks</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: This work develops a novel off‐policy safe reinforcement learning (RL) approach for optimal tracking of continuous‐time nonlinear systems, affine in control input. The main contribution consists of the synthesis of an optimal tracker under safety guarantees. A novel formulation is developed enabling optimal tracking of references while satisfying state‐based safety constraints. The tracking error and the state dynamics are considered to form an augmented system, facilitating this dual objective with the primary goal being to guarantee the safety without compromising the system performance. To this end, the safety is achieved during the exploration phase, by dynamically adjusting control inputs that are solutions of quadratic programming (QP) problem that incorporates zeroing control barrier function (ZCBF) conditions. Additionally, the safety during exploitation (operational phase) of the learned policy is strengthened by integrating a reciprocal control barrier function (RCBF) into the cost function, leading to an effective trade‐off between safety and system performance. Neural networks are employed to approximate the optimal control law, and novel mathematically rigorous proofs are developed to guarantee the safety, the stability, and the convergence towards optimality. Finally, the effectiveness of the approach is assessed using a simulation example. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of International Journal of Robust & Nonlinear Control is the property of Wiley-Blackwell and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=194235571 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1002/rnc.70518 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 14 StartPage: 5378 Subjects: – SubjectFull: Continuous time systems Type: general – SubjectFull: Reinforcement learning Type: general – SubjectFull: Quadratic programming Type: general – SubjectFull: Feedback control systems Type: general – SubjectFull: Stability of linear systems Type: general – SubjectFull: Safety regulations Type: general – SubjectFull: Artificial neural networks Type: general Titles: – TitleFull: Safe Reinforcement Learning for Optimal Tracking of Continuous‐Time Nonlinear Systems. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Kanso, Soha – PersonEntity: Name: NameFull: Shekhar Jha, Mayank – PersonEntity: Name: NameFull: Theilliol, Didier IsPartOfRelationships: – BibEntity: Dates: – D: 10 M: 07 Text: Jul2026 Type: published Y: 2026 Identifiers: – Type: issn-print Value: 10498923 Numbering: – Type: volume Value: 36 – Type: issue Value: 10 Titles: – TitleFull: International Journal of Robust & Nonlinear Control Type: main |
| ResultId | 1 |