Safe Reinforcement Learning for Optimal Tracking of Continuous‐Time Nonlinear Systems.

Saved in:
Bibliographic Details
Title: Safe Reinforcement Learning for Optimal Tracking of Continuous‐Time Nonlinear Systems.
Authors: Kanso, Soha1 (AUTHOR), Shekhar Jha, Mayank1 (AUTHOR) mayank-shekhar.jha@univ-lorraine.fr, Theilliol, Didier1 (AUTHOR)
Source: International Journal of Robust & Nonlinear Control. Jul2026, Vol. 36 Issue 10, p5378-5391. 14p.
Subjects: Continuous time systems, Reinforcement learning, Quadratic programming, Feedback control systems, Stability of linear systems, Safety regulations, Artificial neural networks
Abstract: This work develops a novel off‐policy safe reinforcement learning (RL) approach for optimal tracking of continuous‐time nonlinear systems, affine in control input. The main contribution consists of the synthesis of an optimal tracker under safety guarantees. A novel formulation is developed enabling optimal tracking of references while satisfying state‐based safety constraints. The tracking error and the state dynamics are considered to form an augmented system, facilitating this dual objective with the primary goal being to guarantee the safety without compromising the system performance. To this end, the safety is achieved during the exploration phase, by dynamically adjusting control inputs that are solutions of quadratic programming (QP) problem that incorporates zeroing control barrier function (ZCBF) conditions. Additionally, the safety during exploitation (operational phase) of the learned policy is strengthened by integrating a reciprocal control barrier function (RCBF) into the cost function, leading to an effective trade‐off between safety and system performance. Neural networks are employed to approximate the optimal control law, and novel mathematically rigorous proofs are developed to guarantee the safety, the stability, and the convergence towards optimality. Finally, the effectiveness of the approach is assessed using a simulation example. [ABSTRACT FROM AUTHOR]
Copyright of International Journal of Robust & Nonlinear Control is the property of Wiley-Blackwell and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Description
Abstract:This work develops a novel off‐policy safe reinforcement learning (RL) approach for optimal tracking of continuous‐time nonlinear systems, affine in control input. The main contribution consists of the synthesis of an optimal tracker under safety guarantees. A novel formulation is developed enabling optimal tracking of references while satisfying state‐based safety constraints. The tracking error and the state dynamics are considered to form an augmented system, facilitating this dual objective with the primary goal being to guarantee the safety without compromising the system performance. To this end, the safety is achieved during the exploration phase, by dynamically adjusting control inputs that are solutions of quadratic programming (QP) problem that incorporates zeroing control barrier function (ZCBF) conditions. Additionally, the safety during exploitation (operational phase) of the learned policy is strengthened by integrating a reciprocal control barrier function (RCBF) into the cost function, leading to an effective trade‐off between safety and system performance. Neural networks are employed to approximate the optimal control law, and novel mathematically rigorous proofs are developed to guarantee the safety, the stability, and the convergence towards optimality. Finally, the effectiveness of the approach is assessed using a simulation example. [ABSTRACT FROM AUTHOR]
ISSN:10498923
DOI:10.1002/rnc.70518