Electric Vehicle Charging and Discharging Scheduling Method Based on Clustering and Deep Reinforcement Learning.
Saved in:
| Title: | Electric Vehicle Charging and Discharging Scheduling Method Based on Clustering and Deep Reinforcement Learning. |
|---|---|
| Authors: | He, Chunqi1 (AUTHOR) hcq561@mail.shiep.edu.cn, Li, Jiang1 (AUTHOR) |
| Source: | Energies (19961073). May2026, Vol. 19 Issue 9, p2238. 36p. |
| Subject Terms: | *Reinforcement learning, *Clustering algorithms, *Scheduling, *Mixed integer linear programming, *Electric power system management, *Electric vehicle charging stations, *Markov processes |
| Abstract: | With the large-scale integration of electric vehicles (EVs) into the power grid, uncoordinated charging behavior has aggravated load fluctuations in the power system. Deep reinforcement learning can optimize EV charging and discharging strategies through dynamic decision-making, thereby alleviating the operational pressure imposed on the grid by load variations. However, under large-scale EV integration scenarios, challenges still remain, including the excessively high dimensionality of the state space and the resulting decline in training efficiency. In addition, the coupling between existing clustering methods and dynamic scheduling mechanisms is still insufficiently tight. To address these issues, this study proposes a cluster-based deep reinforcement learning method for EV charging and discharging scheduling, referred to as CDRL. First, a probabilistic behavioral model is constructed based on EV charging transaction data to characterize the stochasticity of user charging behavior. A Density–Centroid Hybrid Clustering (DCHC) method is then adopted to cluster the charging behavior characteristics of EVs. Subsequently, at the cluster level, a day-ahead base load forecasting model is introduced, and the forecasting results are fed into a mixed-integer linear programming (MILP) model to generate the charging and discharging power allocation tasks for each cluster. At the individual level, the EV charging and discharging process is formulated as a Markov decision process (MDP), and a deep Q-network (DQN) is employed for policy learning, thereby achieving the decomposition of cluster-level tasks into individual scheduling decisions. The simulation results demonstrate that the proposed method can effectively reduce charging costs and smooth system load fluctuations while improving training convergence speed and policy stability. [ABSTRACT FROM AUTHOR] |
| Database: | Energy & Power Source |
|
Full text is not displayed to guests.
Login for full access.
|
|
| Abstract: | With the large-scale integration of electric vehicles (EVs) into the power grid, uncoordinated charging behavior has aggravated load fluctuations in the power system. Deep reinforcement learning can optimize EV charging and discharging strategies through dynamic decision-making, thereby alleviating the operational pressure imposed on the grid by load variations. However, under large-scale EV integration scenarios, challenges still remain, including the excessively high dimensionality of the state space and the resulting decline in training efficiency. In addition, the coupling between existing clustering methods and dynamic scheduling mechanisms is still insufficiently tight. To address these issues, this study proposes a cluster-based deep reinforcement learning method for EV charging and discharging scheduling, referred to as CDRL. First, a probabilistic behavioral model is constructed based on EV charging transaction data to characterize the stochasticity of user charging behavior. A Density–Centroid Hybrid Clustering (DCHC) method is then adopted to cluster the charging behavior characteristics of EVs. Subsequently, at the cluster level, a day-ahead base load forecasting model is introduced, and the forecasting results are fed into a mixed-integer linear programming (MILP) model to generate the charging and discharging power allocation tasks for each cluster. At the individual level, the EV charging and discharging process is formulated as a Markov decision process (MDP), and a deep Q-network (DQN) is employed for policy learning, thereby achieving the decomposition of cluster-level tasks into individual scheduling decisions. The simulation results demonstrate that the proposed method can effectively reduce charging costs and smooth system load fluctuations while improving training convergence speed and policy stability. [ABSTRACT FROM AUTHOR] |
|---|---|
| ISSN: | 19961073 |
| DOI: | 10.3390/en19092238 |