
Editorial
Physics-Informed Deep Reinforcement Learning with Adaptive Action Masking for Secure Second-Level Emergency Scheduling in Energy Networks
@ARTICLE{10.4108/ew.14204, author={Bochun Zhan and Zhengbo Shan and Ke Wang and Rong Yan and Shengmin Qiu and Qingbiao Lin and Zhantao Fan and Nan Lou and Xixi Zhang}, title={Physics-Informed Deep Reinforcement Learning with Adaptive Action Masking for Secure Second-Level Emergency Scheduling in Energy Networks}, journal={EAI Endorsed Transactions on Energy Web}, volume={13}, number={1}, publisher={EAI}, journal_a={EW}, year={2026}, month={9}, keywords={real-time market scheduling, situational risk, data-knowledge driven, secure decision-making, physics-informed deep reinforcement learning, adaptive action masking}, doi={10.4108/ew.14204} }- Bochun Zhan
Zhengbo Shan
Ke Wang
Rong Yan
Shengmin Qiu
Qingbiao Lin
Zhantao Fan
Nan Lou
Xixi Zhang
Year: 2026
Physics-Informed Deep Reinforcement Learning with Adaptive Action Masking for Secure Second-Level Emergency Scheduling in Energy Networks
EW
EAI
DOI: 10.4108/ew.14204
Abstract
With the integration of a high proportion of volatile renewable energy into modern power grids, systems frequently face severe supply-demand imbalances and escalated situational risks under extreme events. Addressing the high-dimensional risks and stringent response timing requirements in real-time markets of new-type power systems, this paper proposes a Physics-Informed Deep Reinforcement Learning (PI-DRL) framework with adaptive action masking for secure energy network scheduling and emergency control. First, to mitigate the “curse of dimensionality” inherent in massive heterogeneous resources, a situational risk-driven dynamic feature selection and action masking technique is established to exponentially compress the agent’s exploration space, solving the convergence bottleneck of large-scale resource dispatch. Second, targeting low-probability, high-risk scenarios, a knowledge-driven physical security shield is constructed, utilizing local Jacobian sensitivity matrices to achieve near-real-time physical correction of actions. This overcomes the lack of boundary awareness in traditional black-box models. Furthermore, a Security Reward Shaping mechanism is introduced to foster endogenous security awareness within the agent through closed-loop training, ensuring millisecond-level decision safety during online deployment. Validation on a modified IEEE 118-bus system using high-fidelity meteorological big data from China demonstrates that the proposed method accelerates decision-making by 400-fold compared to traditional MILP algorithms, enabling second-level response. Compared to pure data-driven models, it reduces operational costs by 12.4% and ensures absolute compliance with physical constraints while suppressing risk spikes within 3 minutes during an extreme 450 MW power drop event.
Copyright © 2026 Bochun Zhan et al., licensed to EAI. This is an open access article distributed under the terms of the CC BY-NCSA 4.0, which permits copying, redistributing, remixing, transformation, and building upon the material in any medium so long as the original work is properly cited.

