
Research Article
Deep Reinforcement Learning-Based Intelligent Control for Efficiency Enhancement in Thermal Power Plant Fuel Management
@ARTICLE{10.4108/ew.11848, author={Rui Zhu and Qiang Liu and Guofeng Li and Xufeng Hong and Zhenlu Tian and Yanjun Guo and Shengju Hao}, title={Deep Reinforcement Learning-Based Intelligent Control for Efficiency Enhancement in Thermal Power Plant Fuel Management}, journal={EAI Endorsed Transactions on Energy Web}, volume={13}, number={1}, publisher={EAI}, journal_a={EW}, year={2026}, month={5}, keywords={Deep Reinforcement Learning, Thermal Power Plants, Fuel Management, Efficiency Optimization, Multi‑Objective Control, CO₂ Emissions}, doi={10.4108/ew.11848} }- Rui Zhu
Qiang Liu
Guofeng Li
Xufeng Hong
Zhenlu Tian
Yanjun Guo
Shengju Hao
Year: 2026
Deep Reinforcement Learning-Based Intelligent Control for Efficiency Enhancement in Thermal Power Plant Fuel Management
EW
EAI
DOI: 10.4108/ew.11848
Abstract
Thermal power plants remain a significant component of global power generation; however, several limitations persist. Hence, this research work has been developed on the basis of a proposed intelligent fuel management system based on Deep Reinforcement Learning techniques with a Proximal Policy Optimization (PPO) algorithm as a step toward increasing efficiency and sustainability of operation of thermal power plants. In this work, a fuel management problem has been formulated as a Markov Decision Process (MDP) environment within which a Deep Reinforcement Learning agent interacts with the boiler–turbine and condenser system using real efficiency data from a thermal power plant. of a thermal power plant. A multi-objective reward function was formulated using a reward shaping strategy, whereby the reward signal is explicitly structured to guide the reinforcement learning agent toward thermodynamically efficient and emission-aware plant operation. The reward formulation maximizes thermal efficiency while penalizing higher heat rate, auxiliary power consumption, and CO₂ emissions. Experimental results demonstrate that the proposed Deep Reinforcement Learning approach outperforms conventional control models. The efficiency level of this system raises from 33.68% to 35.72%, marking a relative improvement of 2.04%, with a lowered auxiliary power demand from 6.08% to 5.73%. More significantly, this optimized policy provides an expected 15-20% reduction in CO₂ emissions and lowers the heat rate from 14,000 kJ/kWh down to 11,000–12,000 kJ/kWh,000 kJ/kWh from previous levels. Convergence has been observed in the rise of episode reward values and reducing loss values during training. The current work marks a fresh start utilizing the power of PPO-Based Deep RL with Multiple Reward design in real-time closed-loop fuel management operations as a highly scalable and adaptable alternative compared to rule-set and traditional supervised learning methods.


