
Editorial
A Transformer-Enhanced Multi-Agent Reinforcement Learning Model for Resilience Optimization in Educational Equipment Manufacturing Supply Chains
@ARTICLE{10.4108/eetsis.14151, author={Yiquan Kong}, title={A Transformer-Enhanced Multi-Agent Reinforcement Learning Model for Resilience Optimization in Educational Equipment Manufacturing Supply Chains}, journal={EAI Endorsed Transactions on Scalable Information Systems}, volume={13}, number={2}, publisher={EAI}, journal_a={SIS}, year={2026}, month={8}, keywords={Transformer, Multi-agent reinforcement learning, Supply chain resilience, Disruption-risk}, doi={10.4108/eetsis.14151} }- Yiquan Kong
Year: 2026
A Transformer-Enhanced Multi-Agent Reinforcement Learning Model for Resilience Optimization in Educational Equipment Manufacturing Supply Chains
SIS
EAI
DOI: 10.4108/eetsis.14151
Abstract
INTRODUCTION: Educational equipment manufacturing supply chains are vulnerable to demand fluctuations, equipment failures, logistics disruptions, and cross-node risk propagation, while conventional approaches often separate state prediction from recovery decision-making. OBJECTIVES: This study proposes TMARL-ESCR, a Transformer-enhanced multi-agent reinforcement learning framework for supply chain resilience optimization. METHODS: The model represents suppliers, manufacturers, logistics providers, and distribution centers as a dynamic network and uses a Transformer to capture long-range temporal dependencies and cross-node interactions. Multitask prediction heads estimate future demand, logistics lead time, available capacity, and disruption risk, and these predictions are fused with current states to support coordinated recovery decisions. Under a centralized training and decentralized execution framework, multiple agents jointly optimize procurement, production, transportation, and inventory reallocation through local-global rewards and explicit operational constraints. RESULTS:The prediction module achieves a demand WMAPE of 14.38% and a disruption-risk AUC of 0.941. The complete model reaches a 95.2% order fulfillment rate, a 7.1-day recovery time, a 0.5% constraint violation rate, and a resilience score of 0.892. Compared with Transformer-MAPPO, TMARL-ESCR reduces recovery time by 19.3% and improves the resilience score by 5.9%. CONCLUSION: Experiments integrating M5 Forecasting, AI4I 2020, LaDe, and an SCML-based simulation environment evaluate predictive accuracy and resilience optimization under multiple disruption scenarios, demonstrating improved proactive recovery and coordinated resilience under complex disruptions.
Copyright © 2026 Yiquan Kong, licensed to EAI. This is an open access article distributed under the terms of the CC BY-NC-SA 4.0, which permits copying, redistributing, remixing, transformation, and building upon the material in any medium so long as the original work is properly cited.


