
Editorial
Reinforcement Learning Method for Collaborative Multi-Objective Path Planning of IIoT Drones and Unmanned Vehicles in Random Demand and Dynamic Delivery Environments
@ARTICLE{10.4108/eetsis.13700, author={Li Qiao and Changkai Xu and Bin Ni}, title={Reinforcement Learning Method for Collaborative Multi-Objective Path Planning of IIoT Drones and Unmanned Vehicles in Random Demand and Dynamic Delivery Environments}, journal={EAI Endorsed Transactions on Scalable Information Systems}, volume={13}, number={2}, publisher={EAI}, journal_a={SIS}, year={2026}, month={8}, keywords={Industrial Internet of Things, Drone-Unmanned Vehicle Collaboration, Stochastic Demand, Multi-Objective Path Planning, MADDPG, Guidance Factor}, doi={10.4108/eetsis.13700} }- Li Qiao
Changkai Xu
Bin Ni
Year: 2026
Reinforcement Learning Method for Collaborative Multi-Objective Path Planning of IIoT Drones and Unmanned Vehicles in Random Demand and Dynamic Delivery Environments
SIS
EAI
DOI: 10.4108/eetsis.13700
Abstract
INTRODUCTION: Dynamic Industrial Internet of Things (IIoT) delivery is characterized by stochastic order arrivals, heterogeneous UAV–UGV capabilities, and real-time demand updates, making conventional static routing methods insufficient for collaborative delivery planning. Existing UAV trajectory-planning and truck–drone routing studies usually focus on either single-platform trajectory control or offline vehicle routing, and they rarely integrate stochastic IIoT demand updating, UAV–UGV rendezvous coordination, operational constraints, and multi-objective decision-making into a unified reinforcement-learning framework. OBJECTIVES: To address this limitation, this study proposes a guidance-factor-enhanced MADDPG framework for collaborative multi-objective path planning of UAVs and UGVs in random-demand and dynamic delivery environments. METHODS: Order arrivals are modeled by a Poisson process, while order location, payload, priority, and time-window attributes are generated through multivariate random distributions and updated by IIoT terminals every 5 s. Payload, endurance, speed, rendezvous, and task-sequencing constraints are incorporated into a normalized multi-objective function considering delivery time, delivery cost, demand satisfaction, and UAV endurance loss. A distance–demand–priority guidance factor and a reconstructed reward function are further introduced to alleviate sparse-reward effects and guide agents toward high-value demand regions. RESULTS: Simulation and ablation results show that the reconstructed reward is essential for stable learning, while the guidance factor significantly accelerates convergence. Under 40 IIoT nodes, the proposed method converges after approximately 2,000 episodes, whereas the model without guidance information requires about 5,000 episodes. Comparative experiments further show that the proposed method improves system throughput and collaborative delivery efficiency while maintaining feasible UAV trajectories. CONCLUSION: The proposed framework provides an adaptive reinforcement-learning solution for UAV–UGV cooperative logistics in dynamic IIoT environments.
Copyright © 2026 Li Qiao et al., licensed to EAI. This is an open access article distributed under the terms of the CC BY-NC-SA 4.0, which permits copying, redistributing, remixing, transformation, and building upon the material in any medium so long as the original work is properly cited.


