
Editorial
Reference-Assisted Learner-Aligned Distributed Stacking for Edge Intrusion Detection in the Industrial Internet of Things
@ARTICLE{10.4108/eetsis.14107, author={Jing Li and Shuhao Shen and Kangrui Xu and Lin Cui}, title={Reference-Assisted Learner-Aligned Distributed Stacking for Edge Intrusion Detection in the Industrial Internet of Things}, journal={EAI Endorsed Transactions on Scalable Information Systems}, volume={13}, number={3}, publisher={EAI}, journal_a={SIS}, year={2026}, month={8}, keywords={Industrial Internet of Things, edge intrusion detection, distributed stacking, non-IID data, reference data, model fusion}, doi={10.4108/eetsis.14107} }- Jing Li
Shuhao Shen
Kangrui Xu
Lin Cui
Year: 2026
Reference-Assisted Learner-Aligned Distributed Stacking for Edge Intrusion Detection in the Industrial Internet of Things
SIS
EAI
DOI: 10.4108/eetsis.14107
Abstract
Networked manufacturing couples sensors, programmable logic controllers, and industrial gateways to production that cannot simply be paused. An intrusion detector in this setting must control false alarms and edge-resource use as well as detect attacks. Industrial Internet of Things (IIoT) edge nodes add three practical constraints: training traffic cannot be centralized, local statistical distributions differ, and the deployed models may be structurally heterogeneous. We address this setting with Reference-Assisted Learner-Aligned Distributed Stacking (RA-LADS). It evaluates node models on mutually exclusive reference pools, groups LightGBM and XGBoost responses by learner family, and represents each family by its mean, standard deviation, and five quantiles in a fixed, permutation-invariant 7M-dimensional vector. Across 20 matched partitioning and training seeds, the full 14-dimensional summary attained a mean F1 of 0.984514. Its gain over a global 7-dimensional summary without family semantics was 0.000756 (95% CI: 0.000559–0.000962; Holm-adjusted p = 1.08 × 10−5). By comparison, the difference from node-wise raw prediction concatenation was −0.000009, with an interval spanning zero. Random balanced grouping was not significantly different from true learner-family grouping after multiplicity correction; family identity is therefore a useful grouping prior, but not the only one. Mean F1 scores for a compact reference-pool LightGBM, a 41-parameter Tiny Deep Sets model, and a heterogeneous one-model-per-node summary were 0.983887, 0.983659, and 0.983865. The full method delivered small yet reproducible paired gains over all three controls. At fixed false-alarm-rate (FAR) budgets, no stable advantage over a single LightGBM appeared between 0.1% and 2% FAR; at 5% FAR, F1 increased by 0.000417 (Holm-adjusted p = 0.017). Five-seed retraining on two external NetFlow datasets did not show a general advantage. The claim is therefore limited to settings in which node responses carry exploitable distributional structure. Within that boundary, RA-LADS offers a lightweight interface for heterogeneous tree collaboration when labeled reference data and explicit alert operating points are available, although system cost still grows linearly.
Copyright © 2026 Jing Li et al., licensed to EAI. This is an open access article distributed under the terms of the CC BY-NC-SA 4.0, which permits copying, redistributing, remixing, transformation, and building upon the material in any medium so long as the original work is properly cited.


