
Editorial
Relative Position Encoding-Enhanced Vision Transformer for Image-Level Defect Classification in Manufacturing Textured-Surface Inspection
@ARTICLE{10.4108/eetsis.13596, author={Jiamin Liu and Yuhui Sun}, title={Relative Position Encoding-Enhanced Vision Transformer for Image-Level Defect Classification in Manufacturing Textured-Surface Inspection}, journal={EAI Endorsed Transactions on Scalable Information Systems}, volume={13}, number={3}, publisher={EAI}, journal_a={SIS}, year={2026}, month={9}, keywords={Vision transformer, relative position encoding, surface defect classification, industrial visual inspection, manufacturing quality inspection}, doi={10.4108/eetsis.13596} }- Jiamin Liu
Yuhui Sun
Year: 2026
Relative Position Encoding-Enhanced Vision Transformer for Image-Level Defect Classification in Manufacturing Textured-Surface Inspection
SIS
EAI
DOI: 10.4108/eetsis.13596
Abstract
INTRODUCTION: In textile and textured-surface manufacturing, visual defects in woven fabrics, leather-like materials, tiles, and grid products directly influence product grading, downstream cutting, assembly reliability, after-sales repair, and supply-chain delivery stability. Weak defects embedded in repetitive textures are still difficult for manual inspection and conventional handcrafted vision methods. OBJECTIVES: This study constructs an image-level defective/non-defective classification model for manufacturing quality inspection by enhancing Vision Transformer with relative spatial modeling, so that subtle texture disruption can be identified before products enter subsequent processing or distribution stages. METHODS: A ViT-Small backbone was combined with learnable relative position encoding. Input images were resized, normalized, divided into 16 × 16 patches, and transformed into token embeddings. Relative spatial bias was inserted into multi-head self-attention to encode patch-to-patch displacement while retaining absolute positional information. The model was evaluated with precision, accuracy, recall, F1-score, and AUROC on a supervised MVTec-derived texture subset. RESULTS: ViT-RPE outperformed ResNet-18, an activation-embedded CNN, EfficientFormer-L1, Swin-Tiny, and the baseline ViT under the same supervised image-level classification protocol. It achieved an accuracy of 0.954 ± 0.004 and an AUROC of 0.977 ± 0.003 on DAGM, and an accuracy of 0.939 ± 0.005 and an AUROC of 0.969 ± 0.004 on the MVTec-derived subset. Ablation results showed that combining absolute and relative positional information produced the best performance, with limited additional computational cost. CONCLUSION: The method provides a compact image-level pre-screening solution for automated product inspection in manufacturing lines, especially where repetitive texture, material appearance consistency, and rapid quality sorting are required.

