
Research Article
NLP-Based Robust Object Detection and Recognition Through Multimodal Relation Graph Construction Using MPST AND DSWDCNN
@ARTICLE{10.4108/eetsis.11813, author={Bo Feng and Jiayi Yang and Jingyue Xue }, title={NLP-Based Robust Object Detection and Recognition Through Multimodal Relation Graph Construction Using MPST AND DSWDCNN}, journal={EAI Endorsed Transactions on Scalable Information Systems}, volume={12}, number={5}, publisher={EAI}, journal_a={SIS}, year={2026}, month={4}, keywords={Object Detection and Recognition, Deep Learning, Natural Language Processing, Human Computer Interface (HCI) Applications, Multimodal Relation Graph Construction, Entity Relation Identification, Artificial Intelligence (AI)}, doi={10.4108/eetsis.11813} }- Bo Feng
Jiayi Yang
Jingyue Xue
Year: 2026
NLP-Based Robust Object Detection and Recognition Through Multimodal Relation Graph Construction Using MPST AND DSWDCNN
SIS
EAI
DOI: 10.4108/eetsis.11813
Abstract
INTRODUCTION: Nowadays, object detection and recognition play an important role in various applications such as surveillance, autonomous driving, robotics, and medical imaging. However, none of the traditional works focuses on analyzing the explicit relationships, logical dependencies, and semantic conflicts in text-rich or complex scenes, affecting object recognition accuracy. OBJECTIVES: A Natural Language Processing (NLP)-based robust object detection and recognition framework is developed through multimodal relation graph construction using Minimum Persistence Spanning Tree (MPST) and Deep Swim Wishart Distribution Convolutional Neural Network (DSWDCNN). METHODS: Initially, image with their corresponding captions is collected. Then, the image preprocessing and text preprocessing are done independently. From the preprocessed image, the object is detected using You Aspect-ratio Adaptive Anchors Only Look Once version-8 (YAAAOLOv8), followed by visualization. Meanwhile, from the preprocessed text, entity relations are identified. The multimodal relation graph is constructed using MPST. Further, the features from preprocessed text, relation graphs, detected objects, and visualized-images are extracted. Next, the multimodal analysis is carried out. RESULTS: In the meantime, the word embedding is performed on the preprocessed texts. Finally, the object recognition is carried out using DSWDCNN. CONCLUSION: The proposed framework achieves an object recognition accuracy of 98.8569%, demonstrating its effectiveness under weakly supervised conditions.
Copyright © 2026 Bo Feng et al., licensed to EAI. This is an open access article distributed under the terms of the CC BY-NC-SA 4.0, which permits copying, redistributing, remixing, transformation, and building upon the material in any medium so long as the original work is properly cited.


