About | Contact Us | Register | Login
ProceedingsSeriesJournalsSearchEAI
sis 26(7):

Editorial

Research on the Application of Large Language Model in Data Integration

Download3 downloads
Cite
BibTeX Plain Text
  • @ARTICLE{10.4108/eetsis.10245,
        author={Zhanfang Chen and Yuan Ren and Xiaoming Jiang and Ruipeng Qi},
        title={Research on the Application of Large Language Model in Data Integration},
        journal={EAI Endorsed Transactions on Scalable Information Systems},
        volume={12},
        number={7},
        publisher={EAI},
        journal_a={SIS},
        year={2026},
        month={1},
        keywords={large language models (LLMs), metadata fine-tuning, Transformer, data governance, reinforcement learning, symbolic regression},
        doi={10.4108/eetsis.10245}
    }
    
  • Zhanfang Chen
    Yuan Ren
    Xiaoming Jiang
    Ruipeng Qi
    Year: 2026
    Research on the Application of Large Language Model in Data Integration
    SIS
    EAI
    DOI: 10.4108/eetsis.10245
Zhanfang Chen1,*, Yuan Ren1, Xiaoming Jiang1, Ruipeng Qi1
  • 1: Changchun University of Science and Technology
*Contact email: chenzhanfang@cust.edu.cn

Abstract

INTRODUCTION: Large Language Models (LLMs), a major breakthrough in artificial intelligence, have been widely applied across various domains in recent years. Their powerful capabilities in language comprehension and generation enable effective handling of diverse natural language processing tasks, such as text generation, question answering, machine translation, and information retrieval. This paper investigates the application of LLM technology in data integration, a core aspect of data governance. In contrast to end-to-end black-box approaches, we reframe data integration as a problem of discovering interpretable mapping rules through symbolic regression. OBJECTIVES: We begin by defining the fundamental problem of data integration. We then propose a general-purpose large model framework for data governance, built on a deep symbolic regression foundation. The framework comprises a symbolic expression generator and a metadata-enhanced executor, aiming to achieve both high accuracy and interpretability. METHODS: The model is trained using a combination of recurrent neural networks and reinforcement learning techniques, for expression generation and the execution of the discovered rules is structured based on a Transformer encoder architecture enhanced with a dedicated metadata embedding layer. To enhance performance, we incorporate metadata fine-tuning, where the generated symbolic expressions serve as key metadata to guide the integration process. RESULTS: Finally, the proposed model is evaluated on two representative data integration tasks, with experimental results demonstrating its effectiveness. CONCLUSION: The results validate its practical quality and highlight the advantage of the symbolic regression paradigm in enhancing interpretability.

Keywords
large language models (LLMs), metadata fine-tuning, Transformer, data governance, reinforcement learning, symbolic regression
Received
2025-09-11
Accepted
2026-01-13
Published
2026-01-26
Publisher
EAI
http://dx.doi.org/10.4108/eetsis.10245

Copyright © 2026 Zhanfang Chen et al., licensed to EAI. This is an open access article distributed under the terms of the CC BY-NCSA 4.0, which permits copying, redistributing, remixing, transformation, and building upon the material in any medium so long as the original work is properly cited.

EBSCOProQuestDBLPDOAJPortico
EAI Logo

About EAI

  • Who We Are
  • Leadership
  • Research Areas
  • Partners
  • Media Center
  • Cookie Preferences

Community

  • Membership
  • Conference
  • Recognition
  • Sponsor Us

Publish with EAI

  • Publishing
  • Journals
  • Proceedings
  • Books
  • EUDL