About | Contact Us | Register | Login
ProceedingsSeriesJournalsSearchEAI
Proceedings of the 4th International Conference on Image, Algorithms, and Artificial Intelligence, ICIAAI 2026, 22-24 May 2026, Singapore, Singapore

Research Article

Medical Image Segmentation Based on the CLIP and SAM Base Models

Download11 downloads
Cite
BibTeX Plain Text
  • @INPROCEEDINGS{10.4108/eai.22-5-2026.2365106,
        author={Jingyi  Yu},
        title={Medical Image Segmentation Based on the CLIP and SAM Base Models},
        proceedings={Proceedings of the 4th International Conference on Image, Algorithms, and Artificial Intelligence, ICIAAI 2026, 22-24 May 2026, Singapore, Singapore},
        publisher={EAI},
        proceedings_a={ICIAAI},
        year={2026},
        month={8},
        keywords={Base model Medical image segmentation CLIP SAM Multimodal fusion},
        doi={10.4108/eai.22-5-2026.2365106}
    }
    
  • Jingyi Yu
    Year: 2026
    Medical Image Segmentation Based on the CLIP and SAM Base Models
    ICIAAI
    EAI
    DOI: 10.4108/eai.22-5-2026.2365106
Jingyi Yu1,*
  • 1: Computer Engineering Department, TAIYUAN INSTITUTE OF TECHNOLOGY, Taiyuan, Shanxi, China
*Contact email: yv105086@163.com

Abstract

Medical image segmentation is a key technology for achieving precise medical care. Traditional methods rely on a large amount of labeled data and have limited generalization capabilities. Visual foundation models represented by CLIP and SAM have obtained general capabilities through large-scale pre-training, providing a new paradigm for medical segmentation. CLIP achieves semantic-guided segmentation through image-text alignment, while SAM achieves general segmentation through prompt interaction. Both can effectively reduce the reliance on labeled data. This paper systematically reviews the applications of these two types of models in medical segmentation: first, analyzes the necessity of their applications; second, separately summarizes the technical routes and adaptation progress of semantic-guided methods based on CLIP and prompt-interaction methods based on SAM; finally, discusses their advantages and challenges. The review shows that these models significantly improve data efficiency and cross-domain generalization capabilities, but still need further exploration in medical specificity adaptation and fine-grained segmentation.

Keywords
Base model, Medical image segmentation, CLIP, SAM, Multimodal fusion
Published
2026-08-31
Publisher
EAI
http://dx.doi.org/10.4108/eai.22-5-2026.2365106
Copyright © 2026–2026 EAI
EBSCOProQuestDBLPDOAJPortico
EAI Logo

About EAI

  • Who We Are
  • Leadership
  • Research Areas
  • Partners
  • Media Center
  • Cookie Preferences

Community

  • Membership
  • Conference
  • Recognition
  • Sponsor Us

Publish with EAI

  • Publishing
  • Journals
  • Proceedings
  • Books
  • EUDL