About | Contact Us | Register | Login
ProceedingsSeriesJournalsSearchEAI
Proceedings of the 4th International Conference on Image, Algorithms, and Artificial Intelligence, ICIAAI 2026, 22-24 May 2026, Singapore, Singapore

Research Article

Multi-Armed Bandit Algorithms and Their Applications in Online Recommendation Systems

Download8 downloads
Cite
BibTeX Plain Text
  • @INPROCEEDINGS{10.4108/eai.22-5-2026.2365249,
        author={Zijie  Zhao},
        title={Multi-Armed Bandit Algorithms and Their Applications in Online Recommendation Systems},
        proceedings={Proceedings of the 4th International Conference on Image, Algorithms, and Artificial Intelligence, ICIAAI 2026, 22-24 May 2026, Singapore, Singapore},
        publisher={EAI},
        proceedings_a={ICIAAI},
        year={2026},
        month={8},
        keywords={Multi-armed bandit recommendation systems exploration--exploitation},
        doi={10.4108/eai.22-5-2026.2365249}
    }
    
  • Zijie Zhao
    Year: 2026
    Multi-Armed Bandit Algorithms and Their Applications in Online Recommendation Systems
    ICIAAI
    EAI
    DOI: 10.4108/eai.22-5-2026.2365249
Zijie Zhao1,*
  • 1: Polytechnic Institute, Purdue University, Lafayette, 47901, United States of America
*Contact email: zhaozijie_2024@sina.com

Abstract

Recommendation systems are used in various online services like streaming media services, e-commerce sites and news apps to enable users to find content of their choice quickly. This often involves sequential decision-making without knowing the rewards or pay-offs that the recommendations yield. A classic example of such decision-making is the multi-armed bandit (MAB) framework. This paper gives an overview of the basic MAB algorithms and their use cases for online recommendation systems. The paper covered some of the popular bandit algorithms, such as Upper Confidence Bound (UCB), Thompson Sampling, contextual bandits, off-policy evaluation, etc. Bandit algorithms are instrumental in solving the exploration-exploitation dilemma in the recommendation problem. They can learn the user preference and meanwhile, keep optimizing the quality of the item list. However, there are still a few practical problems, such as noisy feedback, large action space, and non-stationarity of user behavior, that remain to be solved.

Keywords
Multi-armed bandit, recommendation systems, exploration–exploitation
Published
2026-08-31
Publisher
EAI
http://dx.doi.org/10.4108/eai.22-5-2026.2365249
Copyright © 2026–2026 EAI
EBSCOProQuestDBLPDOAJPortico
EAI Logo

About EAI

  • Who We Are
  • Leadership
  • Research Areas
  • Partners
  • Media Center
  • Cookie Preferences

Community

  • Membership
  • Conference
  • Recognition
  • Sponsor Us

Publish with EAI

  • Publishing
  • Journals
  • Proceedings
  • Books
  • EUDL