About | Contact Us | Register | Login
ProceedingsSeriesJournalsSearchEAI
Proceedings of the 4th International Conference on Image, Algorithms, and Artificial Intelligence, ICIAAI 2026, 22-24 May 2026, Singapore, Singapore

Research Article

Off-Policy Evaluation of Contextual Bandit Algorithms for Personalized News Recommendation Using the MIND Dataset

Download13 downloads
Cite
BibTeX Plain Text
  • @INPROCEEDINGS{10.4108/eai.22-5-2026.2365082,
        author={Yikai  Zhao},
        title={Off-Policy Evaluation of Contextual Bandit Algorithms for Personalized News Recommendation Using the MIND Dataset},
        proceedings={Proceedings of the 4th International Conference on Image, Algorithms, and Artificial Intelligence, ICIAAI 2026, 22-24 May 2026, Singapore, Singapore},
        publisher={EAI},
        proceedings_a={ICIAAI},
        year={2026},
        month={8},
        keywords={Contextual bandits off-policy evaluation inverse propensity weighting SNIPS MIND dataset},
        doi={10.4108/eai.22-5-2026.2365082}
    }
    
  • Yikai Zhao
    Year: 2026
    Off-Policy Evaluation of Contextual Bandit Algorithms for Personalized News Recommendation Using the MIND Dataset
    ICIAAI
    EAI
    DOI: 10.4108/eai.22-5-2026.2365082
Yikai Zhao1,*
  • 1: University of Utah, Salt Lake City, UT, USA
*Contact email: U1555016@utah.edu

Abstract

News recommendation systems rely on user feedback to improve ranking decisions. Running online tests adds cost. It can also interrupt the user experience. Off-policy evaluation (OPE) gives a safer option for estimating new policies. The paper models news recommendation as a contextual bandit task. The paper constructs TF-IDF features for news articles. The paper also trains a logistic regression model as the target policy. Using the logged data, the paper compute inverse propensity weighting (IPS) and its self-normalized version (SNIPS) to evaluate policy performance. The paper examines the distribution of importance weights and uses bootstrap sampling to study the stability of the IPS estimator. The results show that IPS is sensitive to variation in its weights. SNIPS produces estimates that change less across samples. These patterns are clear. They show that weight normalization matters. They also show that simple diagnostic checks are important when applying OPE to real news recommendation logs.

Keywords
Contextual bandits, off-policy evaluation, inverse propensity weighting, SNIPS, MIND dataset
Published
2026-08-31
Publisher
EAI
http://dx.doi.org/10.4108/eai.22-5-2026.2365082
Copyright © 2026–2026 EAI
EBSCOProQuestDBLPDOAJPortico
EAI Logo

About EAI

  • Who We Are
  • Leadership
  • Research Areas
  • Partners
  • Media Center
  • Cookie Preferences

Community

  • Membership
  • Conference
  • Recognition
  • Sponsor Us

Publish with EAI

  • Publishing
  • Journals
  • Proceedings
  • Books
  • EUDL