
Research Article
Off-Policy Evaluation of Contextual Bandit Algorithms for Personalized News Recommendation Using the MIND Dataset
@INPROCEEDINGS{10.4108/eai.22-5-2026.2365082, author={Yikai Zhao}, title={Off-Policy Evaluation of Contextual Bandit Algorithms for Personalized News Recommendation Using the MIND Dataset}, proceedings={Proceedings of the 4th International Conference on Image, Algorithms, and Artificial Intelligence, ICIAAI 2026, 22-24 May 2026, Singapore, Singapore}, publisher={EAI}, proceedings_a={ICIAAI}, year={2026}, month={8}, keywords={Contextual bandits off-policy evaluation inverse propensity weighting SNIPS MIND dataset}, doi={10.4108/eai.22-5-2026.2365082} }- Yikai Zhao
Year: 2026
Off-Policy Evaluation of Contextual Bandit Algorithms for Personalized News Recommendation Using the MIND Dataset
ICIAAI
EAI
DOI: 10.4108/eai.22-5-2026.2365082
Abstract
News recommendation systems rely on user feedback to improve ranking decisions. Running online tests adds cost. It can also interrupt the user experience. Off-policy evaluation (OPE) gives a safer option for estimating new policies. The paper models news recommendation as a contextual bandit task. The paper constructs TF-IDF features for news articles. The paper also trains a logistic regression model as the target policy. Using the logged data, the paper compute inverse propensity weighting (IPS) and its self-normalized version (SNIPS) to evaluate policy performance. The paper examines the distribution of importance weights and uses bootstrap sampling to study the stability of the IPS estimator. The results show that IPS is sensitive to variation in its weights. SNIPS produces estimates that change less across samples. These patterns are clear. They show that weight normalization matters. They also show that simple diagnostic checks are important when applying OPE to real news recommendation logs.


