
Research Article
A Comparative Analysis of Multi-Armed Bandit Algorithms in the Research of Optimal Product Discovery in E-Commerce
@INPROCEEDINGS{10.4108/eai.22-5-2026.2365147, author={Tailin Su}, title={A Comparative Analysis of Multi-Armed Bandit Algorithms in the Research of Optimal Product Discovery in E-Commerce}, proceedings={Proceedings of the 4th International Conference on Image, Algorithms, and Artificial Intelligence, ICIAAI 2026, 22-24 May 2026, Singapore, Singapore}, publisher={EAI}, proceedings_a={ICIAAI}, year={2026}, month={8}, keywords={Recommender Systems Multi-Armed Bandit Thompson Sampling Cold Start E-commerce}, doi={10.4108/eai.22-5-2026.2365147} }- Tailin Su
Year: 2026
A Comparative Analysis of Multi-Armed Bandit Algorithms in the Research of Optimal Product Discovery in E-Commerce
ICIAAI
EAI
DOI: 10.4108/eai.22-5-2026.2365147
Abstract
This paper investigates how the Multi-Armed Bandit (MAB) model can maximize the discovery of products of instant noodles based on the Ramen Rating data. The paper will treat recommendation as a sequential decision-making process and will compare three particular algorithms, namely, Explore-Then-Commit (ETC), Upper Confidence Bound (UCB), and Thompson Sampling. Cumulative Regret, Optimal Action Accuracy, and Average Reward were used as measures of performance. Experimental evidence shows that Thompson Sampling is significantly more effective compared to the deterministic strategies. It had the least regret and correctly detected the top-performing brand with a high degree of accuracy in less than 1,000 steps. These results imply that probabilistic methods are a strong and economical method of controlling long-tail inventory in sparse-data settings.


