
Research Article
A Mechanism-Level Comparison of ETC, UCB, and Thompson Sampling Under Finite-Horizon Constraints
@INPROCEEDINGS{10.4108/eai.22-5-2026.2365150, author={Yichen Fan}, title={A Mechanism-Level Comparison of ETC, UCB, and Thompson Sampling Under Finite-Horizon Constraints}, proceedings={Proceedings of the 4th International Conference on Image, Algorithms, and Artificial Intelligence, ICIAAI 2026, 22-24 May 2026, Singapore, Singapore}, publisher={EAI}, proceedings_a={ICIAAI}, year={2026}, month={8}, keywords={Multi-armed bandits Explore-Then-Commit (ETC) Upper Confidence Bound (UCB) Thompson Sampling}, doi={10.4108/eai.22-5-2026.2365150} }- Yichen Fan
Year: 2026
A Mechanism-Level Comparison of ETC, UCB, and Thompson Sampling Under Finite-Horizon Constraints
ICIAAI
EAI
DOI: 10.4108/eai.22-5-2026.2365150
Abstract
Multi-armed bandit (MAB) algorithms are commonly used in sequential decision tasks such as online recommendation and advertising. In real-world systems, algorithms often have limited time and data to learn from, and poor decisions can be costly. From a finite-horizon perspective, this paper presents a mechanism-level comparison of Explore-Then-Commit (ETC), Upper Confidence Bound (UCB), and Thompson Sampling (TS). The findings suggest ETC is highly sensitive to the length of the exploration phase and does not allow parameter adjustment once it enters the commit stage. UCB tends to incur relatively high exploration costs in short horizons and is also sensitive to parameter settings. In contrast, Thompson Sampling usually exhibits smoother behavior in the early stages and more stable performance under finite horizons, with less reliance on parameter tuning, although some randomness across runs remains. Overall, algorithm selection in real-world systems should consider early-stage behavior and risk under finite-horizon constraints.


