
Research Article
Improvements of Multi-Armed Bandit Algorithms
@INPROCEEDINGS{10.4108/eai.22-5-2026.2365102, author={Hengnian Pan}, title={Improvements of Multi-Armed Bandit Algorithms}, proceedings={Proceedings of the 4th International Conference on Image, Algorithms, and Artificial Intelligence, ICIAAI 2026, 22-24 May 2026, Singapore, Singapore}, publisher={EAI}, proceedings_a={ICIAAI}, year={2026}, month={8}, keywords={Reinforcement learning; Multi-armed Bandit; recommendation; clinical trials}, doi={10.4108/eai.22-5-2026.2365102} }- Hengnian Pan
Year: 2026
Improvements of Multi-Armed Bandit Algorithms
ICIAAI
EAI
DOI: 10.4108/eai.22-5-2026.2365102
Abstract
As a branch of reinforcement learning, Multi-Armed Bandit(MAB) is used in problems of making optimal decisions such as recommendation and clinical trials. It has the ability of gathering information from unknown environments and using this information to maximize the reward in a process. This paper overviews the paths to further improve this algorithm and identifies two major methods of improvement. The internal method focuses on changing the inner structure of the algorithm to make it more generalizable in various circumstances; the external method instead focuses on constructing multi-phase processes with MABs and other algorithms to utilize MAB in more complex tasks. Both paths let the resulting algorithms show better performances than those before improvements in many different realms. Depending on the application, there are many other factors in specific occasions that may influence the performance of MAB, and targeted adjustments may also be needed.


