
Research Article
The Bandit Learning beyond Stationarity under the Game Theory
@INPROCEEDINGS{10.4108/eai.22-5-2026.2365325, author={Beishen Guo}, title={The Bandit Learning beyond Stationarity under the Game Theory}, proceedings={Proceedings of the 4th International Conference on Image, Algorithms, and Artificial Intelligence, ICIAAI 2026, 22-24 May 2026, Singapore, Singapore}, publisher={EAI}, proceedings_a={ICIAAI}, year={2026}, month={8}, keywords={Endogenous Non-stationarity; Game-theoretic perspective; Strategic Regret}, doi={10.4108/eai.22-5-2026.2365325} }- Beishen Guo
Year: 2026
The Bandit Learning beyond Stationarity under the Game Theory
ICIAAI
EAI
DOI: 10.4108/eai.22-5-2026.2365325
Abstract
It is a universal truth that classical multi-armed bandit (MAB) theory relies on the stationary assumption of reward distributions. However, this assumption is not viable in systems with multiple adaptive interacting agents where interactions endogenously redefine rewards. This paper explores bandit learning in non-stationary systems from a game-theoretic perspective, where learning in game theory, modern bandit theory, and recent advances in multi-player bandits, cooperative learning, federated decision systems, incentive design, adversarial robustness, and teamwork modeling are leveraged to propose a unifying theory of endogenous non-stationarity in bandit learning. The paper provides a conceptual framework and synthesis of strategic regret in interacting systems, stability regimes of learning systems, and how exploration-exploitation trade-offs are transformed in interactive systems.


