Paper: Strategic Arms with Side Communication Prevail Over Low-Regret MAB Algorithms
Authors: Ahmed Ben Yahmed, Clément Calauzènes, Vianney Perchet
Category: Learning Theory, Sequential Learning
Revue: ICASSP 2024
Why did we work on this topic (the problem we want to solve)?
We chose to investigate this topic because we aimed to address a fundamental problem within the strategic multi-armed bandit (MAB) setting. In contrast to the traditional non-strategic MAB, strategic arms are rational entities that strategically adjust their offers by reserving a portion of observed rewards for their own utility and employing communication to influence player decisions. This strategic framework introduces a game-like dynamic, engendering a competition of objectives between the player, striving to minimize regret, and the arms, motivated to maximize their own utilities. Specifically, when arms have perfect information about the player’s behavior, they can establish an equilibrium where they retain nearly all of their value while leaving the player with significant (linear) regret. This study demonstrates that even if complete information is not openly accessible to all arms but is instead shared among them, it is still feasible to achieve a comparable equilibrium. Through this research, we sought to shed light on how strategic arms with side communication can effectively surpass low-regret multi-armed bandit algorithms, providing insights into optimizing decision-making processes in such strategic contexts.
How did we proceed?
Leveraging communication schemes and techniques used in distributed learning, we demonstrate that even though the full historical information is not publicly available for all arms, they can iteratively estimate the true total information at a good rate, enabling them to establish an efficient collusion-style equilibrium, regardless of the low-regret Multi-Armed Bandit (MAB) algorithm used by the player.
What did we find? What did we achieve?
In scenarios involving repeated interactions, converting a single-step collusion scenario into an equilibrium within the cumulative game requires each participant’s ability to identify instances where others deviate from collusive behavior. This study illustrates that even when not all historical information is publicly accessible, in the presence of side communication, even with a minimal scheme, the arms can implement a communication strategy enabling each of them to detect deviations. These may involve falsifying player reports or manipulating shared information. The established equilibrium surpasses any low-regret MAB algorithm, ultimately resulting in reduced player revenues.

What is the originality here?
The originality lies in understanding that the absence of complete information to all arms is not sufficient to avoid bad arms’ equilibrium. In certain settings where side communication exists, arms may effectively collude even though the information is shared among them and not completely known by any of them. This justifies the necessity of developing mechanisms that encompass not only traditional sequential learning in the classical Multi-Armed Bandit (MAB) style but also integrate incentive mechanisms to effectively address the challenges highlighted in this paper.
Check all our research cards 👇
https://medium.com/criteo-engineering/research-cards/home




