- A. Kalanther, S. Bharvirkar, S. Sastry, C. Maheshwari
- IEEE Conference on Decision and Control (CDC), 2026 [Accepted]
- arXiv
- We propose NePPO, a multi-agent RL algorithm for approximating Nash equilibria in general-sum games. It learns a player-independent potential function by minimizing a novel objective with zeroth-order gradient descent.