- A. Kalanther, S. Bharvirkar, D. Bostwick, C. Maheshwari, S. Sastry
- Under review, 2026
- arXiv
- We study how to deploy pretrained level-K policies against an opponent whose reasoning level is unknown. In pursuit-evasion, an orchestrator trained with RL to choose among them earns higher return than classifiers that estimate the opponent's level and play the matching response.