![]() |
Robust and Communication-Efficient Multi-Agent Sequential Decision-Making (Xuchuang Wang et al.)
The rapid deployment of decentralized AI, edge computing, and autonomous systems has made multi-agent sequential decision-making under uncertainty a cornerstone of modern machine learning. Toward this end, multi-agent bandit frameworks and reinforcement learning methods have shown promising results in decentralized load balancing, recommender systems, and network routing. However, most existing methods share major limitations when deployed in real-world environments: they rely on synchronized communication, assume homogeneous agent behavior, or fail to withstand strongly adversarial manipulated feedback. As a result, maintaining optimal coordination and communication efficiency in practical, asynchronous settings remains a critical bottleneck. Furthermore, the theoretical understanding of how heterogeneous agents can effectively cooperate under complex feedback structures (such as fusing absolute rewards with relative dueling feedback) is still largely unexplored.
Building on our ongoing research in online learning theory and multi-agent sequential decision-making, this project aims to make a major step forward by developing a novel, robust multi-agent online learning framework. Specifically, we will design algorithms that optimize both decision regret and communication overhead, while providing theoretical guarantees against adversarial attacks and asynchronous delays.
Objectives:
Asynchronous Frameworks: Design and demonstrate a novel asynchronous multi-agent learning framework that optimally balances the trade-offs between fully distributed and leader-coordinated paradigms, addressing communication delays.
Theoretical Foundations: Formulate the lower and upper bounds for communication complexity and regret under non-stationary and adversarial environments, particularly for multi-fidelity and fused feedback structures.
System Validation: Validate the designed learning models and theoretical analyses by systematically conducting experiments on decentralized resource allocation and distributed routing using synthetic data and real-world network simulators.
Findings So Far:
Our research has recently addressed the critical "sync-gap" in decentralized systems. We developed a framework for "Asynchronous Multi-Agent Bandits" [1] that quantifies the regret penalty incurred by communication delays. Furthermore, we have pioneered methods for "Fusing Reward and Dueling Feedback" [2], allowing agents to learn simultaneously from direct scores and preference-based comparisons. This is vital for systems where absolute metrics are noisy, but relative rankings are clear. To protect these systems from external interference, we also introduced "Stochastic Bandits Robust to Adversarial Attacks" [3], which uses a robust estimator to maintain sublinear regret even when observations are maliciously corrupted.
Selected Publications:
[1] X. Wang, Y.-Z. J. Chen, L. Yang, X. Liu, M. Hajiesmaili, D. Towsley, and J. C.S. Lui, "Asynchronous Multi-Agent Bandits: Fully Distributed vs. Leader-Coordinated Algorithms," ACM International Conference on Measurement and Modeling of Computer Systems (SIGMETRICS), 2025.
[2] X. Wang, Q. Zeng, J. Zuo, X. Liu, M. Hajiesmaili, J. C.S. Lui, and A. Wierman, "Fusing Reward and Dueling Feedback in Stochastic Bandits," Forty-second International Conference on Machine Learning (ICML), 2025.
[3] X. Wang, M. Liu, J. Zuo, X. Liu, J. C.S. Lui, and M. Hajiesmaili, "Stochastic Bandits Robust to Adversarial Attacks," International Conference on Learning Representations (ICLR), 2025.
For further information on this research topic, please contact Prof. Xuchuang Wang.
![]() |