The project aims to develop and evaluate a new method for training world models for multi-agent reinforcement learning (MARL). The primary objective is to improve the robustness of learned agents under test-time distribution shifts, with particular emphasis on observation noise and adversarial perturbations. The proposed approach will use learned world models to support robust policy learning and decision-making in multi-agent environments. GPU resources will be used for training world models and reinforcement-learning agents, as well as for systematic evaluation across different levels and types of test-time perturbations.
Main Supervisor: György Dán (Department of Network and Systems Engineering, KTH Royal Institute of Technology)