Distributed Multi Robot Lunar Cargo Transportation via Phase Decomposed Reinforcement Learning

2026-06-30Robotics

Robotics
AI summary

The authors developed a way for multiple robots to work together to carry cargo on the moon. They broke the task into three parts—lifting, moving, and placing the cargo—and trained each part separately to handle the challenges of teamwork and safety. Their method uses smart decision-making models trained together but run independently on each robot, with a system to keep the robots coordinated and stop if something goes wrong. They tested their approach both in computer simulations and real experiments at a Japanese space facility, showing that their system reliably handled all parts of the task.

modular robotsreinforcement learningcooperative transportmulti-agent systemspayload couplingphase decompositionMarkov state representationcentralized trainingproprioceptionsynchronization mechanism
Authors
Ashutosh Mishra, Elian Neppel, Shreya Santra, Antoine Jonquières, Muhammad Athallah Naufal, Kentaro Uno, Kazuya Yoshida
Abstract
Modular reconfigurable robotic systems provide a scalable solution for cooperative surface operations in future lunar missions. However, cooperative cargo transportation remains challenging due to morphology-dependent topology changes, strong payload-induced coupling, long-horizon decision making, and safety constraints. This paper proposes a phase-decomposed reinforcement learning framework for cooperative cargo transport with distributed robotic units. The task is decomposed into lifting, transportation, and placement, each optimized with a dedicated joint-state policy capturing inter-agent coupling. Centralized training promotes stable convergence, while deployment uses onboard proprioception for control and OptiTrack motion capture for ground-truth evaluation and post-processed metrics. A deterministic phase controller expressed in Markov state representation regulates transitions between stages, and a failure-sensitive synchronization mechanism ensures coordinated progression and safety-aware halting during real-world execution. The framework is evaluated in simulation and through controlled field experiments at a JAXA space exploration test facility. Results demonstrate reliable cooperative transport across all stages in both simulation and hardware experiments.