ArXiv paper proposes milestone-based training for long-horizon LLM agents
MileGPO uses milestone discovery and local evidence to improve credit assignment when training language-model agents on long, multi-step tasks. The authors report state-of-the-art results on ALFWorld and WebShop, but the claims remain limited to the paper’s experiments.