LURE paper proposes pursuit-evasion self-play to train LLM reasoning without task data
A new arXiv paper presents LURE, a zero-data self-play method in which one language model sets task difficulty while another solves verifiable reasoning challenges. The authors report stronger out-of-distribution zero-shot accuracy than trained baselines across nine held-out benchmarks, but the abstract does not…