Paper Argues Evolution Strategies Beat RL at Keeping LLM Answer Sets Diverse
A new arXiv preprint argues that post-training LLMs with evolution strategies — a population-based, gradient-free method that perturbs weights directly — beats reinforcement learning on pass@k and solution coverage. The abstract cites better math-benchmark results but names no models, benchmarks, or numbers.