NC-GRPO paper reports stronger vision-language reasoning from latent-space rollout diversity
An arXiv preprint introduces NC-GRPO, a reinforcement-learning method that perturbs a vision-language model’s hidden representation rather than its input image. The authors report improved out-of-domain mathematical reasoning and hallucination robustness on Qwen2.5-VL-7B, while noting tradeoffs between…