RecurSE proposes bounded self-evaluation for LLM judges
A new arXiv paper proposes training language models to evaluate their own judgments without external gold labels, while using validation checks to identify when recursive improvement should stop.