DirEAG paper proposes a better way to calibrate AI confidence in math answers
A new arXiv paper introduces DirEAG, a method that combines confidence reports from multiple prompts into calibrated evidence about possible answers, including the possibility that none is correct.