Preprint proposes continuous difficulty estimates to reduce LLM self-consistency costs
A new arXiv preprint proposes Flexible Self-Consistency, a method that estimates how uncertain a language model is about a question and adjusts the number of sampled reasoning paths. The authors report token savings of up to 76% while maintaining accuracy comparable to standard self-consistency across various…