Paper proposes a two-step test for choosing LLMs under uncertain evaluations
A new preprint argues that companies can sometimes certify the best assignment of large language models to recurring workloads even when model-quality estimates remain uncertain. It proposes a two-solve test and an evidence-gathering method called CASE.