Study finds tool-equipped language models can make unsupported claims even when checking is possible
A new preprint reports that one tested language model sometimes made unsupported final claims despite having access to evidence-resolving tools, while an automatic checking rule corrected the errors in a small synthetic evaluation.