Prospective study finds LLM accuracy falls sharply on ambiguous clinical registry questions
A multi-site prospective study reports that a large language model matched 87% of human-consensus answers when extracting information from unprocessed medical records, but accuracy fell to 62% on questions requiring event timing and greater clinical reasoning.