MTDiag tests whether medical LLMs can reason across multi-turn diagnostic conversations
A new dataset evaluates large language models as diagnostic agents in interactive clinical encounters rather than through isolated questions. MTDiag combines cases from emergency medicine, clinical records and published case reports, and proposes knowledge-grounded metrics beyond diagnostic accuracy.