MC-CXR benchmark finds misleading context can overturn vision-language model chest X-ray judgments
A new benchmark reports that vision-language models often changed a correct chest X-ray answer after receiving conflicting text or prior-image context, with misleading text exerting a stronger pull than misleading visual context.