Pada si Iroyin
AtunseAI Understanding finifini

Iwadi ti ifojusọna rii pe deede LLM ṣubu ni didasilẹ lori awọn ibeere iforukọsilẹ ile-iwosan aibikita

Iwadii ti ifojusọna ti ọpọlọpọ-ojula ṣe ijabọ pe awoṣe ede nla kan baamu 87% ti awọn idahun ifọkanbalẹ eniyan nigba yiyọ alaye jade lati awọn igbasilẹ iṣoogun ti ko ṣiṣẹ, ṣugbọn deede ṣubu si 62% lori awọn ibeere ti o nilo akoko iṣẹlẹ ati imọran ile-iwosan ti o tobi julọ.

5 min readRead the primary source
Source-page capture accompanying Prospective study finds LLM accuracy falls sharply on ambiguous clinical registry questions
Iwe aṣẹ orisun akọkọOrisun ti o gbasilẹ
Olutẹwe
arxiv.org
Orisun ọna asopọ
arxiv.orghttps://arxiv.org/abs/2608.20373
Orisun iru
Iwe akọkọ - ikede osise, iwe, iforukọsilẹ, tabi oju-iwe ẹgbẹ akọkọ ti a ka taara.
AtokọLoye eyi ni iṣẹju 60

Bẹrẹ nibi

Awọn ofin bọtini

Awoṣe Ede nla (LLM)
Awoṣe ede ti a ṣe ikẹkọ lori titobi ọrọ corpora lati ṣe ipilẹṣẹ ati itupalẹ ọrọ.
Otitọ ilẹ
Awọn aami itọkasi igbẹkẹle ti a lo lati ṣe ikẹkọ tabi ṣe iṣiro awọn abajade awoṣe.
Ṣe idanwo fun ara rẹAwọn awoṣe AI ti ṣalaye adanwo

Kini o ṣẹlẹ

Researchers evaluated a large language model on clinical registry abstraction using unprocessed electronic medical record data from two American College of Cardiology National Cardiovascular Data Registry settings. The study organized registry questions into six categories based on ambiguity and the clinical reasoning needed to answer them.

The paper, revised on arXiv on Aug. 25, reports a pilot and a validation study involving clinical registry questions from the American College of Cardiology’s National Cardiovascular Data Registry. In the pilot at an academic medical center, the model identified candidate data sources for each registry question. Experienced abstractors then used those results to define question-specific document sets. In the validation study at a second center, using a second ACC NCDR registry, the model answered questions from those sets. This design focused the evaluation on extracting and interpreting information from existing medical-record material rather than on a simplified, preselected data table.

Before reviewing model output, two abstractors independently established the and assigned each question to one of six categories: Medication/Event Flag, Binary Clinical Presence, Administrative, Quantitative Laboratory/Physiologic, Clinical Interpretation, and Event Timing. The categories were ordered by the ambiguity and clinical reasoning required to resolve each question. The paper reports 9,430 abstractor answers reconciled into 4,715 consensus answers, including 501 pilot answers and 4,214 validation answers. In the pilot, the average number of candidate data sources ranged from 14.6 for demographics to 89.2 for history and risk factors, with substantial variation reflected in the reported standard deviations.

In validation, human inter-rater agreement was approximately 98%, according to the study. The LLM’s answers exactly matched consensus in 87% of cases, were classified as partial matches in 2%, and did not match in 9%. The paper reports a mean question-level accuracy of 91.5%, with a standard deviation of 13.4%, across 157 questions that each had at least 20 answers. Accuracy varied by ambiguity category, declining from 96% for Medication/Event Flag questions to 62% for Event Timing questions. The source identifies the study as a preprint and does not name the model evaluated in the abstract.

Awọn alaye orisun: arxiv.org ↗

Kini idi ti o ṣe pataki

The results suggest that average accuracy can conceal important weaknesses in healthcare AI. The model performed best on relatively direct medication and event-flag questions and worst when it had to interpret clinical context or determine when an event occurred.

The central finding is not simply that an LLM made errors. It is that performance declined in a structured way as the questions became more ambiguous and demanded more clinical reasoning. A single overall accuracy figure of 91.5% could therefore give an incomplete picture of operational risk. A registry workflow containing many straightforward fields may appear reliable while still performing poorly on a smaller set of questions that require reconstructing context, interpreting clinical evidence, or placing an event in time.

Clinical registries support research, quality measurement, benchmarking, and other forms of healthcare reporting. Errors in abstraction can affect the quality of those downstream records even when the system is not making a diagnosis or treatment recommendation. The reported gap between approximately 98% human agreement and the model’s lower exact-match rate indicates that human review remains important for this task, particularly where the answer depends on multiple documents or on interpreting the sequence of events.

The study also offers a practical way to evaluate healthcare language models: divide questions according to the type of ambiguity they contain, rather than treating all extraction tasks as equivalent. That approach could help organizations identify which fields are suitable for assistance and which require closer review. The source does not show that the model caused patient harm, improved registry operations, reduced costs, or was deployed in routine care. Those outcomes remain unestablished.

Interactive Mechanism

Ibaraẹnisọrọ Mechanism: Bii O Ṣe Nṣiṣẹ Lootọ

Ṣawari imọ-ẹrọ abẹlẹ lẹhin idagbasoke yii ni ibaraenisọrọ.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Ibanisọrọ Erongba Ṣayẹwo+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

Kini lati wo tókàn

The study does not establish whether the findings generalize to other models, hospitals, medical specialties, registries, or clinical decision-making. Further work should test broader settings, compare model versions, examine the consequences of partial errors, and assess whether human review can reliably catch the hardest mistakes.

A key unknown is how broadly the results apply. The source describes a pilot at one academic medical center and validation at a second center using another ACC NCDR registry, but it does not provide evidence in the abstract about community hospitals, other specialties, different record systems, or other registry designs. The evaluated model is also not identified in the abstract. Comparisons across models and model versions would be needed before drawing conclusions about LLM performance generally.

Future evaluations should report more than aggregate exact-match accuracy. Important questions include whether partial answers are clinically usable, which types of mistakes are most common, how often errors involve dates or contradictory records, and whether reviewers can detect model mistakes consistently. The study’s category-level results make Event Timing a particularly important area for follow-up, but the source does not specify the individual questions, error examples, or severity of the incorrect answers.

It is also unknown whether the model’s candidate-source selection in the pilot affected the later validation workflow or whether similar performance would occur when document sets are not curated by experienced abstractors. Further studies could test fully unprocessed records, different levels of human assistance, and prospective monitoring in real registry operations. Until such evidence is available, the findings support targeted human oversight for ambiguous abstraction tasks rather than a conclusion that LLMs are ready to replace clinical abstractors. The source therefore leaves open how the same approach would perform when document selection is not guided by experienced abstractors, when records contain different levels of structure, or when reviewers enter the process at different points.

Awọn itọsọna ti o jọmọ & awọn ibeere

Awọn awoṣe AI ti ṣalayeÌlànà Ìwà AIAI IkẹkọChatGPT & LLMsṢe idanwo ohun ti o mọ — gbiyanju idanwo AI ọfẹ kanWa ọrọ AI kan ninu iwe-itumọ waTẹle olutọpa idasilẹ awoṣe AI
Ṣe eyi wulo?