在本页4 分钟阅读
概述
They use many electronic health record variables and update continuously. They build on traditional scores such as MEWS and NEWS2, which add points for abnormal vital signs. They matter because deterioration is often preceded by hours of subtle changes that busy ward teams miss. Their real value, though, depends on local validation and on the response workflow behind each alert.
深入探讨
Traditional early warning scores add points for abnormal vital signs. The Modified Early Warning Score (MEWS) spread in the early 2000s. The UK's Royal College of Physicians published the National Early Warning Score (NEWS) in 2012 and an update, NEWS2, in 2017. NEWS2 scores respiratory rate, oxygen saturation, use of supplemental oxygen, systolic blood pressure, pulse, level of consciousness including new confusion, and temperature. These scores are transparent, can be calculated at the bedside and use the same thresholds everywhere. But they rely on a few variables and fixed cutoffs. AI deterioration indexes add labs, nursing assessments, diagnoses, medications and trends. Examples include the Epic Deterioration Index built into Epic's EHR, eCART developed at the University of Chicago, and the Rothman Index, which relies heavily on nursing assessments. Most are designed to trigger a rapid response evaluation above a chosen threshold. The strongest evidence ties the model to a workflow. Kaiser Permanente Northern California deployed its Advance Alert Monitor in stages across its hospitals. A 2020 study in the New England Journal of Medicine linked it to lower mortality. A defining feature was that trained remote nurses screened every alert before it reached the bedside team. Validation is the recurring weak point. When University of Michigan researchers tested Epic's sepsis model independently in 2021, it performed substantially worse than the developer had reported. That is a warning for any proprietary score. Models can also learn from clinicians' own worry. Orders for lactate or blood cultures signal that someone already suspects trouble, so the model may flag deterioration that staff have already recognized. A common misconception is that a higher AUROC means better care. What matters at the bedside is how often an alert is right at the chosen threshold, how many alerts each unit receives per shift, how much lead time an alert gives, and whether the response actually changes treatment.
战略影响
背景与规则
行业背景决定了人工智能创意能否与现实接触。
质量控制
领域约束会影响可接受的错误率和监督模型。
构建选择
成功的部署使技术能力与一线工作流程保持一致。
The Future of AI Early Warning Scores for Patient Deterioration
Continuous monitoring with wearable sensors on general wards could give models denser vital sign data, though it also raises questions about false alarms. US federal health IT rules have begun requiring more transparency about predictive decision support in certified EHRs, and many health systems now expect local validation before deployment. The field still needs more prospective and randomized studies that measure mortality, ICU use and staff workload together. Hospitals will likely keep treating these scores as prompts for clinical review, not as automatic triggers for treatment.
现实世界的实施
A ward nurse sees a patient's deterioration score climb over six hours even though the NEWS2 score is only 3. The rise is driven by increasing respiratory rate, falling oxygen saturation and new abnormal labs, so she calls the rapid response team.
Kaiser Permanente's Advance Alert Monitor sends high-risk predictions to remote nurses. They review the chart before contacting the floor team, which filters alerts before they reach bedside staff.
A hospital runs a vendor deterioration model against a year of its own data. At the planned alert threshold it adds little over NEWS2, so the hospital changes the threshold before going live.
A step-down unit uses a deterioration model to decide which patients recently moved out of the ICU get more frequent overnight vital sign checks.
风险与防护栏
监管要求可能会使原本强大的原型失效。
历史数据可能会编码损害特定社区的偏见。
遗留系统可能会造成集成瓶颈和隐性成本。
实施路线图
让领域专家参与从问题框架到评估的整个过程。
在启动前设计审计跟踪和文档。
尽早验证合规性和安全义务。
分阶段推出,并具有明确的停止和回滚标准。
不断探索
Free newsletter
Get the daily AI briefing
Three verified AI stories every weekday morning, written in plain English. Free forever, no ads.
One email each weekday. Unsubscribe in one click. We never sell or share your address.
Test yourself
Take the AI Early Warning Scores for Patient Deterioration quiz
Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.
Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation
常见问题
What is AI Early Warning Scores for Patient Deterioration?
AI early warning scores estimate a hospitalized patient's risk of serious deterioration, such as ICU transfer, cardiac arrest or death, over the next several hours. They use many electronic health record variables and update continuously. They build on traditional scores such as MEWS and NEWS2, which add points for abnormal vital signs. They matter because deterioration is often preceded by hours of subtle changes that busy ward teams miss. Their real value, though, depends on local validation and on the response workflow behind each alert.
Which of these is one of the parameters scored in NEWS2?
NEWS2 scores respiratory rate, oxygen saturation, supplemental oxygen use, systolic blood pressure, pulse, consciousness and temperature. Lab values are not included.
What was a defining feature of Kaiser Permanente's Advance Alert Monitor deployment?
Remote nurse review filtered alerts and tied predictions to a clear response, which the guide calls central to its evidence.
Why does the guide warn that a high AUROC does not guarantee better care?
AUROC summarizes ranking across all thresholds, but bedside value depends on performance at one threshold and on what staff do with the alert.
How can orders for lactate or blood cultures cause a deterioration model to overstate its usefulness?
Features that reflect clinician suspicion let the model echo what staff already know rather than give early warning.
Why does computing AUROC over every observation inflate a deterioration model's performance?
Many repeated easy negatives from stable patients push the metric up. Encounter-level or first-alert analyses are more honest.
继续学习
相关指南
为此主题精选的更多指南