返回新闻
安全AI Understanding 简报

《卫报》报道 NHS 人工智能抄写员记录中存在患者安全错误

英国健康观察告诉《卫报》,英国医疗保健中使用的人工智能抄写员生成了不正确的药物名称、诊断和咨询摘要,而患者和医生可能无法检测到。

5 min readRead the original reporting
Source-provided image accompanying The Guardian reports patient-safety errors in NHS AI scribe records
归因报告来源记录
出版商
theguardian.com
来源链接
theguardian.comhttps://www.theguardian.com/society/2026/aug/31/doctors-ai-scribes-get-names-of-drugs-and-diagnoses-wrong-nhs-watchdog-warns
来源类型
新闻媒体的报道——不是第一方文件。

我们无法独立确认的内容: 此声明归因于指定的商店。我们没有根据第一方文件对其进行验证。 (theguardian.com)

背景60 秒内了解这一点

从这里开始

关键术语

数据集
用于训练、验证或测试的结构化或非结构化示例的集合。
测试一下自己什么是人工智能?测验

发生了什么

The Guardian reports that Healthwatch England has received multiple accounts of AI scribing tools producing inaccurate consultation transcripts and summaries. Examples include an incorrect diagnosis, a confused medication name and a missing instruction about obtaining a repeat prescription.

The Guardian reported on 31 August 2026 that Healthwatch England had warned that AI systems used to listen to and transcribe medical consultations can misidentify drug names and illnesses. The report describes AI scribes as a rapidly expanding technology in England, where doctors and hospitals are already using 27 different systems. The government’s 10-year NHS health plan presents the tools as a way to reduce administrative work and give clinicians more time with patients.

One case involved a woman whose AI-generated summary incorrectly stated that she had demyelination, a serious condition that can be associated with multiple sclerosis. According to The Guardian, the woman was an NHS health professional and noticed the problem when she checked the tool’s account of her MRI result. The hospital later corrected the record to “null demyelination.” The woman told the newspaper that receiving an incorrect diagnosis and then being told it was a typo was traumatising.

The Guardian also reported that patients identified other errors: an AI scribe confused a prescribed drug with another medication with a similar name, and an automatically generated letter omitted a consultant’s instruction that a patient should seek a repeat migraine prescription from their GP. Healthwatch said it had heard multiple stories in which patients noticed mistakes that health professionals had not. The supplied source is a reported secondary account; these individual cases, the number of affected patients and the precise error rates of the tools were not independently confirmed here through a primary regulator report or NHS .

来源详情: theguardian.com ↗

为什么这很重要

Errors in records used for patient care could affect treatment, medication access and trust in NHS automation. The Guardian reports that England has no nationwide oversight framework for AI scribes after the MHRA decided not to classify them as medical devices.

The practical risk is that an incorrect transcript can become part of a patient’s medical record and influence later care. A wrong diagnosis may cause distress or trigger inappropriate follow-up, while a confused medication name could contribute to a prescribing or treatment error. A missing instruction can also create an access problem if a patient cannot obtain medication when expected. The Guardian’s examples show that the harm does not depend on an AI system making a dramatic clinical decision; a small transcription or summarisation error can matter when it enters a clinical workflow.

Healthwatch England told The Guardian that mistakes may persist when patients do not catch them. That places part of the quality-control burden on people who may not know what was said, may be unwell or may not have access to the generated record. It also raises a workflow question for clinicians: if doctors must check every transcript carefully, the promised time savings may be reduced. Dr Shier Ziser Dawood, a London GP cited by The Guardian, previously described an AI scribe recording that she had told a patient to continue Prozac even though she had not prescribed or discussed it.

The report also highlights uneven performance across consultations. Dr Charlotte Blease, an AI-in-healthcare researcher at Uppsala University, told The Guardian that errors are more likely when several people are present, when a patient has a complex medical history or when English is not the patient’s first language. More than half of 1,003 UK GPs in Blease’s survey reportedly considered their ambient-AI records more accurate than records they produced themselves, but she also said the error rate could be worse with AI. The survey result does not establish the clinical safety of any particular tool, and the article does not provide comparative error measurements or independent validation.

Interactive Mechanism

互动机制:它实际上是如何运作的

以交互方式探索这一发展背后的基础技术。

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
交互式概念检查+10 Points
What is AI? Quiz

A route planner searches possible journeys using explicit rules. What does this illustrate about AI?

接下来看什么

Watch for clearer NHS procedures for reviewing, correcting and reporting AI-generated records; evidence on error rates across different tools and patient groups; and whether regulators reconsider the oversight status of AI scribes.

A central issue is whether patients will have a simple, visible way to review AI-generated notes, report errors and obtain corrections. Healthwatch told The Guardian that there is an urgent need for clarity about how mistakes made by scribing tools or by the professionals using them should be reported and corrected. Useful safeguards would need to define who is responsible for checking the record, how quickly corrections must be made and how amendments are communicated to later clinicians.

The Guardian reported that Healthwatch was concerned by the Medicines and Healthcare products Regulatory Agency’s decision not to classify AI scribes as medical devices. According to the report, that decision means there will be no England-wide oversight specifically intended to ensure that the systems are safe and effective. The supplied article does not include the MHRA’s reasoning, the NHS’s response to Healthwatch’s warning or details of any national testing regime, so the regulatory consequences remain important unknowns.

Further reporting should establish how the 27 systems differ, whether they are used for recording only or also for drafting clinical documents, and how frequently clinicians review their output before it reaches a formal record. It should also examine performance for accents, multilingual consultations, complex histories and multi-person appointments, following earlier complaints in Rotherham about an AI receptionist struggling with strong Yorkshire accents. The Guardian’s report does not show how widespread those problems are or whether any tool has been withdrawn, so conclusions about the technology as a whole should remain limited.

相关指南和测验

什么是人工智能?AI 伦理人工智能模型解释人工智能培训测试你所知道的——尝试免费的人工智能测验在我们的词汇表中查找人工智能术语关注AI监管追踪器
觉得这有用吗?