返回新聞
安全性AI Understanding 簡報

《衛報》報道 NHS 人工智慧抄寫員記錄中存在病患安全錯誤

英國健康觀察告訴《衛報》,英國醫療保健中使用的人工智慧抄寫員產生了不正確的藥物名稱、診斷和諮詢摘要,而患者和醫生可能無法檢測到。

5 min readRead the original reporting
Source-provided image accompanying The Guardian reports patient-safety errors in NHS AI scribe records
歸因報告來源記錄
出版商
theguardian.com
來源連結
theguardian.comhttps://www.theguardian.com/society/2026/aug/31/doctors-ai-scribes-get-names-of-drugs-and-diagnoses-wrong-nhs-watchdog-warns
來源類型
新聞媒體的報道-不是第一方文件。

我們無法獨立確認的內容: 此聲明歸因於指定的商店。我們沒有根據第一方文件對其進行驗證。 (theguardian.com)

背景60 秒內了解這一點

從這裡開始

關鍵術語

數據集
用於訓練、驗證或測試的結構化或非結構化範例的集合。
測試一下自己什麼是人工智慧?測驗

發生了什麼事

The Guardian reports that Healthwatch England has received multiple accounts of AI scribing tools producing inaccurate consultation transcripts and summaries. Examples include an incorrect diagnosis, a confused medication name and a missing instruction about obtaining a repeat prescription.

The Guardian reported on 31 August 2026 that Healthwatch England had warned that AI systems used to listen to and transcribe medical consultations can misidentify drug names and illnesses. The report describes AI scribes as a rapidly expanding technology in England, where doctors and hospitals are already using 27 different systems. The government’s 10-year NHS health plan presents the tools as a way to reduce administrative work and give clinicians more time with patients.

One case involved a woman whose AI-generated summary incorrectly stated that she had demyelination, a serious condition that can be associated with multiple sclerosis. According to The Guardian, the woman was an NHS health professional and noticed the problem when she checked the tool’s account of her MRI result. The hospital later corrected the record to “null demyelination.” The woman told the newspaper that receiving an incorrect diagnosis and then being told it was a typo was traumatising.

The Guardian also reported that patients identified other errors: an AI scribe confused a prescribed drug with another medication with a similar name, and an automatically generated letter omitted a consultant’s instruction that a patient should seek a repeat migraine prescription from their GP. Healthwatch said it had heard multiple stories in which patients noticed mistakes that health professionals had not. The supplied source is a reported secondary account; these individual cases, the number of affected patients and the precise error rates of the tools were not independently confirmed here through a primary regulator report or NHS .

來源詳情: theguardian.com ↗

為什麼這很重要

Errors in records used for patient care could affect treatment, medication access and trust in NHS automation. The Guardian reports that England has no nationwide oversight framework for AI scribes after the MHRA decided not to classify them as medical devices.

The practical risk is that an incorrect transcript can become part of a patient’s medical record and influence later care. A wrong diagnosis may cause distress or trigger inappropriate follow-up, while a confused medication name could contribute to a prescribing or treatment error. A missing instruction can also create an access problem if a patient cannot obtain medication when expected. The Guardian’s examples show that the harm does not depend on an AI system making a dramatic clinical decision; a small transcription or summarisation error can matter when it enters a clinical workflow.

Healthwatch England told The Guardian that mistakes may persist when patients do not catch them. That places part of the quality-control burden on people who may not know what was said, may be unwell or may not have access to the generated record. It also raises a workflow question for clinicians: if doctors must check every transcript carefully, the promised time savings may be reduced. Dr Shier Ziser Dawood, a London GP cited by The Guardian, previously described an AI scribe recording that she had told a patient to continue Prozac even though she had not prescribed or discussed it.

The report also highlights uneven performance across consultations. Dr Charlotte Blease, an AI-in-healthcare researcher at Uppsala University, told The Guardian that errors are more likely when several people are present, when a patient has a complex medical history or when English is not the patient’s first language. More than half of 1,003 UK GPs in Blease’s survey reportedly considered their ambient-AI records more accurate than records they produced themselves, but she also said the error rate could be worse with AI. The survey result does not establish the clinical safety of any particular tool, and the article does not provide comparative error measurements or independent validation.

Interactive Mechanism

互動機制:它實際上是如何運作的

以互動方式探索這項發展背後的基礎技術。

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
互動式概念檢查+10 Points
What is AI? Quiz

A route planner searches possible journeys using explicit rules. What does this illustrate about AI?

接下來看什麼

Watch for clearer NHS procedures for reviewing, correcting and reporting AI-generated records; evidence on error rates across different tools and patient groups; and whether regulators reconsider the oversight status of AI scribes.

A central issue is whether patients will have a simple, visible way to review AI-generated notes, report errors and obtain corrections. Healthwatch told The Guardian that there is an urgent need for clarity about how mistakes made by scribing tools or by the professionals using them should be reported and corrected. Useful safeguards would need to define who is responsible for checking the record, how quickly corrections must be made and how amendments are communicated to later clinicians.

The Guardian reported that Healthwatch was concerned by the Medicines and Healthcare products Regulatory Agency’s decision not to classify AI scribes as medical devices. According to the report, that decision means there will be no England-wide oversight specifically intended to ensure that the systems are safe and effective. The supplied article does not include the MHRA’s reasoning, the NHS’s response to Healthwatch’s warning or details of any national testing regime, so the regulatory consequences remain important unknowns.

Further reporting should establish how the 27 systems differ, whether they are used for recording only or also for drafting clinical documents, and how frequently clinicians review their output before it reaches a formal record. It should also examine performance for accents, multilingual consultations, complex histories and multi-person appointments, following earlier complaints in Rotherham about an AI receptionist struggling with strong Yorkshire accents. The Guardian’s report does not show how widespread those problems are or whether any tool has been withdrawn, so conclusions about the technology as a whole should remain limited.

相關指引和測驗

什麼是人工智慧?AI 倫理人工智慧模型解釋人工智慧培訓測試你所知道的—嘗試免費的人工智慧測驗在我們的詞彙表中尋找人工智慧術語關注AI監管追蹤器
覺得有用嗎?