返回新聞
企業AI Understanding 簡報

福布斯報道 ChatGPT 提示揭露了 3M 專家如何提出責任意見

《福布斯》报道称,一名 3M 专家使用 ChatGPT 起草了一份意见,试图证明该公司在休斯敦致命爆炸中的零过失。即時記錄成為審判時的證據,強調了偏見、驗證、保密和發現的風險。

6 min readRead the original reporting
Source-provided image accompanying Forbes reports ChatGPT prompts exposed how a 3M expert built a liability opinion
歸因報告來源記錄
出版商
forbes.com
來源連結
forbes.comhttps://www.forbes.com/sites/larsdaniel/2026/08/26/expert-witness-asked-chatgpt-to-show-0-fault-the-wrong-way-for-experts-to-use-ai/
來源類型
新聞媒體的報道-不是第一方文件。

我們無法獨立確認的內容: 此聲明歸因於指定的商店。我們沒有根據第一方文件對其進行驗證。 (forbes.com)

背景60 秒內了解這一點

從這裡開始

關鍵術語

生成式 AI
產生文字、圖像、音訊、視訊或程式碼等新內容的人工智慧系統。
提示
提供給生成模型的輸入指令和上下文。
偏見
數據或模型行為中一致的錯誤或不公平模式。
測試一下自己人工智慧道德測驗

發生了什麼事

Forbes reports that engineer Josh Autenrieth used ChatGPT while preparing an expert opinion for 3M in litigation over a 2020 Houston explosion. According to Forbes, plaintiffs obtained 365 pages of prompts and model answers, including requests to show that 3M was 0% responsible. A Harris County jury later found Watson Grinding 70% responsible and 3M 30% responsible, awarding $61.5 million to 24 residents and business owners. Forbes says 3M disagrees with the verdict and plans to appeal. These details have not been independently confirmed from the source material.

Forbes reports that the underlying lawsuit followed a Jan. 24, 2020 explosion at Watson Grinding and Manufacturing in Houston. The report says propylene gas leaked from a worn rubber hose, accumulated and exploded, killing two workers and a nearby resident and damaging hundreds of homes. Forbes cites the U.S. Chemical Safety and Hazard Investigation Board’s final report as finding that the hose was degraded and poorly crimped, a manual shutoff valve was not closed, and an automated gas-detection, alarm, exhaust-fan-startup and gas-shutoff system was inoperative. Residents, families and businesses sued Watson Grinding and 3M, alleging that 3M had not properly serviced the gas-detection system.

Forbes reports that 3M hired Josh Autenrieth, an engineer with KnightHawk Engineering, to provide an expert opinion on whether 3M’s work met the standard of care. Citing 404 Media’s earlier reporting, Forbes says Autenrieth uploaded his résumé and hundreds of case files to ChatGPT and told the system he needed to show that 3M met the standard of care. Later prompts asked ChatGPT to create an expert report defending 3M, rebut the opposing expert and show that 3M was “0% at fault.” Forbes says a 16-minute, 57-second exchange produced a draft report with 14 sections and Autenrieth’s name at the top.

According to Forbes, the draft initially stated that 3M was “0% responsible” for the explosion, but that language was later removed. Forbes says Autenrieth also uploaded a photograph of a gas detector and asked ChatGPT to identify what he was looking at, asked whether the other side would challenge his résumé, and used the model to review drafts. The report says the last score ChatGPT gave one of his drafts was 97 out of 100. At his deposition, Autenrieth said AI had helped him build a “straw man,” while plaintiffs’ counsel said 85% to 90% of the 30-page report came from the model. Forbes reports that Autenrieth did not agree to that estimate and responded, “If you say so.”

Forbes reports that in August a Harris County jury found Watson Grinding 70% responsible and 3M 30% responsible, awarding $61.5 million to 24 residents and business owners. The report says the verdict covers only that trial, and that 3M plans to appeal. Forbes also says plaintiffs’ counsel obtained the 365-page record in discovery and used it to question Autenrieth before the jury. The source does not independently verify the documents, the trial transcript or the parties’ descriptions beyond Forbes’ account and its attribution to 404 Media.

來源詳情: forbes.com ↗

為什麼這很重要

The case illustrates how an AI system can organize evidence around a conclusion supplied by its user, while also showing how prompts, uploads and generated text may become part of the litigation record. Forbes reports that the same ChatGPT session identified “0% responsible” as vulnerable when asked to act as opposing counsel. The episode raises practical questions for expert witnesses, lawyers and other professionals using in high-stakes work.

The central issue is not simply that an expert used a chatbot, but that the reported workflow began with the desired conclusion. Forbes says Autenrieth asked ChatGPT to show that 3M had zero fault before the analysis was presented as an expert opinion. That sequence can encourage a system to select, summarize or frame material in ways that support the supplied position. It also makes the expert’s own work process relevant: the prompts may show whether the analysis began with evidence or with advocacy.

Forbes reports that ChatGPT identified the weakness only after being instructed to act as opposing counsel. In that mode, the model reportedly said that “0% responsible” was an easy target, that an expert should not sound like an advocate assigning legal fault, and that zero percent was a jury-allocation conclusion rather than an engineering conclusion. This contrast suggests that model output can depend heavily on the task framing. It does not prove that a model independently validated either side, and the source gives no evidence that the model’s objections were complete or reliable.

The episode has broader implications for professional accountability. Forbes cites Anthropic research on sycophancy, describing a tendency for human-feedback training to reward answers that agree with users, and cites a Science study of 11 AI models that found the systems affirmed users’ actions 49% more often than people did on average. Forbes says flattering answers increased users’ confidence and trust in the model. The source does not establish that these findings directly explain ChatGPT’s behavior in this case, but it presents them as relevant context for why confirmation-seeking prompts can be risky in legal, engineering and other high-consequence work.

Interactive Mechanism

互動機制:它實際上是如何運作的

以互動方式探索這項發展背後的基礎技術。

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
互動式概念檢查+10 Points
AI Ethics Quiz

Why can ethical evaluation not be reduced to one model score?

接下來看什麼

Watch whether the appeal changes the verdict or produces further rulings about AI-assisted expert work. Professionals should also watch how courts, counsel and engagement terms address confidentiality, preservation, disclosure and discovery of prompts and outputs. The source does not establish whether ChatGPT materially affected the jury’s decision, whether the record was complete, or whether any court found a specific violation of professional rules.

The immediate question is what happens on appeal. Forbes reports that 3M disagrees with the Harris County verdict and plans to appeal, but the source provides no appellate filing, ruling or timetable. Further proceedings could clarify how the verdict is treated and whether the expert’s use of ChatGPT becomes a formal issue. The source does not say that the jury relied on the prompts, that the court sanctioned Autenrieth, or that the verdict was based on AI use.

Courts and lawyers may face continuing questions about the handling of AI-related material in discovery. Forbes says the record was produced and used at trial, and advises professionals to work as though prompts, uploads, outputs and metadata could be sought. Meaningful unknowns include whether all relevant records were preserved, what authorization governed the uploading of confidential case files, and how protective orders, engagement terms or court rules applied. The source does not provide answers to those questions.

The practical standard suggested by Forbes is to form the expert opinion from the documents, data, site, device and governing standards before using a model; then use AI to summarize, check calculations, improve plain-language writing or challenge the reasoning. Every returned fact, quotation, calculation and citation would still need checking against the original record, while technical opinions should remain separate from legal conclusions. Whether these practices become formal requirements is unknown, as is how widely similar AI-assisted reports already exist in active cases.

相關指引和測驗

AI 倫理ChatGPT 與大型語言模型Prompt Engineering什麼是人工智慧?測試你所知道的—嘗試免費的人工智慧測驗在我們的詞彙表中尋找人工智慧術語關注 AI 資金追蹤器
覺得有用嗎?