語言人工智慧指南

命名實體識別

Named entity recognition, or NER, identifies spans of text that refer to categories such as people, organizations, and places.

閱讀時間約2分鐘最後更新

概述

It finds mentions under a chosen schema. Linking a mention to a particular real-world record is a separate entity-linking task.

重點摘要

  • Define types and span boundaries.
  • Preserve offsets into the original text.
  • Keep recognition separate from identity linking.

深入探討

Define the entity types and span rules before training or evaluation. Should a company suffix be included? Is a product an organization, a separate type, or outside the schema? Inconsistent annotation rules can make a dataset internally contradictory. NER systems may assign token-level labels and combine adjacent tokens into spans. Subword tokenization requires care when aligning labels with the original text. Preserve character offsets so applications can show exactly which passage produced an extracted value. Evaluate both boundaries and types. Identifying only “Northstar” when the annotated organization is “Northstar Research Labs” may count as a span error even if the general category is correct. Report the matching convention with precision and recall so scores can be interpreted. Context can change the label. “Jordan” might identify a person, country, or organization in different passages. A recognized name is not verified identity information. When using extraction for redaction, search, or record matching, test the downstream outcome and handle ambiguous or missed mentions explicitly.

技術洞察

NER and redaction are not equivalent. A system that misses a private name or identifier can leave sensitive information visible even when its average recognition score is high.

Recognize a mention without inventing an identity

  1. Use the invented sentence “Jordan joined Northstar Research Labs in June.”
  2. Mark Jordan as a person mention and Northstar Research Labs as an organization mention under a documented schema.
  3. Do not attach a particular biography or company registration unless a separate linking step has evidence for that match.

The constructed example separates locating a name from resolving who or what it identifies.

戰略影響

速度與規模

語言工作流程可以在不犧牲一致性的情況下更快地移動。

交通與覆蓋範圍

它擴展了跨語言和溝通方式的訪問。

更明確的決策

團隊可以花更多時間進行判斷,而自動化則可以處理重複。

現實世界的實施

Highlight organizations mentioned in a news article with original text offsets.

Build a review queue for possible names before approving a redacted document.

風險與防護欄

幻覺的事實可以悄悄地進入報告、支持流程或研究成果。

及時的敏感性可能會在類似的請求中產生不一致的結果。

如果存取控制薄弱,敏感文字資料可能會暴露。

實施路線圖

1

在推出之前定義輸出格式、語氣和品質標準。

2

當準確性很重要時,請使用可信任來源進行地面回應。

3

為高風險輸出保留人工審查檢查點。

4

追蹤故障模式並定期重新訓練提示或工作流程。

資料來源與延伸閱讀

不斷探索

Free newsletter

Get the daily AI briefing

Three verified AI stories every weekday morning, written in plain English. Free forever, no ads.

One email each weekday. Unsubscribe in one click. We never sell or share your address.

Test yourself

Take the Named Entity Recognition quiz

Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.

開始測驗

Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation

下一步指南

實體連結和消歧

常見問題

Does finding a name prove who the person is?

No. A text mention can be ambiguous. Resolving it to a particular person requires additional evidence and a separate linking process.