返回新聞
企業AI Understanding 簡報

OpenAI 發現隨著代理商進入日常工作,企業人工智慧差距正在擴大

OpenAI 8 月 12 日的企業訊號更新表示,前沿公司正在擴大其在人工智慧深度方面的領先優勢,代理使用已超越了工程領域,但該公司警告說,代幣數量只是工作的代表,而不是商業價值。

6 min readRead the primary source
主要來源文件來源記錄
出版商
OpenAI's Enterprise Signals report and August 12 publication
來源連結
openai.comhttps://openai.com/index/how-enterprises-put-ai-to-work
來源類型
主要文件-我們直接閱讀的官方公告、文件、文件或第一方頁面。
背景60 秒內了解這一點

從這裡開始

關鍵術語

記憶體(代理記憶體)
AI 代理程式跨步驟或會話使用儲存的上下文來提高連續性。
分類
模型將輸入分配給一個或多個預定義類別的任務。
基準測試
用於測量和比較模型性能的標準化測試或資料集。
測試一下自己AI 代理測驗

發生了什麼事

OpenAI published two complementary reports on August 12, including an updated Enterprise Signals analysis of aggregated, de-identified customer usage and a working paper on adoption across firms and workers. The company says enterprise AI is shifting from answering questions toward delegated, tool-using work. Its most striking comparison is between frontier firms, defined as the top 10% of monthly output-token intensity, and typical firms near the middle of the distribution: the gap grew from 2.6× in January to 8.3× in June. OpenAI presents these as usage signals, not proof that one group creates eight times more value.

The report’s first measure is not model quality or revenue; it is output tokens per active user. OpenAI uses that volume as a rough proxy for the depth and duration of AI-assisted work, reasoning that longer, multi-step tasks tend to generate more output than a short answer. The company also states the limitation plainly: a brief response can be valuable, while a long response can add little. The 8.3× comparison therefore describes a difference in observed usage intensity inside OpenAI’s customer base, not an audited productivity multiplier.

The shift is visible in Codex’s share of enterprise activity. As of June, OpenAI says Codex generated 64% of combined Codex and ChatGPT output tokens among enterprise customers. Its description of agentic work is concrete: tools can help a worker find information, edit files, create deliverables, and carry out multi-step tasks autonomously or under supervision. The result is a change in the unit of work being delegated—from asking an assistant for guidance to asking an agent to complete a reviewable task.

The frontier definition matters because it is relative, not a permanent league table. OpenAI ranks customers each month by output tokens per active user, labels the top 10% frontier firms, and compares them with companies between the 45th and 55th percentiles. The same analysis says the gap is visible across industries, ranging from 11.7× in information and technology to 5.3× in manufacturing. Those comparisons show a distribution of usage within the studied customer population; they do not identify which companies are named or explain every cause of the difference.

The report also tracks where adoption is spreading. Since February, OpenAI says weekly active enterprise Codex users grew 108× in legal, 41× in sales, 41× in recruiting, and 26× in marketing, compared with 5× in engineering. Frontier users were more likely to use advanced capabilities: 21% used Plugins weekly and 19% used skills, versus 9% and 3% at typical firms. OpenAI says the analysis covers more than 10 million messages and uses automated ; no OpenAI employee reviewed customer messages, according to the disclosure.

來源詳情: OpenAI's Enterprise Signals report and August 12 publication ↗

為什麼這很重要

The update changes the enterprise AI question from who has access to who has built the operating conditions for delegated work. The leading firms in this dataset are not described as using a different class of model; they are connecting models to context, tools, repeatable workflows, and review processes more deeply.

That distinction is practical for organizations that cannot buy their way to the frontier. A team may have the same model access as a larger company and still see little benefit if employees must copy context manually, agents cannot reach approved systems, or every output is trapped in a one-off chat. OpenAI’s own explanation points toward an adoption stack: shared workflows, data infrastructure, continuous learning, permissions, governance, and human review. Those are organizational investments, not just product settings.

The cross-function numbers also challenge the idea that agentic AI is mainly a software-engineering story. OpenAI says legal, sales, recruiting, and marketing are among the fastest-growing areas for weekly Codex users, while system and agent operations make up meaningful shares of agentic messages in recruiting, sales, policy, and communications. For a nonprofit, public agency, or small business, the relevant question is not whether it can automate everything. It is whether a narrow process—research, reporting, case preparation, or document review—can be made faster while keeping a person accountable for the result.

The adoption gap may also become a governance gap. OpenAI says agents need access to company systems and that frontier firms set rules for where agents can operate, what they can access, when they can act, and how people review higher-risk decisions. Deeper usage can therefore increase both capability and exposure: a connected agent may save time, but a mistaken permission, stale data source, or unreviewed action can travel farther than a bad chat answer. The report’s emphasis on controls is a reminder that deployment maturity includes limits and auditability.

The evidence remains a company-produced view of its own enterprise ecosystem. OpenAI says the data are aggregated and de-identified, and its automated systems classify message content, but readers cannot independently inspect the underlying customer records from the public page. The measures also privilege activity that produces tokens and may miss value created through short interactions, offline work, or tools outside OpenAI’s products. The strongest supported claim is about a widening usage pattern—not a universal ranking of business performance.

Interactive Mechanism

互動機制:它實際上是如何運作的

以互動方式探索這項發展背後的基礎技術。

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
互動式概念檢查+10 Points
AI Agents Quiz

What most distinguishes an AI agent from a basic chatbot?

接下來看什麼

The next test is whether the usage gap predicts durable outcomes outside OpenAI’s measurement system. Watch for independent evidence connecting agent deployment to verified task completion, quality, cost, worker experience, and incident rates rather than treating more tokens as the destination.

Independent researchers should test whether output-token intensity remains useful when compared with operational measures such as cycle time, error correction, customer outcomes, or revenue per employee. A fair comparison would separate the effect of model access from the effect of workflow design, training, data quality, and employee selection. It should also report failures and rework, because a long agent run that produces a polished but unusable deliverable is not the same as completed work.

Organizations should look inside their own distributions instead of copying a top-10% label. The useful unit may be a team, process, or task family: how often an agent can complete a bounded workflow, how much context it needs, how many handoffs require human intervention, and what kinds of errors recur. OpenAI’s monthly percentile method provides a way to describe adoption depth, but it does not by itself reveal whether the frontier firms have better data, more permissive budgets, different staffing, or more mature internal controls.

The safety signal to track is the relationship between autonomy and review. As agents gain access to files, browsers, CRM systems, code repositories, and financial or legal workflows, teams will need explicit permission boundaries, logging, rollback paths, and escalation rules. The practical is not maximum autonomy. It is whether a worker can see what the system did, verify the important steps, and stop or correct an action before it creates material harm.

Finally, watch how OpenAI’s product vocabulary becomes measurable in practice. The report groups Plugins, skills, memory, app access, computer use, and persistence into a broader model of agent capability, but adoption percentages do not tell readers which combinations deliver reliable value. Future updates will be more useful if they disclose task-level denominators, uncertainty, model and product changes, and outcomes over time. Until then, this August 12 release is a timely signal that enterprise AI is becoming more agentic, with meaningful uncertainty about how much of that activity translates into durable public benefit.

相關指引和測驗

人工智慧代理人工智慧模型解釋AI 的未來AI 倫理測試你所知道的—嘗試免費的人工智慧測驗在我們的詞彙表中尋找人工智慧術語關注 AI 資金追蹤器
覺得有用嗎?