Torna alle notizie
ImpresaAI Understanding briefing

OpenAI rileva che il divario nell'intelligenza artificiale aziendale si sta ampliando man mano che gli agenti entrano nel lavoro quotidiano

L'aggiornamento Enterprise Signals del 12 agosto di OpenAI afferma che le aziende di frontiera stanno ampliando il loro vantaggio nella profondità dell'intelligenza artificiale, con l'uso degli agenti che va oltre l'ingegneria, ma l'azienda avverte che il volume dei token è solo un indicatore del lavoro, non del valore aziendale.

6 min readRead the primary source
Documento di origine primariaFonte registrata
Editore
OpenAI's Enterprise Signals report and August 12 publication
Collegamento alla fonte
openai.comhttps://openai.com/index/how-enterprises-put-ai-to-work
Tipo di fonte
Documento principale: un annuncio ufficiale, un documento, un documento o una pagina proprietaria che leggiamo direttamente.
ContestoComprendilo in 60 secondi

Inizia qui

Termini chiave

Memoria (memoria dell'agente)
Contesto archiviato che un agente AI utilizza attraverso passaggi o sessioni per migliorare la continuità.
Classificazione
Un'attività in cui un modello assegna un input a una o più categorie predefinite.
Punto di riferimento
Un test o un set di dati standardizzato utilizzato per misurare e confrontare le prestazioni del modello.
Mettiti alla provaQuiz sugli agenti IA

Cosa è successo

OpenAI published two complementary reports on August 12, including an updated Enterprise Signals analysis of aggregated, de-identified customer usage and a working paper on adoption across firms and workers. The company says enterprise AI is shifting from answering questions toward delegated, tool-using work. Its most striking comparison is between frontier firms, defined as the top 10% of monthly output-token intensity, and typical firms near the middle of the distribution: the gap grew from 2.6× in January to 8.3× in June. OpenAI presents these as usage signals, not proof that one group creates eight times more value.

The report’s first measure is not model quality or revenue; it is output tokens per active user. OpenAI uses that volume as a rough proxy for the depth and duration of AI-assisted work, reasoning that longer, multi-step tasks tend to generate more output than a short answer. The company also states the limitation plainly: a brief response can be valuable, while a long response can add little. The 8.3× comparison therefore describes a difference in observed usage intensity inside OpenAI’s customer base, not an audited productivity multiplier.

The shift is visible in Codex’s share of enterprise activity. As of June, OpenAI says Codex generated 64% of combined Codex and ChatGPT output tokens among enterprise customers. Its description of agentic work is concrete: tools can help a worker find information, edit files, create deliverables, and carry out multi-step tasks autonomously or under supervision. The result is a change in the unit of work being delegated—from asking an assistant for guidance to asking an agent to complete a reviewable task.

The frontier definition matters because it is relative, not a permanent league table. OpenAI ranks customers each month by output tokens per active user, labels the top 10% frontier firms, and compares them with companies between the 45th and 55th percentiles. The same analysis says the gap is visible across industries, ranging from 11.7× in information and technology to 5.3× in manufacturing. Those comparisons show a distribution of usage within the studied customer population; they do not identify which companies are named or explain every cause of the difference.

The report also tracks where adoption is spreading. Since February, OpenAI says weekly active enterprise Codex users grew 108× in legal, 41× in sales, 41× in recruiting, and 26× in marketing, compared with 5× in engineering. Frontier users were more likely to use advanced capabilities: 21% used Plugins weekly and 19% used skills, versus 9% and 3% at typical firms. OpenAI says the analysis covers more than 10 million messages and uses automated ; no OpenAI employee reviewed customer messages, according to the disclosure.

Dettagli della fonte: OpenAI's Enterprise Signals report and August 12 publication ↗

Perché è importante

The update changes the enterprise AI question from who has access to who has built the operating conditions for delegated work. The leading firms in this dataset are not described as using a different class of model; they are connecting models to context, tools, repeatable workflows, and review processes more deeply.

That distinction is practical for organizations that cannot buy their way to the frontier. A team may have the same model access as a larger company and still see little benefit if employees must copy context manually, agents cannot reach approved systems, or every output is trapped in a one-off chat. OpenAI’s own explanation points toward an adoption stack: shared workflows, data infrastructure, continuous learning, permissions, governance, and human review. Those are organizational investments, not just product settings.

The cross-function numbers also challenge the idea that agentic AI is mainly a software-engineering story. OpenAI says legal, sales, recruiting, and marketing are among the fastest-growing areas for weekly Codex users, while system and agent operations make up meaningful shares of agentic messages in recruiting, sales, policy, and communications. For a nonprofit, public agency, or small business, the relevant question is not whether it can automate everything. It is whether a narrow process—research, reporting, case preparation, or document review—can be made faster while keeping a person accountable for the result.

The adoption gap may also become a governance gap. OpenAI says agents need access to company systems and that frontier firms set rules for where agents can operate, what they can access, when they can act, and how people review higher-risk decisions. Deeper usage can therefore increase both capability and exposure: a connected agent may save time, but a mistaken permission, stale data source, or unreviewed action can travel farther than a bad chat answer. The report’s emphasis on controls is a reminder that deployment maturity includes limits and auditability.

The evidence remains a company-produced view of its own enterprise ecosystem. OpenAI says the data are aggregated and de-identified, and its automated systems classify message content, but readers cannot independently inspect the underlying customer records from the public page. The measures also privilege activity that produces tokens and may miss value created through short interactions, offline work, or tools outside OpenAI’s products. The strongest supported claim is about a widening usage pattern—not a universal ranking of business performance.

Interactive Mechanism

Meccanismo interattivo: come funziona realmente

Esplora la tecnologia alla base di questo sviluppo in modo interattivo.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Verifica concettuale interattiva+10 Points
AI Agents Quiz

What most distinguishes an AI agent from a basic chatbot?

Cosa guardare dopo

The next test is whether the usage gap predicts durable outcomes outside OpenAI’s measurement system. Watch for independent evidence connecting agent deployment to verified task completion, quality, cost, worker experience, and incident rates rather than treating more tokens as the destination.

Independent researchers should test whether output-token intensity remains useful when compared with operational measures such as cycle time, error correction, customer outcomes, or revenue per employee. A fair comparison would separate the effect of model access from the effect of workflow design, training, data quality, and employee selection. It should also report failures and rework, because a long agent run that produces a polished but unusable deliverable is not the same as completed work.

Organizations should look inside their own distributions instead of copying a top-10% label. The useful unit may be a team, process, or task family: how often an agent can complete a bounded workflow, how much context it needs, how many handoffs require human intervention, and what kinds of errors recur. OpenAI’s monthly percentile method provides a way to describe adoption depth, but it does not by itself reveal whether the frontier firms have better data, more permissive budgets, different staffing, or more mature internal controls.

The safety signal to track is the relationship between autonomy and review. As agents gain access to files, browsers, CRM systems, code repositories, and financial or legal workflows, teams will need explicit permission boundaries, logging, rollback paths, and escalation rules. The practical is not maximum autonomy. It is whether a worker can see what the system did, verify the important steps, and stop or correct an action before it creates material harm.

Finally, watch how OpenAI’s product vocabulary becomes measurable in practice. The report groups Plugins, skills, memory, app access, computer use, and persistence into a broader model of agent capability, but adoption percentages do not tell readers which combinations deliver reliable value. Future updates will be more useful if they disclose task-level denominators, uncertainty, model and product changes, and outcomes over time. Until then, this August 12 release is a timely signal that enterprise AI is becoming more agentic, with meaningful uncertainty about how much of that activity translates into durable public benefit.

Guide e quiz correlati

Agenti dell'intelligenza artificialeSpiegazione dei modelli di intelligenza artificialeFuturo dell'IAEtica dell'IAMetti alla prova ciò che sai: prova un quiz gratuito sull'intelligenza artificialeCerca un termine AI nel nostro glossarioSegui il tracker dei finanziamenti AI
Lo hai trovato utile?