返回新聞
產品展示AI Understanding 簡報

36氪報道OpenAI為代理應用開啟了Codex Harness

36氪報道,OpenAI 將 Codex Harness 定位為開發者建立代理應用程式的開放運行時,將 Codex 擴展到自己的應用程式、命令列工具和 IDE 整合之外。該報告稱運行時管理會話、上下文、工具呼叫、沙箱和人工批准,但聲明並未…

6 min readRead the linked source
Source-provided image accompanying 36Kr reports OpenAI opened Codex Harness for agent applications
來源參考來源記錄
出版商
eu.36kr.com
來源連結
eu.36kr.comhttps://eu.36kr.com/en/p/3952749463895174
來源類型
連結來源-主要來源狀態尚未確定。
背景60 秒內了解這一點

從這裡開始

關鍵術語

AGI(通用人工智慧)
一個假設的人工智慧系統,可以在許多領域以人類層級執行大多數智力任務。
嵌入
擷取文字、影像或其他資料語意的數位向量表示。
代幣
由語言模型處理的文字區塊,例如單字或符號。
測試一下自己AI 代理測驗

發生了什麼事

36Kr reports that OpenAI released a developer post on August 20 describing Codex as a platform built on an open agent harness. According to the report, developers can use codex exec, an SDK or an app-server to embed Codex capabilities into enterprise consoles, customer-service systems, security tools and internal applications. The report presents this as a shift from Codex being primarily a programming assistant to serving as reusable execution infrastructure.

36Kr reports that OpenAI’s August 20 developer post presented Codex as a platform and described an open-source Codex Harness underneath the Codex App, CLI and IDE extensions. The report says the harness manages session state, context, tool calls, sandboxes and human approvals. It also says developers can connect those capabilities to existing products through codex exec, an SDK or an app-server that supports threads, turns, event streams and approval protocols. No primary OpenAI document was supplied with the source, so these details are attributed to 36Kr and are not independently confirmed here.

The report says the capabilities were introduced in stages. It describes Codex SDK availability alongside the general release of Codex in October 2025, later use of codex exec for scripts and continuous-integration tasks, and the app-server’s support for longer-running sessions. It also says OpenAI published an explanation of the Codex agent loop in January 2026, covering how the model receives instructions, calls tools, reads results and continues to the next step. 36Kr’s central interpretation is that OpenAI has now grouped these functions under the Codex Harness name and is encouraging developers to build products around them.

According to 36Kr, OpenAI cited several applications in its own blog post. The report says Cisco uses the Codex SDK in App Builder for Cisco Cloud Control, while GitHub and JetBrains integrate Codex into existing development environments. It also says Thrive Holdings and Crete use Codex for tax preparation and that one pilot handled 7,000 tax returns while reducing preparation time by about one-third. These examples and measurements are presented by 36Kr as claims from OpenAI; the source does not provide independent testing, methodology or documentation from the organizations involved.

The report compares Codex Harness with DeepSeek Harness, which it says DeepSeek open-sourced on August 14 under the DSH abbreviation. 36Kr says DSH treats models, tools, skills, sessions, sandboxes, storage, the agent loop, scheduling and the user interface as plugins managed through a system called Cordis. It characterizes Codex as a more integrated runtime whose non-core components must be adapted to its engine, while describing DSH as more modular. The article also discusses Kimi Code and ZCode as related but differently structured agent systems, rather than presenting them as the same product.

來源詳情: eu.36kr.com ↗

為什麼這很重要

The reported change could give OpenAI a larger role in the software layer that determines how AI agents use tools, preserve context, request approval and complete multistep work. That layer can affect reliability, cost, observability and user control independently of the underlying model. The report also places OpenAI’s move in competition with more open agent runtimes, including DeepSeek Harness and other open-source projects.

The reported move matters because an agent runtime controls the operational steps between a model’s output and a completed task. According to 36Kr, Codex Harness can maintain context, call tools, execute work in sandboxes and route actions through human approvals. For organizations agents into internal software, those functions can determine whether an agent is observable, interruptible and compatible with existing permissions. The report does not establish that Codex Harness meets any particular enterprise security or compliance standard.

36Kr reports that OpenAI presented an ARC-AGI-3 comparison in which GPT-5.6 Sol scored 13.3% alone and 38.3% with Codex Harness’s continuous reasoning and context-compression capabilities. The article also says the output volume fell to about one-sixth of the original amount. Those figures, if reproduced, would suggest that orchestration and context handling can materially change the performance and cost of the same model. However, the source provides no test protocol, baseline details, independent replication or information about statistical significance, so the numbers should be treated as reported company results rather than settled evidence.

The article argues that model companies have incentives to own or distribute the runtime around their models. 36Kr says a runtime can expose where tasks fail, where tool calls stall, how often retries occur and which context strategies consume fewer tokens. This could help a model developer improve both its models and its execution system. The source also notes that any use of task records for training would depend on applicable privacy policies; it does not establish what Codex or DeepSeek collect, retain or use.

A more open runtime could also change the balance between model vendors and application developers. 36Kr says DSH’s plugin-based architecture is designed to let developers replace models, tools, storage, sandboxes, and loop components without maintaining a long-lived fork of the entire project. That flexibility may appeal to teams that want to change providers or execution strategies. At the same time, the article reports that DSH is still an early preview and that community feedback raised concerns about speed, consumption, documentation, compatibility and uneven plugin quality. Those observations are reported feedback, not a systematic evaluation.

Interactive Mechanism

互動機制:它實際上是如何運作的

以互動方式探索這項發展背後的基礎技術。

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
互動式概念檢查+10 Points
AI Agents Quiz

An agent must create a draft calendar event for Tuesday at 2 p.m. Which evidence would establish the requested result?

接下來看什麼

The most important open questions are whether Codex Harness is genuinely usable outside OpenAI’s model ecosystem, what components developers can replace, how data from embedded workflows is handled, and whether the reported performance gains hold under independent testing. The supplied report does not establish broad availability, pricing, contractual terms, privacy practices or production reliability.

First, watch for evidence that Codex Harness can operate with models and tools beyond OpenAI’s preferred stack. 36Kr describes Codex as open and embeddable, but the supplied report does not specify the license, the exact replaceable components, supported model adapters, deployment options or whether critical parts remain closed. Those details will determine whether developers receive a portable runtime or a tightly coupled OpenAI integration.

Second, independent tests should examine the reported ARC-AGI-3 improvement and reduction. Useful verification would include the exact model versions, prompts, harness settings, context limits, tool configuration, failure handling and cost assumptions. Without those details, the comparison cannot show whether the gain comes from a generally useful runtime design or from a particular proprietary setup.

Third, organizations considering embedded agents will need clearer information about data governance and control. The report describes enterprise code, customer data, internal tools, approvals and task records moving through a shared runtime. It does not say where those records are processed, how long they are retained, whether administrators can audit or delete them, or whether they may be used for model improvement. Those unknowns are material for deployments involving confidential or regulated work.

Finally, track whether the growing collection of agent runtimes becomes a durable infrastructure market or remains a set of model-specific developer tools. 36Kr reports that DeepSeek Harness attracted substantial early developer attention and that Kimi Code and ZCode already provide related execution capabilities. Adoption in production will depend on stability, rollback mechanisms, permission boundaries, documentation, cost predictability and independent security review. The supplied report establishes interest and product positioning, but not broad enterprise adoption or reliable superiority.

相關指引和測驗

人工智慧代理人工智慧模型解釋Prompt Engineering測試你所知道的—嘗試免費的人工智慧測驗在我們的詞彙表中尋找人工智慧術語關注 AI 模型發布追蹤器
覺得有用嗎?