返回新聞
創新AI Understanding 簡報

36氪报道DeepSeek推出了用于AI自我提升研究的开发者预览版工具

36氪報道,DeepSeek 發布了 DeepSeek Harness,這是一個開發者預覽環境,旨在讓 AI 代理程式修改、評估和回溯軟體元件。該報告稱,該系統使用一種名為 Cordis 的機制來管理不斷變化的插件和依賴項,但其發布狀態、功能和更廣泛的…

6 min readRead the linked source
Source-provided image accompanying 36Kr reports DeepSeek launched a developer-preview harness for AI self-improvement research
來源參考來源記錄
出版商
eu.36kr.com
來源連結
eu.36kr.comhttps://eu.36kr.com/en/p/3952918063774852
來源類型
連結來源-主要來源狀態尚未確定。
背景60 秒內了解這一點

從這裡開始

關鍵術語

AGI(通用人工智慧)
一個假設的人工智慧系統,可以在許多領域以人類層級執行大多數智力任務。
記憶體(代理記憶體)
AI 代理程式跨步驟或會話使用儲存的上下文來提高連續性。
思想鏈
一種推理風格,人工智慧模型將問題分解為中間步驟。
測試一下自己AI 代理測驗

發生了什麼事

According to 36Kr, DeepSeek released DeepSeek Harness, or DSH, as a developer preview the previous week. The report describes it as an operating environment for AI agents that can install, modify and remove components while preserving the ability to recover from failed changes. 36Kr says DSH is part of DeepSeek’s broader exploration of recursive self-improvement, rather than evidence that AI systems can independently create more capable AI.

36Kr’s Aug. 24 report says DeepSeek released DeepSeek Harness, or DSH, as a developer preview after other AI companies and projects had introduced comparable “harness” systems. The article places the launch within DeepSeek co-founder Liang Wenfeng’s reported progression from large language models to reasoning, agents, continuous learning, self-iteration and embodied intelligence. The source presents recursive self-improvement, or RSI, as a means toward artificial general intelligence, not as a synonym for AGI. It also says DSH is an early milestone and remains far from demonstrating a technological singularity.

The source does not provide an exact release date, public availability conditions, performance measurements or a primary DeepSeek announcement. The report identifies a core mechanism in DSH called Cordis. According to 36Kr, Cordis is intended to address a limitation of conventional plugin systems: removing a plugin may not fully remove its effects without restarting the surrounding process. That limitation matters for an agent operating with tools, permissions, sandboxes, memories and session state, because repeated modifications could otherwise interrupt a running system or leave it in an unrecoverable intermediate condition.

36Kr says Cordis tracks and revokes the effects of components when they exit and automatically re-coordinates dependencies when relationships between components change. The article connects those capabilities to the concepts of temporal and spatial composability in a paper that it says DeepSeek and Peking University published alongside the release. 36Kr reports that the underlying approach has been used for four years in Koishi, an open-source chatbot framework with more than 4,000 community plugins. The article treats that history as evidence of engineering experience, while acknowledging that the mechanism still needs testing in larger plugin ecosystems.

The supplied source does not include the paper’s title in a form that can be independently checked beyond its reference to “a programming paradigm for spatiotemporal composability,” nor does it provide benchmarks, failure rates, reproducible test procedures or evidence that Cordis has enabled DSH to improve a frontier model. The article also describes user criticism following the release, saying some users found deployment cumbersome, configuration complex and the overall experience unfinished. It then surveys related efforts at Google DeepMind, Anthropic, Tencent and MiniMax. Those examples include automated program search, agent research loops and reported internal evaluations. They are contextual claims made by 36Kr, not independently verified findings in the supplied material. The central concrete development remains the reported release of DSH and its proposed software environment for iterative agent modification.

來源詳情: eu.36kr.com ↗

為什麼這很重要

The reported launch points to a shift from evaluating AI mainly through periodic model releases toward building environments in which models can repeatedly act, receive feedback and modify their own supporting software. If the system works as described, the practical challenge is less autonomous invention than reliable experimentation: changes must be testable, reversible and judged against useful outcomes. The report does not independently establish that DSH improves a model’s core capabilities or brings AI closer to artificial general intelligence.

The significance of the reported launch lies in infrastructure for iteration. A conventional model release creates a visible comparison point: a new model is evaluated, ranked and deployed. The system described by 36Kr instead emphasizes what happens between releases. An agent operates in an environment, changes a component, runs an evaluation, observes the result and decides whether to keep or reverse the change. That approach could make model development more continuous, but only if the environment provides reliable tests and the feedback signal reflects real usefulness rather than a narrow metric.

The report’s emphasis on rollback and dependency management highlights a practical safety and reliability problem. An AI agent with authority to alter tools or scaffolding can fail in ways that are difficult to diagnose if each change contaminates the next experiment. A mechanism that records effects, removes them and restores a known state could make experimentation more controlled. The source does not establish that DSH provides comprehensive security, prevents unauthorized actions, or protects data and permissions. Those remain important unknowns, especially because the described environment may include memory, tools, sandboxes and session state.

The article also argues that access to real work and high-quality feedback could become a competitive advantage. In its account, coding environments are valuable because they generate structured activity traces and rapid signals about where an AI system succeeds or fails. This is a plausible strategic implication of the report, but it should not be treated as proof that data from user activity will produce better models. The source gives no independent evidence about data governance, consent, privacy protections, training use, or whether any reported gains generalize beyond internal tasks. Nor does it show that recursive loops produce increasing returns rather than diminishing returns.

Interactive Mechanism

互動機制:它實際上是如何運作的

以互動方式探索這項發展背後的基礎技術。

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
互動式概念檢查+10 Points
AI Agents Quiz

An agent must create a draft calendar event for Tuesday at 2 p.m. Which evidence would establish the requested result?

接下來看什麼

The main questions are whether DeepSeek publishes technical documentation, source code, evaluations or demonstrations that substantiate the report’s claims; how reliably DSH handles failed modifications and changing dependencies; and whether it is usable outside DeepSeek’s own engineering environment. Readers should also distinguish a tool-assisted improvement loop from recursive self-improvement in the stronger sense of an AI designing and training substantially more capable successors.

Verification should begin with the release itself. A meaningful next step would be public technical documentation or a repository that specifies DSH’s supported models, operating requirements, permissions model, sandbox boundaries and recovery behavior. Independent users would need to reproduce the reported plugin installation, removal and dependency-reconciliation functions. Without such material, the public cannot determine whether DSH is broadly usable, a limited internal tool, or primarily a research prototype.

Evaluators should look for evidence that separates engineering automation from stronger claims about recursive self-improvement. Useful tests would show whether an agent can propose changes, run objective evaluations, recover from failed changes and improve performance across tasks that were not used to guide the loop. The supplied report provides no such independent evaluation. It also does not establish whether any gains affect model training, inference, tool use, software scaffolding or only a narrowly defined benchmark.

The broader claims in 36Kr’s article require particular caution. The report cites forecasts about when RSI might occur and describes efforts by several major AI companies, but predictions and company statements are not evidence that a self-improvement threshold has been reached. Future reporting should examine human oversight, shutdown authority, auditability, data provenance and the possibility that optimization loops exploit their evaluators. It should also track whether DSH receives stable public support, whether deployment complaints are resolved, and whether independent researchers can verify any practical capability gains.

相關指引和測驗

人工智慧代理人工智慧模型解釋人工智慧培訓AI 的未來測試你所知道的—嘗試免費的人工智慧測驗在我們的詞彙表中尋找人工智慧術語關注 AI 模型發布追蹤器
覺得有用嗎?