返回新闻
创新AI Understanding 简报

36氪报道DeepSeek推出了用于AI自我提升研究的开发者预览版工具

36氪报道,DeepSeek 发布了 DeepSeek Harness,这是一个开发者预览环境,旨在让 AI 代理修改、评估和回滚软件组件。该报告称,该系统使用一种名为 Cordis 的机制来管理不断变化的插件和依赖项,但其发布状态、功能和更广泛的……

6 min readRead the linked source
Source-provided image accompanying 36Kr reports DeepSeek launched a developer-preview harness for AI self-improvement research
来源参考来源记录
出版商
eu.36kr.com
来源链接
eu.36kr.comhttps://eu.36kr.com/en/p/3952918063774852
来源类型
链接来源——主要来源状态尚未确定。
背景60 秒内了解这一点

从这里开始

关键术语

AGI(通用人工智能)
一个假设的人工智能系统,可以在许多领域以人类水平执行大多数智力任务。
内存(代理内存)
AI 代理跨步骤或会话使用存储的上下文来提高连续性。
思想链
一种推理风格,人工智能模型将问题分解为中间步骤。
测试一下自己AI 代理测验

发生了什么

According to 36Kr, DeepSeek released DeepSeek Harness, or DSH, as a developer preview the previous week. The report describes it as an operating environment for AI agents that can install, modify and remove components while preserving the ability to recover from failed changes. 36Kr says DSH is part of DeepSeek’s broader exploration of recursive self-improvement, rather than evidence that AI systems can independently create more capable AI.

36Kr’s Aug. 24 report says DeepSeek released DeepSeek Harness, or DSH, as a developer preview after other AI companies and projects had introduced comparable “harness” systems. The article places the launch within DeepSeek co-founder Liang Wenfeng’s reported progression from large language models to reasoning, agents, continuous learning, self-iteration and embodied intelligence. The source presents recursive self-improvement, or RSI, as a means toward artificial general intelligence, not as a synonym for AGI. It also says DSH is an early milestone and remains far from demonstrating a technological singularity.

The source does not provide an exact release date, public availability conditions, performance measurements or a primary DeepSeek announcement. The report identifies a core mechanism in DSH called Cordis. According to 36Kr, Cordis is intended to address a limitation of conventional plugin systems: removing a plugin may not fully remove its effects without restarting the surrounding process. That limitation matters for an agent operating with tools, permissions, sandboxes, memories and session state, because repeated modifications could otherwise interrupt a running system or leave it in an unrecoverable intermediate condition.

36Kr says Cordis tracks and revokes the effects of components when they exit and automatically re-coordinates dependencies when relationships between components change. The article connects those capabilities to the concepts of temporal and spatial composability in a paper that it says DeepSeek and Peking University published alongside the release. 36Kr reports that the underlying approach has been used for four years in Koishi, an open-source chatbot framework with more than 4,000 community plugins. The article treats that history as evidence of engineering experience, while acknowledging that the mechanism still needs testing in larger plugin ecosystems.

The supplied source does not include the paper’s title in a form that can be independently checked beyond its reference to “a programming paradigm for spatiotemporal composability,” nor does it provide benchmarks, failure rates, reproducible test procedures or evidence that Cordis has enabled DSH to improve a frontier model. The article also describes user criticism following the release, saying some users found deployment cumbersome, configuration complex and the overall experience unfinished. It then surveys related efforts at Google DeepMind, Anthropic, Tencent and MiniMax. Those examples include automated program search, agent research loops and reported internal evaluations. They are contextual claims made by 36Kr, not independently verified findings in the supplied material. The central concrete development remains the reported release of DSH and its proposed software environment for iterative agent modification.

来源详情: eu.36kr.com ↗

为什么这很重要

The reported launch points to a shift from evaluating AI mainly through periodic model releases toward building environments in which models can repeatedly act, receive feedback and modify their own supporting software. If the system works as described, the practical challenge is less autonomous invention than reliable experimentation: changes must be testable, reversible and judged against useful outcomes. The report does not independently establish that DSH improves a model’s core capabilities or brings AI closer to artificial general intelligence.

The significance of the reported launch lies in infrastructure for iteration. A conventional model release creates a visible comparison point: a new model is evaluated, ranked and deployed. The system described by 36Kr instead emphasizes what happens between releases. An agent operates in an environment, changes a component, runs an evaluation, observes the result and decides whether to keep or reverse the change. That approach could make model development more continuous, but only if the environment provides reliable tests and the feedback signal reflects real usefulness rather than a narrow metric.

The report’s emphasis on rollback and dependency management highlights a practical safety and reliability problem. An AI agent with authority to alter tools or scaffolding can fail in ways that are difficult to diagnose if each change contaminates the next experiment. A mechanism that records effects, removes them and restores a known state could make experimentation more controlled. The source does not establish that DSH provides comprehensive security, prevents unauthorized actions, or protects data and permissions. Those remain important unknowns, especially because the described environment may include memory, tools, sandboxes and session state.

The article also argues that access to real work and high-quality feedback could become a competitive advantage. In its account, coding environments are valuable because they generate structured activity traces and rapid signals about where an AI system succeeds or fails. This is a plausible strategic implication of the report, but it should not be treated as proof that data from user activity will produce better models. The source gives no independent evidence about data governance, consent, privacy protections, training use, or whether any reported gains generalize beyond internal tasks. Nor does it show that recursive loops produce increasing returns rather than diminishing returns.

Interactive Mechanism

互动机制:它实际上是如何运作的

以交互方式探索这一发展背后的基础技术。

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
交互式概念检查+10 Points
AI Agents Quiz

An agent must create a draft calendar event for Tuesday at 2 p.m. Which evidence would establish the requested result?

接下来看什么

The main questions are whether DeepSeek publishes technical documentation, source code, evaluations or demonstrations that substantiate the report’s claims; how reliably DSH handles failed modifications and changing dependencies; and whether it is usable outside DeepSeek’s own engineering environment. Readers should also distinguish a tool-assisted improvement loop from recursive self-improvement in the stronger sense of an AI designing and training substantially more capable successors.

Verification should begin with the release itself. A meaningful next step would be public technical documentation or a repository that specifies DSH’s supported models, operating requirements, permissions model, sandbox boundaries and recovery behavior. Independent users would need to reproduce the reported plugin installation, removal and dependency-reconciliation functions. Without such material, the public cannot determine whether DSH is broadly usable, a limited internal tool, or primarily a research prototype.

Evaluators should look for evidence that separates engineering automation from stronger claims about recursive self-improvement. Useful tests would show whether an agent can propose changes, run objective evaluations, recover from failed changes and improve performance across tasks that were not used to guide the loop. The supplied report provides no such independent evaluation. It also does not establish whether any gains affect model training, inference, tool use, software scaffolding or only a narrowly defined benchmark.

The broader claims in 36Kr’s article require particular caution. The report cites forecasts about when RSI might occur and describes efforts by several major AI companies, but predictions and company statements are not evidence that a self-improvement threshold has been reached. Future reporting should examine human oversight, shutdown authority, auditability, data provenance and the possibility that optimization loops exploit their evaluators. It should also track whether DSH receives stable public support, whether deployment complaints are resolved, and whether independent researchers can verify any practical capability gains.

相关指南和测验

人工智能代理人工智能模型解释人工智能培训AI 的未来测试你所知道的——尝试免费的人工智能测验在我们的词汇表中查找人工智能术语关注 AI 模型发布跟踪器
觉得这有用吗?