Zpět na Novinky
InovaceInstruktáž AI Understanding

36Kr uvádí, že DeepSeek spustilo vývojářskou ukázku pro výzkum sebezdokonalování AI

36Kr uvádí, že DeepSeek vydalo DeepSeek Harness, vývojářské prostředí pro náhled, které umožňuje agentům umělé inteligence upravovat, vyhodnocovat a vracet softwarové komponenty. Zpráva říká, že systém používá mechanismus nazvaný Cordis ke správě měnících se pluginů a závislostí, ale jeho stav vydání, možnosti a širší…

6 min readRead the linked source
Source-provided image accompanying 36Kr reports DeepSeek launched a developer-preview harness for AI self-improvement research
Odkaz na zdrojZdroj zaznamenán
Vydavatel
eu.36kr.com
Odkaz na zdroj
eu.36kr.comhttps://eu.36kr.com/en/p/3952918063774852
Typ zdroje
Propojený zdroj — stav primárního zdroje nebyl stanoven.
KontextPochopte to za 60 sekund

Začněte zde

Klíčové pojmy

AGI (obecná umělá inteligence)
Hypotetický systém umělé inteligence, který může provádět většinu intelektuálních úkolů na lidské úrovni v mnoha oblastech.
Paměť (paměť agenta)
Uložený kontext, který agent AI používá v krocích nebo relacích ke zlepšení kontinuity.
Řetězec myšlení
Styl uvažování, kdy model AI rozkládá problém na mezikroky.
Otestujte seKvíz AI agentů

Co se stalo

According to 36Kr, DeepSeek released DeepSeek Harness, or DSH, as a developer preview the previous week. The report describes it as an operating environment for AI agents that can install, modify and remove components while preserving the ability to recover from failed changes. 36Kr says DSH is part of DeepSeek’s broader exploration of recursive self-improvement, rather than evidence that AI systems can independently create more capable AI.

36Kr’s Aug. 24 report says DeepSeek released DeepSeek Harness, or DSH, as a developer preview after other AI companies and projects had introduced comparable “harness” systems. The article places the launch within DeepSeek co-founder Liang Wenfeng’s reported progression from large language models to reasoning, agents, continuous learning, self-iteration and embodied intelligence. The source presents recursive self-improvement, or RSI, as a means toward artificial general intelligence, not as a synonym for AGI. It also says DSH is an early milestone and remains far from demonstrating a technological singularity.

The source does not provide an exact release date, public availability conditions, performance measurements or a primary DeepSeek announcement. The report identifies a core mechanism in DSH called Cordis. According to 36Kr, Cordis is intended to address a limitation of conventional plugin systems: removing a plugin may not fully remove its effects without restarting the surrounding process. That limitation matters for an agent operating with tools, permissions, sandboxes, memories and session state, because repeated modifications could otherwise interrupt a running system or leave it in an unrecoverable intermediate condition.

36Kr says Cordis tracks and revokes the effects of components when they exit and automatically re-coordinates dependencies when relationships between components change. The article connects those capabilities to the concepts of temporal and spatial composability in a paper that it says DeepSeek and Peking University published alongside the release. 36Kr reports that the underlying approach has been used for four years in Koishi, an open-source chatbot framework with more than 4,000 community plugins. The article treats that history as evidence of engineering experience, while acknowledging that the mechanism still needs testing in larger plugin ecosystems.

The supplied source does not include the paper’s title in a form that can be independently checked beyond its reference to “a programming paradigm for spatiotemporal composability,” nor does it provide benchmarks, failure rates, reproducible test procedures or evidence that Cordis has enabled DSH to improve a frontier model. The article also describes user criticism following the release, saying some users found deployment cumbersome, configuration complex and the overall experience unfinished. It then surveys related efforts at Google DeepMind, Anthropic, Tencent and MiniMax. Those examples include automated program search, agent research loops and reported internal evaluations. They are contextual claims made by 36Kr, not independently verified findings in the supplied material. The central concrete development remains the reported release of DSH and its proposed software environment for iterative agent modification.

Podrobnosti o zdroji: eu.36kr.com ↗

Proč na tom záleží

The reported launch points to a shift from evaluating AI mainly through periodic model releases toward building environments in which models can repeatedly act, receive feedback and modify their own supporting software. If the system works as described, the practical challenge is less autonomous invention than reliable experimentation: changes must be testable, reversible and judged against useful outcomes. The report does not independently establish that DSH improves a model’s core capabilities or brings AI closer to artificial general intelligence.

The significance of the reported launch lies in infrastructure for iteration. A conventional model release creates a visible comparison point: a new model is evaluated, ranked and deployed. The system described by 36Kr instead emphasizes what happens between releases. An agent operates in an environment, changes a component, runs an evaluation, observes the result and decides whether to keep or reverse the change. That approach could make model development more continuous, but only if the environment provides reliable tests and the feedback signal reflects real usefulness rather than a narrow metric.

The report’s emphasis on rollback and dependency management highlights a practical safety and reliability problem. An AI agent with authority to alter tools or scaffolding can fail in ways that are difficult to diagnose if each change contaminates the next experiment. A mechanism that records effects, removes them and restores a known state could make experimentation more controlled. The source does not establish that DSH provides comprehensive security, prevents unauthorized actions, or protects data and permissions. Those remain important unknowns, especially because the described environment may include memory, tools, sandboxes and session state.

The article also argues that access to real work and high-quality feedback could become a competitive advantage. In its account, coding environments are valuable because they generate structured activity traces and rapid signals about where an AI system succeeds or fails. This is a plausible strategic implication of the report, but it should not be treated as proof that data from user activity will produce better models. The source gives no independent evidence about data governance, consent, privacy protections, training use, or whether any reported gains generalize beyond internal tasks. Nor does it show that recursive loops produce increasing returns rather than diminishing returns.

Interactive Mechanism

Interaktivní mechanismus: Jak to vlastně funguje

Interaktivně prozkoumejte základní technologii tohoto vývoje.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Interaktivní kontrola konceptu+10 Points
AI Agents Quiz

An agent must create a draft calendar event for Tuesday at 2 p.m. Which evidence would establish the requested result?

Na co se dále dívat

The main questions are whether DeepSeek publishes technical documentation, source code, evaluations or demonstrations that substantiate the report’s claims; how reliably DSH handles failed modifications and changing dependencies; and whether it is usable outside DeepSeek’s own engineering environment. Readers should also distinguish a tool-assisted improvement loop from recursive self-improvement in the stronger sense of an AI designing and training substantially more capable successors.

Verification should begin with the release itself. A meaningful next step would be public technical documentation or a repository that specifies DSH’s supported models, operating requirements, permissions model, sandbox boundaries and recovery behavior. Independent users would need to reproduce the reported plugin installation, removal and dependency-reconciliation functions. Without such material, the public cannot determine whether DSH is broadly usable, a limited internal tool, or primarily a research prototype.

Evaluators should look for evidence that separates engineering automation from stronger claims about recursive self-improvement. Useful tests would show whether an agent can propose changes, run objective evaluations, recover from failed changes and improve performance across tasks that were not used to guide the loop. The supplied report provides no such independent evaluation. It also does not establish whether any gains affect model training, inference, tool use, software scaffolding or only a narrowly defined benchmark.

The broader claims in 36Kr’s article require particular caution. The report cites forecasts about when RSI might occur and describes efforts by several major AI companies, but predictions and company statements are not evidence that a self-improvement threshold has been reached. Future reporting should examine human oversight, shutdown authority, auditability, data provenance and the possibility that optimization loops exploit their evaluators. It should also track whether DSH receives stable public support, whether deployment complaints are resolved, and whether independent researchers can verify any practical capability gains.

Související průvodci a kvízy

Agenti AIVysvětlení modelů AIŠkolení AIBudoucnost AIOtestujte si, co víte – vyzkoušejte bezplatný kvíz AIVyhledejte si termín AI v našem slovníkuPostupujte podle sledování vydání modelu AI
Považujete to za užitečné?