返回新闻
安全AI Understanding 简报

OpenAI 破坏协调活动以提取与 Moonshot AI 相关的模型推理

OpenAI 表示,它停止了一项利用其模型重现隐藏推理的大规模努力,并将该活动的核心集群归因于中国初创公司 Moonshot AI。

4 min readRead the original reporting
Source-page capture accompanying OpenAI disrupts coordinated campaign to extract model reasoning linked to Moonshot AI
归因报告来源记录
出版商
cnbc.com
来源链接
cnbc.comhttps://www.cnbc.com/2026/10/01/openai-chinas-moonshot-ai-kimi.html
来源类型
新闻媒体的报道——不是第一方文件。

我们无法独立确认的内容: 此声明归因于指定的商店。我们没有根据第一方文件对其进行验证。 (cnbc.com)

背景60 秒内了解这一点

从这里开始

关键术语

人工智能治理
指导人工智能如何在社会中开发和使用的政策、标准和监督机制。
蒸馏
将知识从大型教师模型压缩到较小的学生模型中。
测试一下自己人工智能道德测验

发生了什么

OpenAI identified a coordinated campaign that began in early July and surged to 16,000 requests from more than 4,000 users over two days, aiming to extract protected reasoning from its AI models. The company traced related activity to a broader cluster of over 15,000 users and says it fully disrupted the campaign by July 28. OpenAI attributes a core cluster of the effort to individuals associated with Chinese startup Moonshot AI, the maker of the Kimi model. The operators did not breach OpenAI’s encryption or databases; instead they manipulated model interactions to surface hidden reasoning. OpenAI shared its findings with peers in the Frontier Model Forum and with government information‑sharing channels.

OpenAI’s internal security team detected a surge of requests in early July that appeared to be probing its models for hidden reasoning patterns. Over a two‑day period, more than 4,000 distinct users submitted roughly 16,000 requests designed to coax out internal chain‑of‑thought explanations that are not normally exposed to end‑users.

Analysis of the request logs revealed a larger network of activity involving over 15,000 user accounts. OpenAI concluded that the campaign was coordinated, with many requests sharing similar prompts and timing, suggesting an organized effort rather than isolated curiosity.

OpenAI attributes a core cluster of the activity to actors linked to Moonshot AI, the Chinese startup behind the Kimi chatbot. While the company cannot confirm that all participants were directly coordinated by Moonshot, the pattern of requests points to a group associated with that firm.

The operators did not breach OpenAI’s encryption, databases, or stored conversation logs. Instead, they leveraged the model’s interactive interface to extract reasoning that the model internally generates but typically hides from users. OpenAI describes this technique as "adversarial ," a form of model‑output exploitation.

By July 28, OpenAI says it had fully disrupted the campaign, cutting off the malicious requests and implementing additional safeguards. The findings have been shared with other AI developers through the Frontier Model Forum and with relevant government channels to help prevent similar attacks.

来源详情: cnbc.com ↗

为什么这很重要

The incident illustrates a growing threat vector known as "adversarial ," where attackers coax a model’s internal reasoning into a form that can be reused to train competing systems. If successful, such extraction could let rivals replicate advanced capabilities without the costly research and safety safeguards required to develop frontier models, raising both safety and national‑security concerns. The link to Moonshot AI underscores the geopolitical dimension of AI model theft, as U.S. firms fear that Chinese developers may accelerate their own AI programs using stolen insights. OpenAI’s rapid disruption demonstrates the importance of monitoring model‑output abuse, but the episode also reveals gaps in current defenses, as the attackers bypassed traditional data‑security measures by exploiting the model’s interactive interface.

Adversarial represents a novel attack surface that bypasses traditional data‑security controls, highlighting the need for new defensive measures that monitor and limit the type of reasoning a model can reveal during user interactions.

If attackers can reliably extract hidden reasoning, they could use that information to train rival models that mimic the capabilities of frontier systems without incurring the same research and development costs, potentially eroding the competitive advantage of firms that invest heavily in safety and alignment.

The involvement of a Chinese AI firm adds a geopolitical layer, as U.S. policymakers have expressed concern about technology transfer that could accelerate AI development in rival nations, raising national‑security implications.

OpenAI’s rapid response and information‑sharing demonstrate an emerging industry practice of collective threat intelligence, which could become a critical component of as similar attacks become more sophisticated.

Interactive Mechanism

互动机制:它实际上是如何运作的

以交互方式探索这一发展背后的基础技术。

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
交互式概念检查+10 Points
AI Ethics Quiz

Impossibility results in algorithmic fairness (e.g. Kleinberg et al., Chouldechova) show what?

接下来看什么

Future disclosures from OpenAI about additional mitigation techniques for adversarial , and any policy responses from U.S. or Chinese regulators. Watch for follow‑up statements from Moonshot AI and potential legal or diplomatic actions. Monitor the Frontier Model Forum for coordinated industry guidelines on protecting model reasoning and for any broader collaboration on threat intelligence sharing.

OpenAI may publish additional technical details or mitigation strategies for adversarial , which could influence how other AI labs design their model interfaces and output controls.

Regulatory bodies in the United States and China might consider new rules or guidelines addressing model‑output abuse, especially if the technique is deemed a national‑security risk.

Moonshot AI’s response, if any, could signal whether the company acknowledges involvement or disputes the attribution, potentially leading to diplomatic or legal repercussions.

The Frontier Model Forum may develop industry‑wide standards for monitoring and limiting model reasoning exposure, and future collaborations could include joint research on detection algorithms.

相关指南和测验

AI 伦理人工智能模型解释人工智能培训测试你所知道的——尝试免费的人工智能测验在我们的词汇表中查找人工智能术语关注AI监管追踪器
觉得这有用吗?