Volver a Noticias
SeguridadAI Understanding sesión informativa

OpenAI interrumpe la campaña coordinada para extraer el razonamiento del modelo vinculado a Moonshot AI

OpenAI dice que detuvo un esfuerzo a gran escala que utilizaba sus modelos para reproducir razonamientos ocultos, atribuyendo un grupo central de la actividad a la startup china Moonshot AI.

4 min readRead the original reporting
Source-page capture accompanying OpenAI disrupts coordinated campaign to extract model reasoning linked to Moonshot AI
Informes atribuidosFuente registrada
Editor
cnbc.com
Enlace fuente
cnbc.comhttps://www.cnbc.com/2026/10/01/openai-chinas-moonshot-ai-kimi.html
Tipo de fuente
Informe de un medio de comunicación, no un documento propio.

Lo que no pudimos confirmar de forma independiente: Este reclamo se atribuye al medio mencionado. No lo verificamos con un documento de origen. (cnbc.com)

ContextoEntiende esto en 60 segundos

Empieza aquí

Términos clave

Gobernanza de la IA
Políticas, estándares y mecanismos de supervisión que guían cómo se desarrolla y utiliza la IA en la sociedad.
Destilación
Comprimir el conocimiento de un modelo de profesor grande a un modelo de estudiante más pequeño.
Ponte a pruebaPrueba de ética de la IA

que paso

OpenAI identified a coordinated campaign that began in early July and surged to 16,000 requests from more than 4,000 users over two days, aiming to extract protected reasoning from its AI models. The company traced related activity to a broader cluster of over 15,000 users and says it fully disrupted the campaign by July 28. OpenAI attributes a core cluster of the effort to individuals associated with Chinese startup Moonshot AI, the maker of the Kimi model. The operators did not breach OpenAI’s encryption or databases; instead they manipulated model interactions to surface hidden reasoning. OpenAI shared its findings with peers in the Frontier Model Forum and with government information‑sharing channels.

OpenAI’s internal security team detected a surge of requests in early July that appeared to be probing its models for hidden reasoning patterns. Over a two‑day period, more than 4,000 distinct users submitted roughly 16,000 requests designed to coax out internal chain‑of‑thought explanations that are not normally exposed to end‑users.

Analysis of the request logs revealed a larger network of activity involving over 15,000 user accounts. OpenAI concluded that the campaign was coordinated, with many requests sharing similar prompts and timing, suggesting an organized effort rather than isolated curiosity.

OpenAI attributes a core cluster of the activity to actors linked to Moonshot AI, the Chinese startup behind the Kimi chatbot. While the company cannot confirm that all participants were directly coordinated by Moonshot, the pattern of requests points to a group associated with that firm.

The operators did not breach OpenAI’s encryption, databases, or stored conversation logs. Instead, they leveraged the model’s interactive interface to extract reasoning that the model internally generates but typically hides from users. OpenAI describes this technique as "adversarial ," a form of model‑output exploitation.

By July 28, OpenAI says it had fully disrupted the campaign, cutting off the malicious requests and implementing additional safeguards. The findings have been shared with other AI developers through the Frontier Model Forum and with relevant government channels to help prevent similar attacks.

Detalles de la fuente: cnbc.com ↗

Por qué es importante

The incident illustrates a growing threat vector known as "adversarial ," where attackers coax a model’s internal reasoning into a form that can be reused to train competing systems. If successful, such extraction could let rivals replicate advanced capabilities without the costly research and safety safeguards required to develop frontier models, raising both safety and national‑security concerns. The link to Moonshot AI underscores the geopolitical dimension of AI model theft, as U.S. firms fear that Chinese developers may accelerate their own AI programs using stolen insights. OpenAI’s rapid disruption demonstrates the importance of monitoring model‑output abuse, but the episode also reveals gaps in current defenses, as the attackers bypassed traditional data‑security measures by exploiting the model’s interactive interface.

Adversarial represents a novel attack surface that bypasses traditional data‑security controls, highlighting the need for new defensive measures that monitor and limit the type of reasoning a model can reveal during user interactions.

If attackers can reliably extract hidden reasoning, they could use that information to train rival models that mimic the capabilities of frontier systems without incurring the same research and development costs, potentially eroding the competitive advantage of firms that invest heavily in safety and alignment.

The involvement of a Chinese AI firm adds a geopolitical layer, as U.S. policymakers have expressed concern about technology transfer that could accelerate AI development in rival nations, raising national‑security implications.

OpenAI’s rapid response and information‑sharing demonstrate an emerging industry practice of collective threat intelligence, which could become a critical component of as similar attacks become more sophisticated.

Interactive Mechanism

Mecanismo interactivo: cómo funciona realmente

Explore la tecnología subyacente detrás de este desarrollo de forma interactiva.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Verificación interactiva del concepto+10 Points
AI Ethics Quiz

Impossibility results in algorithmic fairness (e.g. Kleinberg et al., Chouldechova) show what?

Qué ver a continuación

Future disclosures from OpenAI about additional mitigation techniques for adversarial , and any policy responses from U.S. or Chinese regulators. Watch for follow‑up statements from Moonshot AI and potential legal or diplomatic actions. Monitor the Frontier Model Forum for coordinated industry guidelines on protecting model reasoning and for any broader collaboration on threat intelligence sharing.

OpenAI may publish additional technical details or mitigation strategies for adversarial , which could influence how other AI labs design their model interfaces and output controls.

Regulatory bodies in the United States and China might consider new rules or guidelines addressing model‑output abuse, especially if the technique is deemed a national‑security risk.

Moonshot AI’s response, if any, could signal whether the company acknowledges involvement or disputes the attribution, potentially leading to diplomatic or legal repercussions.

The Frontier Model Forum may develop industry‑wide standards for monitoring and limiting model reasoning exposure, and future collaborations could include joint research on detection algorithms.

Guías y cuestionarios relacionados

Ética de la IAModelos de IA explicadosEntrenamiento de IAPon a prueba lo que sabes: prueba un cuestionario gratuito sobre IABusque un término de IA en nuestro glosarioSiga el rastreador de regulaciones de IA
¿Encontró esto útil?