返回新闻
产品展示AI Understanding 简报

SpaceX launches Grok 4.7 with new safety and agent features

SpaceX introduced Grok 4.7, its latest large language model, featuring improved cost-efficiency, multi-agent capabilities, and enhanced safety benchmarks for malicious request blocking.

4 min readRead the linked source
Source-provided image accompanying SpaceX launches Grok 4.7 with new safety and agent features
来源参考来源记录
出版商
siliconangle.com
来源链接
siliconangle.comhttps://siliconangle.com/2026/09/21/spacex-launches-grok-4-7-with-long-horizon-processing-safety-upgrades/
来源类型
链接来源——主要来源状态尚未确定。
背景60 秒内了解这一点

从这里开始

关键术语

大语言模型(LLM)
在海量文本语料库上训练来生成和分析文本的语言模型。
强化学习
通过奖励信号进行训练,代理学习能够最大化长期回报的行动。
XAI(可解释的人工智能)
使人工智能预测更加透明和易于理解的技术和实践。
测试一下自己AI 模型解释测验

发生了什么

SpaceX launched Grok 4.7, a new large language model that the company describes as its most capable to date. The release includes a new base model, enhanced , and compatibility with the Grok Bot harness for parallel multi-agent tasks. The model is available for purchase with specific pricing tiers and offers a faster, higher-cost variant for latency-sensitive workloads.

SpaceX Corp. introduced Grok 4.7, identifying it as its most capable large language model to date. The model series originated from xAI Corp., which was merged with xAI Inc. and subsequently folded into SpaceX, bringing the Grok models and associated AI data centers under the SpaceX umbrella.

The company evaluated Grok 4.7 using CursorBench 4.0, a benchmark developed by Cursor, a recent SpaceX acquisition. According to SiliconANGLE, Grok 4.7 completed these challenges at an average cost of $4.69 per task, outperforming GPT-5.6 Sol and Fable 5.1 in cost-efficiency. SpaceX noted that it evaluated hardware-intensive versions of the competing models that prioritize output quality over cost.

On more widely used benchmarks, Grok 4.7 outperformed Fable 5.1 on the Harvey Legal Agent Benchmark and EEBench, which covers chip design tasks. However, the model scored lower than GPT-6 Astra, OpenAI’s latest LLM, on the EEBench. SpaceX attributes the performance gains to a new base model and an enhanced workflow that involved tougher training tasks and longer processing times.

The model is built to work with the Grok Bot harness, a set of resources allowing Grok 4.7 to split complex work among multiple AI agents. These agents can perform tasks in parallel and verify each other’s output. SpaceX also reported that Grok 4.7 set records on LatchBio and HackerBench, benchmarks testing the model's ability to block malicious biology research and cybersecurity requests.

Grok 4.7 is available starting at $2 per million input tokens and $6 per million output tokens. A version for latency-sensitive workloads is offered that processes prompts twice as fast but costs twice as much. This launch occurred less than a week after SpaceX released Grok Voice Transcribe 2.0, a text-to-speech model with improved accuracy and lower cost.

来源详情: siliconangle.com

为什么这很重要

The launch represents a significant consolidation of AI capabilities within SpaceX, following the merger of xAI and the acquisition of Cursor. By integrating advanced multi-agent processing and specific safety benchmarks for biology and cybersecurity, SpaceX is positioning Grok 4.7 as a competitive alternative to leading models from OpenAI and other providers. The explicit pricing structure and availability of a speed-optimized variant provide concrete data points for enterprise adoption decisions, while the reported performance on specialized benchmarks like Harvey Legal Agent and EEBench suggests targeted improvements in professional and technical domains.

The integration of xAI into SpaceX has resulted in a unified AI strategy that combines model development with significant infrastructure and acquisition assets, such as Cursor. This consolidation allows for tighter integration between model capabilities and specialized tools, potentially creating a more cohesive ecosystem for enterprise users.

The specific focus on cost-efficiency in the CursorBench 4.0 evaluation highlights a shift in competitive dynamics where operational cost per task is becoming a key differentiator alongside raw performance. The reported $4.69 per task average provides a concrete metric for businesses evaluating the total cost of ownership for AI-driven workflows.

The introduction of the Grok Bot harness for multi-agent parallel processing addresses a common limitation in single-agent LLMs, which often struggle with complex, multi-step tasks. By enabling agents to verify each other’s output, SpaceX is attempting to improve reliability and speed, which is critical for high-stakes applications in legal and technical fields.

The safety upgrades, specifically the record-setting performance on LatchBio and HackerBench, are significant for organizations concerned with AI misuse. Demonstrating the ability to block malicious requests in biology and cybersecurity domains is a practical safety feature that may influence procurement decisions for regulated industries.

Interactive Mechanism

互动机制:它实际上是如何运作的

以交互方式探索这一发展背后的基础技术。

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
交互式概念检查+10 Points
AI Models Explained Quiz

In AI, what are a model's "parameters"?

接下来看什么

Independent verification of the benchmark claims, particularly regarding the cost-efficiency comparison with GPT-5.6 Sol and Fable 5.1, is currently lacking. Users should monitor for third-party evaluations that confirm the model's performance on legal and chip design tasks. Additionally, the practical impact of the Grok Bot harness on real-world multi-agent workflows and the effectiveness of the new safety safeguards against evolving malicious prompts will be key indicators of the model's long-term viability and safety profile.

Independent third-party benchmarks are needed to verify SpaceX’s claims regarding cost-efficiency and performance on specialized tasks like legal and chip design. Current data relies solely on SpaceX’s internal evaluations and the specific benchmark developed by its acquisition, Cursor.

The practical utility of the Grok Bot harness in real-world scenarios will determine if the multi-agent approach offers a tangible advantage over existing single-agent or simpler multi-agent frameworks. User feedback on the reliability of the cross-verification process will be a key indicator.

The effectiveness of the new safety safeguards against novel or evolving malicious prompts will be tested over time. While current benchmarks show strong performance, the dynamic nature of AI security threats means that ongoing monitoring and updates will be necessary to maintain these safety records.

相关指南和测验

人工智能模型解释人工智能代理AI 伦理测试你所知道的——尝试免费的人工智能测验在我们的词汇表中查找人工智能术语
觉得这有用吗?