返回新聞
產品展示AI Understanding 簡報

SpaceX launches Grok 4.7 with new safety and agent features

SpaceX introduced Grok 4.7, its latest large language model, featuring improved cost-efficiency, multi-agent capabilities, and enhanced safety benchmarks for malicious request blocking.

4 min readRead the linked source
Source-provided image accompanying SpaceX launches Grok 4.7 with new safety and agent features
來源參考來源記錄
出版商
siliconangle.com
來源連結
siliconangle.comhttps://siliconangle.com/2026/09/21/spacex-launches-grok-4-7-with-long-horizon-processing-safety-upgrades/
來源類型
連結來源-主要來源狀態尚未確定。
背景60 秒內了解這一點

從這裡開始

關鍵術語

大語言模型(LLM)
在海量文本語料庫上訓練來產生和分析文本的語言模型。
強化學習
透過獎勵訊號進行訓練,代理學習能夠最大化長期回報的行動。
XAI(可解釋的人工智慧)
使人工智慧預測更加透明和易於理解的技術和實踐。
測試一下自己AI 模型解釋測驗

發生了什麼事

SpaceX launched Grok 4.7, a new large language model that the company describes as its most capable to date. The release includes a new base model, enhanced , and compatibility with the Grok Bot harness for parallel multi-agent tasks. The model is available for purchase with specific pricing tiers and offers a faster, higher-cost variant for latency-sensitive workloads.

SpaceX Corp. introduced Grok 4.7, identifying it as its most capable large language model to date. The model series originated from xAI Corp., which was merged with xAI Inc. and subsequently folded into SpaceX, bringing the Grok models and associated AI data centers under the SpaceX umbrella.

The company evaluated Grok 4.7 using CursorBench 4.0, a benchmark developed by Cursor, a recent SpaceX acquisition. According to SiliconANGLE, Grok 4.7 completed these challenges at an average cost of $4.69 per task, outperforming GPT-5.6 Sol and Fable 5.1 in cost-efficiency. SpaceX noted that it evaluated hardware-intensive versions of the competing models that prioritize output quality over cost.

On more widely used benchmarks, Grok 4.7 outperformed Fable 5.1 on the Harvey Legal Agent Benchmark and EEBench, which covers chip design tasks. However, the model scored lower than GPT-6 Astra, OpenAI’s latest LLM, on the EEBench. SpaceX attributes the performance gains to a new base model and an enhanced workflow that involved tougher training tasks and longer processing times.

The model is built to work with the Grok Bot harness, a set of resources allowing Grok 4.7 to split complex work among multiple AI agents. These agents can perform tasks in parallel and verify each other’s output. SpaceX also reported that Grok 4.7 set records on LatchBio and HackerBench, benchmarks testing the model's ability to block malicious biology research and cybersecurity requests.

Grok 4.7 is available starting at $2 per million input tokens and $6 per million output tokens. A version for latency-sensitive workloads is offered that processes prompts twice as fast but costs twice as much. This launch occurred less than a week after SpaceX released Grok Voice Transcribe 2.0, a text-to-speech model with improved accuracy and lower cost.

來源詳情: siliconangle.com

為什麼這很重要

The launch represents a significant consolidation of AI capabilities within SpaceX, following the merger of xAI and the acquisition of Cursor. By integrating advanced multi-agent processing and specific safety benchmarks for biology and cybersecurity, SpaceX is positioning Grok 4.7 as a competitive alternative to leading models from OpenAI and other providers. The explicit pricing structure and availability of a speed-optimized variant provide concrete data points for enterprise adoption decisions, while the reported performance on specialized benchmarks like Harvey Legal Agent and EEBench suggests targeted improvements in professional and technical domains.

The integration of xAI into SpaceX has resulted in a unified AI strategy that combines model development with significant infrastructure and acquisition assets, such as Cursor. This consolidation allows for tighter integration between model capabilities and specialized tools, potentially creating a more cohesive ecosystem for enterprise users.

The specific focus on cost-efficiency in the CursorBench 4.0 evaluation highlights a shift in competitive dynamics where operational cost per task is becoming a key differentiator alongside raw performance. The reported $4.69 per task average provides a concrete metric for businesses evaluating the total cost of ownership for AI-driven workflows.

The introduction of the Grok Bot harness for multi-agent parallel processing addresses a common limitation in single-agent LLMs, which often struggle with complex, multi-step tasks. By enabling agents to verify each other’s output, SpaceX is attempting to improve reliability and speed, which is critical for high-stakes applications in legal and technical fields.

The safety upgrades, specifically the record-setting performance on LatchBio and HackerBench, are significant for organizations concerned with AI misuse. Demonstrating the ability to block malicious requests in biology and cybersecurity domains is a practical safety feature that may influence procurement decisions for regulated industries.

Interactive Mechanism

互動機制:它實際上是如何運作的

以互動方式探索這項發展背後的基礎技術。

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
互動式概念檢查+10 Points
AI Models Explained Quiz

In AI, what are a model's "parameters"?

接下來看什麼

Independent verification of the benchmark claims, particularly regarding the cost-efficiency comparison with GPT-5.6 Sol and Fable 5.1, is currently lacking. Users should monitor for third-party evaluations that confirm the model's performance on legal and chip design tasks. Additionally, the practical impact of the Grok Bot harness on real-world multi-agent workflows and the effectiveness of the new safety safeguards against evolving malicious prompts will be key indicators of the model's long-term viability and safety profile.

Independent third-party benchmarks are needed to verify SpaceX’s claims regarding cost-efficiency and performance on specialized tasks like legal and chip design. Current data relies solely on SpaceX’s internal evaluations and the specific benchmark developed by its acquisition, Cursor.

The practical utility of the Grok Bot harness in real-world scenarios will determine if the multi-agent approach offers a tangible advantage over existing single-agent or simpler multi-agent frameworks. User feedback on the reliability of the cross-verification process will be a key indicator.

The effectiveness of the new safety safeguards against novel or evolving malicious prompts will be tested over time. While current benchmarks show strong performance, the dynamic nature of AI security threats means that ongoing monitoring and updates will be necessary to maintain these safety records.

相關指引和測驗

人工智慧模型解釋人工智慧代理AI 倫理測試你所知道的—嘗試免費的人工智慧測驗在我們的詞彙表中尋找人工智慧術語
覺得有用嗎?