Komawa Labarai
SamfuraAI Understanding takaitaccen bayani

SpaceX launches Grok 4.7 with new safety and agent features

SpaceX introduced Grok 4.7, its latest large language model, featuring improved cost-efficiency, multi-agent capabilities, and enhanced safety benchmarks for malicious request blocking.

4 min readRead the linked source
Source-provided image accompanying SpaceX launches Grok 4.7 with new safety and agent features
Tushen tusheAn rubuta tushen tushe
Mawallafi
siliconangle.com
Tushen hanyar haɗin gwiwa
siliconangle.comhttps://siliconangle.com/2026/09/21/spacex-launches-grok-4-7-with-long-horizon-processing-safety-upgrades/
Nau'in tushe
Tushen da aka haɗa - ba a kafa matsayin tushen farko ba.
MaganaFahimtar wannan a cikin daƙiƙa 60

Fara a nan

Mabuɗin sharuddan

Babban Samfurin Harshe (LLM)
Samfurin harshe da aka horar akan babban haɗin gwiwar rubutu don samarwa da tantance rubutu.
Ƙarfafa Koyo
Horowa ta siginar lada inda wakili ke koyon ayyuka waɗanda ke haɓaka dawowa na dogon lokaci.
XAI (Bayyana AI)
Dabaru da ayyuka don yin hasashen AI mafi fahimi da fahimta.
Gwada kankaAI Model An Bayyana Tambayoyi

Me ya faru

SpaceX launched Grok 4.7, a new large language model that the company describes as its most capable to date. The release includes a new base model, enhanced , and compatibility with the Grok Bot harness for parallel multi-agent tasks. The model is available for purchase with specific pricing tiers and offers a faster, higher-cost variant for latency-sensitive workloads.

SpaceX Corp. introduced Grok 4.7, identifying it as its most capable large language model to date. The model series originated from xAI Corp., which was merged with xAI Inc. and subsequently folded into SpaceX, bringing the Grok models and associated AI data centers under the SpaceX umbrella.

The company evaluated Grok 4.7 using CursorBench 4.0, a benchmark developed by Cursor, a recent SpaceX acquisition. According to SiliconANGLE, Grok 4.7 completed these challenges at an average cost of $4.69 per task, outperforming GPT-5.6 Sol and Fable 5.1 in cost-efficiency. SpaceX noted that it evaluated hardware-intensive versions of the competing models that prioritize output quality over cost.

On more widely used benchmarks, Grok 4.7 outperformed Fable 5.1 on the Harvey Legal Agent Benchmark and EEBench, which covers chip design tasks. However, the model scored lower than GPT-6 Astra, OpenAI’s latest LLM, on the EEBench. SpaceX attributes the performance gains to a new base model and an enhanced workflow that involved tougher training tasks and longer processing times.

The model is built to work with the Grok Bot harness, a set of resources allowing Grok 4.7 to split complex work among multiple AI agents. These agents can perform tasks in parallel and verify each other’s output. SpaceX also reported that Grok 4.7 set records on LatchBio and HackerBench, benchmarks testing the model's ability to block malicious biology research and cybersecurity requests.

Grok 4.7 is available starting at $2 per million input tokens and $6 per million output tokens. A version for latency-sensitive workloads is offered that processes prompts twice as fast but costs twice as much. This launch occurred less than a week after SpaceX released Grok Voice Transcribe 2.0, a text-to-speech model with improved accuracy and lower cost.

Bayanan tushe: siliconangle.com

Me ya sa yake da mahimmanci

The launch represents a significant consolidation of AI capabilities within SpaceX, following the merger of xAI and the acquisition of Cursor. By integrating advanced multi-agent processing and specific safety benchmarks for biology and cybersecurity, SpaceX is positioning Grok 4.7 as a competitive alternative to leading models from OpenAI and other providers. The explicit pricing structure and availability of a speed-optimized variant provide concrete data points for enterprise adoption decisions, while the reported performance on specialized benchmarks like Harvey Legal Agent and EEBench suggests targeted improvements in professional and technical domains.

The integration of xAI into SpaceX has resulted in a unified AI strategy that combines model development with significant infrastructure and acquisition assets, such as Cursor. This consolidation allows for tighter integration between model capabilities and specialized tools, potentially creating a more cohesive ecosystem for enterprise users.

The specific focus on cost-efficiency in the CursorBench 4.0 evaluation highlights a shift in competitive dynamics where operational cost per task is becoming a key differentiator alongside raw performance. The reported $4.69 per task average provides a concrete metric for businesses evaluating the total cost of ownership for AI-driven workflows.

The introduction of the Grok Bot harness for multi-agent parallel processing addresses a common limitation in single-agent LLMs, which often struggle with complex, multi-step tasks. By enabling agents to verify each other’s output, SpaceX is attempting to improve reliability and speed, which is critical for high-stakes applications in legal and technical fields.

The safety upgrades, specifically the record-setting performance on LatchBio and HackerBench, are significant for organizations concerned with AI misuse. Demonstrating the ability to block malicious requests in biology and cybersecurity domains is a practical safety feature that may influence procurement decisions for regulated industries.

Interactive Mechanism

Ingantacciyar hanyar sadarwa: Yadda A zahiri yake Aiki

Bincika fasahar da ke bayan wannan ci gaban ta hanyar mu'amala.

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
Duba ra'ayi na hulɗa+10 Points
AI Models Explained Quiz

In AI, what are a model's "parameters"?

Abin kallo na gaba

Independent verification of the benchmark claims, particularly regarding the cost-efficiency comparison with GPT-5.6 Sol and Fable 5.1, is currently lacking. Users should monitor for third-party evaluations that confirm the model's performance on legal and chip design tasks. Additionally, the practical impact of the Grok Bot harness on real-world multi-agent workflows and the effectiveness of the new safety safeguards against evolving malicious prompts will be key indicators of the model's long-term viability and safety profile.

Independent third-party benchmarks are needed to verify SpaceX’s claims regarding cost-efficiency and performance on specialized tasks like legal and chip design. Current data relies solely on SpaceX’s internal evaluations and the specific benchmark developed by its acquisition, Cursor.

The practical utility of the Grok Bot harness in real-world scenarios will determine if the multi-agent approach offers a tangible advantage over existing single-agent or simpler multi-agent frameworks. User feedback on the reliability of the cross-verification process will be a key indicator.

The effectiveness of the new safety safeguards against novel or evolving malicious prompts will be tested over time. While current benchmarks show strong performance, the dynamic nature of AI security threats means that ongoing monitoring and updates will be necessary to maintain these safety records.

Jagorori masu alaƙa & tambayoyin tambayoyi

AI Model ya bayyanaWakilan AIƊa'a ta AIGwada abin da kuka sani - gwada gwajin AI kyautaNemo kalmar AI a cikin ƙamus ɗin mu
An sami wannan yana da amfani?