Haberlere Geri Dön
ÜrünAI Understanding brifing

SpaceX launches Grok 4.7 with new safety and agent features

SpaceX introduced Grok 4.7, its latest large language model, featuring improved cost-efficiency, multi-agent capabilities, and enhanced safety benchmarks for malicious request blocking.

4 min readRead the linked source
Source-provided image accompanying SpaceX launches Grok 4.7 with new safety and agent features
Kaynak referansıKaynak kaydedildi
Yayıncı
siliconangle.com
Kaynak bağlantısı
siliconangle.comhttps://siliconangle.com/2026/09/21/spacex-launches-grok-4-7-with-long-horizon-processing-safety-upgrades/
Kaynak türü
Bağlantılı kaynak — birincil kaynak durumu belirlenmedi.
Bağlam60 saniyede bunu anlayın

Buradan başlayın

Anahtar terimler

Büyük Dil Modeli (LLM)
Metin oluşturmak ve analiz etmek için çok büyük metin toplulukları üzerinde eğitilmiş bir dil modeli.
Takviyeli Öğrenme
Bir aracının uzun vadeli getiriyi en üst düzeye çıkaracak eylemleri öğrendiği ödül sinyalleriyle eğitim.
XAI (Açıklanabilir Yapay Zeka)
Yapay zeka tahminlerini daha şeffaf ve anlaşılır hale getirmeye yönelik teknikler ve uygulamalar.
Kendinizi test edinYapay Zeka Modelleri Açıklaması Testi

Ne oldu?

SpaceX launched Grok 4.7, a new large language model that the company describes as its most capable to date. The release includes a new base model, enhanced , and compatibility with the Grok Bot harness for parallel multi-agent tasks. The model is available for purchase with specific pricing tiers and offers a faster, higher-cost variant for latency-sensitive workloads.

SpaceX Corp. introduced Grok 4.7, identifying it as its most capable large language model to date. The model series originated from xAI Corp., which was merged with xAI Inc. and subsequently folded into SpaceX, bringing the Grok models and associated AI data centers under the SpaceX umbrella.

The company evaluated Grok 4.7 using CursorBench 4.0, a benchmark developed by Cursor, a recent SpaceX acquisition. According to SiliconANGLE, Grok 4.7 completed these challenges at an average cost of $4.69 per task, outperforming GPT-5.6 Sol and Fable 5.1 in cost-efficiency. SpaceX noted that it evaluated hardware-intensive versions of the competing models that prioritize output quality over cost.

On more widely used benchmarks, Grok 4.7 outperformed Fable 5.1 on the Harvey Legal Agent Benchmark and EEBench, which covers chip design tasks. However, the model scored lower than GPT-6 Astra, OpenAI’s latest LLM, on the EEBench. SpaceX attributes the performance gains to a new base model and an enhanced workflow that involved tougher training tasks and longer processing times.

The model is built to work with the Grok Bot harness, a set of resources allowing Grok 4.7 to split complex work among multiple AI agents. These agents can perform tasks in parallel and verify each other’s output. SpaceX also reported that Grok 4.7 set records on LatchBio and HackerBench, benchmarks testing the model's ability to block malicious biology research and cybersecurity requests.

Grok 4.7 is available starting at $2 per million input tokens and $6 per million output tokens. A version for latency-sensitive workloads is offered that processes prompts twice as fast but costs twice as much. This launch occurred less than a week after SpaceX released Grok Voice Transcribe 2.0, a text-to-speech model with improved accuracy and lower cost.

Kaynak ayrıntıları: siliconangle.com

Neden önemli?

The launch represents a significant consolidation of AI capabilities within SpaceX, following the merger of xAI and the acquisition of Cursor. By integrating advanced multi-agent processing and specific safety benchmarks for biology and cybersecurity, SpaceX is positioning Grok 4.7 as a competitive alternative to leading models from OpenAI and other providers. The explicit pricing structure and availability of a speed-optimized variant provide concrete data points for enterprise adoption decisions, while the reported performance on specialized benchmarks like Harvey Legal Agent and EEBench suggests targeted improvements in professional and technical domains.

The integration of xAI into SpaceX has resulted in a unified AI strategy that combines model development with significant infrastructure and acquisition assets, such as Cursor. This consolidation allows for tighter integration between model capabilities and specialized tools, potentially creating a more cohesive ecosystem for enterprise users.

The specific focus on cost-efficiency in the CursorBench 4.0 evaluation highlights a shift in competitive dynamics where operational cost per task is becoming a key differentiator alongside raw performance. The reported $4.69 per task average provides a concrete metric for businesses evaluating the total cost of ownership for AI-driven workflows.

The introduction of the Grok Bot harness for multi-agent parallel processing addresses a common limitation in single-agent LLMs, which often struggle with complex, multi-step tasks. By enabling agents to verify each other’s output, SpaceX is attempting to improve reliability and speed, which is critical for high-stakes applications in legal and technical fields.

The safety upgrades, specifically the record-setting performance on LatchBio and HackerBench, are significant for organizations concerned with AI misuse. Demonstrating the ability to block malicious requests in biology and cybersecurity domains is a practical safety feature that may influence procurement decisions for regulated industries.

Interactive Mechanism

İnteraktif Mekanizma: Aslında Nasıl Çalışıyor?

Bu gelişmenin arkasında yatan teknolojiyi etkileşimli olarak keşfedin.

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
İnteraktif Konsept Kontrolü+10 Points
AI Models Explained Quiz

In AI, what are a model's "parameters"?

Bundan sonra ne izlenecek?

Independent verification of the benchmark claims, particularly regarding the cost-efficiency comparison with GPT-5.6 Sol and Fable 5.1, is currently lacking. Users should monitor for third-party evaluations that confirm the model's performance on legal and chip design tasks. Additionally, the practical impact of the Grok Bot harness on real-world multi-agent workflows and the effectiveness of the new safety safeguards against evolving malicious prompts will be key indicators of the model's long-term viability and safety profile.

Independent third-party benchmarks are needed to verify SpaceX’s claims regarding cost-efficiency and performance on specialized tasks like legal and chip design. Current data relies solely on SpaceX’s internal evaluations and the specific benchmark developed by its acquisition, Cursor.

The practical utility of the Grok Bot harness in real-world scenarios will determine if the multi-agent approach offers a tangible advantage over existing single-agent or simpler multi-agent frameworks. User feedback on the reliability of the cross-verification process will be a key indicator.

The effectiveness of the new safety safeguards against novel or evolving malicious prompts will be tested over time. While current benchmarks show strong performance, the dynamic nature of AI security threats means that ongoing monitoring and updates will be necessary to maintain these safety records.

İlgili kılavuzlar ve testler

Yapay Zeka Modellerinin AçıklamasıYapay Zeka AracılarıYapay Zeka EtiğiBildiklerinizi test edin; ücretsiz bir yapay zeka testini deneyinSözlüğümüzde bir yapay zeka terimine bakın
Bunu yararlı buldunuz mu?