العودة إلى الأخبار
المنتجAI Understanding إحاطة

تطلق SpaceXAI إصدار Grok 4.6 لوكلاء الترميز طويلي التشغيل

تقول SpaceXAI إن Grok 4.6 متاح الآن مع نافذة سياق مكونة من 500000 رمز مميز، وضوابط تفكير موسعة، وتدريب يهدف إلى الترميز المستمر والبحث وعمل المشروع التفاعلي.

6 min readRead the primary source
وثيقة المصدر الأساسيتم تسجيل المصدر
الناشر
SpaceXAI's August 12 Grok 4.6 announcement and API documentation
رابط المصدر
docs.x.aihttps://docs.x.ai/developers/grok-4-6
نوع المصدر
المستند الأساسي - إعلان رسمي أو ورقة أو ملف أو صفحة الطرف الأول التي نقرأها مباشرة.
السياقافهم هذا في 60 ثانية

ابدأ هنا

المصطلحات الرئيسية

API (واجهة برمجة التطبيقات)
طريقة منظمة لنظام برمجي واحد لإرسال الطلبات إلى نظام آخر وتلقي الاستجابات منه.
XAI (الذكاء الاصطناعي القابل للتفسير)
تقنيات وممارسات لجعل تنبؤات الذكاء الاصطناعي أكثر شفافية وقابلية للفهم.
مصدر البيانات
الأصل الموثق والملكية والتاريخ لمجموعة البيانات أو القطعة الأثرية النموذجية.
اختبر نفسكمسابقة وكلاء الذكاء الاصطناعي

ماذا حدث

أصدرت SpaceXAI Grok 4.6 في 12 أغسطس كنموذج حدودي يهدف إلى البرمجة والمهام الوكيلة والعمل المعرفي. وتقول الشركة إن النموذج مصمم ليظل فعالاً عبر وظائف أطول ومتعددة الخطوات: البحث عن موضوع غير مألوف، وتحليل المعلومات، وتغيير قاعدة التعليمات البرمجية، وتحويل الفكرة إلى تطبيق عملي أو أي قطعة أثرية مصقولة أخرى. الإصدار متاح من خلال xAI API وGrok Build وCursor والبوابات النموذجية بما في ذلك OpenRouter وVercel وCloudflare. إنه منتج تم شحنه وليس إعلانًا لخارطة الطريق، على الرغم من أن معظم أدلة الأداء المنشورة حتى الآن تأتي من SpaceXAI نفسها.

تحدد وثائق واجهة برمجة التطبيقات (API) الرسمية النموذج على أنه `grok-4.6` وتمنحه نافذة سياق مكونة من 500000 رمز مميز. فهو يقبل إدخالات النص والصور وينتج إخراج النص، بدون حد محدد لإخراج النص. يمكن للمطورين اختيار جهد تفكير منخفض أو متوسط ​​أو مرتفع أو مرتفع جدًا. يسرد SpaceXAI التسعير الأساسي أقل من 200000 رمز سريع بسعر 2 دولار لكل مليون رمز إدخال، و0.50 دولار لكل مليون رمز إدخال مخبأ، و6 دولارات لكل مليون رمز إخراج؛ تكلف المطالبات التي تتجاوز هذا الحد 4 دولارات و1 دولارًا و12 دولارًا على التوالي. البديل السريع يكلف ضعف المعدلات الأساسية. هذه هي أسعار الإطلاق ومواصفات المنتج، وليست ضمانًا لزمن الوصول أو التوفر أو إجمالي تكلفة عبء العمل.

يقول SpaceXAI إن Grok 4.6 تلقى تدريبًا تكميليًا أطول من Grok 4.5. تصف الشركة المواد المنسقة التي تم إنشاؤها للنموذج للاستدلال والمفاهيم التقنية المتقدمة، والبيانات الهندسية عالية الجودة، والتغييرات على المُحسِّن ووصفة التدريب. ثم استخدمت Grok 4.5 لتجديد مسارات الضبط الدقيق الخاضعة للإشراف عبر إعدادات الاستدلال، وأدوات الوكلاء، والعلوم والتكنولوجيا والهندسة والرياضيات، وهندسة البرمجيات، والعمل المعرفي، مع عمليات التحقق القائمة على النموذج التي تعمل على تصفية الآثار الإشكالية. يقال إن بيئات التعلم المعزز غطت الترميز العام، والعمل المعرفي، وتحسين النواة، وتطوير الويب، والتصميم بمساعدة الكمبيوتر. لا يكشف الإعلان عن عدد المعلمات أو حساب التدريب أو مصدر البيانات الكامل أو تفاصيل التنفيذ الكافية لمجموعة خارجية لإعادة إنتاج عملية التدريب.

The launch emphasizes sustained work and product creation rather than one isolated coding answer. SpaceXAI says internal projects showed stronger first passes on visual and interactive applications than Grok 4.5, followed by iterative refinement and more self-testing on longer trajectories. That is a useful description of the intended behavior, but the public examples are selected demonstrations. They do not establish how often the model detects its own errors, whether verification catches subtle regressions, or how performance changes when an agent has restricted tools, incomplete context, a large legacy repository, or a task that runs for hours.

The company reports an Artificial Analysis Intelligence Index score of 61, matching the score it lists for GPT-5.6 Sol and one point behind Fable 5 Max. Its table also reports 69.9% on CursorBench 3.2, 65.9% on DeepSWE 1.1, 61.3% on FrontierCode 1.1 Extended, 57.5% on APEX-Agents, and 26% on Terminal-Bench 3.0. The results are mixed rather than a universal lead: the table places other models ahead on several coding and terminal tests. SpaceXAI says third-party figures use the best self-reported or publicly available results, so differences in harnesses, reasoning budgets, and test settings remain important limitations.

تفاصيل المصدر: SpaceXAI's August 12 Grok 4.6 announcement and API documentation ↗

لماذا يهم

ينضم Grok 4.6 إلى سباق نموذجي يتم تنظيمه بشكل متزايد حول وكلاء يمكنهم تنفيذ مشروع بدلاً من مجرد الاستجابة للمطالبة. إن نافذة السياق الكبيرة والتفكير القابل للتعديل واستخدام الأدوات والتدريب على المسارات الطويلة يمكن أن تسهل على المطور أو الفريق الصغير تفويض جزء محدد من البحث والترميز والاختبار والمراجعة. ستعتمد القيمة العملية بشكل أقل على موضع واحد في لوحة المتصدرين بقدر ما تعتمد على ما إذا كان النموذج يمكنه الحفاظ على المتطلبات، واستخدام الأدوات بأمان، وإنتاج عمل يمكن للشخص فحصه وتصحيحه.

For software teams, continuity is the central product claim. Long-running agents must remember architectural constraints, understand changes made earlier in a session, avoid undoing correct work, and verify that a fix does not break another part of the system. A 500,000-token window gives the model room for more repository context and tool history, but context capacity is not the same as dependable recall or reasoning. Large prompts can include irrelevant or conflicting material, cost more above xAI's 200,000-token pricing threshold, and still fail to contain the one file or requirement that determines the right answer.

The training description also shows how frontier labs are using earlier models to manufacture and filter trajectories for later ones. Regenerating supervised examples across several reasoning efforts and agent harnesses may improve consistency between the model and the environments in which it will operate. It also raises unanswered questions about error inheritance and evaluation independence. If one model creates training traces and model-based checks decide which traces survive, developers need evidence that the resulting system is not simply becoming better at satisfying the same automated judges while retaining blind spots those judges miss.

Availability across the API, Cursor, Grok Build, and gateways lowers the friction of testing the release in existing workflows. The introductory offer of twice the included usage in Cursor and Grok Build for one week may accelerate real-world trials. Teams should still compare total task cost, completion time, correction effort, and failure recovery rather than token prices alone. The fast variant may help interactive work, but doubling the per-token price creates a tradeoff that only task-level testing can resolve. A slower model that completes correctly on the first attempt can be cheaper than a fast model that requires repeated repair.

The public-interest boundary is just as important as coding performance. An agent operating across files, browsers, terminals, or business systems can propagate a mistaken assumption farther than a conventional chatbot answer. SpaceXAI says Grok 4.6 received its widest pre-deployment testing suite and expanded post-deployment and third-party testing, but the announcement provides only a high-level safety description. It does not publish a detailed system card, disaggregated refusal and misuse results, incident thresholds, or evidence for every domain in which the company says safeguards were calibrated. Users should treat the safety language as a vendor claim pending fuller documentation and independent testing.

Interactive Mechanism

الآلية التفاعلية: كيف تعمل فعليًا

استكشف التكنولوجيا الأساسية وراء هذا التطور بشكل تفاعلي.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
التحقق من المفهوم التفاعلي+10 Points
AI Agents Quiz

What most distinguishes an AI agent from a basic chatbot?

ماذا تشاهد بعد ذلك

سيأتي الدليل الحاسم من الاختبارات القابلة للتكرار والمشاريع العادية بعد الإطلاق. شاهد ما إذا كان Grok 4.6 يمكنه إنهاء المهام الطويلة دون الانجراف، وما إذا كان اختباره الذاتي يرصد عيوبًا حقيقية، وكيف تتغير تكلفته مع السياقات الكبيرة وإعادة المحاولة، وما إذا كانت SpaceXAI تنشر ما يكفي من تفاصيل السلامة والتقييم ليقوم الغرباء بفحص المطالبات. التوفر المبكر له معنى؛ ويظل الحكم الذاتي الذي يمكن الاعتماد عليه سؤالا تجريبيا.

First, compare the model under matched conditions. Independent evaluators should hold the agent harness, tools, repository snapshot, time limit, token budget, and reasoning effort constant when comparing Grok 4.6 with Grok 4.5 and competing systems. A useful report should include completed tasks, partial successes, regressions, invalid tool calls, human interventions, wall-clock time, and total cost. A single benchmark percentage cannot show whether failures are easy to repair or whether an agent silently changes unrelated files while obtaining a passing test result.

Second, test the long-context claim as a systems question. Developers should vary repository size and context quality, then measure whether the model retrieves the right constraints, maintains a plan through compaction, and notices contradictions introduced earlier in the trajectory. The documentation recommends a prompt cache key so requests route consistently and cache hits remain reliable, and it points long loops toward context compaction. Those features can improve economics and continuity, but teams need to monitor what compaction removes and whether a cached conversation preserves obsolete assumptions after the underlying project changes.

Third, look for fuller safety disclosure. SpaceXAI says its safeguards cover legitimate vulnerability patching, engineering design, and AI research, with broad pre-deployment, post-deployment, and third-party testing. The next useful publication would identify threat models, evaluation sets, pass thresholds, high-risk capabilities, known failure modes, and mitigations at both model and product layers. It should distinguish what the base model learned from what Grok Build, Cursor, an API gateway, or a customer's own sandbox enforces. Without that separation, users cannot tell which safety property travels with the model and which depends on the surrounding application.

Finally, watch adoption beyond launch-week incentives. Cursor and Grok Build users can test Grok 4.6 immediately, while API customers can choose ordinary or faster inference. Sustained usage, public postmortems, and task-level comparisons will show whether the model's stronger interactive first passes translate into maintainable software and useful research artifacts. SpaceXAI has shipped a material new option with concrete specifications. What remains unknown is how reliably it sustains autonomous work in messy production environments, where permissions, incomplete requirements, changing files, and human review matter as much as raw benchmark capability.

الأدلة والاختبارات ذات الصلة

وكلاء الذكاء الاصطناعيشرح نماذج الذكاء الاصطناعيترميز الذكاء الاصطناعيسلامة الذكاء الاصطناعياختبر ما تعرفه – جرّب اختبارًا مجانيًا للذكاء الاصطناعيابحث عن مصطلح الذكاء الاصطناعي في قاموسنااتبع أداة تعقب إصدار نموذج الذكاء الاصطناعي
وجدت هذا مفيدا؟