العودة إلى الأخبار
المنتجAI Understanding إحاطة

توضح شركة IBM كيفية بناء نماذج الاستدلال المفتوحة Granite 4.2

قامت شركة IBM بإصدار Granite 4.2، وهي عائلة مكونة من نماذج لغوية كثيفة 3B و8B و30B تم تدريبها باستخدام التدريب المسبق طويل السياق والبيانات المنطقية والتعلم المعزز في بيئات استخدام الأدوات الحقيقية.

6 min readRead the primary source
Primary-source image accompanying IBM details how it built the open Granite 4.2 reasoning models
وثيقة المصدر الأساسيتم تسجيل المصدر
الناشر
huggingface.co
رابط المصدر
huggingface.cohttps://huggingface.co/blog/ibm-granite/granite-4-2
نوع المصدر
المستند الأساسي - إعلان رسمي أو ورقة أو ملف أو صفحة الطرف الأول التي نقرأها مباشرة.
السياقافهم هذا في 60 ثانية

ابدأ هنا

المصطلحات الرئيسية

تعزيز التعلم من ردود الفعل البشرية (RLHF)
طريقة تدريب تستخدم إشارات التفضيل البشري لتشكيل سلوك النموذج.
التعلم المعزز
التدريب من خلال إشارات المكافأة حيث يتعلم الوكيل الإجراءات التي تزيد من العائد على المدى الطويل.
الذاكرة (ذاكرة الوكيل)
السياق المُخزن الذي يستخدمه وكيل الذكاء الاصطناعي عبر الخطوات أو الجلسات لتحسين الاستمرارية.
اختبر نفسكوأوضح نماذج الذكاء الاصطناعي مسابقة

ماذا حدث

IBM’s Granite team published a technical account of Granite 4.2, a new family of dense, decoder-only reasoning language models available in 3B, 8B and 30B sizes under the Apache 2.0 license. The models support thinking and non-thinking modes, low-effort reasoning, native tool calling and a context window extended to 512K tokens during training. IBM says the 8B and 30B versions were additionally trained to use tools, edit and run code, operate terminals and search the web inside sandboxed environments.

In a Hugging Face article published August 25, 2026, IBM’s Granite team described Granite 4.2 as the first Granite family built as dense, decoder-only reasoning models. The family has three sizes: 3B, 8B and 30B parameters. All use the same broad architecture, including grouped-query attention, rotary position embeddings, SwiGLU feed-forward layers, RMSNorm and bfloat16 precision. The article says the models were pretrained from scratch on approximately 15 trillion tokens through five phases, with the final phase extending the context window to 512K tokens. The architecture table lists a 131,072-token sequence length for the models, while the training strategy is described as extending context to 512K; the source does not explain that distinction in detail.

The post-training recipe is the central change IBM describes. Supervised fine-tuning used about 7.2 million samples, or roughly 100 billion tokens, combining agentic and non-agentic material. IBM says the data included software engineering, tool calling, terminal use, search, mathematics, multilingual instruction following, science, reasoning and safety examples. The company says it normalized the data into a common chat format, used GPT-OSS-120B and Gemma 4 as language-model judges, removed low-quality or invalid examples, and applied heuristic filtering and SHA-256-based deduplication. For the 30B model, IBM added a second fine-tuning phase that increased the share of agentic coding data while retaining about 16% replay data from the original mixture.

After fine-tuning, IBM applied a staged reinforcement-learning pipeline. All three models received foundational for verifiable tasks and a final RLHF stage for preference and safety. The 8B and 30B models also received agentic reinforcement learning in three stages: software engineering, terminal operation and web search. IBM says those stages used real repositories, live shell environments and browsing tools, with rewards based on whether tasks were completed. The training used asynchronous GRPO, with separate generation and training workers, and relied on NeMo-RL and NeMo-Gym. The 3B model did not receive the agentic-RL block. The source also describes quantized releases in FP8, NVFP4, MXFP4 and multiple GGUF formats.

تفاصيل المصدر: huggingface.co ↗

لماذا يهم

The release provides unusually detailed visibility into how an openly licensed model family combines conventional pretraining with staged for tool use. It also gives developers smaller models, quantized variants and OpenAI-compatible serving options that could make local or self-hosted reasoning and agentic workflows more practical. The reported benchmark results are IBM’s own evaluations, however, and the source does not establish independent replication, real-world reliability or broad availability beyond the described model releases.

The release matters because it makes the training process more inspectable than a typical model announcement. IBM provides stage-by-stage descriptions of the data, reward signals, rollout environments and optimization settings, including the distinction between verifiable rewards, judge-based rewards and agentic outcome rewards. That documentation is useful to researchers and developers evaluating whether tool use should be learned through ordinary instruction tuning, or a combination. It also makes clear that the three model sizes are not simply scaled versions of one another: the 8B and 30B versions receive additional training intended to teach actions in environments, while the 3B version follows a shorter path.

The Apache 2.0 license and the release of quantized variants may broaden the practical options for organizations that want to run models under their own infrastructure. The article describes support for Transformers, vLLM and SGLang, an OpenAI-compatible endpoint, and integration instructions for OpenCode, Pi and OpenHands. These features could reduce adaptation work for teams already using compatible serving and agent frameworks. They do not, by themselves, show that the models can run economically on ordinary consumer hardware. The source identifies large-scale training on an NVIDIA GB200 NVL72 cluster hosted by CoreWeave and gives extensive distributed-training details, but it does not provide complete inference-cost comparisons or hardware requirements for each quantized model.

IBM’s reported results suggest capability increases with model size, especially on the listed reasoning, long-context and agentic coding evaluations. The source reports, for example, SWE-Bench Verified scores of 47.67 for 8B and 57.00 for 30B, RULER 128K scores of 71.41 and 81.38, and AIME25 scores of 86.67 and 89.17. These numbers are useful as a record of the company’s evaluation claims, but they are not independent evidence. The article does not describe a third-party audit, confidence intervals, contamination analysis, comparative testing against current alternatives or the operational failure modes encountered in the environments. It also does not establish how the models behave when tools return misleading information or when tasks have safety consequences.

Interactive Mechanism

الآلية التفاعلية: كيف تعمل فعليًا

استكشف التكنولوجيا الأساسية وراء هذا التطور بشكل تفاعلي.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
التحقق من المفهوم التفاعلي+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

ماذا تشاهد بعد ذلك

The key questions are whether Granite 4.2’s reported capabilities hold up in independent testing, how much performance is lost in its FP8, FP4 and GGUF variants, and how reliably the 8B and 30B models act in less controlled environments. Users should also examine licensing and deployment details, hardware requirements, tool-call safety and the limits of the models’ long-context and agentic performance before treating benchmark scores as production evidence.

وينبغي أن تكون التقييمات المستقلة هي المتابعة الأولى. سيحتاج المراجعون إلى إعادة إنتاج المعايير المدرجة حيثما أمكن ذلك، ومقارنة النماذج بأنظمة مماثلة الحجم، وفحص مطالبات التقييم وتحديد ما إذا كانت بيانات التدريب تتداخل مع مجموعات الاختبار. تستحق النتائج الوكيلة تدقيقًا دقيقًا بشكل خاص لأن معدلات النجاح يمكن أن تعتمد بشكل كبير على الأدوات واختيار المستودع والاختبارات المخفية وأدوات المتصفح وحدود المهام. يقوم مصدر IBM بالإبلاغ عن البيئات وبعض تفاصيل التكوين، ولكنه لا يوفر معلومات كافية هنا لتحديد مدى تمثيلها لأعمال الإنتاج.

قيود النشر هي أمر مهم آخر غير معروف. يقول المصدر أنه يمكن تقديم Granite 4.2 من خلال البنية التحتية المتوافقة مع OpenAI ويقدم العديد من تنسيقات القياس الكمي، لكنه لا يحدد متطلبات الذاكرة أو الإنتاجية أو زمن الوصول أو استخدام الطاقة أو تدهور الجودة لكل متغير. يحتاج مطالبة سياق 512K أيضًا إلى اختبار عملي: سعة السياق الطويلة ليست مثل الاسترجاع الموثوق أو التفكير في كل جزء من المدخلات الطويلة جدًا. يجب على المطورين قياس الأداء على أعباء العمل الخاصة بهم والتحقق من كيفية تصرف عناصر التحكم في التفكير واقتطاع السجل وتحليل استدعاءات الأداة في حزمة الخدمة التي اختاروها.

تظل أسئلة السلامة والحوكمة مفتوحة أيضًا. تقول شركة IBM إن مرحلة RLHF النهائية تتضمن تحسين التفضيلات، ومقاومة كسر الحماية، والرفض المناسب، وعقوبة الاستدلال المفرط. لا تقدم المقالة نتائج مفصلة تتعلق بالسلامة، أو معدلات خطأ الرفض، أو تحليل الخصوصية، أو اختبار الأمان، أو أدلة حول السلوك خارج بيئات التدريب. كما أنها لا توضح ما إذا كان محتوى سلسلة الأفكار يتم عرضه دائمًا أو تصفيته أو التعامل معه بشكل مختلف عبر الواجهات. قبل استخدام Granite 4.2 في التطبيقات اللاحقة، ستحتاج المؤسسات إلى اختبارها الخاص لأذونات الأداة، ومعالجة البيانات، والحقن الفوري، وإمكانية التدقيق، والمراجعة البشرية. وبالتالي فإن الأهمية المباشرة للإصدار تكمن في النماذج نفسها والوصف الأكثر شفافية للخيارات الهندسية التي تقف وراءها؛ لا يزال يتعين تحديد مدى تفوقهم في العالم الحقيقي.

الأدلة والاختبارات ذات الصلة

شرح نماذج الذكاء الاصطناعيتدريب الذكاء الاصطناعيوكلاء الذكاء الاصطناعيالمحولاتاختبر ما تعرفه – جرّب اختبارًا مجانيًا للذكاء الاصطناعيابحث عن مصطلح الذكاء الاصطناعي في قاموسنااتبع أداة تعقب إصدار نموذج الذكاء الاصطناعي
وجدت هذا مفيدا؟