Back to News
ProductAI Understanding briefing

Microsoft launches Decision-1 model for fast structured decisions

Microsoft has released Microsoft-Decision-1, a specialized AI model for structured decision-making that claims 35x faster latency than GPT-6 Sol and top accuracy on 36 internal benchmarks.

4 min readRead the linked source
Source-provided image accompanying Microsoft launches Decision-1 model for fast structured decisions
Source referenceSource recorded
Publisher
eu.36kr.com
Source type
Linked source — primary-source status has not been established.
Also cited

Story last revised

ContextUnderstand this in 60 seconds

Key terms

Classification
A task where a model assigns an input to one or more predefined categories.
Post-training
Training steps applied after pretraining, such as instruction tuning, preference optimization, and safety tuning.
Benchmark
A standardized test or dataset used to measure and compare model performance.
Test yourselfAI Agents Quiz

What changed since publication

  1. First published
  2. This source provides detailed technical specifications, pricing, and internal testing results for Microsoft-Decision-1, including its Qwen3.5-9B base, $0.042 per million input token pricing, and specific performance comparisons against GPT-6 Sol and GPT-5.6 Luna.

What happened

Microsoft unveiled Microsoft-Decision-1, a new AI model designed specifically for structured decision-making tasks such as routing, , and workflow control. The model is built on the Qwen3.5-9B base with specialized . According to 36Kr, the P50 version is 35 times faster than OpenAI's GPT-6 Sol and ranks first in accuracy across 36 internal tests. The model is available to developers via Microsoft Foundry, with OpenRouter support planned.

Microsoft announced the release of Microsoft-Decision-1, a model tailored for structured decision-making rather than general text generation. The model selects optimal solutions from preset options and assigns probability scores to each alternative, enabling downstream applications to determine next steps such as retrying, escalating, or submitting for manual review.

According to 36Kr, the P50 version of the model is 35 times faster than OpenAI's GPT-6 Sol and 4.5 times faster than Quyet-1.0-Large. Microsoft claims it ranks first in accuracy across 36 tests covering nearly 150,000 questions. The model is built on the Qwen3.5-9B base with specialized focused on single decision scoring.

Pricing is set at $0.042 per million input tokens, with output tokens being free. This structure is designed for high-frequency, repetitive decision-making applications. The model is currently available to developers through Microsoft Foundry, with support for OpenRouter planned.

Internal testing by Microsoft's Xbox Research team showed the model classified over 10,000 pieces of game feedback with quality comparable to GPT-6 Sol but at 14 times the speed and 200 times lower cost. The Microsoft Copilot team found similar performance to GPT-5.6 Luna in AI response evaluation tasks.

Microsoft reported that the model maintains stability under perturbation, with an average decision flip rate of only 1.3% across eight forms of perturbation. Security testing across 11 benchmarks involving 5,250 requests showed the model successfully rejects harmful requests while maintaining utility for normal use.

Source details: eu.36kr.com ↗

Why it matters

This launch addresses a specific bottleneck in AI agent workflows: the latency and cost of making high-frequency, repetitive decisions. By offering a model that is significantly faster and cheaper than general-purpose LLMs for these specific tasks, Microsoft provides a practical tool for enterprises building complex automation pipelines. The specialized architecture allows for deterministic outputs with probability scores, which is crucial for reliable agent orchestration. However, the performance claims are based on Microsoft's internal testing, and independent verification is pending.

The model targets a specific niche in AI agent workflows where latency and cost are critical. By offloading structured decision tasks from general-purpose LLMs, developers can build more efficient and cost-effective automation pipelines.

The specialized design, which outputs probability scores for preset options, provides a more deterministic and controllable interface for workflow orchestration compared to free-text generation.

The pricing model, with free output tokens, makes it economically viable for high-volume, low-complexity decision tasks that would be prohibitively expensive with standard LLMs.

The reliance on internal benchmarks means the performance claims are not yet independently verified, which is a significant limitation for enterprise adoption decisions.

Interactive Mechanism

Interactive Mechanism: How It Actually Works

Explore the underlying technology behind this development interactively.

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
Interactive Concept Check+10 Points
AI Agents Quiz

An agent must create a draft calendar event for Tuesday at 2 p.m. Which evidence would establish the requested result?

What to watch next

Independent third-party benchmarks to verify the claimed 35x speed advantage and accuracy rankings. The rollout of OpenRouter support, which would make the model accessible to a broader developer ecosystem outside the Microsoft stack. Future migrations of the model to other base architectures, including Microsoft's MAI series and OpenAI models, as mentioned in the source.

Independent third-party evaluations will be crucial to validate Microsoft's claims about speed and accuracy, particularly the 35x latency advantage over GPT-6 Sol.

The expansion of access to OpenRouter will determine how widely the model is adopted outside the Microsoft ecosystem.

Future updates to the model's base architecture, including potential migrations to Microsoft's MAI series or OpenAI models, could affect its performance and compatibility.

Related guides & quizzes

Updates and corrections

This canonical story is updated in place when the developing event materially changes. Its URL and original publication date never change.

  • This source provides detailed technical specifications, pricing, and internal testing results for Microsoft-Decision-1, including its Qwen3.5-9B base, $0.042 per million input token pricing, and specific performance comparisons against GPT-6 Sol and GPT-5.6 Luna.
See the public corrections log
Found this useful?