返回新聞
產品展示AI Understanding 簡報

AWS 發布開源 AI 代理程式決策模型 Strands Decider 2B

AWS 於 2026 年 10 月 1 日發布了 Strands Decider 2B,這是一個擁有 20 億參數的開源模型,旨在為 AI 代理程式做出快速的本地路由和策略決策,而無需產生文字。

5 min readRead the linked source
Source-provided image accompanying AWS releases Strands Decider 2B, an open-source AI agent decision model
來源參考來源記錄
出版商
shattered.io
來源連結
shattered.iohttps://shattered.io/aws-strands-decider-2b-open-source-ai-agent-model-2026/
來源類型
連結來源-主要來源狀態尚未確定。
背景60 秒內了解這一點

從這裡開始

關鍵術語

人工智慧代理
一種可以觀察、推理並採取行動來實現目標的軟體系統,通常使用工具和記憶體。
API(應用程式介面)
一種軟體系統向另一個系統發送請求並接收回應的結構化方式。
大語言模型(LLM)
在海量文本語料庫上訓練來產生和分析文本的語言模型。
測試一下自己AI 代理測驗

發生了什麼事

Amazon Web Services released Strands Decider 2B, a specialized open-source AI model intended to handle routine decision-making tasks within workflows. Unlike general-purpose large language models that generate text, this 2-billion-parameter model evaluates predefined options and returns a selection or confidence score in a single forward pass. Released under the Apache 2.0 license with full training materials, the model is designed to run locally on consumer hardware, aiming to reduce the latency and token costs associated with using frontier models for simple routing and guardrail checks.

AWS released Strands Decider 2B on October 1, 2026, as part of its strands-labs initiative. The model is distinct from traditional chatbots because it does not generate free-form text. Instead, it functions as a decision engine that takes a set of predefined options and context, then returns a specific choice, a yes/no probability, or a confidence score. This design allows it to operate within the control loop of an , handling tasks such as tool selection, model routing, and guardrail enforcement without the overhead of text generation.

The model contains approximately 2 billion parameters, a size that AWS states allows it to run on laptop CPUs, consumer GPUs, or Apple silicon without requiring a cloud API connection. It is released under the Apache 2.0 license, and AWS has made the weights, training recipes, and evaluation scripts available for download. This open-source approach enables developers to inspect, retrain, and deploy the model on their own infrastructure, which is particularly relevant for organizations handling regulated data or requiring offline capabilities.

According to AWS Newsroom, the model is optimized for fast experimentation and local development, with a claimed latency of under 100 milliseconds for local execution. Secondary reporting from SiliconANGLE and TechTimes highlights the cost-efficiency of this approach, noting that eliminating text generation for routine decisions reduces token consumption and response latency. While a specific parameter count of 1.9 billion and a median latency of 115 milliseconds on an Nvidia RTX 3090 have appeared in community benchmarks, these figures are not confirmed in AWS's official specifications.

The release positions AWS against other cloud providers and AI labs that have focused on scaling cloud-hosted agent components. By providing a local-first, open-source tool for agent control, AWS is targeting the operational overhead of running agents at scale. The model is intended to complement, not replace, larger frontier models, which would still handle complex reasoning and creative writing tasks, while Strands Decider 2B manages the high-frequency, low-complexity decisions that occur between those larger steps.

來源詳情: shattered.io ↗

為什麼這很重要

The release addresses a significant operational inefficiency in current architectures, where complex, expensive language models are often used for simple binary or multiple-choice decisions. By providing a lightweight, locally runnable tool, AWS offers developers a way to lower inference costs and improve response times for high-volume agent deployments. This move also signals a shift toward modular agent infrastructure, where specific components are optimized for narrow tasks rather than relying on a single general-purpose model for all functions.

The primary significance of Strands Decider 2B lies in its potential to reduce the cost and latency of deployments. In many current agent architectures, every routing decision or policy check is sent to a large language model, incurring the cost of full text generation even when the output is a simple selection. By offloading these tasks to a smaller, specialized model, developers can significantly lower their inference bills and improve system responsiveness.

The open-source nature of the release is also a strategic move. By providing the full training materials and code under a permissive license, AWS allows developers to customize the model for their specific use cases. This transparency is a differentiator compared to hosted decision APIs, which lock developers into a vendor's infrastructure and pricing model. For enterprises concerned with data privacy or compliance, the ability to run the model entirely offline is a critical feature.

This release reflects a broader trend in the AI industry toward modularization. Rather than relying on a single, massive model for all tasks, developers are increasingly looking for specialized components that can handle specific functions more efficiently. Strands Decider 2B is a concrete example of this shift, offering a practical tool for optimizing the 'connective tissue' of agent systems.

Interactive Mechanism

互動機制:它實際上是如何運作的

以互動方式探索這項發展背後的基礎技術。

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
互動式概念檢查+10 Points
AI Agents Quiz

An agent must create a draft calendar event for Tuesday at 2 p.m. Which evidence would establish the requested result?

接下來看什麼

Developers should monitor the model's real-world performance in production environments, particularly regarding latency consistency across different hardware configurations. Additionally, the adoption of this model by other cloud providers or the emergence of competing open-source decision models will indicate whether this modular approach becomes a standard practice in agent development.

The real-world performance of the model will be a key factor in its adoption. While AWS claims sub-100-millisecond latency, actual performance will vary based on hardware, batch size, and integration complexity. Developers will need to conduct their own benchmarks to determine if the model meets their specific latency and throughput requirements.

The competitive landscape will also be important to monitor. If other cloud providers or AI labs release similar open-source decision models, it could lead to a standardization of this approach in agent development. Conversely, if the model fails to gain traction, it may indicate that the market is not yet ready for such specialized, modular components.

Finally, the evolution of agent frameworks will be a factor. As agent orchestration tools become more sophisticated, they may incorporate specialized models like Strands Decider 2B as standard components. This could lead to a new category of 'agent infrastructure' models that are optimized for specific, narrow tasks rather than general-purpose intelligence.

相關指引和測驗

人工智慧代理人工智慧模型解釋AI 的未來測試你所知道的—嘗試免費的人工智慧測驗在我們的詞彙表中尋找人工智慧術語關注 AI 模型發布追蹤器
覺得有用嗎?