뉴스로 돌아가기
제품AI Understanding 브리핑

AWS, 오픈 소스 AI 에이전트 결정 모델인 Strands Decider 2B 출시

AWS는 2026년 10월 1일에 텍스트를 생성하지 않고 AI 에이전트에 대한 로컬 라우팅 및 정책 결정을 빠르게 내리도록 설계된 20억 매개변수 오픈 소스 모델인 Strands Decider 2B를 출시했습니다.

5 min readRead the linked source
Source-provided image accompanying AWS releases Strands Decider 2B, an open-source AI agent decision model
소스 참조녹음된 소스
출판사
shattered.io
소스 링크
shattered.iohttps://shattered.io/aws-strands-decider-2b-open-source-ai-agent-model-2026/
소스 유형
연결된 소스 — 기본 소스 상태가 설정되지 않았습니다.
맥락60초 안에 이해하세요

여기서 시작하세요

주요 용어

AI 에이전트
종종 도구와 메모리를 사용하여 목표를 달성하기 위해 관찰하고, 추론하고, 조치를 취할 수 있는 소프트웨어 시스템입니다.
API(애플리케이션 프로그래밍 인터페이스)
한 소프트웨어 시스템이 다른 시스템에 요청을 보내고 응답을 받는 구조화된 방식입니다.
대형 언어 모델(LLM)
텍스트를 생성하고 분석하기 위해 대규모 텍스트 말뭉치를 학습한 언어 모델입니다.
자신을 테스트해 보세요AI 에이전트 퀴즈

무슨 일이 일어났나요?

Amazon Web Services released Strands Decider 2B, a specialized open-source AI model intended to handle routine decision-making tasks within workflows. Unlike general-purpose large language models that generate text, this 2-billion-parameter model evaluates predefined options and returns a selection or confidence score in a single forward pass. Released under the Apache 2.0 license with full training materials, the model is designed to run locally on consumer hardware, aiming to reduce the latency and token costs associated with using frontier models for simple routing and guardrail checks.

AWS released Strands Decider 2B on October 1, 2026, as part of its strands-labs initiative. The model is distinct from traditional chatbots because it does not generate free-form text. Instead, it functions as a decision engine that takes a set of predefined options and context, then returns a specific choice, a yes/no probability, or a confidence score. This design allows it to operate within the control loop of an , handling tasks such as tool selection, model routing, and guardrail enforcement without the overhead of text generation.

The model contains approximately 2 billion parameters, a size that AWS states allows it to run on laptop CPUs, consumer GPUs, or Apple silicon without requiring a cloud API connection. It is released under the Apache 2.0 license, and AWS has made the weights, training recipes, and evaluation scripts available for download. This open-source approach enables developers to inspect, retrain, and deploy the model on their own infrastructure, which is particularly relevant for organizations handling regulated data or requiring offline capabilities.

According to AWS Newsroom, the model is optimized for fast experimentation and local development, with a claimed latency of under 100 milliseconds for local execution. Secondary reporting from SiliconANGLE and TechTimes highlights the cost-efficiency of this approach, noting that eliminating text generation for routine decisions reduces token consumption and response latency. While a specific parameter count of 1.9 billion and a median latency of 115 milliseconds on an Nvidia RTX 3090 have appeared in community benchmarks, these figures are not confirmed in AWS's official specifications.

The release positions AWS against other cloud providers and AI labs that have focused on scaling cloud-hosted agent components. By providing a local-first, open-source tool for agent control, AWS is targeting the operational overhead of running agents at scale. The model is intended to complement, not replace, larger frontier models, which would still handle complex reasoning and creative writing tasks, while Strands Decider 2B manages the high-frequency, low-complexity decisions that occur between those larger steps.

소스 세부정보: shattered.io ↗

왜 중요한가요?

The release addresses a significant operational inefficiency in current architectures, where complex, expensive language models are often used for simple binary or multiple-choice decisions. By providing a lightweight, locally runnable tool, AWS offers developers a way to lower inference costs and improve response times for high-volume agent deployments. This move also signals a shift toward modular agent infrastructure, where specific components are optimized for narrow tasks rather than relying on a single general-purpose model for all functions.

The primary significance of Strands Decider 2B lies in its potential to reduce the cost and latency of deployments. In many current agent architectures, every routing decision or policy check is sent to a large language model, incurring the cost of full text generation even when the output is a simple selection. By offloading these tasks to a smaller, specialized model, developers can significantly lower their inference bills and improve system responsiveness.

The open-source nature of the release is also a strategic move. By providing the full training materials and code under a permissive license, AWS allows developers to customize the model for their specific use cases. This transparency is a differentiator compared to hosted decision APIs, which lock developers into a vendor's infrastructure and pricing model. For enterprises concerned with data privacy or compliance, the ability to run the model entirely offline is a critical feature.

This release reflects a broader trend in the AI industry toward modularization. Rather than relying on a single, massive model for all tasks, developers are increasingly looking for specialized components that can handle specific functions more efficiently. Strands Decider 2B is a concrete example of this shift, offering a practical tool for optimizing the 'connective tissue' of agent systems.

Interactive Mechanism

대화형 메커니즘: 실제로 작동하는 방식

이 개발의 이면에 있는 기본 기술을 대화식으로 살펴보세요.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
대화형 개념 확인+10 Points
AI Agents Quiz

An agent must create a draft calendar event for Tuesday at 2 p.m. Which evidence would establish the requested result?

다음에 무엇을 볼 것인가

Developers should monitor the model's real-world performance in production environments, particularly regarding latency consistency across different hardware configurations. Additionally, the adoption of this model by other cloud providers or the emergence of competing open-source decision models will indicate whether this modular approach becomes a standard practice in agent development.

The real-world performance of the model will be a key factor in its adoption. While AWS claims sub-100-millisecond latency, actual performance will vary based on hardware, batch size, and integration complexity. Developers will need to conduct their own benchmarks to determine if the model meets their specific latency and throughput requirements.

The competitive landscape will also be important to monitor. If other cloud providers or AI labs release similar open-source decision models, it could lead to a standardization of this approach in agent development. Conversely, if the model fails to gain traction, it may indicate that the market is not yet ready for such specialized, modular components.

Finally, the evolution of agent frameworks will be a factor. As agent orchestration tools become more sophisticated, they may incorporate specialized models like Strands Decider 2B as standard components. This could lead to a new category of 'agent infrastructure' models that are optimized for specific, narrow tasks rather than general-purpose intelligence.

관련 가이드 및 퀴즈

AI 에이전트AI 모델 설명AI의 미래알고 있는 내용을 테스트해 보세요. 무료 AI 퀴즈를 시도해 보세요.용어집에서 AI 용어를 찾아보세요.AI 모델 출시 추적기를 따르세요.
이것이 유용하다고 생각하시나요?