뉴스로 돌아가기
기업AI Understanding 브리핑

Salesforce는 SageMaker 컨트롤을 사용하여 Agentforce 모델을 가용성 영역 전체에 확산시킵니다.

AWS에 따르면 Salesforce는 새로운 SageMaker 추론 구성 요소 배치 기능을 사용하여 GPU 공유 비용 절감을 유지하면서 가용 영역 전체에서 Agentforce 모델 복사본의 균형을 유지했습니다.

5 min readRead the primary source
Primary-source image accompanying Salesforce uses SageMaker controls to spread Agentforce models across availability zones
기본 소스 문서녹음된 소스
출판사
aws.amazon.com
소스 링크
aws.amazon.comhttps://aws.amazon.com/blogs/machine-learning/spreading-the-load-how-salesforce-met-multi-az-ha-with-sagemaker-inference-components/
소스 유형
기본 문서 — 우리가 직접 읽는 공식 발표, 논문, 서류 또는 자사 페이지입니다.
맥락60초 안에 이해하세요

여기서 시작하세요

주요 용어

API(애플리케이션 프로그래밍 인터페이스)
한 소프트웨어 시스템이 다른 시스템에 요청을 보내고 응답을 받는 구조화된 방식입니다.
알고리즘
문제를 해결하거나 작업을 완료하기 위해 컴퓨터가 따르는 정의된 규칙 또는 단계 세트입니다.
추론
훈련된 모델이 예측 또는 출력을 생성하는 런타임 단계입니다.
자신을 테스트해 보세요AI 에이전트 퀴즈

무슨 일이 일어났나요?

AWS says Salesforce deployed Agentforce models with SageMaker Components and used the new SchedulingConfig parameter to distribute model copies across Availability Zones and instances. Salesforce reported an 8x reduction in infrastructure costs from co-hosting multiple models on shared GPUs, but the default placement behavior did not guarantee the two-zone resilience required for its production models.

AWS and Salesforce describe a production-serving problem involving Agentforce, Salesforce’s AI foundation for agents. SageMaker Components allow multiple models to share GPU-backed infrastructure, which the source says reduced Salesforce’s infrastructure costs by eight times. The tradeoff was that SageMaker’s default placement evaluated each deployment operation independently. Copies of a particular model could therefore be unevenly distributed even when the endpoint itself used multiple Availability Zones. This left endpoint-level multi-zone configuration insufficient to guarantee model-level distribution. The issue was specifically the relationship between shared infrastructure and the placement of individual model copies.

The new control is exposed through the SchedulingConfig parameter in the CreateInferenceComponent API. Its AvailabilityZoneBalance setting governs how evenly copies are distributed between zones, while PlacementStrategy controls placement within each zone. SPREAD distributes copies across as many instances as possible to improve fault isolation; BINPACK puts copies on fewer instances to improve utilization. The source says Salesforce selected SPREAD for its production high-availability requirement. These settings make the intended placement behavior explicit at deployment time and connect the choice of distribution to the operational goal being pursued.

AWS gives a two-zone example with four instances and four copies of a model. With SPREAD and a maximum imbalance of one copy, the intended result is two copies in each zone. For a model requiring only two copies, a maximum imbalance of zero targets exactly one copy per zone. The same scheduling settings are intended to remain in effect during scale-out, scale-in, and endpoint or -component updates. The source also recommends a consolidation strategy for longer-term cleanup after repeated scaling operations. Together, these details describe placement as an ongoing scheduling concern rather than a setting applied only once at initial deployment.

소스 세부정보: aws.amazon.com ↗

왜 중요한가요?

The deployment addresses a practical problem in serving AI systems: a multi-zone endpoint can still leave all copies of one model concentrated in a single zone or instance. The configuration gives enterprise teams explicit controls for balancing availability-zone placement and choosing between fault isolation and higher utilization.

The distinction between endpoint-level and model-level resilience is important for organizations operating many AI models on shared infrastructure. A multi-zone endpoint does not automatically ensure that every model has a copy in every zone. If a model’s copies are concentrated, an instance failure or zone outage can remove that model even though other workloads on the endpoint remain available. This means a broad availability-zone design can appear resilient while leaving a particular model exposed. The placement question therefore has to be evaluated at the level of each model’s copies.

The placement feature connects reliability requirements with an explicit resource tradeoff. SPREAD can reduce the number of model copies lost in an instance failure, while BINPACK can improve accelerator utilization by concentrating workloads. For Salesforce, the source presents the choice as a way to retain the economic benefit of multi-model GPU hosting while meeting an internal requirement that every production model have two-zone support. The configuration does not remove the tradeoff; it gives teams a direct way to choose how their copies occupy the available instances and zones.

The reported result is consequential for enterprise AI operations because it moves high availability from a general architecture goal into deployment settings that can be inspected and managed. The source says Salesforce achieved two-zone compliance for its model fleet, preserved its co-hosting savings, maintained distribution during scaling, and avoided breaking zone balance during model updates. Those are claims from the AWS customer-success account, not independently audited performance findings. The post does not establish how the system behaved during a real zone outage or whether every model and region had identical conditions. The implementation account is therefore useful for understanding the control and its stated outcome, while leaving independent validation open.

Interactive Mechanism

대화형 메커니즘: 실제로 작동하는 방식

이 개발의 이면에 있는 기본 기술을 대화식으로 살펴보세요.

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
대화형 개념 확인+10 Points
AI Agents Quiz

An agent must create a draft calendar event for Tuesday at 2 p.m. Which evidence would establish the requested result?

다음에 무엇을 볼 것인가

The source does not provide independent uptime measurements, outage-test results, latency data, or a complete accounting of the reported cost reduction. Teams adopting the pattern will need to verify capacity in each target zone, monitor placement after scaling, and determine whether best-effort placement meets their own compliance requirements.

Capacity remains a central limitation. AWS recommends On-Demand Capacity Reservations in each target Availability Zone for zone-constrained regions, and the source says Salesforce pre-provisioned reserved GPU capacity to help obtain balanced placement. Without sufficient capacity, SageMaker may partially deploy copies on available instances because the enforcement mode described is permissive. That makes the feature usable under constraint, but it can leave the final distribution less balanced than intended. The desired configuration and the placement that is actually achieved can consequently differ when the required capacity is not available in every target zone.

Operators will need to monitor whether the desired placement persists over time. The source points to SageMaker AI Insights and CloudWatch metrics covering availability-zone skew, -component copy counts by zone, rebalancing events and duration, and insufficient-capacity errors. It also warns that an HA-critical component should not be reduced to one copy, because one copy cannot span two zones. The practical question is how quickly teams detect and remediate an imbalance before it becomes an availability problem. Monitoring must cover both the number of copies and their distribution, especially as scaling and updates change the deployment.

Important details remain unknown from the source. It does not state the geographic regions involved, the number of production endpoints or models covered, the cost of reserved capacity, the effect on latency and throughput, or the availability target Salesforce was trying to meet. It also does not provide comparative failure testing against the default . Enterprise teams should therefore treat the post as an implementation pattern and customer-reported result, then validate capacity, failover behavior, monitoring, and total cost in their own environments. Those checks are necessary to determine whether the reported balance between resilience and utilization applies to their own workloads and operating conditions.

관련 가이드 및 퀴즈

AI 에이전트AI 모델 설명AI의 미래알고 있는 내용을 테스트해 보세요. 무료 AI 퀴즈를 시도해 보세요.용어집에서 AI 용어를 찾아보세요.AI 자금 추적기를 팔로우하세요
이것이 유용하다고 생각하시나요?