Home / Blog / Spreading the load: How Salesforce met Multi-AZ HA with…
Tech News

Spreading the load: How Salesforce met Multi-AZ HA with SageMaker

Learn how Salesforce used Amazon SageMaker AI Inference Component placement (the SchedulingConfig parameter) to distribute model copies across multiple.

By Dillip Chowdary • Aug 29, 2026 • Source: AWS Machine Learning Blog

Spreading the load: How Salesforce met Multi-AZ HA with SageMaker

What happened

Spreading the load: How Salesforce met Multi-AZ HA with SageMaker Inference Components

Salesforce has met its Multi-AZ high availability compliance requirements using Amazon SageMaker AI Inference Components. They distributed model copies across Availability Zones with SchedulingConfig.

This article covers the mechanism of model placement and how it ensures cost efficiency. It is for machine learning engineers and cloud infrastructure builders.

How it works

Salesforce needed to meet strict compliance rules for high availability across multiple Availability Zones. They solved this challenge by implementing Amazon SageMaker AI Inference Components to deploy their machine learning models. Using this technology, Salesforce successfully distributed model copies across different Availability Zones. This specific deployment strategy allowed them to maintain constant service availability even if one zone experienced an outage. The transition to this new system was documented as a successful use case on the AWS Machine Learning Blog.

The core of this achievement lies in the balance between system reliability and infrastructure cost. Previously, achieving high availability compliance often required dedicating separate hardware resources to each model copy, which increased overall expenses. Salesforce bypassed this financial penalty by using multi-model co-hosting features. They did not have to compromise on cost efficiency to ensure applications remained fully compliant with enterprise standards. This deployment shows that large scale machine learning services can achieve high reliability without incurring prohibitive infrastructure costs.

Spreading the load: How Salesforce met Multi-AZ HA with SageMaker
Illustration · Pexels

The technical solution centers on placing Amazon SageMaker AI Inference Components across cloud infrastructure. Salesforce configured their setup using the SchedulingConfig parameter to control how model copies are assigned to instances. This parameter works by instructing the scheduler to distribute model instances across multiple Availability Zones. By actively managing these placement decisions, the system prevents all copies of a single model from hosting within the same zone. This separation ensures that a localized hardware failure does not disable the service.

Advertisement

Tech Pulse Daily

Get tomorrow's pulse first

Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.

Why it matters

At the same time, this method maintains benefits of multi-model co-hosting on shared instances. Co-hosting allows multiple machine learning models to share the same underlying compute resources, reducing idle capacity and lowering operation fees. By using the SchedulingConfig parameter, Salesforce could guarantee that their models were spread out for safety while still sharing host resources within each zone. This dual approach provides high availability and cost efficiency, proving that secure hosting does not require dedicated servers for every model copy.

Meeting Multi-AZ high availability compliance is critical for enterprise platforms like Salesforce that support global operations. If a service goes offline, it can cause immediate disruption to business workflows and client services. Historically, achieving this level of protection required significant financial investments in redundant, isolated hardware platforms. The ability to distribute model copies using the SchedulingConfig parameter changes this equation. It shows that companies can meet rigorous compliance rules while keeping their infrastructure spend at a reasonable and manageable level.

Who is affected

Furthermore, this deployment establishes a practical blueprint for hosting large volumes of machine learning models. As organizations scale their use of artificial intelligence, the cost of running independent endpoints for every single model becomes unsustainable. Co-hosting models on shared resources is the primary way to control these host operation expenses. By demonstrating that co-hosted models can be safely distributed across Availability Zones, Salesforce has shown a path forward for other enterprises facing similar scaling challenges and strict service level agreements.

This development directly affects systems builders and machine learning engineers who design and maintain large scale deployment pipelines. These professionals must regularly balance the conflicting demands of strict compliance audits and tight infrastructure budgets. The availability of the SchedulingConfig parameter in Amazon SageMaker AI provides them with a concrete tool to achieve both goals simultaneously. Systems engineers can now design architectures that are resilient to zone outages without having to request any additional budget for redundant specific hardware hosting instances.

Enterprise customers who rely on Salesforce services also benefit from this infrastructure improvement. They receive more stable and reliable applications that are less likely to experience downtime during localized cloud disruptions. At the same time, because Salesforce can run these models efficiently, the overall cost of providing these host services is kept under control. Ultimately, any organization using Amazon SageMaker AI for hosting multiple models can apply these same design techniques to improve their own operational resilience and compliance posture.

What to watch next

Moving forward, builders should verify how their own model hosting configurations handle zone outages. It is important to test the behavior of the SchedulingConfig parameter under simulated infrastructure failures to ensure model copies are distributed correctly. Teams should verify that their compliance monitoring tools can track model placement across multiple Availability Zones in real time. Regularly auditing these placements will ensure that subsequent configuration updates or scaling activities do not accidentally consolidate model instances back into a single Availability Zone.

In addition, organizations should watch for further updates from cloud providers regarding new automated placement controls. As machine learning workloads become more complex, the platform tools for managing resource allocation will likely become more sophisticated. Builders should monitor the development of scheduling algorithms that automatically optimize both infrastructure cost and service availability without manual configuration. Keeping track of these changes will help teams maintain efficient operations as they scale their enterprise artificial intelligence services to meet evolving global business needs.

Developer Action Items

  • Verify the claim on the official Amazon / AWS / Salesforce page (or AWS Machine Learning Blog), not from this recap alone.
  • Name the surface that moved — API, policy, model, hardware, or commercial terms — before you Slack the thread.
  • Assign one owner a day to read the primary material and decide: this-sprint, this-quarter, or noise.
  • Do not change production on day-one coverage. Watch the vendor changelog and one independent write-up first.

Advertisement

🔎 More interesting news

5-min tech signal

Weekday briefing for engineers who skip the noise.

No spam · Unsubscribe anytime

Advertisement

✈️ CareerPilot

Your AI job-search copilot

Match your resume against live Ashby, Greenhouse & Lever openings — fit scores, job-specific resume optimization and email alerts.

Find matching jobs →

Free Tools

Browse all tools →