AWS and OpenAI announce $38 billion 7-year strategic partnership. OpenAI gets access to hundreds of thousands of NVIDIA GPUs and tens of millions of CPUs for...

What the Partnership Covers

AWS and OpenAI have committed to a $38 billion, seven-year strategic partnership — the largest cloud-AI deal on record. Under the agreement, OpenAI gains access to hundreds of thousands of NVIDIA GPUs alongside tens of millions of CPUs running on Amazon's infrastructure. The GPU pool handles the training and inference of large models, while the CPU fleet covers the surrounding work: data preprocessing, orchestration, serving lighter workloads, and the countless supporting services that sit around a model in production.

A multi-year term matters here as much as the dollar figure. Training frontier models is not a one-time purchase of compute; it is a continuous pipeline where each generation feeds the next. Locking in capacity over seven years gives OpenAI predictable access to hardware that is otherwise supply-constrained, and gives AWS a large, committed anchor tenant for its accelerated-computing capacity.

Why Both GPUs and CPUs Matter

It is easy to focus only on the GPU count, since accelerators do the heavy math behind model training. But the ratio of tens of millions of CPUs to hundreds of thousands of GPUs reflects how real AI systems actually run. Much of the total cost and complexity lives outside the accelerator itself.

  • Data pipelines: cleaning, tokenizing, and staging training data before it ever reaches a GPU.
  • Orchestration: scheduling jobs, managing checkpoints, and keeping thousands of accelerators busy without idle gaps.
  • Serving: routing requests, handling retries, and running the API and safety layers that wrap each model call.
  • Everything else: logging, monitoring, storage, and the general-purpose services that a large application depends on.

Keeping expensive GPUs saturated depends on having enough conventional compute to feed them. A deal that pairs both at scale is a sign of planning for the full workload, not just the headline training runs.

The Tradeoffs of Committing at This Scale

A commitment of this size buys reliability but trades away flexibility. The main advantage is guaranteed capacity: when compute is scarce, having reserved hardware means training and product roadmaps are not blocked waiting for availability. It also tends to lower the effective unit cost compared with buying capacity on demand.

The cost is concentration. Standardizing on one provider's infrastructure means aligning deeply with that provider's networking, storage, and operational model, which raises the effort required to move workloads elsewhere later. For a project running at this scale, that is usually an accepted tradeoff — the operational gains from a single, well-integrated environment outweigh the theoretical benefit of staying provider-agnostic.

What Teams Building on AI Can Take From This

Most organizations will never negotiate a deal of this magnitude, but the underlying reasoning transfers directly. When planning AI infrastructure, budget for the supporting compute, not just the accelerators — the CPU-side work often determines whether your GPUs are actually being used efficiently. If your workloads are steady and predictable, reserved or committed capacity typically costs less than paying on demand; if they are spiky or experimental, keep the flexibility to scale down.

The broader signal is that access to compute has become a strategic asset worth securing years in advance. For anyone whose product depends on running models continuously, treating capacity planning as a first-class concern — rather than an afterthought once the model works — is the practical lesson to carry forward.

Automate Your Content with AI Video Generator

Try it Free →