AT&T and Microsoft scale trillion-token workloads with Microsoft Foundry and AMD
AT&T processed approximately one trillion tokens while building OTel2.0 on Microsoft Foundry Managed Compute, open AI models, and a mix of AMD and NVIDIA GPU…
By Dillip Chowdary • Aug 05, 2026 • Source: Microsoft Azure Blog
AT&T processed approximately one trillion tokens while building OTel2.0 on Microsoft Foundry Managed Compute, open AI models, and a mix of AMD and NVIDIA GPU infrastructure. The work is described as a production-scale telecom AI effort, not a lab experiment, and is framed around flexible model choice plus infrastructure that can absorb that volume of token traffic.
The stack combines Foundry Managed Compute for orchestration and capacity, open models for the model layer, and heterogeneous AMD and NVIDIA GPUs for the hardware path. That setup implies the workload was not locked to a single accelerator or vendor SKU. Managed compute handles provisioning and scale; dual-GPU support lets operators route jobs where capacity, cost, or model support fits best rather than forcing one GPU family for every stage of training or inference.
Advertisement
Tech Pulse Daily
Get tomorrow's pulse first
Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.
For engineers and builders, the signal is that trillion-token runs are being treated as normal production traffic for a large telecom application, not as one-off research batches. Teams designing similar systems need to plan for managed compute that can grow with token load, for model selection that can change without re-architecting the control plane, and for GPU fleets that are not single-vendor. OTel2.0 is the concrete product context where those constraints were stress-tested at that scale.
In market terms, the pairing of AT&T’s telecom AI work with Microsoft Foundry and multi-vendor GPUs sits in a broader push for production-scale AI on cloud managed platforms rather than pure DIY clusters. Open models plus Foundry Managed Compute give a path that does not require a single closed model stack, while AMD and NVIDIA together reduce dependence on one accelerator supplier. That combination is aimed at carriers and other heavy operators who need both scale and the ability to switch models or hardware as demand shifts.
The practical takeaway is to treat token volume, managed compute, model flexibility, and multi-GPU capacity as co-equal design inputs when scoping production AI for telecom-class workloads. Watch whether similar Foundry Managed Compute deployments publish more detail on how open models and AMD/NVIDIA mix under sustained trillion-token load, and how that pattern extends beyond OTel2.0 to other carrier AI systems.
Advertisement