Home / Blog / AT&T and Microsoft scale trillion-token workloads with…
Engineering

AT&T and Microsoft scale trillion-token workloads with Microsoft Foundry and AMD

AT&T processed approximately one trillion tokens while building OTel2.0 on Microsoft Foundry Managed Compute, open AI models, and a mix of AMD and NVIDIA GPU…

By Dillip Chowdary • Aug 05, 2026 • Source: Microsoft Azure Blog

AT&T and Microsoft scale trillion-token workloads with Microsoft Foundry and AMD

AT&T processed approximately one trillion tokens while building OTel2.0 on Microsoft Foundry Managed Compute, open AI models, and a mix of AMD and NVIDIA GPU infrastructure. The work is described as a production-scale telecom AI effort, not a lab experiment, and is framed around flexible model choice plus infrastructure that can absorb that volume of token traffic.

The stack combines Foundry Managed Compute for orchestration and capacity, open models for the model layer, and heterogeneous AMD and NVIDIA GPUs for the hardware path. That setup implies the workload was not locked to a single accelerator or vendor SKU. Managed compute handles provisioning and scale; dual-GPU support lets operators route jobs where capacity, cost, or model support fits best rather than forcing one GPU family for every stage of training or inference.

Advertisement

Tech Pulse Daily

Get tomorrow's pulse first

Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.

For engineers and builders, the signal is that trillion-token runs are being treated as normal production traffic for a large telecom application, not as one-off research batches. Teams designing similar systems need to plan for managed compute that can grow with token load, for model selection that can change without re-architecting the control plane, and for GPU fleets that are not single-vendor. OTel2.0 is the concrete product context where those constraints were stress-tested at that scale.

In market terms, the pairing of AT&T’s telecom AI work with Microsoft Foundry and multi-vendor GPUs sits in a broader push for production-scale AI on cloud managed platforms rather than pure DIY clusters. Open models plus Foundry Managed Compute give a path that does not require a single closed model stack, while AMD and NVIDIA together reduce dependence on one accelerator supplier. That combination is aimed at carriers and other heavy operators who need both scale and the ability to switch models or hardware as demand shifts.

The practical takeaway is to treat token volume, managed compute, model flexibility, and multi-GPU capacity as co-equal design inputs when scoping production AI for telecom-class workloads. Watch whether similar Foundry Managed Compute deployments publish more detail on how open models and AMD/NVIDIA mix under sustained trillion-token load, and how that pattern extends beyond OTel2.0 to other carrier AI systems.

Advertisement

🔎 More interesting news

5-min tech signal

Weekday briefing for engineers who skip the noise.

No spam · Unsubscribe anytime

Advertisement

✈️ CareerPilot

Your AI job-search copilot

Match your resume against live Ashby, Greenhouse & Lever openings — fit scores, job-specific resume optimization and email alerts.

Find matching jobs →

Free Tools

Browse all tools →