As AI clusters scale to hundreds of thousands of GPUs, traditional packet switching is hitting a "power wall." The live demonstration of Optical Circuit Sw...

The power wall in packet-switched AI fabrics

Every large training cluster relies on a network that moves gradients and activations between accelerators. In a conventional packet-switched fabric, each switch parses, buffers, and forwards traffic packet by packet. That work is dominated by the optical transceivers and switch ASICs converting between light and electrons at every hop. As clusters grow toward hundreds of thousands of GPUs, the number of hops and transceivers climbs faster than the useful compute, and the fabric starts to consume a disproportionate share of the power and cooling budget. That is the "power wall": the point where adding more network to connect more GPUs costs more energy per delivered bit than the accelerators can justify.

The pressure is worse because AI traffic is not like general datacenter traffic. Collective operations such as all-reduce produce large, predictable flows between fixed groups of accelerators. Paying for the flexibility of per-packet routing on traffic that is this structured is expensive in exactly the resource that has become scarce.

How optical circuit switching changes the tradeoff

Optical circuit switching (OCS) takes a different approach. Instead of inspecting packets, it establishes a direct light path between two endpoints and holds that connection open, steering the optical signal without converting it back to electrical form at the switch. Once a circuit is set up, data passes through as photons, so the switch itself adds little latency and consumes little power regardless of how much traffic crosses it.

The tradeoff is that circuits must be provisioned before they carry data, and reconfiguring them is slower than forwarding an individual packet. OCS therefore fits workloads whose communication pattern is known ahead of time and stays stable for the duration of a job — which is precisely the shape of large-scale model training and its recurring collective operations.

What Marvell and Lumentum bring to the fabric

Delivering an optical fabric at cluster scale requires two distinct pieces of the stack. Lumentum contributes the optical components — the light sources, switching elements, and photonic building blocks that physically route and preserve signal integrity across a circuit. Marvell contributes the connectivity silicon and interconnect technology that terminates those optical links at the accelerators and integrates the fabric with the rest of the system. A live demonstration matters here because optical switching has long been credible in principle; showing the pieces working together as an AI fabric is the step that moves it from research toward something an operator can plan around.

Practical questions before adopting OCS

Teams evaluating an optical circuit switching fabric should reason about fit rather than raw speed. Useful questions include:

  • Is the workload's communication pattern stable enough that slower circuit setup is amortized over long-running jobs?
  • How does the scheduler map job topology onto available circuits, and what happens when a job needs to reconfigure mid-run?
  • Where does packet switching still belong — for control traffic, storage, or bursty east-west flows that OCS handles poorly?
  • What are the operational implications of holding light paths open, including failure recovery and how a broken circuit reroutes?

The likely near-term outcome is not that optical circuit switching replaces packet switching outright, but that large clusters become hybrids: OCS carries the heavy, predictable collective traffic to stay under the power wall, while packet switches continue to handle everything that needs per-flow flexibility.

Automate Your Content with AI Video Generator

Try it Free →