Marvell and Google announce multi-year pact for custom AI silicon targeting Gemini 3.5. Analysis of the 50x performance-per-watt leap. Read more.

What the partnership covers

Marvell and Google have announced a multi-year agreement to build custom AI silicon for Google’s Gemini line, with work aimed at the Gemini 3.5 generation. The chips are planned on a 5nm process and are not off-the-shelf accelerators rebranded for the cloud. They are purpose-built for Google’s models, data movement patterns, and power envelopes in its own data centers.

Custom silicon matters when you already know the workload shape. Google can fix the model architecture, serving topology, and memory hierarchy first, then ask Marvell to implement the math, interconnect, and packaging around those constraints. That is a different problem from shipping a general-purpose GPU that must run every framework and every model family reasonably well.

Reading the 50x performance-per-watt claim

The partnership’s headline claim is a large jump in performance per watt relative to prior generations of hardware used for the same class of work. Treat that figure as a system-level ratio, not a single FLOPS number on a slide. Performance per watt usually folds in effective throughput on the target model, sustained power under realistic batching, and how much energy is spent moving data rather than computing.

When you evaluate claims like this, separate three layers: peak silicon efficiency, rack-level efficiency after networking and cooling, and end-to-end tokens or inferences per joule in production. A 50x leap is only meaningful if the baseline, workload, and measurement boundary are the same. For operators outside Google, the useful takeaway is directional: model-specific accelerators can beat general silicon on energy cost when utilization is high and the software stack is co-designed with the hardware.

  • Ask what baseline the comparison uses (previous custom chip, commercial GPU, or CPU path).
  • Ask whether the metric is measured on Gemini 3.5 training, inference, or both.
  • Ask how much of the gain comes from process (5nm) versus architecture and compiler work.

Why multi-year custom AI silicon is hard

A multi-year pact exists because custom AI silicon is a long feedback loop. Specs freeze early, process nodes have long lead times, and compilers, kernels, and serving stacks must mature on silicon that does not exist yet. Miss the model shift and you ship an efficient chip for last year’s architecture.

Marvell’s role is typically the full stack around the compute die: high-speed SerDes, packaging, power delivery, and the interfaces that let racks scale without the interconnect becoming the bottleneck. Google’s role is the model roadmap, the runtime, and the fleet constraints. Success depends less on a single clever block and more on whether those two roadmaps stay aligned as Gemini evolves.

Practical implications for teams watching the space

If you buy cloud inference rather than design chips, this kind of deal is a signal about where large model providers will put capital: fewer generic FLOPS, more specialized paths for their flagship models. That can improve cost and latency for Gemini users over time, while leaving multi-model platforms more dependent on flexible accelerators.

If you build systems or procure hardware, use the announcement as a checklist, not a product roadmap. Prefer vendors who document power at the rack, not only the die. Plan capacity around energy and cooling, not peak TOPS alone. Keep software portable enough that a custom path for one model family does not lock your entire stack. Custom 5nm AI silicon is a bet that co-design beats generality for a known workload; your job is to decide whether your workload is stable and large enough to justify the same tradeoff.

Automate Your Content with AI Video Generator

Try it Free →