Home / Blog / GEM Training: How Meta Doubled the Efficiency of Its…
Engineering

GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model

Meta’s Generative Ads Recommendation Model (GEM) is the foundation model behind ads recommendations on Instagram and Facebook. Meta…

By Dillip Chowdary • Aug 03, 2026 • Source: Meta Engineering

GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model

Meta’s Generative Ads Recommendation Model (GEM) is the foundation model behind ads recommendations on Instagram and Facebook. Meta Engineering reports that GEM now trains at LLM scale on several thousand latest-generation GPUs, with end-to-end training efficiency doubled to 20–25% Model FLOPs Utilization (MFU) while training FLOPs scaled 4x.

The operational claim is joint: more compute and better use of that compute. MFU measures how much of theoretical model FLOPs are realized in training. Reaching 20–25% MFU at multi-thousand-GPU, LLM-scale ads training means the stack is converting a larger share of peak FLOPs into useful work rather than leaving capacity idle on communication, pipeline bubbles, or underutilized kernels.

Advertisement

Tech Pulse Daily

Get tomorrow's pulse first

Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.

For engineers building large recommendation or foundation-model systems, the numbers matter more than the product label. Ads ranking at this scale is not a small-model fine-tune: it is multi-thousand-GPU training with the same efficiency pressure as LLM pretraining. Doubling E2E efficiency while increasing training FLOPs 4x is a concrete capacity lever—more model scale or more experiment throughput without a proportional jump in cluster hours.

In market terms, this sits inside Meta’s core ads stack on Instagram and Facebook, where recommendation quality is revenue-critical. An LLM-scale ads foundation model trained more efficiently is a competitive training-system story as much as a modeling one: whoever runs larger recommendation models at higher MFU can iterate faster or train richer models on the same hardware budget.

Practical takeaway: treat MFU and E2E training efficiency as first-class metrics when scaling recsys foundation models, not only peak GPU count. Watch whether Meta publishes more on the systems work behind the 2x efficiency gain and the 4x FLOPs scale-up—those details are what other teams can reuse when moving ads or ranking models into the multi-thousand-GPU regime.

Advertisement

🔎 More interesting news

5-min tech signal

Weekday briefing for engineers who skip the noise.

No spam · Unsubscribe anytime

Advertisement

✈️ CareerPilot

Your AI job-search copilot

Match your resume against live Ashby, Greenhouse & Lever openings — fit scores, job-specific resume optimization and email alerts.

Find matching jobs →

Free Tools

Browse all tools →