Home / Blog / GEM Training: How Meta Doubled the Efficiency of Its…
Engineering

GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model

**Meta**’s **Generative Ads Recommendation Model (GEM)** is the foundation model behind ads recommendations on **Instagram** and **Facebook**. Meta…

By Dillip Chowdary • Aug 03, 2026 • Source: Meta Engineering

GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model

**Meta**’s **Generative Ads Recommendation Model (GEM)** is the foundation model behind ads recommendations on **Instagram** and **Facebook**. Meta Engineering reports that GEM now trains at **LLM scale** on several thousand latest-generation GPUs, with end-to-end training efficiency doubled to **20–25% Model FLOPs Utilization (MFU)** while training FLOPs scaled **4x**.

The operational claim is joint: more compute and better use of that compute. **MFU** measures how much of theoretical model FLOPs are realized in training. Reaching **20–25% MFU** at multi-thousand-GPU, LLM-scale ads training means the stack is converting a larger share of peak FLOPs into useful work rather than leaving capacity idle on communication, pipeline bubbles, or underutilized kernels.

Advertisement

Tech Pulse Daily

Get tomorrow's pulse first

Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.

For engineers building large recommendation or foundation-model systems, the numbers matter more than the product label. Ads ranking at this scale is not a small-model fine-tune: it is multi-thousand-GPU training with the same efficiency pressure as LLM pretraining. Doubling E2E efficiency while increasing training FLOPs **4x** is a concrete capacity lever—more model scale or more experiment throughput without a proportional jump in cluster hours.

In market terms, this sits inside Meta’s core ads stack on **Instagram** and **Facebook**, where recommendation quality is revenue-critical. An LLM-scale ads foundation model trained more efficiently is a competitive training-system story as much as a modeling one: whoever runs larger recommendation models at higher MFU can iterate faster or train richer models on the same hardware budget.

Practical takeaway: treat **MFU** and E2E training efficiency as first-class metrics when scaling recsys foundation models, not only peak GPU count. Watch whether Meta publishes more on the systems work behind the **2x** efficiency gain and the **4x** FLOPs scale-up—those details are what other teams can reuse when moving ads or ranking models into the multi-thousand-GPU regime.

Advertisement

🔎 More interesting news

5-min tech signal

Weekday briefing for engineers who skip the noise.

No spam · Unsubscribe anytime

Advertisement

✈️ CareerPilot

Your AI job-search copilot

Match your resume against live Ashby, Greenhouse & Lever openings — fit scores, job-specific resume optimization and email alerts.

Find matching jobs →

Free Tools

Browse all tools →