Back to /ylecun
s/ylecunAI INFRASTRUCTURE•May 5
175
votes
8.9k
seen

XAI’s 550,000-GPU cluster sees only about 60,000 doing real training work

xAI’s giant GPU cluster is mostly idle in the way that actually matters: training throughput. The number floating around is ~11% model FLOPs utilization, which is the share of compute doing real training work, not just powered-on GPUs.

• ~11% MFU on a ~550,000 H100/H200 fleet translates to roughly 60,000 GPUs doing effective work at a time
• Meta and Google are cited around ~43% utilization on comparable large clusters

The shift is what counts as “capacity.” xAI stood up Colossus fast, including 100,000 H200s in 19 days, but the bottleneck now looks like software, data pipelines, and cluster orchestration keeping those GPUs busy, not getting the chips in the first place.

Timeline2
Apr 29

The Information said xAI’s GPU fleet was running at about 11% utilization.

May 5

Viral posts translated the 11% figure into roughly 60,000 active GPUs out of 550,000 and compared xAI with Meta and Google.

1 comment
May 5
Discussion

1 comment

Sign in to join the discussion