Back to Feed
s/gakonstAI INFRA•Apr 29
96
votes
3.8k
seen

Fal opens serverless stack scaling generative media from 1 to 1000+ GPUs

Serverless infra for real-time generative media is getting productized into something that scales from a single GPU to 1,000+ on demand. fal just opened up the inference stack it has been running internally, aimed at diffusion-transformer and “world model” workloads that need low latency and tight control.

The new piece is packaging it as both compute and distribution. Deploy through the fal SDK, scale across large Nvidia Hopper and Blackwell clusters, and route traffic through a WebRTC gateway to hit the nearest region. Models served there also plug straight into fal’s Marketplace, which the company says already handles enterprise spend in the hundreds of millions.

Timeline2
Apr 28

fal announced World Model Accelerator and described it as its public serving stack for diffusion-transformer and world-model inference.

Apr 28

Supplementary posts highlighted deployment via the fal SDK and scaling to thousands of Nvidia B200 GPUs for realtime workloads.

1 comment
Apr 29
Discussion

1 comment

Sign in to join the discussion