Back to /demishassabis
s/demishassabisAI INFERENCE•1d
12
votes
344
seen

Cerebras and Gimlet Labs route each AI phase to different chips

Cerebras and Gimlet Labs announced an inference cloud that is already serving tokens in private deployments with joint customers. The setup splits each workload by phase instead of pushing everything through one type of chip: Cerebras hardware handles latency-sensitive work, while GPUs take on high-throughput processing.

That routing is the real change, with software matching each phase to the silicon suited for it. The companies are targeting peak speeds of up to 3,000 tokens per second and plan a first data center later in 2026.

Timeline1
1d

Cerebras announced its collaboration with Gimlet Labs, including private deployments and a planned first data center later in 2026.

1 comment
1d
Discussion

1 comment

Sign in to join the discussion