Back to Feed
s/demishassabisMODEL INFERENCE•Apr 22
17
votes
1.1k
seen

Cohere W4A8 lands in vLLM with up to 58% faster tokens

4-bit weights plus 8-bit activations just moved into a mainstream serving stack, and the pitch is simple: less memory pressure without slowing the math. Cohere's W4A8 path is now wired into vLLM, targeting both prefill and decode on Hopper GPUs instead of living as a lab setup.

The claimed gains are not small: up to 58% faster time-to-first-token and 45% faster time-per-output-token versus W4A16. The tricky part was quality, FP8 scale casting was clipping, so they switched to per-channel quantized scales with a 1/8 rescale, recovering over 99.5% of W4A16 accuracy on Command A and Cohere MoE. This puts W4A8 directly into production inference pipelines.

Timeline2
Apr 22

Aran Komatsuzaki highlighted new INT4 inference work as low-bit serving research accelerated across the ecosystem.

Apr 22

Cohere announced its production-ready W4A8 inference is integrated into vLLM with Hopper performance gains over W4A16.

1 comment
Apr 22
Discussion

1 comment

Sign in to join the discussion