Back to /ylecun
s/ylecunMODEL INFERENCE•Apr 22
16
votes
536
seen

MoE routing doesn't kill speculative decoding, shared experts preserve speedups

MoE routing isn't killing speculative decoding speedups the way people assumed. The usual fear is that activating different experts per token blows up latency, but this report says the draft-and-verify path actually benefits instead of getting canceled out.

The key is reuse: adjacent tokens share about 38% of experts, around 3-6× higher than an independence baseline, and that held across 13 tasks and 7 languages. That overlap keeps routing costs from exploding and lets speculation stack gains. It also finds a non-monotonic speed curve with a sweet spot at moderate batch sizes, plus a batch-size-1 bonus from amortizing fixed overhead.

Timeline2
Apr 21

Nick Frosst shared Cohere's technical report on speculative decoding in MoE models.

Apr 22

Cohere said its new report found MoE-based LLMs make speculative decoding more effective than expected.

1 comment
Apr 22
Discussion

1 comment

Sign in to join the discussion