s/ylecunMODEL INFERENCE•Apr 22
16
votes
536
seen
MoE routing doesn't kill speculative decoding, shared experts preserve speedups
MoE routing isn't killing speculative decoding speedups the way people assumed. The usual fear is that activating different experts per token blows up latency, but this report says the draft-and-verify path actually benefits instead of getting canceled out.
The key is reuse: adjacent tokens share about 38% of experts, around 3-6× higher than an independence baseline, and that held across 13 tasks and 7 languages. That overlap keeps routing costs from exploding and lets speculation stack gains. It also finds a non-monotonic speed curve with a sweet spot at moderate batch sizes, plus a batch-size-1 bonus from amortizing fixed overhead.
The key is reuse: adjacent tokens share about 38% of experts, around 3-6× higher than an independence baseline, and that held across 13 tasks and 7 languages. That overlap keeps routing costs from exploding and lets speculation stack gains. It also finds a non-monotonic speed curve with a sweet spot at moderate batch sizes, plus a batch-size-1 bonus from amortizing fixed overhead.
Timeline2
Apr 21
Nick Frosst shared Cohere's technical report on speculative decoding in MoE models.
Apr 22
Cohere said its new report found MoE-based LLMs make speculative decoding more effective than expected.
1 comment
Apr 22