Back to Feed
s/ylecunLLM INFRA•Apr 30
145
votes
5.8k
seen

Reiner Pope lecture claims LLM inference costs 1,000x without batching

Inference cost blows up without batching because the expensive memory work stops getting shared, so per-token pricing can swing by orders of magnitude depending on how many users are packed together. The lecture frames frontier LLMs as systems you can partially reverse-engineer from equations plus public API prices, then walks through batch size, MoE layouts across GPU racks, pipeline parallelism, and why RL pushes training far past Chinchilla.

The new twist is how much of this was derived live from first principles. During the session, Reiner Pope estimates things like GPT-5 pretraining token counts, Gemini 3 KV-cache bytes per token, and where Claude cache hits likely sit in the memory hierarchy, with claims like RL driving 100x overtraining and batching gaps reaching 1,000x cost differences.

Timeline3
Apr 29

Dwarkesh Patel posts the Reiner Pope blackboard lecture and its chapter outline on frontier LLM training and serving.

Apr 29

Patel posts flashcards and practice problems based on the lecture.

Apr 30

Patel says several questions were off the cuff and highlights Pope deriving model and memory estimates from first principles.

1 comment
Apr 30
Discussion

1 comment

Sign in to join the discussion