The Life of an LLM Inference — A Prompt's 28 Stops Inside llama.cpp
After you press enter on "hello, llama", those 5 tokens travel 28 stops inside llama.cpp before they come back as an answer — the tokenizer splits bytes into BPE ids, the embedding table fetches 4096-dim vectors, the KV cache decides the memory ceiling of the whole run, attention fuses Q·Kᵀ through FlashAttention's online softmax into a single SRAM kernel, the MoE router lights up only 8 of 256 experts, MLA collapses the entire KV table into latent form, quantisation squeezes a 16 GB model into 5 GB, speculative decoding shaves 60% off decode time, continuous batching keeps the GPU busy between requests, TP/PP/EP splits 405B across 8 cards, GBNF forces output to obey a JSON schema, the vision encoder turns an image into 256 tokens, reasoning models "*think*" 10000 internal tokens, Blackwell fp4 fits 405B into a single node, and finally SSE pushes every character one frame at a time to the user's screen. Every stop maps to a real file and a real function in llama.cpp / vLLM, read line by line.