FlashAttention 1–4: How IO-Awareness Reshaped the Attention Kernel
Prerequisite: GPU 与 Triton 入门 (in Chinese) covers the GPU memory hierarchy and Triton basics; 谁偷走了 5090 的算力和显存 (in Chinese) covers FLOPs, arithmetic intensity, the roofline model and peak memory, and measures on an RTX 5090 why the $N\times N$ score matrix dominates both the runtime and the activation memory of standard attention — the two problems this post starts from.
评论
?
参与讨论