FlashAttention 1–4: How IO-Awareness Reshaped the Attention Kernel

Prerequisite: GPU 与 Triton 入门 (in Chinese) covers the GPU memory hierarchy and Triton basics; 谁偷走了 5090 的算力和显存 (in Chinese) covers FLOPs, arithmetic intensity, the roofline model and peak memory, and measures on an RTX 5090 why the $N\times N$ score matrix dominates both the runtime and the activation memory of standard attention — the two problems this post starts from.
添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论