Pushing memory bound kernels beyond the speed of light with lossless decompression

The fastest a memory bound kernel can go is set by the time required to transfer the data to the SMs. How can we do better?

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论