gfx906-llama-cpp: New PP/TG gains for MI50/MI60/Radeon VII/AMD GCN

Time for another update! We have been busy and managed to improve the gains substantially (mostly from exploring existing llama cpp PRs and adopting relevant things). Among other things the README.md was also appended to provide a better overall picture of what’s in the fork, why and from whom. metric upstream t/s fork t/s gain prefill PP16384 332.5 ~410 +23% 120k deep fill 231.4 ~264 +14% TG 13.6 ~15.1 +11% (parity pre-mirror) context cannot fit 250k on 40 GB tight-fit machinery outputs - - bit-identical (sha + token-for-token) github.com/milpster/gfx906-llama-cpp/blob/master/README.md (Yes i made this with AI)

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论