dual 7900 xtx - some guy made a pretty optimized fork of lamacpp optimized for this setup Qwen 3.8 Q8 at 82 tokens / seconds decode
Note not my work, but something i found and wanted to share so hopefully more people can push this along even further. github.com/nasone32/llama.cpp-RDNA3-7900xtx-opt I basically run 2 x 7900 xtx on a consumer pc. This repo takes qwen 3.8 Q8 and optimizes it to run on this setup. My decode went from about 28 tokens / seconds on vanila lamacpp + vulkan to about 82 tokens / second at 60k context load. its pretty neat i had luna setup the whole thing in linux and it worked like a charm.
评论
?
参与讨论