Is RX 6800 + 6800 XT a sensible upgrade from 2x RTX 2060 OC 12GB for llama.cpp?

I’m currently running llama.cpp on two RTX 2060 12GB cards, so 24GB total VRAM. With Qwen3.8 27B IQ4_XS at 131k context I’m getting around 45 tok/s, which is actually pretty good for this setup. I found a deal on an RX 6800 16GB and an RX 6800 XT 16GB, so I’d be going from 24GB to 32GB total VRAM. On paper the AMD cards are obviously much stronger and have higher memory bandwidth, but I’m not sure how well that translates to llama.cpp, especially in a mixed multi-GPU AMD setup. Would this actually be a meaningful upgrade, or would I mostly just be gaining more VRAM and the ability to run higher quants / longer context? I’m also curious how good ROCm or Vulkan is these days on the 6800 series for llama.cpp, because Ive seen pretty mixed reports. If anyone here is running a 6800 / 6800 XT setup, especially with Qwen3.8 27B, I’d love to hear what kind of tok/s you’re getting and whether youd consider it worth switching from CUDA.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论