1x32GB V100 vs 2x16GB V100 vs 5060ti 16GB for QWEN 3.8

Hi All, I am currently contemplating an upgrade from my 5060ti 16gb. I am getting ~40t/s with 130k context on Qwen 3.8 IQ3_S HF quant. I am running llama.cpp on linux. Objective is to increase context and use a better quant and also free up 5060 for other tasks. The options I am considering are 1x32GB V100 and 2x16GB V100. Theoretically, 2x16GB should be superior to 5060 and 1x32gb. One issue I have to deal with is that I am limited in terms of CPU to GPU comms - I only have 2x x4 lines available. Any other good options in the same price range?

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论