4 GPUs (MI50) llama.cpp or vLLM?

Hi, I've been running vLLM on my MI50 because of tensor-parallel support. It works, but I have some complaints. For one, the quants seem much harder to find than GGUFs. Also, model switching is a pain, especially with the insanely long startup time. I recently discovered that llama.cpp supports tensor-parallel (I've been away for a long while). Is there any reason I should stick with vLLM? submitted by /u/FrozenAptPea [link] [comments]

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论