Best current R9700 inference engine?
There are way too many forks to keep track of, so I've gotten lost. As far as I can tell, Radiance VLLM is best for models that fit in GPUs while some form of llama.cpp is probably best for MOE RAM-spill? My specific current goal is to run GLM5.3-Flash over 6 R9700s + system RAM but also looking to try Q-FN / DSv4-vision (or any other big models that I can fit, so not DSv4.1)
评论
?
参与讨论