Deepseek V4 Flash on a single RTX 6000 Pro - vLLM-Moet
Wow... github.com/kacper-daftcode/vLLM-Moet Using this customized vllm provided as a docker, I'm able to run DS V4 Flash on a single RTX 6000 Pro (apparently it also works on a single 5090 - check his readme, but I haven't tried). Apparently this also works with GLM 5.2 (though you need at least two 6000 pros, which is still amazing). Setting 130K context, I needed around 150 GB of RAM to get past the safetensor sharding, but once it is fully loaded in VRAM I am able to fit it all in the GPU (If you
评论
?
参与讨论