Production Qwen 3.6-27B VLLM config?
Hi everyone, I've spent about four days now trying to find the eight configuration for running Qwen 27b in production using VLLM but have been getting significant decreases in performance with the FP8 safetensors in comparison to the llama.cpp variants, especially under load/high concurrency. Our quality benchmarks drop by almost 20% in absolute terms. Would some mind sharing their VLLM config they use in production that execute correctly? I'm at a loss at this point :) Thanks in advance! (Update) Current c
评论
?
参与讨论